A smart management system for smart medical treatment and a management method thereof
By integrating genetic, environmental, and lifestyle data through an intelligent management system, a disease association network is constructed and dynamically assessed. This solves the problem of capturing the interaction of multiple factors in existing technologies, enabling personalized disease risk assessment and intervention recommendations, and improving the scientific rigor and real-time nature of the assessment.
Patent Information
- Application Number
- CN202511349325.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-09-22
AI Technical Summary
Existing disease risk assessment technologies struggle to fully capture the interactions of multi-dimensional factors and lack a unified data processing framework, leading to distorted risk assessment results that fail to meet personalized needs and universal requirements. Furthermore, the lack of dynamic update capabilities limits the application of precision medicine.
The system employs an intelligent management system to acquire genetic, environmental, and lifestyle data through a data collection module. It then constructs a disease association network, processes the data, and partitions it into risk categories. AI algorithms are used for data cleaning, standardization, and interaction analysis to generate individual disease risk assessment results and intervention recommendations.
It enables multi-dimensional and dynamic assessment of disease risk, improves the scientific rigor and interpretability of the assessment, supports the development of personalized intervention strategies, and meets the needs of precision medicine.
Smart Images

Figure CN120853952B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical systems, and more particularly to a smart management system and its management method for smart healthcare. Background Technology
[0002] In the field of disease risk assessment, traditional methods have long relied on independent analysis of single-dimensional factors, making it difficult to comprehensively capture the multi-dimensional mechanisms of disease development. For example, existing technologies often use linear superposition to assess single-factor risk, ignoring the amplifying or inhibiting effects of non-linear interactions such as gene-environment, gene-lifestyle, and environment-lifestyle on disease risk. This simplification can lead to distorted risk assessment results, failing to accurately reflect an individual's actual health risk under complex exposure backgrounds. Existing technologies face significant difficulties in integrating the three categories of risk factors: genes, environment, and lifestyle. The sources, formats, and quantification standards of genetic data, environmental exposure data, and lifestyle data vary greatly, lacking a unified standardized processing framework. Furthermore, the raw data often contains redundant records, mapping errors, or missing values, further exacerbating the difficulty of data quality control. The lack of medical knowledge mapping also prevents the accurate interpretation of the biological significance of risk factors, limiting the scientific rigor and interpretability of the model.
[0003] Traditional risk assessment models are mostly based on static data for prediction, making it difficult to capture the dynamic evolution of risk factors. For example, the immediate impact of sudden environmental exposure, behavioral adjustments, or gene expression fluctuations on disease risk is not effectively incorporated into the analytical framework of existing technologies. Furthermore, interaction modeling is limited to single-factor or pairwise interactions, lacking a systematic quantification of the synergistic effects of genes, environment, and lifestyle habits, leading to inaccurate identification of key risk pathways. In addition, model outputs are mostly static risk probabilities, without clearly defined intervention targets, limiting the operability of clinical decision-making. Existing disease risk assessment technologies struggle to balance personalized needs with universality requirements. On the one hand, most models are designed based on average population levels, failing to fully consider the uniqueness of individual gene-environment-lifestyle combinations; on the other hand, models for specific diseases lack a unified assessment framework across disease types, limiting the technology's broad applicability. Moreover, insufficient model interpretability and weak dynamic update capabilities further hinder the application of these technologies in precision medicine. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention discloses a smart management system and its management method for smart healthcare that integrates three types of risk factors—genes, environment, and lifestyle habits—and analyzes their interactions through dynamic modeling.
[0005] This invention discloses a smart management system for smart healthcare, comprising:
[0006] The data acquisition module is used to collect an individual's genetic data, environmental exposure data, and lifestyle data.
[0007] The disease association network construction module is used to construct a disease association network based on genetic data, environmental exposure data, and lifestyle data, which includes the associations between nodes of three types of risk factors.
[0008] The data processing module processes the factor nodes and calculates the disease type corresponding to each processed node to generate the risk probability of different diseases.
[0009] The risk partitioning module is used to partition the disease association network according to the type of risk factor and output multiple risk subnetworks, where each risk subnetwork corresponds to the same type of risk factor;
[0010] The risk assessment module is used to generate individual disease risk assessment results and intervention recommendations based on multiple risk subnetworks;
[0011] The disease association network construction module is connected to the data acquisition module, the data processing module is connected to the disease association network construction module, the risk zoning module is connected to the data processing module, and the risk assessment module is connected to the risk zoning module.
[0012] Furthermore, the data processing module includes preprocessing, medical knowledge mapping, and standardization.
[0013] Preprocessing:
[0014] Data cleaning was performed on the collected genetic data, environmental exposure data, and lifestyle data to eliminate mapping errors or redundancy;
[0015] Medical knowledge mapping:
[0016] Gene loci are mapped to disease-related genes in OMIM or GWASCatalog, and different gene risk weights for different diseases are obtained based on the disease-related genes.
[0017] Environmental factors are linked to the list of pathogenic factors in the WHO's "Environment and Health Guidelines," and environmental exposure indices for different diseases are obtained based on the data in the list of pathogenic factors.
[0018] Lifestyle habits are converted into ICD-10 behavioral risk factor codes, and the hazard ratio (HR) values for different diseases are obtained based on these codes.
[0019] Standardization process:
[0020] This is used to normalize the risk weights of genes, environmental exposure indices, and lifestyle risk HR values, and to determine the independent risk contribution of each factor to a specific disease.
[0021] Furthermore, the risk subnetwork generated by the risk partitioning module when partitioning by risk factor type includes a genetic subnetwork, an environmental subnetwork, and a lifestyle subnetwork.
[0022] When partitioning by risk factor type, the risk partitioning module retains cross-type association edges between gene subnetwork, environment subnetwork, and lifestyle subnetwork. These cross-type association edges include gene-environment interaction edges, gene-lifestyle interaction edges, and environment-lifestyle interaction edges.
[0023] Furthermore, the risk assessment module is used to convert the independent risk contribution values of multiple risk sub-networks into standardized probabilities respectively;
[0024] The risk assessment module also includes an interaction risk calculation submodule, used to generate risk probabilities and the highest-risk path:
[0025] It is used to analyze the interactions between genes and environment, genes and lifestyle, and environment and lifestyle, and calculates the joint risk contribution value after the interaction based on the preset synergistic and antagonistic rules.
[0026] For the synergistic effect of the three factors, the highest risk combination among the pairwise interactions should be calculated first, and then the influence of the third factor should be added.
[0027] All possible risk transmission paths are ranked according to the rules that the risk of interaction is greater than the risk of a single factor and the path supported by high scientific evidence is greater than the path supported by low scientific evidence.
[0028] The highest-risk path is selected based on the highest risk contribution value and the biological mechanism. The risk contribution value of the highest-risk path is converted into a standardized probability and combined with the baseline risk of the population to generate the individual absolute risk probability.
[0029] Output the absolute risk probability of an individual and the corresponding key risk factors and their interaction combinations.
[0030] Furthermore, the data types of genetic data, environmental exposure data, and lifestyle data are as follows:
[0031] Genetic data includes at least one of the following: single nucleotide polymorphisms, gene expression profiles, and epigenetic data;
[0032] Environmental exposure data includes:
[0033] Physical environment: at least one of PM2.5 concentration, noise level, and history of heavy metal exposure;
[0034] Chemical environment: Indoor pollutant exposure levels of at least one of formaldehyde and benzene;
[0035] Lifestyle data includes:
[0036] Behavioral data: average daily smoking amount, alcohol consumption and frequency, exercise duration, and sleep quality score;
[0037] Dietary data: frequency of nutrient intake, dietary structure score.
[0038] Furthermore, the intelligent management system also includes:
[0039] The health data monitoring module is used to monitor the health indicator data corresponding to each sub-network in multiple risk sub-networks in real time, and obtain a multi-dimensional health dataset.
[0040] The risk source identification module is used to trace risks based on multi-dimensional health datasets and detect the initial risk source corresponding to each risk sub-network.
[0041] The risk propagation network construction module is used to build a corresponding risk propagation network based on each initial risk source. The network includes the risk source, the nodes affected by the risk propagation, and the propagation path. The risk propagation network construction module allows the risk source to propagate across sub-networks.
[0042] The critical risk path acquisition module is used to analyze the risk propagation network through the network center analysis model and obtain the critical risk path corresponding to each risk sub-network.
[0043] The risk assessment module is used to generate individual disease risk assessment results and intervention recommendations based on key risk pathways;
[0044] The health data monitoring module is connected to the risk assessment module, the risk source identification module is connected to the health data monitoring module, the risk propagation network construction module is connected to the risk source identification module, and the key risk path acquisition module is connected to the risk propagation network construction module.
[0045] Furthermore, the health data monitoring module is used for real-time monitoring:
[0046] Gene subnetwork: monitoring gene expression levels related to environmental factors and lifestyle habits;
[0047] Environmental subnetwork: Monitors environmental exposure indicators affected by lifestyle habits;
[0048] Lifestyle subnetwork: monitoring behavioral data related to genes and environment.
[0049] Furthermore, the risk source identification module includes:
[0050] The hybrid risk source detection unit is used to determine whether a risk is caused by two or three types of factors in the gene-environment-lifestyle combination. The judgment criteria include the spatiotemporal correlation or interaction strength threshold of cross-sub-network data.
[0051] For the identified mixed risk sources, a joint risk label containing multiple types of factors is generated and associated with its corresponding disease type.
[0052] This invention discloses a smart management method for smart healthcare, which uses any of the smart management systems for smart healthcare described above, including:
[0053] S1: Data collection and standardization processing, acquiring genetic data, environmental exposure data, and lifestyle data of the target individual. Genetic data includes gene locus information related to the target disease, environmental exposure data covers physical, chemical, and biological environmental factors, and lifestyle data includes records of diet, exercise, rest, and medical behavior. The three types of data are standardized to eliminate differences in data dimensions and establish a unified data format.
[0054] S2: Multi-source risk feature extraction. Based on biological mechanisms and epidemiological evidence, it extracts susceptibility site features from gene data, key pollutant features from environmental exposure data, and high-risk behavioral features from lifestyle data to form three sets of risk features, and marks the association strength between each feature and the target disease.
[0055] S3: Dynamic interaction modeling, constructing a gene-environment-lifestyle ternary interaction model, quantifying the synergistic effects of gene-environment interaction, gene-lifestyle interaction and environment-lifestyle interaction through nonlinear algorithms, generating dynamic risk transmission paths, and simulating the cumulative impact of different combinations of risk factors on the probability of occurrence of target diseases;
[0056] S4: Risk propagation network construction, which integrates three types of risk feature sets with dynamic risk propagation paths into a multi-dimensional risk propagation network, where nodes correspond to risk factors, edges represent the strength and direction of the interaction between risk factors, and network weights are determined by biological pathway validation and statistical significance.
[0057] S5: Screening of key risk paths. Based on the topological characteristics of the risk propagation network, the centrality analysis and sensitivity analysis methods are used to identify the key paths that contribute the most to the risk of the target disease, eliminate non-significantly associated paths, and form a simplified risk-driven sub-network.
[0058] S6: Personalized risk assessment. Standardized data of the target individual is input into the risk-driven sub-network. Combined with the output of the dynamic interaction model, the disease risk score of the target individual within a specific time window is calculated, and a risk assessment report containing risk level, key driving factors and potential intervention targets is generated.
[0059] S7: Results output and update maintenance. Output the risk assessment report in a visual form, including risk trend charts, heat maps of key risk factors, and a list of intervention recommendations; regularly update the risk transmission network and dynamic interaction model based on newly acquired genomic, environmental, or lifestyle data to achieve dynamic iterative optimization of risk assessment results.
[0060] The beneficial effects of this invention are:
[0061] The beneficial effects of this invention are as follows: It comprehensively captures the biological basis, external exposure, and behavioral driving factors of disease occurrence through integrated data collection of genes, environment, and lifestyle habits; it transforms risk factors into a structured association network, intuitively displaying the causal relationships or synergistic pathways between factors, providing visual support for disease analysis and risk assessment; through a three-level processing flow of data cleaning, medical mapping, and standardization, it solves the problems of quality control, semantic interpretation, and quantitative calculation of multi-source heterogeneous data, ensuring that subsequent modules output scientifically reliable results; it divides the disease association network into sub-networks according to the type of risk factor and retains cross-type association edges, achieving hierarchical management of risk factors while capturing the synergistic or antagonistic relationships between different types of factors, breaking through the limitations of single-factor independent assessment, and truly reflecting the complex formation mechanism of individual disease risk, providing support for accurately identifying key risk pathways and formulating multi-dimensional intervention strategies; the risk assessment module will integrate independent risks... Contribution values are converted into standardized probabilities, enabling unified measurement and comparable analysis of different types of risk factors. Interactive influences are quantified through pre-defined rules, and the most biologically plausible and risk-contributing transmission paths are selected. Individual absolute risk probabilities are generated by combining baseline population risk, outputting accurate assessment results including key factors and interaction combinations. The dynamic module upgrades from periodic static assessment to real-time dynamic monitoring through real-time monitoring, risk tracing, dynamic modeling, and path analysis. It captures the instantaneous changes and evolutionary logic of risk, supports closed-loop feedback between risk and intervention, and improves the timeliness and targeting of interventions. The static and dynamic modules work together to form a continuous spatiotemporal assessment of baseline prediction plus real-time early warning, covering the entire health to disease cycle. Real-time data is used to verify and correct static models, improving the accuracy of risk assessment in complex scenarios. This meets the requirements of the medical field for the scientific rigor and interpretability of risk assessment, while also adapting to the real-time and personalized needs of precision medicine. Attached Figure Description
[0062] Figure 1 This is a flowchart illustrating the system module connection of a smart management system for smart healthcare according to an embodiment of this application. Detailed Implementation
[0063] To enable those skilled in the art to better understand the present invention, the technical solutions in the specific embodiments of the present invention will be clearly and completely described below.
[0064] This invention discloses a smart management system for smart healthcare, comprising: a data acquisition module for collecting individual genetic data, environmental exposure data, and lifestyle data; a disease association network construction module for constructing a disease association network based on genetic data, environmental exposure data, and lifestyle data, containing the relationships between nodes of three types of risk factors; a data processing module for processing factor nodes and calculating the disease type corresponding to each processed node to generate the risk probability of different diseases; a risk partitioning module for partitioning the disease association network according to risk factor type, outputting multiple risk subnetworks, where each risk subnetwork corresponds to the same type of risk factor; and a risk assessment module for generating individual disease risk assessment results and intervention suggestions based on multiple risk subnetworks. Risk factor type refers to the essential attribute category of the risk factor, used to distinguish whether the source of risk is genetic inheritance, environmental exposure, or lifestyle habits. Classification structurates complex risk factors, facilitating subsequent independent analysis or interaction studies of different types of factors. The classification is based on the biological attributes of risk factors—genes, exposure pathways—environment, and behavioral autonomy—lifestyle habits—rather than risk level or disease type. This invention comprehensively captures the biological basis, external exposure, and behavioral driving factors of disease occurrence through a three-pronged data collection approach encompassing genes, environment, and lifestyle habits. It transforms risk factors into a structured network of relationships, visually demonstrating the causal relationships or synergistic pathways between genes, environment, and lifestyle habits, providing visualization support for disease analysis and risk assessment.
[0065] This smart healthcare management system deeply integrates artificial intelligence technology to achieve intelligent upgrades and improved efficiency in disease risk assessment. In the data processing stage, AI algorithms efficiently clean and standardize multi-source heterogeneous data. Machine learning models identify redundant information and mapping errors in genetic, environmental, and lifestyle data, significantly improving the efficiency and accuracy of data preprocessing. During medical knowledge mapping, AI can connect with authoritative databases such as the OMIM online human Mendelian genetics database and the GWASCatalog genome-wide association studies catalog. Through natural language processing technology, it quickly matches gene loci with disease associations, dynamically updating gene risk weights, environmental exposure indices, and lifestyle hazard ratios, making the retrieval of medical knowledge more immediate and comprehensive.
[0066] When constructing disease association networks, AI-driven graph neural networks can deeply mine the potential associations between nodes of three types of risk factors. They not only capture known interactions but also identify new synergistic or antagonistic pathways, making the network structure more closely aligned with complex biological mechanisms. In the risk assessment phase, AI algorithms can optimize interaction quantification models, using deep learning to analyze the ternary synergistic effects of genes, environment, and lifestyle in real time. This accurately identifies the highest-risk pathways and dynamically adjusts risk probability calculations based on real-time monitoring data, making the assessment results more timely and personalized.
[0067] In the dynamic monitoring module, AI's real-time data processing capabilities can efficiently integrate multi-dimensional health indicators. Through time-series analysis, it can quickly locate initial risk sources across sub-networks. By leveraging dynamic graph neural networks to construct a risk propagation network, it can track the propagation path of risks between different sub-networks in real time, providing immediate evidence for the generation of intervention recommendations. Simultaneously, AI can periodically iterate and optimize the risk propagation network and interaction model based on new data, continuously improving the system's assessment accuracy and achieving end-to-end intelligent processing from data collection to intervention recommendations. This better meets the needs of precision medicine for efficient, accurate, and dynamic risk assessment.
[0068] The genetic data of this invention are collected through gene sequencers or gene chips, with priority given to clinically validated testing equipment to ensure data accuracy. Environmental data includes physical environment data collected through real-time monitoring sensors deployed in individual living areas, such as PM2.5 sensors and sound level meters; heavy metal exposure history obtained by detecting biological samples using inductively coupled plasma mass spectrometry; and chemical environment data collected through electrochemical sensors or gas chromatography-mass spectrometry. Lifestyle data includes behavioral data such as daily smoking volume, exercise duration, and sleep quality scores collected in real-time by wearable devices such as smart bracelets and smartwatches, and supplemented by user self-reporting via a mobile app. Dietary data includes nutrient intake frequency and dietary structure scores recorded in a dietary log via an app, and nutrient intake and dietary structure scores are automatically calculated using a food composition database such as the Chinese Food Composition Table.
[0069] In one implementation, the data processing module includes preprocessing, medical knowledge mapping, and standardization. Preprocessing involves cleaning the collected genetic, environmental exposure, and lifestyle data to eliminate mapping errors and redundancy. Medical knowledge mapping maps gene loci to disease-related genes in OMIM or GWASCatalog, deriving different genetic risk weights for different diseases based on these genes. Environmental factors are associated with the pathogenic factor list in the WHO's "Environment and Health Guidelines," deriving different environmental exposure indices for different diseases based on the data in the pathogenic factor list. Lifestyle habits are converted into ICD-10 behavioral risk factor codes, deriving different lifestyle risk hazard values for different diseases based on these codes. Standardization normalizes the genetic risk weights, environmental exposure indices, and lifestyle hazard values to obtain the independent risk contribution of each factor to a specific disease. Preprocessing cleanses the data, removing mapping errors such as mislabeled gene loci and redundant records. Redundant records include duplicate environmental monitoring data and logical inconsistencies, avoiding misjudgments of risk due to junk data. Gene data cleaning removes sequencing error sites, ensuring the accurate detection of gene mutations, etc. For lifestyle habits, cleaning was used to correct for missing values in the sleep quality score (0-10 points within the normal range of 12), avoiding interference from outliers in the model. For missing data, such as weekly PM2.5 monitoring values, domain knowledge was used for imputation, such as interpolation using the average of the same period or data from nearby stations, ensuring no gaps in subsequent analysis. Environmental data imputation used WHO air quality guidelines limits combined with historical data to fill in noise levels. Dietary data was standardized, with daily vegetable intake of 300g and weekly vegetable intake of 2.1kg standardized to a daily average vegetable intake of 300g. Medical knowledge mapping mapped raw data, such as gene loci and lifestyle habits, to authoritative medical databases, giving them biological significance related to diseases, transforming data from numbers into interpretable risk signals.
[0070] Taking an average daily alcohol consumption of 50ml as an example, the type of drinking behavior is first determined based on the amount and frequency of alcohol consumption, such as long-term drinking or occasional drinking. Then, the corresponding category in the ICD-10 code is matched. For example, an average daily alcohol consumption of 50ml for more than one year corresponds to the code Z72.1 alcohol abuse. The correlation between the code and the HR value is based primarily on cohort study data published in the International Agency for Research on Cancer (IARC) Carcinogenicity Assessment Report or the China Chronic Disease and Risk Factor Surveillance Report. For example, the liver disease HR value of 2.0 corresponding to the Z72.1 code comes from the risk ratio of alcohol abuse to non-abuse in a national liver disease cohort study.
[0071] Genetic data mapping, for example, maps rs7903146 to the TCF7L2 gene, a diabetes risk gene included in OMIM, with an OR value of 1.3, generating gene risk weights. Lifestyle mapping converts 50ml of daily alcohol consumption into the ICD-10 code Z72.1, with a HR value of 2.0 associated with liver disease. Standardized risk indicators make risk factors for different diseases comparable. In lung cancer risk, the environmental exposure index for PM2.5 exposure is 0.8. In cardiovascular disease risk, the environmental exposure index for PM2.5 exposure is 0.6. Both can be calculated based on the WHO guidelines' unified exposure-risk model, facilitating cross-sectional assessment of the differences in the impact of PM2.5 on different diseases. Known associations in the database are used to discover implicit risk pathways in user data. Even if a user did not mention a family history of diabetes, a PPARGPro12Ala mutation in their genetic data, mapped to a diabetes association in the GWASCatalog, was automatically included in the risk assessment. Standardization processes normalize the heterogeneous data of genes, environment, and lifestyle into dimensionless values in the 0-1 range. To balance feature contributions, avoid model bias, and prevent features with large numerical ranges from dominating the model, the risk contributions of genes, environment, and lifestyle habits are determined by actual biological significance, rather than data scale. A three-level processing flow—data cleaning, medical mapping, and standardization—solves the problems of quality control, semantic interpretation, and quantitative calculation for multi-source heterogeneous data. This is the transitional stage from raw data collection to accurate risk assessment, ensuring that subsequent modules such as risk zoning and interactive analysis can output scientific and reliable results. For example, gene mutation risk sources spread to the environmental subnetwork through the gene-enzyme activity-environmental pollutant metabolism pathway, or lifestyle risk sources spread to the gene subnetwork through the behavior-physiological indicator-gene expression pathway. For cross-subnetwork propagation pathways, multi-factor synergistic effect values are calculated, based on studies of biological mechanisms or epidemiological interactions between factors. After mapping gene loci to OMIM or GWASCatalog, risk weights are assigned based on odds ratio (OR): OR ≤ 1.2 corresponds to 0.2-0.4, 1.2-1.5 corresponds to 0.4-0.6, 1.5-2.0 corresponds to 0.6-0.8, and > 2.0 corresponds to 0.8-1.0. For environmental factors associated with WHO guidelines, indexes are assigned based on the ratio of actual exposure concentration to limit: ratio ≤ 0.5 corresponds to 0.2-0.4, 0.5-1.0 corresponds to 0.4-0.6, 1.0-2.0 corresponds to 0.6-0.8, and > 2.0 corresponds to 0.8-1.0. After converting lifestyle habits to ICD-10 codes, risk weights are assigned based on risk hazard ratio (HR): HR ≤ 1.1 corresponds to 0.2-0.4, 1.1-1.3 corresponds to 0.4-0.6, 1.3-1.8 corresponds to 0.6-0.8, and > 1.8 corresponds to 0.8-1.0.
[0072] As one implementation method, the risk partitioning module generates risk subnetworks including genetic, environmental, and lifestyle subnetworks when partitioning by risk factor type. The module retains cross-type association edges between these subnetworks, including gene-environment interaction edges, gene-lifestyle interaction edges, and environment-lifestyle interaction edges. The weights of these cross-type association edges are determined based on normalized interaction ratios or hazard ratios from epidemiological studies. Subnetwork partitioning decomposes the complex disease association network into three independent subnetworks: genetic, environmental, and lifestyle. Each subnetwork contains only risk factor nodes of the same type and internal association edges. This avoids mixing different types of factors and facilitates in-depth analysis of single-type risks, such as genetic risks. The contribution of different types of factors to the disease is calculated through independent subnetworks, achieving hierarchical risk attribution. This allows for precise identification of risk sources; for example, if an individual's diabetes risk is found to be 10% genetic, 50% environmental, and 40% lifestyle-related, environmental and lifestyle factors are identified as the root causes. By employing three types of interaction edges—gene-environment, gene-lifestyle, and environment-lifestyle—the synergistic or antagonistic relationships between different risk factors are preserved. The synergistic or antagonistic effects between risk factors are determined by the type and intensity of their interactions. Synergistic effects may manifest as risk superposition (e.g., the simple addition of the independent risks of two factors) or amplification (e.g., multiplying risk contribution values to reflect synergistic enhancement). Antagonistic effects may manifest as risk offsetting (e.g., one factor reducing the influence coefficient of another factor) or inhibition (e.g., weakening the overall risk through negative regulation). Based on authoritative medical research evidence, such as synergistic coefficients of gene-environment interactions derived from GWAS studies, antagonistic relationships of environment-lifestyle interactions referenced from WHO guidelines, or pre-defined biological mechanism models, the quantitative assessment of interactions ensures that it conforms to scientific principles and clinical practice, thereby accurately reflecting the real risk changes under the combined effects of multiple factors. Cross-type association edges can reveal hidden risks. For example, GSTM1 gene deletion combined with smoking increases the risk of lung cancer fourfold, significantly higher than the sum of the individual effects of the two. Based on an individual's unique combination of genes, environment, and lifestyle habits, specific risk contributions can be calculated; for instance, individuals carrying the ALDH2 mutation have a significantly higher risk of liver cancer from alcohol consumption than ordinary drinkers. Cross-type association edges allow risks to flow between different sub-networks, forming cross-type risk propagation paths. These edges can also simulate risk evolution processes, such as a long-term high-sugar diet leading to insulin resistance, activation of the TCF7L2 gene, and increased risk of diabetes. By analyzing bottleneck nodes in the propagation path, more effective intervention strategies can be designed. The existence of interacting edges must be based on authoritative medical evidence to ensure that the network structure conforms to biological mechanisms. False causality based on statistical associations should be avoided. When calculating the synergy coefficient, a weighting coefficient must first be determined based on the level of evidence supporting the interaction, and then multiplied by a base synergy coefficient obtained from a database to obtain the final synergy coefficient used for calculation.
[0073] Example: Lung Cancer Risk Assessment Scenario
[0074] 1. Subnetwork Independent Analysis
[0075] Gene subnetwork:
[0076] Nodes: EGFR mutation (risk weight 0.7), KRAS wild-type (risk weight 0.3);
[0077] Intra-gene network association: EGFR mutation - tyrosine kinase pathway activation - increased risk of lung cancer.
[0078] Environmental subnetwork:
[0079] Nodes: PM2.5 exposure (exposure index 0.8), radon concentration exceeding the standard (exposure index 0.6);
[0080] Related: PM2.5 - DNA damage - increased risk of lung cancer.
[0081] Lifestyle Habits Sub-Network:
[0082] Key indicators: Smoking (HR=2.5), lack of exercise (HR=1.3);
[0083] Related: Smoking - Bronchial mucosal damage - Increased risk of lung cancer.
[0084] 2. Cross-type interaction analysis
[0085] Gene-environment interactions:
[0086] GSTM1 deletion (gene) - PM2.5 exposure (environment) - decreased detoxification capacity - increased DNA damage (cooperation coefficient 1.5).
[0087] Gene-lifestyle interaction:
[0088] CHRNA5 variant (gene) - smoking (lifestyle) - increased nicotine addiction - increased smoking amount (coefficient of 2.0);
[0089] Interaction between environment and lifestyle:
[0090] High-temperature fried food (lifestyle habits) - polycyclic aromatic hydrocarbon exposure (environment) - lung cancer risk superposition (HR=3.0).
[0091] 3. Comprehensive Risk Calculation
[0092] By integrating risks within subnetworks and cross-type interaction risks, the individual lung cancer risk probability is ultimately generated, and key driving pathways are identified.
[0093] The risk zoning module divides the disease association network into gene, environment, and lifestyle sub-networks according to the type of risk factor. Its core value lies in the fact that while realizing hierarchical management of risk factors, it retains cross-type association edges between genes and environment, genes and lifestyle, and environment and lifestyle. This enables the system to capture the synergistic or antagonistic relationships between different types of risk factors, thereby breaking through the limitations of single-factor independent assessment and more realistically reflecting the complex formation mechanism of individual disease risk. It provides structured support for accurately identifying key risk paths and formulating multi-dimensional intervention strategies, avoiding risk misjudgment or one-sided intervention plans caused by ignoring the interaction between factors.
[0094] As one implementation method, the risk assessment module converts the independent risk contribution values of multiple risk sub-networks into standardized probabilities. The risk assessment module also includes an interaction risk calculation submodule, used to generate risk probabilities and the highest-risk path, and to analyze gene-environment, gene-lifestyle, and environment-lifestyle interactions. Based on preset synergistic and antagonistic rules, it calculates the joint risk contribution value after the interactions. For the synergistic effect of the three factors, the combination with the highest risk among pairwise interactions is calculated first, and then the influence of the third factor is added.
[0095] When there is clear medical evidence, such as authoritative literature reports and clinical research data, indicating that a disease requires the simultaneous presence of gene mutation, specific environmental exposure, and unhealthy lifestyle habits to trigger a sudden increase in risk, it is necessary to first verify the necessity of the three-way interaction through biological pathways. Then, a nonlinear algorithm, such as a gradient boosting model based on machine learning, should be used to quantify the synergistic effect value of the three parties. This effect value is independent of the summation results of pairwise interactions and participates in the calculation of the risk contribution value independently. Taking the synergistic effect of lung cancer gene defect-pollutant-smoking as an example, if clinical research confirms that TP53 mutation only significantly enhances the carcinogenic effect when PM2.5 and smoking coexist, it is necessary to first extract the epidemiological association data of this three-way combination, fit the three-way interaction coefficient through a logistic regression model, and include it as an independent parameter in the calculation of the joint risk contribution value to ensure that it conforms to the multifactorial pathogenesis mechanism of the disease.
[0096] Specific synergistic effects also include three-way nonlinear interactions, where the risk increases exponentially when all three factors act together, rather than being a sum or product of pairwise interactions. For three-factor synergistic effects, we prioritize screening the combination with the highest risk among pairwise interactions based on medical evidence, and then superimpose the independent risk contribution value of the third factor. This process assumes that the three-factor synergistic effect is a linear sum of pairwise interaction effects and single-factor effects, and does not currently include quantitative analysis of direct three-factor interactions. Three-way nonlinear interactions may underestimate or overestimate actual risks and are applicable to scenarios where the three-factor interaction mechanism is not yet clearly defined.
[0097] To assess the synergistic effect of the three factors, we first calculate the combination with the highest risk in each pairwise interaction, and then incorporate the interaction effect of the third factor in the following manner:
[0098] If there is evidence of tripartite synergy, such as the gene-environment-lifestyle interaction coefficient reported in the literature, the joint risk contribution value is calculated according to the pre-defined tripartite synergy rules. If there is no tripartite evidence, the independent risk contribution value of the third factor is added.
[0099] The system treats the synergistic effect of the three factors based on a sufficiency grading system of medical evidence:
[0100] If there is clear evidence of pairwise interactions, such as the synergy coefficients of gene-environment and environment-lifestyle, which are supported by literature, then the combination with the highest risk in the pairwise interactions should be calculated first, and then the independent risk contribution value of the third factor should be added. In this process, the synergistic effect of the three factors is linearly superimposed.
[0101] If there is evidence of tripartite interaction, such as the joint involvement of genes, environment, and lifestyle in a pathogenic pathway, the joint risk contribution value is calculated using a pre-defined tripartite synergy rule. For interaction combinations with insufficient evidence, only single-factor and pairwise interaction analyses are retained to avoid unfounded inferences.
[0102] All possible risk propagation paths are ranked according to the rules that interaction risk is greater than single-factor risk and paths supported by high scientific evidence are greater than paths supported by low scientific evidence. Path finding can be performed using either the multi-standard Dijkstra algorithm or the A* algorithm. The path with the highest risk contribution value and consistent with biological mechanisms is selected as the highest-risk path. The risk contribution value of the highest-risk path is converted into a standardized probability, and combined with the baseline risk of the population to generate the individual's absolute risk probability. The individual's absolute risk probability and the corresponding key risk factors and interaction combinations are output.
[0103] The normalized independent risk contribution value, ranging from 0 to 1, is used as an input variable and converted into a relative risk probability through a logistic regression model. The formula is: Relative Risk Probability = 1 / (1+e^(-(a×x+b))), where x is the independent risk contribution value, and a and b are model parameters obtained by fitting the model using maximum likelihood estimation based on population epidemiological data. Different diseases correspond to different values of a and b, such as a=2.5 and b=-1.2 for lung cancer, and a=2.1 and b=-1.0 for diabetes. The relative risk probability needs to be combined with the baseline risk of the population to calculate the individual's absolute risk probability. The baseline risk of the population is obtained through national or regional disease surveillance data. For example, if the baseline risk of lung cancer in a certain region is 1%, the formula is: Individual Absolute Risk Probability = Relative Risk Probability × Population Baseline Risk. Taking lung cancer as an example, if the independent risk contribution of gene risk is 0.6, the relative risk probability calculated by the formula is 1 / (1+e^(-(2.5×0.6-1.2)))=1 / (1+e^(-0.3))≈0.575. Combined with the baseline risk of 1% in the population, the individual absolute risk probability is 0.575×1%=0.575%.
[0104] The independent risk contribution values of the gene, environment, and lifestyle subnetworks are transformed into standardized probabilities in the 0-1 range. Through normalization, Min-Max standardization or logarithmic transformation, the dimensional differences between different factors are eliminated. Based on population epidemiological data, the normalized values are mapped to the relative probability of disease occurrence. This allows for direct comparison of gene, environment, and lifestyle risks. A unified input format is provided for subsequent interaction calculations, supporting multi-factor integration in machine learning or statistical models. The interaction risk calculation submodule can capture synergistic and antagonistic effects of multiple factors, analyze gene-environment, gene-lifestyle, and environment-lifestyle interactions, and quantify joint risk contribution values. Different algorithms vary greatly in their ability to capture interactions; therefore, selection should be based on data type. Genetic data is suitable for Bayesian networks, while behavioral data is suitable for time series models.
[0105] Synergistic rules: Risk amplification, such as genetic sensitivity traits doubling the risk of environmental exposure;
[0106] Antagonistic rule: risk offsetting, such as regular exercise reducing the risk of diabetes from a high-sugar diet, and exercise improving insulin sensitivity;
[0107] Prioritize pairwise interactions: First calculate the pairwise combinations with the highest risk among gene-environment, gene-lifestyle, and environment-lifestyle;
[0108] Superimposed third factor: On the basis of pairwise interactions, the independent risk or regulatory effect of a third factor is superimposed, such as the risk of reinforcement by gene-environment interaction and lifestyle habits. This avoids the underestimation or overestimation of risk caused by linear superposition of single factors, and truly reflects biological phenomena where 1-1>2 or 1-1<1.
[0109] Risk path ranking:
[0110] All possible risk transmission paths are ranked by single-factor, pairwise, and three-factor synergy to identify the highest-risk path.
[0111] Interaction risk > Single-factor risk:
[0112] Multifactor synergistic effects are common in the occurrence of diseases and are the most important. For example, most cancers, such as lung cancer, are rarely caused by a single gene or environmental factor.
[0113] Prioritize the evaluation of combined pathways such as gene-environment and environment-lifestyle, rather than isolated single-factor pathways.
[0114] Pathways with strong scientific evidence > Pathways with weak scientific evidence
[0115] Logic: Based on the level of evidence, such as RCT > cohort study > case report, a reliable path is selected;
[0116] Application: Exclude low-evidence pathways such as "stress-gene mutation-cancer" and prioritize pathways with clear mechanisms such as "smoking-oxidative stress-lung cancer".
[0117] Path enumeration: Lists all possible risk transmission paths. Examples include gene-disease, gene-environment-disease, environment-lifestyle-disease, etc.
[0118] Evidence filtering: Eliminating paths that lack support from medical literature;
[0119] Risk contribution value sorting: sorted from high to low based on the risk contribution value of interaction;
[0120] Biological rationality verification: Ensure that the pathway conforms to known physiological and pathological mechanisms.
[0121] Avoid misjudgments caused by spurious correlations, such as the coincidental statistical association of "coffee drinking - lung cancer";
[0122] Focus on high-risk, high-evidence pathways, reduce redundant calculations, and improve assessment speed.
[0123] Highest risk path selection and absolute risk generation
[0124] The path with the greatest risk contribution and consistent with biological mechanisms is selected from the sorted paths, and the individual absolute risk probability is generated by combining it with the baseline risk of the population.
[0125] Clinical interpretability: The absolute risk probability directly corresponds to the individual's likelihood of developing the disease, making it easier for doctors to determine the screening frequency;
[0126] Once the highest-risk pathway is identified, intervention measures can be designed for key nodes in the pathway, such as environmental exposure avoidance in the "gene-environment interaction" pathway.
[0127] Output results:
[0128] Individual absolute risk probability: quantifying risk level, such as "25% risk of lung cancer in the next 5 years";
[0129] Key risk factors: Factors that contribute more than 20% to any single factor;
[0130] Interaction combinations: synergistic pairs of multiple factors that lead to increased risk, such as smoking-PM2.5 exposure, GSTM1 deficiency-smoking.
[0131] Intervention priority ranking:
[0132] High-risk interaction combinations should be prioritized for intervention;
[0133] In single-factor studies, behaviors with an HR value greater than 2 are given priority for correction, such as mandatory smoking cessation for those who smoke more than 10 cigarettes per day.
[0134] Numerical definition of cooperation and antagonism rules:
[0135] Gene-environment interaction: Based on GWAS studies, toxicological experimental data and historical data, a pre-set synergy coefficient is used, such as the synergy coefficient of GSTM1 deficiency-PM2.5 exposure = 1.5;
[0136] Environment-lifestyle interaction: Refer to the exposure-behavior interaction model in the WHO's "Environment and Health Guidelines";
[0137] Gene-lifestyle interaction: Based on molecular mechanism studies, such as the CHRNA5 variant-smoking co-coefficient of 2.0, which corresponds to an increase in smoking due to decreased nicotine metabolism enzyme activity.
[0138] The specific synergistic effects between diseases are derived from historical experience, actual research, and existing publicly available data.
[0139] Example: Assessing an individual's risk of lung cancer:
[0140] Genetic data: Detection of GSTM1 gene deletion;
[0141] Environmental data: Long-term exposure to PM2.5 pollution;
[0142] Lifestyle habits: Smoking an average of 10 cigarettes per day for 10 years.
[0143] Preprocessing:
[0144] Genetic data: Sequencing errors were eliminated, confirming that the GSTM1 deletion was a true variant;
[0145] Environmental data: Missing PM2.5 monitoring values were filled in using the average value for the same period;
[0146] Lifestyle habits: Convert "smoking 10 cigarettes / day" into a standard behavior code.
[0147] Medical knowledge mapping
[0148] Gene-disease association: GSTM1 deletion - mapped to GWASCatalog - confirmed to be associated with lung cancer risk, preset gene risk weight = high, corresponding to a pre-standardized value of 0.6.
[0149] Environment-Disease Linkage:
[0150] PM2.5 exposure – mapped to WHO's "Environment and Health Guidelines" – is identified as a causative factor for lung cancer, with an environmental exposure index of extremely high, corresponding to a pre-standardized value of 0.8.
[0151] Lifestyle Habits - Disease Linkage:
[0152] Smoking - converted to ICD-10 code Z72.0 - associated lung cancer HR value = 2.5 (hazard ratio).
[0153] The standardization process involves first normalizing the values to the 0-1 range, and then converting them into probability values.
[0154] Genetic risk: High-risk weights are mapped to 0.6 using Min-Max normalization;
[0155] Environmental exposure: PM2.5 exceedance level corresponds to 0.8;
[0156] Lifestyle habits HR value: 0.55 after logarithmic transformation.
[0157] Datafication of Synergies:
[0158] Preset synergy coefficients based on research data;
[0159] Gene-environment interaction (GSTM1 deletion - PM2.5):
[0160] According to GWAS studies, the synergy coefficient is 1.5, and the risk ratio is increased to 1.5 times that of the individual effect.
[0161] Gene-lifestyle interaction (GSTM1 deletion - smoking):
[0162] Based on toxicological experiments, the synergy coefficient was 2.0, indicating that metabolic defects exacerbated the toxicity of smoking.
[0163] Interaction between environment and lifestyle habits (PM2.5 - smoking):
[0164] WHO guidelines recommend a superposition factor of 3.0, indicating a synergistic effect between pollutants and smoke carcinogens.
[0165] Calculate the joint risk contribution value of each pair of interactions:
[0166] Gene-environment interaction: Joint risk contribution value = (0.6 + 0.8) × 1.5 = 2.1;
[0167] Gene-lifestyle interaction: Joint risk contribution value = (0.6 + 0.55) × 2.0 = 2.3;
[0168] Environment-lifestyle interaction: Joint risk contribution value = (0.8 + 0.55) × 3.0 = 4.05;
[0169] Three-factor synergistic effect: Priority and highest pairwise interaction - superimposed with a third factor
[0170] Filter for the highest value of pairwise interactions: Environment-Lifestyle interaction (4.05);
[0171] Influence of superimposed genetic factors: The combined risk contribution of the three factors = 4.05 × (1 + 0.6) = 6.48;
[0172] Optimal path selection:
[0173] Risk contribution priority:
[0174] Single-factor risk contribution values: Genes (0.6) < Environment (0.8) < Lifestyle (0.55);
[0175] Pairwise interaction risk contribution values: Environment-lifestyle (4.05) > Gene-lifestyle (2.3) > Gene-environment (2.1);
[0176] The three-factor risk contribution value (6.48) is greater than all pairwise interactions, and is therefore determined to be the highest risk path.
[0177] Evidence Level Calibration:
[0178] Environment-lifestyle interaction (PM2.5-smoking): Level of evidence I (RCT study), weight = 1.0;
[0179] Genetic factors combined: Level of evidence II (cohort study), weight = 0.8;
[0180] Final path score: 6.48 × (1.0 × 0.8) = 5.184.
[0181] Disease probability generation: from joint risk contribution value to absolute probability.
[0182] 1. Standardized probability transformation
[0183] The joint risk contribution value of 5.184 is mapped to the 0-1 interval using the Logistic function: standardized probability. Let a=2.5, b=-1.2. Then, when the joint risk contribution value z=5.184, P≈1.0, which means that the relative risk of an individual having the disease under this path is 100%.
[0184] 2. Combined with baseline risk in the population
[0185] Assuming the baseline probability of lung cancer in a certain region is 1%, then the individual absolute risk probability is: Absolute risk = 1% × 0.95 = 0.95%.
[0186] Results and Intervention Recommendations:
[0187] Key risk factors: PM2.5 exposure (0.8), smoking (0.55), GSTM1 deficiency (0.6);
[0188] Interaction combination: PM2.5-smoking (synergistic coefficient 3.0), genetic defect amplification environment-lifestyle risk;
[0189] Intervention priority:
[0190] Quit smoking immediately (eliminate the risk factors in your lifestyle);
[0191] Wear a mask to protect against smog (reduce the environmental exposure index to 0.4);
[0192] Real-time monitoring of lung inflammation markers (to verify changes in pathway risk).
[0193] By simulating the risk changes after removing a factor in a pathway, such as how smoking cessation reduces the risk of smoking-gene interactions, the system demonstrates intervention benefits to patients and improves adherence. The risk assessment module converts independent risk contribution values into standardized probabilities, enabling unified measurement and comparable analysis of different types of risk factors. The interaction risk calculation submodule quantifies the interactive effects of genes, environment, and lifestyle habits through pre-defined synergistic and antagonistic rules, overcoming the limitations of single-factor independent assessment and capturing the risk amplification or inhibition effects under the combined influence of multiple factors. For three-factor synergistic effects, it prioritizes calculating key pairwise interaction combinations and superimposes the influence of the third factor, ensuring full risk coverage in complex interaction scenarios. Using a ranking rule of interaction risk > single-factor risk, and high-evidence pathway > low-evidence pathway, the system can screen out the transmission pathway with the most biological rationality and risk contribution, combine it with the population baseline risk to generate individual absolute risk probabilities, and finally output accurate assessment results including key factors and interaction combinations. Baseline risk for the population was primarily determined using national data, such as the "China Chronic Disease Report" from the National Center for Disease Control and Prevention. Secondary data used was local surveillance data from the target individual's region. If no such data was available, cohort data from a system-built database of over 5000 cases with a follow-up period of over 5 years was used. Matched populations were selected based on individual age, gender, and region. The incidence rate of the target disease within a specific time window for this population was calculated and used as the baseline risk.
[0194] As one implementation method, the data types of genetic data, environmental exposure data, and lifestyle data are as follows: Genetic data includes at least one of single nucleotide polymorphisms, gene expression profiles, and epigenetic data; genetic data originates from gene sequencing, microarray detection, and historical data from hospitals. Environmental exposure data includes: Physical environment: at least one of PM2.5 concentration, noise level, and history of heavy metal exposure. Chemical environment: exposure to at least one indoor pollutant such as formaldehyde and benzene. Lifestyle data includes: Behavioral data: average daily smoking amount, alcohol consumption and frequency, exercise duration, and sleep quality score. Dietary data: nutrient intake frequency and dietary structure score.
[0195] As one implementation method, the intelligent management system also includes a health data monitoring module, used to monitor health indicator data corresponding to each subnetwork in multiple risk subnetworks in real time, obtaining a multi-dimensional health dataset. A risk source identification module is used to trace risks based on the multi-dimensional health dataset, detecting the initial risk source corresponding to each risk subnetwork. A risk propagation network construction module is used to construct a corresponding risk propagation network based on each initial risk source. The network includes the risk source, nodes affected by risk propagation, and propagation paths. The risk propagation network construction module allows risk sources to propagate across subnetworks.
[0196] Dynamic modeling of risk propagation paths across sub-networks can be achieved by establishing a multi-source dynamic data collection system. This system integrates gene expression monitoring data, periodic environmental sensor data, and wearable device behavior logs via standardized data interfaces. After data cleaning and time synchronization, the data is stored in a time-series database. A risk propagation network is constructed based on a dynamic graph database. Starting from the risk source, initial propagation rules are defined using a biological mechanism knowledge base. Combined with periodic data, dynamic graph neural networks or time-series association rule mining algorithms are used to identify dynamic association edges across sub-networks. For example, when the expression level of a certain inflammatory factor suddenly increases in the gene sub-network, a DGNN model is used to periodically search for relevant pollutant exposure data in the environmental sub-network and abnormal activity records in the lifestyle sub-network to construct a cross-sub-network propagation path of abnormal gene expression – increased environmental exposure – altered behavioral patterns.
[0197] Taking a sudden increase in inflammatory cytokine expression in a gene subnetwork as an example of a risk source, its propagation to the environmental subnetwork requires first matching the corresponding biological pathways using a medical knowledge graph. For instance, inflammatory cytokines activate the CYP450 enzyme system in the liver, which participates in the metabolism of environmental pollutants. Then, based on real-time monitored physiological indicators such as enzyme activity levels, the pathway activation status is verified. If enzyme activity is significantly increased, a propagation pathway can be constructed: sudden increase in inflammatory cytokine expression → increased CYP450 enzyme activity → accelerated metabolic rate of environmental pollutants → enhanced exposure effect of pollutants in the environmental subnetwork. For the cross-subnetwork propagation of different types of risk sources, the process of biological pathway matching, physiological indicator verification, and propagation pathway construction must be followed.
[0198] The critical risk path acquisition module analyzes the risk propagation network using a network centrality analysis model to identify the critical risk paths corresponding to each risk sub-network. Network centrality analysis models can employ types such as betweenness centrality, proximity centrality, and PageRank. Betweenness centrality measures the importance of a node as a mediator in the network, reflecting its control over information transmission; proximity centrality assesses a node's reachability by calculating the sum of the shortest paths from a node to all other nodes, reflecting its convenience within the network; and PageRank, using an algorithm that simulates the weight of web page links, measures a node's influence and importance in the network. The system can analyze the topology of the risk propagation network using these models to identify key nodes and paths that contribute significantly to the risk of the target disease, thereby obtaining the critical risk paths corresponding to each risk sub-network. The risk assessment module generates individual disease risk assessment results and intervention recommendations based on the critical risk paths.
[0199] Considering the characteristic of node weights changing over time in dynamic networks, a sliding time window method is used to segment the dynamic network in the traditional static network analysis model. The continuous time series is divided into multiple time windows, with the window length adjusted according to the data type (e.g., 24 hours / window for environmental data, 7 days / window for lifestyle data). Within each time window, the betweenness centrality and proximity centrality of nodes are calculated. Simultaneously, a time decay coefficient is introduced to assign decreasing weights to the centrality indicators of historical time windows; the further away from the current time, the smaller the weight coefficient. Specific values are determined through cross-validation. Finally, the dynamic centrality index of the node is obtained by weighted summation, thereby identifying critical paths at different time stages. For example, when analyzing the risk transmission path related to PM2.5 exposure, if the betweenness centrality of a node is 0.8 in the most recent 24-hour window and 0.6 in the previous 24-hour window, and the time decay coefficient is set to 0.5, then the dynamic centrality of this node is 0.8 + 0.6 × 0.5 = 1.1.
[0200] The aforementioned intelligent management system is a dynamic system, including a risk propagation network construction module. Starting from an initial risk source, it constructs a dynamic network containing cross-sub-network propagation paths, such as gene expression changes – lifestyle adjustments – environmental exposure responses. It reveals the real-time propagation path of risk between genes, environment, and lifestyle. Cross-sub-network propagation paths include, for example, gene mutation risk sources propagating to the environmental sub-network via the "gene-enzyme activity-environmental pollutant metabolism" path, or lifestyle risk sources propagating to the gene sub-network via the "behavior-physiological indicators-gene expression" path. The cross-sub-network propagation path construction starts by identifying abnormal indicators from health monitoring data as initial risk sources, searching the medical knowledge base to match associated biological pathways, collecting relevant physiological indicators to verify pathway activation, and if activated, tracing affected nodes along the pathway, constructing cross-sub-network propagation paths, and recording the node association strength. The risk assessment module generates real-time assessment results and intervention recommendations based on dynamic critical paths, such as immediately avoiding current PM2.5 exposure and adjusting exercise plans. Intervention measures are matched with the real-time status of the risk, such as recommending immediate protection for ongoing "gene-environment interaction risks," improving the timeliness and targeting of interventions. It supports closed-loop feedback of risk and intervention, such as real-time monitoring of the decline in path risk after simulated intervention to verify the effectiveness of measures. It upgrades from periodic static assessment to real-time dynamic monitoring, capturing instantaneous changes in risk; from fixed network models to adaptive propagation modeling, reflecting the dynamic evolution logic of risk; and from historical factor attribution to immediate source tracing and path tracking, enabling precise targeting of intervention plans. The intelligent management system combination gives it dynamic response capabilities of monitoring, source tracing, modeling, and intervention, making it particularly suitable for scenarios requiring real-time prevention and control, such as early warning of acute exacerbations of chronic diseases and emergency response to environmental pollution.
[0201] This invention monitors and predicts disease risk by combining static and dynamic methods. Static key risk factors are not equivalent to dynamic risk attribution. Static factors identify combinations of factors associated with disease risk from large amounts of data, such as smoking-advanced age-increased lung cancer risk, which is attribution at the statistical association level and does not involve the real-time causal chain of risk generation. Dynamic risk attribution locates the source of the current risk in real-time data, such as sudden exposure to secondhand smoke-triggered elevated lung inflammation markers-entering the risk transmission path, which is tracing the real-time causal chain. Static assessment, based on baseline data, shows that someone who is overweight-family history-30% risk of diabetes has weight-genetics as a key factor. Dynamic assessment, through real-time monitoring, detects a sudden increase in dietary calories in the past week-long period-long-term weight gain-elevated insulin resistance markers-triggered the risk transmission path of diet-weight-metabolism, requiring immediate intervention. The dynamic module supplements information on the dynamic risk generation process through real-time data, rather than repeating static factor analysis. To achieve real-time response in risk prevention and control, the health data monitoring module collects multi-dimensional dynamic indicators in real time, such as gene expression fluctuations, real-time PM2.5 concentration, and continuous blood glucose data, forming a dynamic health record and capturing immediate changes in risk factors, such as sudden high-sugar diets and short-term formaldehyde exposure, thus solving the problem of lag in static assessment.
[0202] The risk source identification module locates initial risk sources based on real-time data. For example, excessive daily alcohol consumption might be the trigger for elevated liver enzymes, upgrading from historical correlation analysis to real-time causal tracing and accurately identifying the direct driving factors of risk. The risk transmission network construction module supports dynamic path modeling across sub-networks, such as gene expression abnormalities – lifestyle adjustments – environmental exposure responses, revealing the real-time transmission chain of risk between genes, environment, and lifestyle habits, such as PM2.5 exposure – oxidative stress gene activation – decreased willingness to exercise. The key risk path acquisition module filters core transmission paths in real-time, such as the current environmental exposure – behavioral abnormality path which has the highest risk contribution, adapting to the stage-specific characteristics of risk evolution. The risk assessment module generates real-time intervention suggestions based on dynamic paths, such as immediately wearing an anti-smog mask and adjusting exercise plans. Intervention measures are precisely matched to the current state of risk, improving timeliness, such as triggering immediate protection during PM2.5 peak periods. It supports a closed loop of intervention-monitoring-feedback, such as simulating real-time monitoring of the decrease in gene-smoking interaction risk after smoking cessation.
[0203] Static and dynamic modules work together to construct a full-cycle risk prevention and control system. The static module provides a baseline risk level and potential trends, such as the lower risk limit determined by genetic susceptibility. The dynamic module captures real-time fluctuations and sudden triggers in risk, such as exceeding the upper risk limit due to sudden environmental exposure. This synergy between static and dynamic modules forms a continuous spatiotemporal assessment of baseline prediction and real-time early warning, covering the entire lifecycle from health to sub-health to disease. The dynamic module uses real-time data to validate and correct the static model, such as adjusting preset gene-environment interaction coefficients using real-time gene expression data. This achieves dynamic calibration of data, model, and evidence, improving the accuracy of risk assessment in complex scenarios. This invention, through the organic synergy of static and dynamic modules, not only meets the stringent requirements of the medical field for the scientific rigor and interpretability of risk assessment but also adapts to the real-time and personalized needs of precision medicine, demonstrating significant technological leadership and clinical applicability.
[0204] As one implementation method, the health data monitoring module is used to monitor the gene subnetwork in real time, either weekly or monthly, with the specific frequency set according to needs. It monitors the expression levels of genes related to environmental factors and lifestyle habits, such as the response of antioxidant genes to PM2.5 exposure and the response of appetite regulation genes to dietary patterns. The environmental subnetwork monitors environmental exposure indicators influenced by lifestyle habits, such as the impact of window ventilation time on indoor formaldehyde concentration and the impact of smoking behavior on secondhand smoke pollutant concentration. The lifestyle habit subnetwork monitors behavioral data related to genes and the environment, such as alcohol consumption in individuals carrying alcohol metabolism gene variants and the duration of exercise in individuals with high formaldehyde exposure.
[0205] As one implementation method, the risk source identification module includes a hybrid risk source detection unit, which is used to determine whether the risk is caused by two or three types of factors in the gene-environment-lifestyle categories. The determination criteria include the spatiotemporal correlation of cross-sub-network data and the interaction strength threshold.
[0206] Quantitative criteria for detecting mixed risk sources can be achieved in the following ways: Regarding spatiotemporal correlation, the time window can be set as the time span of exposure to each risk factor, such as the innate existence of gene variations, the cumulative period of environmental exposure over the past year, and the persistence of lifestyle habits over the past three months. It is required that the exposure times of the factors overlap or are sequentially related, such as changes in lifestyle habits occurring within one month after environmental exposure. The spatial scope needs to define the geographical or environmental area where the risk factors act. The threshold for the intensity of interaction can be set based on epidemiological statistical methods, such as using logistic regression to calculate the hazard ratio of pairwise or triadic combinations of genes, environment, and lifestyle habits. A hazard ratio (HR) > 1.5 and a Bonferroni correction p < 0.01 are used as the threshold for significant synergistic effects. Simultaneously, combined with the level of evidence from medical literature, combinations of factors across sub-networks that meet the threshold conditions are identified as mixed risk sources, and joint risk labels containing multiple types of factors are generated.
[0207] The quantitative criteria for identifying mixed risk sources are as follows: temporal correlation requires an overlap of risk factor exposure windows of ≥30 days; spatial correlation is based on an individual's 5-kilometer living circle, with overlapping factor exposure areas constituting the criteria. Interaction strength requires a joint HR value >1.5 and P <0.01. Meeting both the temporal and spatial correlation and interaction strength thresholds simultaneously qualifies a source as a mixed risk source.
[0208] For identified mixed risk sources, a joint risk label containing multiple types of factors is generated and associated with its corresponding disease type. Diseases are often caused by two or three types of factors from genes, environment, and lifestyle habits. For example, lung cancer is synergistically driven by gene mutation, smoking, and air pollution. The mixed risk source detection unit can accurately identify such complex scenarios. The system analyzes the temporal sequence and spatial correlation of factor exposure, such as the time chain of congenital gene mutation - long-term PM2.5 exposure in adulthood - recent smoking, distinguishing between true synergy and accidental co-occurrence. For example, the statistical association between coffee consumption and lung cancer may be confounded by smoking, which can be excluded through spatiotemporal analysis. It supports causal chain modeling, such as the cross-life-cycle path of childhood lead exposure - adult hypertension - insufficient exercise in middle age. Examples of joint risk labels include the "gene mutation - environmental exposure - unhealthy behavior" combined label. The system simplifies the clinical interpretation of multifactorial risks; for example, doctors can quickly identify the need for simultaneous interventions such as gene monitoring, environmental exposure avoidance, and behavioral modification through the labels. It supports cross-system interaction of risk data; for example, electronic health records can directly call labels to generate personalized follow-up plans.
[0209] For each type of mixed risk factor, a combined intervention of gene regulation, environmental improvement, and behavioral modification was designed:
[0210] At the genetic level: Antioxidant supplements are recommended to regulate the Nrf2 pathway;
[0211] Environmental factors: Installing air purifiers reduces PM2.5 exposure;
[0212] Behavioral aspect: Participate in smoking cessation support groups.
[0213] Improve the systematic nature of interventions. Single-factor interventions may fail due to the presence of other factors. For example, quitting smoking alone without improving air pollution may still pose a risk of lung cancer.
[0214] The intervention effect was dynamically verified by monitoring changes in the risk contribution value of the joint label and quantitatively assessing the mitigating effect of smoking cessation-environmental governance on synergistic risks.
[0215] Genes and Environment: Some people are born with a genetic predisposition to environmental sensitivity, carrying a deficiency in the GSTM1 gene, resulting in a poor ability to metabolize PM2.5. When exposed to smog, these individuals are more likely to accumulate toxins, leading to lung inflammation—experiencing symptoms such as coughing and shortness of breath earlier than the average person.
[0216] Genes can double environmental risks – for example, the risk of an ordinary person inhaling smog is 3, while for this type of person it may be 8. The specific difference can be established by statistically linking data from historical databases.
[0217] Genetics and Lifestyle: People carrying the FTO gene mutation are more prone to hunger, leading to a high-calorie diet. The gene makes people carrying the FTO gene mutation want to eat more—and thus have a much higher risk of obesity than the average person.
[0218] Environment and Lifestyle Habits: Pollution makes people lazy; rooms with excessive formaldehyde levels irritate the respiratory tract, causing discomfort and thus reducing physical activity. Polluted environment forces reduced exercise, leading to decreased cardiopulmonary function and increased cardiovascular risk.
[0219] Example:
[0220] Single-factor risk calculation:
[0221] Scenario: Assessing an individual's risk of lung cancer
[0222] Input data:
[0223] Genetic data: Carries a TP53 gene mutation;
[0224] Environmental data: Long-term exposure to PM2.5, annual average 65 μg / m³ 3 ;
[0225] Lifestyle habits: Smoking an average of 10 cigarettes per day for 10 years;
[0226] Data mapping and weight assignment:
[0227] TP53 mutation - mapped to OMIM lung cancer-related genes, with a preset gene risk weight of "high";
[0228] PM2.5 levels exceeding standards - linked to WHO carcinogens, environmental exposure index "extremely high";
[0229] Smoking 10 cigarettes a day is mapped to the ICD-10 smoking code, indicating a "high" HR value for lifestyle habits.
[0230] Single-factor risk calibration:
[0231] Genetic risk: According to population data, people with TP53 mutations have a 5 times higher risk of lung cancer than the general population;
[0232] Environmental risk: The risk increases by 10% for every 100% exceedance of PM2.5 standards;
[0233] Lifestyle risks: Cumulative risk coefficient of smoking for 10 years.
[0234] Single-factor risk output:
[0235] Individual risks associated with genes;
[0236] Individual environmental risks;
[0237] Lifestyle habits pose individual risks.
[0238] Example: Interaction risk calculation:
[0239] Gene-environment interaction
[0240] TP53 mutations weaken cell repair capabilities, making the DNA damage effect of PM2.5 stronger.
[0241] Gene-lifestyle interaction
[0242] Mechanism: Individuals with TP53 mutations are more sensitive to the carcinogenic effects of smoking.
[0243] Environment-Lifestyle Interaction
[0244] Mechanism: Smoke from smoking works synergistically with carcinogens in PM2.5.
[0245] Example: Three-factor synergy:
[0246] Genetic defects, pollution, and smoking form a carcinogenic triangle, with the risk increasing exponentially.
[0247] Take the highest value among the pairwise interactions, and then add another risk.
[0248] Comprehensive risk assessment: Identify the highest-risk path;
[0249] The system automatically filters critical paths based on the following dimensions:
[0250] Risk contribution: Interaction risk > Single factor risk;
[0251] Biological rationale: Prioritize pathways with clear mechanistic evidence, such as the smoking-oxidative stress-lung cancer pathway, which is more reliable than the PM2.5-random mutation-lung cancer pathway;
[0252] Real-time data matching degree: Active pathways in the current monitoring data, such as when an individual's oxidative stress index is elevated, oxidative stress-related pathways are preferentially activated.
[0253] Example: Identification of the highest risk path
[0254] Candidate paths:
[0255] Univariate pathway: TP53 mutation - lung cancer;
[0256] Bivariate pathway: TP53-PM2.5-lung cancer;
[0257] Two-factor pathway: smoking-PM2.5-lung cancer;
[0258] The three-factor pathway is: TP53-PM2.5-smoking-lung cancer.
[0259] System determination:
[0260] The three-factor pathway carries the highest risk and conforms to the typical carcinogenic pattern of gene defect-environmental exposure-adverse behavior, thus being marked as the highest-risk pathway.
[0261] Risk probability generation, overall risk probability:
[0262] The risk contribution value of the highest risk path is standardized to the range of 0-100%.
[0263] Output result:
[0264] Your risk of lung cancer is XX, mainly caused by the combined effects of TP53 gene mutation, PM2.5 exposure, and smoking.
[0265] Verify the best path:
[0266] Intervention targeting the three-factor pathway:
[0267] Quit smoking and eliminate the risks associated with your lifestyle.
[0268] Wearing a mask to protect against smog can reduce PM2.5 exposure by 50%.
[0269] Regularly taking antioxidants can alleviate the oxidative stress effects of TP53 mutations.
[0270] Risk recalculation:
[0271] Smoking risk elimination - the three-factor pathway becomes a two-factor pathway, TP53-PM2.5;
[0272] Result comparison:
[0273] Risk probability before intervention - Risk probability decreases after intervention.
[0274] After intervention, the highest-risk pathway shifted to the single-factor TP53 mutation pathway, and the risk level decreased from high to medium.
[0275] Disease association networks are static maps that capture all possible risk paths.
[0276] The risk propagation network is a real-time navigation system that dynamically plans the most likely risk path and calculates the risk contribution value based on the current situation, such as whether one is smoking or in a contaminated area.
[0277] It should be understood that those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A smart management system for smart healthcare, characterized by, Comprise: a data collection module for collecting individual genetic data, environmental exposure data and lifestyle data; a disease association network construction module for constructing a disease association network containing the association relationships between nodes of the three types of risk factors based on the genetic data, environmental exposure data and lifestyle data; a data processing module for processing the factor nodes and calculating the disease types corresponding to the processed nodes to generate the risk probabilities of different diseases; a risk partitioning module for partitioning the disease association network according to the risk factor types and outputting multiple risk sub-networks, wherein each risk sub-network corresponds to risk factors of the same type; a risk assessment module for generating individual disease risk assessment results and intervention suggestions according to the multiple risk sub-networks; the disease association network construction module is connected to the data collection module, the data processing module is connected to the disease association network construction module, the risk partitioning module is connected to the data processing module, and the risk assessment module is connected to the risk partitioning module; the risk assessment module is used for converting the independent risk contribution values of the multiple risk sub-networks into standardized probabilities; the risk assessment module further comprises an interaction risk calculation sub-module for generating risk probabilities and the highest risk path: for analyzing the interactions of gene-environment, gene-lifestyle, and environment-lifestyle, calculating the joint risk contribution values after the interactions according to preset synergistic and antagonistic rules; for the synergistic effect of the three factors, the combination with the highest risk in the two-way interaction is calculated first, and then the influence of the third factor is superimposed; all possible risk transmission paths are sorted according to the rules that the interaction risk is greater than the single-factor risk and the path supported by high scientific evidence is greater than the path supported by low scientific evidence; the path with the highest risk contribution value and conforming to the biological mechanism is selected as the highest risk path, the risk contribution value of the highest risk path is converted into a standardized probability, and the individual absolute risk probability is generated by combining the population baseline risk; output the individual absolute risk probability and the corresponding key risk factors and interaction combinations.
2. The intelligent management system for intelligent medicine according to claim 1, characterized in that: the data processing module includes preprocessing, medical knowledge mapping and standardization processing: preprocessing: performing data cleaning on the collected genetic data, environmental exposure data and lifestyle data to exclude mapping errors or redundancies; medical knowledge mapping: mapping the genetic loci to the disease-related genes in OMIM or GWAS Catalog, and obtaining the genetic risk weights of different diseases according to the disease-related genes; associating environmental factors with the pathogenic factor list in WHO Environmental Health Guidelines, and obtaining the environmental exposure indexes of different diseases according to the data in the pathogenic factor list; convert the lifestyle habits into ICD-10 behavior risk factor codes, and obtain the lifestyle HR values of different diseases according to the behavior risk factor codes; standardization processing: for normalizing the genetic risk weights, environmental exposure indexes and lifestyle HR values, the independent risk contribution values of each factor to a specific disease.
3. The intelligent management system for intelligent medicine according to claim 2, characterized in that: The risk sub-networks generated by the risk partition module when partitioning by risk factor type include a genetic sub-network, an environmental sub-network, and a lifestyle habit sub-network; The risk partition module retains cross-type association edges between the genetic sub-network, the environmental sub-network, and the lifestyle habit sub-network when partitioning by risk factor type, and the cross-type association edges include gene-environment interaction edges, gene-lifestyle habit interaction edges, and environment-lifestyle habit interaction edges.
4. The intelligent management system for intelligent medicine according to claim 1, characterized in that: The data types of the genetic data, the environmental exposure data, and the lifestyle habit data are respectively: The genetic data include at least one of single nucleotide polymorphisms, gene expression profiles, and epigenetic data; The environmental exposure data include: Physical environment: at least one of PM2.5 concentration, noise level, and heavy metal exposure history; Chemical environment: indoor pollutant exposure of at least one of formaldehyde and benzene; The lifestyle habit data include: Behavioral data: average daily smoking amount, alcohol consumption amount and frequency, exercise duration, and sleep quality score; Dietary data: nutrient intake frequency and dietary structure score.
5. The intelligent management system for intelligent medicine according to claim 1, characterized in that: The intelligent management system further includes: a health data monitoring module configured to periodically monitor health indicator data corresponding to each of the plurality of risk sub-networks to obtain a multi-dimensional health data set; a risk source identification module configured to trace the risk according to the multi-dimensional health data set to detect an initial risk source corresponding to each risk sub-network; a risk propagation network construction module configured to construct a corresponding risk propagation network according to each initial risk source, the network including a risk source, a node affected by risk propagation, and a propagation path, and the risk propagation network construction module allows the risk source to propagate across sub-networks; a key risk path acquisition module configured to analyze the risk propagation network by a network center analysis model to acquire a key risk path corresponding to each risk sub-network; a risk assessment module configured to generate an individual disease risk assessment result and an intervention suggestion according to the key risk path; The health data monitoring module is connected to the risk assessment module, the risk source identification module is connected to the health data monitoring module, the risk propagation network construction module is connected to the risk source identification module, and the key risk path acquisition module is connected to the risk propagation network construction module.
6. The intelligent management system for intelligent medicine according to claim 5, characterized in that: The health data monitoring module is configured to monitor in real time: The genetic sub-network: monitoring the gene expression amount related to environmental factors and lifestyle habits; The environmental sub-network: monitoring environmental exposure indicators affected by lifestyle habits; The lifestyle habit sub-network: monitoring behavioral data associated with genes and the environment.
7. The intelligent management system for intelligent medicine according to claim 5, characterized in that: The risk source identification module includes: a mixed risk source detection unit configured to determine whether the risk is caused by two or three types of factors among genes, environments, and lifestyle habits, and the determination is based on the spatiotemporal correlation or interaction strength threshold of cross-sub-network data. For the identified mixed risk sources, a joint risk label containing multiple types of factors is generated, and its corresponding disease type is associated.
8. A smart management method for smart medicine, using a smart management system for smart medicine according to any one of claims 1-7, characterized in that, Comprise: S1: data acquisition and standardization processing, obtaining the gene data, environmental exposure data and living habit data of the target individual, wherein the gene data includes the gene site information related to the target disease, the environmental exposure data covers physical, chemical and biological environmental factors, and the living habit data includes diet, exercise, work and rest and medical behavior records; standardize the three types of data, eliminate the data dimension difference and establish a unified data format; S2: multi-source risk feature extraction, based on biological mechanism and epidemiological evidence, extract susceptible site features in gene data, key pollution factor features in environmental exposure data and high-risk behavior features in living habit data, form three risk feature sets, and label the association strength of each feature with the target disease; S3: dynamic interaction modeling, construct a gene-environment-living habit ternary interaction model, quantify the synergistic effect of gene-environment interaction, gene-living habit interaction and environment-living habit interaction through nonlinear algorithm, generate dynamic risk transmission path, simulate the cumulative effect of different risk factor combinations on the probability of occurrence of target disease; S4: risk transmission network construction, integrate three risk feature sets and dynamic risk transmission path into a multi-dimensional risk transmission network, wherein the nodes correspond to risk factors, the edges represent the interaction strength and direction between risk factors, and the network weight is determined by biological pathway verification and statistical significance; S5: key risk path screening, based on the topological properties of the risk transmission network, use centrality analysis and sensitivity analysis methods to identify the key path with the highest risk contribution to the target disease, eliminate non-significant associated paths, and form a simplified risk-driven subnetwork; S6: individualized risk assessment, input the standardized data of the target individual into the risk-driven subnetwork, combine the output results of the dynamic interaction model, calculate the disease risk score of the target individual within a specific time window, and generate a risk assessment report containing risk level, key driving factors and potential intervention targets; S7: result output and update maintenance, output the risk assessment report in a visual form, including risk trend chart, key risk factor heat map and intervention suggestion list; Regularly update the risk transmission network and dynamic interaction model according to the newly acquired genomic, environmental or living habit data, realize the dynamic iterative optimization of risk assessment results.
Citation Information
Patent Citations
Children nephropathy health management method and system based on multi-source data fusion
CN120260898A
Disease diagnosis prediction method and system based on graph neural network
CN120340822A