Chronic inflammatory disease risk prediction method and system and storage medium
By constructing a structured variation interaction network and time series analysis, and integrating genetic, environmental, and lifestyle data, the problem of early identification of high-risk patients with chronic inflammatory diseases was solved, enabling personalized disease prediction and dynamic monitoring, and improving the accuracy and personalization of prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AFFILIATED HOSPITAL OF GUANGDONG MEDICAL UNIV
- Filing Date
- 2026-01-07
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies struggle to identify high-risk patients with chronic inflammatory diseases in the early stages and lack personalized disease prediction and intervention methods. They also fail to effectively integrate genetic data with dynamic physiological indicators, leading to prediction results that deviate from the actual biological process and lack temporal dynamism.
By constructing a structured variant interaction network, integrating patients' genetic data, environmental exposure, and lifestyle data, and using feature propagation technology to extract enhanced variant feature sets, and combining time series analysis to generate time-enhanced variant feature sequences, risk factor assessments and personalized health intervention recommendations are conducted.
It enables precise, time-series, and personalized prediction of the risk of chronic inflammatory diseases, improves the accuracy and reliability of prediction, provides dynamic monitoring and early warning, and generates highly customized health recommendations.
Smart Images

Figure CN121922362A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biomedical technology, and in particular to a method, system and storage medium for predicting the risk of chronic inflammatory diseases. Background Technology
[0002] Chronic inflammatory diseases, such as chronic obstructive pulmonary disease (COPD), ankylosing spondylitis, and rheumatoid arthritis, are major causes of death and disease burden worldwide. Traditional diagnostic methods for chronic inflammatory diseases rely primarily on clinical symptoms, laboratory tests, and imaging examinations. These methods typically only detect the disease in its middle to late stages, making accurate prediction of early-stage disease progression difficult. Therefore, identifying high-risk patients early and implementing personalized disease prevention and intervention remains a significant challenge in current medical research.
[0003] Currently, disease risk prediction methods based on genetic data suffer from the following limitations: First, most methods rely on static analysis of single gene variants, making it difficult to effectively capture the complex interactions between different variants and between genes and the environment, leading to predictions that deviate from the actual biological process. Second, existing technologies often separate genetic information from patients' dynamic physiological indicators, historical health records, and other temporal information, failing to reflect the evolution of risk over time and resulting in predictions lacking temporal dynamism. Furthermore, how to construct feature models that truly reflect an individual's genetic background, living environment, and health status based on multi-source heterogeneous data, and generate interpretable and actionable personalized assessment reports, remains a common challenge for the current technological system.
[0004] Therefore, how to achieve deep integration of multi-dimensional information in terms of technology, build an interactive network that can dynamically reflect individual characteristics, and on this basis realize accurate, temporal, and personalized prediction of the risk of chronic inflammatory diseases has become a key problem that urgently needs to be solved in this field. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies, this application provides a method, system, and storage medium for predicting the risk of chronic inflammatory diseases, which can integrate multi-dimensional information to achieve dynamic and accurate prediction of chronic inflammatory disease risk, and generate personalized health intervention suggestions based on the risk prediction results.
[0006] In a first aspect, this application provides a method for predicting the risk of chronic inflammatory diseases, the method comprising: Step 1: Obtain raw sequence data related to chronic inflammatory diseases from the patient gene database, clean and standardize them, and construct a structured variant interaction network containing individualized patient characteristics based on the processed data; Step 2: For the structured variant interaction network, integrate the patient's real-time physiological indicators and dynamic health status data, and extract the set of enhanced variant features that can reflect gene-environment interactions through feature propagation technology. Step 3: Based on the enhanced variant feature set, perform time series analysis by combining the patient's historical health data time series records to generate a time series enhanced variant feature sequence for characterizing dynamic changes in risk; Step 4: Evaluate the risk factors of the time-enhanced variation feature sequence, determine the weight distribution of each risk factor, and generate a preliminary risk score sequence accordingly. Step 5: Based on the preliminary risk score sequence, perform probability distribution modeling and bias correction to obtain a comprehensive risk prediction value for chronic inflammatory diseases; Step 6: Extract key risk information based on the comprehensive risk prediction value, and combine it with hierarchical visualization technology to generate a chronic inflammatory disease risk assessment report that includes personalized health intervention recommendations.
[0007] Secondly, this application provides a chronic inflammatory disease risk prediction system, the system comprising: The network construction module is used to obtain raw sequence data related to chronic inflammatory diseases from patient gene databases, clean and standardize them, and construct a structured variant interaction network containing individualized patient characteristics based on the processed data. The set extraction module is used to extract an enhanced set of variant features that reflect gene-environment interactions by integrating real-time physiological indicators and dynamic health status data of patients for structured variant interaction networks and through feature propagation technology. The time series analysis module is used to perform time series analysis based on the enhanced variant feature set and the patient's historical health data time series records, and generate a time series enhanced variant feature sequence to characterize the dynamic changes in risk. The weighting assessment module is used to assess the risk factors of the time-enhanced variation feature sequence, determine the weight distribution of each risk factor, and generate a preliminary risk score sequence accordingly. The risk prediction module is used to perform probability distribution modeling and bias correction based on the preliminary risk score sequence to obtain a comprehensive risk prediction value for chronic inflammatory diseases. The report generation module is used to extract key risk information based on comprehensive risk prediction values and, combined with hierarchical visualization technology, generate a chronic inflammatory disease risk assessment report that includes personalized health intervention recommendations.
[0008] Thirdly, this application provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the aforementioned method for predicting the risk of chronic inflammatory diseases.
[0009] Compared with the prior art, the beneficial effects of the technical solution of this application are at least as follows: 1. By constructing a structured interaction network that integrates gene variation, environmental exposure, and lifestyle habits, and by using feature propagation technology to capture the synergistic effects among multiple factors, the model can more comprehensively reflect the true pathogenesis of diseases, significantly reduce the prediction bias caused by ignoring key interactions, and improve the accuracy and reliability of risk prediction.
[0010] 2. By introducing time series analysis, static gene data is combined with dynamic changes in health status. The generated time-enhanced feature sequences can characterize the evolution trajectory of risk factors, realizing the dynamism and foresight of risk assessment, and enabling dynamic monitoring and early warning of disease risks.
[0011] 3. Individual medical records and family genetic backgrounds are incorporated from the data cleaning stage, and the final report combines hierarchical visualization and personalized suggestion generation rules, so that the prediction results and health advice are closely aligned with the actual situation of specific patients. This provides a highly customized scientific basis for clinical decision-making and personal health management, enhancing the personalization level and practicality of the model.
[0012] 4. In several key stages such as data processing, feature extraction, risk calculation, and report generation, integrity assessment, multi-source verification, and iterative optimization mechanisms have been set up, forming a closed-loop quality control process. This effectively ensures the accuracy and stability of the entire chain from raw data to final output, and guarantees the overall robustness of the technical solution. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a flowchart of a method for predicting the risk of chronic inflammatory diseases according to this application; Figure 2 This is a comparative diagram of the main performance indicators of the chronic inflammatory disease risk prediction methods in the embodiments of this application; Figure 3 This is a schematic diagram comparing ROC curves in the embodiments of this application; Figure 4 This is a schematic diagram comparing the detection capabilities at different disease stages in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of a chronic inflammatory disease risk prediction system according to this application. Detailed Implementation
[0015] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms “comprising” or “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] For ease of understanding, the specific process of the embodiments of this application is described below. Figure 1 The diagram shows a flowchart of a method for predicting the risk of chronic inflammatory diseases provided by the present invention. The flowchart specifically includes the following steps: Step 1: Obtain raw sequence data related to chronic inflammatory diseases from the patient gene database, clean and standardize them, and construct a structured variant interaction network containing individualized patient characteristics based on the processed data.
[0017] In one specific embodiment, step 1 involves obtaining raw sequence data related to chronic inflammatory diseases from a patient gene database, and then cleaning and standardizing it, including: Raw sequence data containing multiple variant forms were extracted from the patient's gene database, and the corresponding variant site information was obtained. The variant forms included at least single nucleotide variants and insertion / deletion variants. The distribution of gene mutation sites was deeply analyzed, and a standardized filter was used to remove non-mutation-related repetitive segments and segments unrelated to chronic inflammatory diseases. By combining patients' medical history and family genetic background data, cross-analysis and association strength calculation were performed on the retained variant sites, and a basic dataset of gene variants was constructed accordingly. The integrity of the basic dataset of gene variations is assessed and multi-source validation is performed, and inconsistent variant sites are corrected.
[0018] Specifically, to achieve accurate risk prediction for chronic inflammatory diseases, a structured variation interaction network that reflects an individual's genetic background and environmental characteristics was constructed in the initial stage of data processing.
[0019] Raw sequence data was extracted from patient gene databases containing various variants associated with chronic inflammatory diseases, such as single nucleotide variants (SNVs) and insertion / deletion variants (INDELs). Information on the corresponding variant sites was also obtained. SNVs include substitutions between adenine and thymine, and between guanine and cytosine, while insertion / deletion variants involve the insertion or deletion of base segments, ranging in length from one to several bases. The data was sourced from standardized gene databases, such as clinical genetics databases, genome-wide association study (GWAS) databases, or public genomic variant databases. By acquiring this gene sequence data, detailed information based on each patient's genetic background can be provided, enabling personalized risk assessment in disease prediction.
[0020] After obtaining the raw data, systematic cleaning and standardization processing is required to remove noise and standardize the data format. This process first removes invalid information unrelated to the disease, including redundant fragments and missing values. The cleaning process involves the application of a standardization filter to remove all gene variant fragments unrelated to chronic inflammatory diseases. For example, the standardization filter identifies and removes non-variant-related repetitive fragments using a preset base frequency threshold. These fragments typically consist of continuously repeating identical base sequences, appearing too frequently in different patient sequences, and are not directly related to chronic inflammatory diseases. Simultaneously, based on publicly available databases of genes related to chronic inflammatory diseases, gene fragments unrelated to these diseases are screened out and removed, such as gene fragments related to bone development. Retaining these fragments would increase the burden of subsequent data processing and affect prediction accuracy, as they do not participate in the occurrence and development of chronic inflammatory diseases. After initial denoising, the distribution and location information of gene variant sites are obtained.
[0021] Patient history records, particularly information on diagnoses of chronic inflammatory diseases and detailed family genetic background data, were cross-analyzed with filtered and retained variant sites. This analysis calculated the association strength between each variant site and the (disease) clinical phenotype, for example, by assessing the difference in frequency of specific variants between the patient and control groups using statistical tests. For example, the formula for calculating the association strength is as follows: The dataset is defined as follows: N11 represents the number of individuals carrying the variant in the patient group; N10 represents the number of individuals not carrying the variant in the patient group; N01 represents the number of individuals carrying the variant in the control group; and N00 represents the number of individuals not carrying the variant in the control group. The association strength calculation results are used to screen and label variants, identifying variant sites with association strengths higher than a preset threshold, thus forming a basic dataset of gene variants rich in personalized background information. This dataset not only includes the variant sites themselves but also incorporates their quantitative association information with an individual's clinical history.
[0022] To ensure data quality and accuracy, the basic genetic variation dataset needs to undergo a completeness assessment. This assessment checks the completeness, correctness, and consistency of the data by comparing the patient data with public databases or other patient datasets. If, during this step, data on certain variant sites is found to be incomplete or inconsistent with existing standard data, the dataset is deemed incomplete, and the original sequence data must be reacquired to supplement and extract the missing information.
[0023] In multi-source data validation, the variant site information in the patient's genetic variant dataset is compared with standard sequences in public genetic databases (such as the NCBI Genetic Database) to verify the accuracy of the variant sites. If significant discrepancies are found between different databases, the inconsistent variant sites are corrected. The correction process requires re-verification of the original sequence data, medical records, and family genetic data to eliminate errors in data extraction or calculation, ensuring the accuracy of the genetic variant dataset. For example, if a genetic variant site has been repeatedly confirmed to be associated with chronic inflammatory diseases in multiple data sources, but there are missing or erroneous records in one database, other reliable data sources will be used to fill in these gaps, ensuring that accurate genetic variant data is ultimately obtained.
[0024] In one specific embodiment, step 1 involves constructing a structured variational interaction network containing individualized patient characteristics based on the processed data, including: The initial network structure is constructed by using the mutation sites in the basic gene mutation dataset as nodes; The strength of the interaction is assessed by calculating the co-occurrence frequency of the corresponding variants of each node in the patient's medical record information, and the statistical correlation between variants is quantified by combining the Pearson correlation coefficient. The two are then fused to determine the edge weights of the initial network. The patient's environmental exposure data and lifestyle data are quantified into attribute features and incorporated into the initial network as node attributes or edge attributes to generate an initial variant interaction network with individual characteristics. The connectivity of the initial mutated interactive network is evaluated. If the average node connectivity is lower than the preset connectivity threshold, the weight calculation method is adjusted and the edge weights are updated. The structured mutation interaction network is obtained by optimizing the initial mutation interaction network after adjusting the edge weights based on the fusion of multi-source individual feature data.
[0025] Specifically, variant site data from the basic dataset of gene variations are used as nodes in the network. Each variant site represents a specific type of variation that occurs in an individual's genome, such as single nucleotide variants (SNVs) and insertion / deletion variants (INDELs). These variant sites, as nodes in the network, form the initial structure of the network, where each node carries the corresponding gene variation information.
[0026] To further construct and optimize the network structure, the interaction strength between each variant site (node) and other information in the patient's medical records was calculated. Medical records provide data on the patient's historical health status, including disease diagnosis, course of disease, and treatment response. By analyzing the frequency of gene variant sites in different patient medical records, the interaction strength between different variants was assessed; high-frequency co-occurrence suggests potential biological synergies or functional associations. To supplement frequency information and capture linear dependencies, the Pearson correlation coefficient of the incidence of each pair of variant nodes in the patient population was calculated as a quantification of association. The interaction strength and statistical association were fused, for example, through a weighted product operation, to generate a comprehensive edge weight value, thus completing the initial network edge weight assignment. These edge weight values reflect the potential synergistic effects between gene variant sites and connect different nodes in the network as edges.
[0027] To personalize the network, non-genetic factors need to be integrated. Environmental exposure data typically includes the duration and concentration of air pollutant exposure, as well as exposure to dust, chemical reagents, etc., and may also include environmental factors such as temperature and humidity, which are quantified into exposure indices (e.g., multiplying the duration of air pollutant exposure by the concentration to obtain the exposure index as the quantification result). Lifestyle data includes behavioral habits such as diet, exercise, smoking, and drinking, which are converted into behavioral scores according to standard scales (e.g., multiplying the number of years of smoking by the number of cigarettes smoked per day to obtain the smoking impact value, converting the frequency of drinking by the number of times per week to the drinking index, and converting the duration of exercise by the cumulative duration per week to the exercise score). These quantified features are assigned as node attributes to relevant variant nodes. For example, a variant node related to respiratory inflammation is associated with the patient's pack-years of smoking; or, as edge attributes, they affect the connection strength between specific node pairs. For example, for a pair of known co-occurring variants, if the patient has a history of high pollution exposure, the weight of its connection edge will receive an additional adjustment factor. By integrating these environmental and lifestyle data into the network, not only are the personalized characteristics of each node enhanced, but a more detailed mapping of the interaction between genes and the environment is also provided. The initial variation interaction network constructed in this way then possesses the ability to reflect individual differences and can better describe the complex interactions between genes and factors such as environment and lifestyle.
[0028] After generating an initial variant interaction network with individualized characteristics, its structural robustness needs to be evaluated. Network connectivity is an indicator that measures the tightness of relationships between nodes in the network, reflecting the network's structural complexity. The average node connectivity of the network is calculated, which is the average number of edges connecting all nodes. A preset connectivity threshold is set based on the average number of associations with variant sites related to chronic inflammatory diseases. If this average connectivity is lower than the preset threshold, it indicates that the network is too sparse, potentially containing information silos and unable to effectively support subsequent feature propagation. This situation triggers a mechanism to adjust the weight calculation method, for example, switching from Pearson correlation coefficient-based calculation to cosine similarity-based calculation. Cosine similarity measures similarity by calculating the ratio of the inner product of two variant occurrence vectors to their norm product; it is more sensitive to sparse data and can uncover more potential weak connections. After switching the calculation method, the similarity of all node pairs is recalculated and the edge weights are updated to improve network connectivity.
[0029] After obtaining a network with satisfactory connectivity, multi-source individual feature data fusion is performed to optimize network stability. This process takes genetic variation data, environmental exposure data, lifestyle data, and real-time physiological indicators of patients as multi-dimensional inputs. Principal component analysis (PCA) is applied to fuse these data. PCA extracts the main directions of data variation by solving for the eigenvalues and eigenvectors of the data covariance matrix, achieving dimensionality reduction and noise reduction. Based on the extracted principal components, the edge weights in the network are globally adjusted; for example, the weights are recalibrated based on the loads of the nodes connected by the edges on important principal components. This weight optimization based on data fusion enables the network structure to more robustly reflect the complex interactions between genes, environment, and lifestyle, thereby outputting a stable and highly individualized structured variation interaction network.
[0030] Step 2: For the structured variant interaction network, integrate the patient's real-time physiological indicators and dynamic health status data, and extract the set of enhanced variant features that can reflect gene-environment interactions through feature propagation technology.
[0031] In one specific embodiment, the process of performing step 2 may specifically include the following steps: Real-time physiological indicator data and dynamic health status data are superimposed as dynamic node attributes onto the structured variation interaction network. A graph convolutional network is applied to propagate features of the superimposed network to obtain the propagated node embedding vectors. The proportion of interaction effects between mutated nodes is calculated and defined as the first interaction effect proportion. It is determined whether the first interaction effect proportion is higher than the first preset proportion threshold. If it is, a feature subset related to the interaction is extracted from the node embedding vector. If not, a single mutation feature is extracted while extracting the feature subset. The single mutation feature is a feature description of a single gene mutation site itself. Integrate feature subsets and / or single variant features to generate an initial set of individualized variant features; The initial individualized variant feature set is validated using multi-source data. Then, the feature coverage rate is calculated. If the feature coverage rate is lower than the preset coverage rate threshold, the propagation parameters of the graph convolutional network are adjusted and iteratively optimized until an enhanced variant feature set that meets the coverage requirements is generated.
[0032] Specifically, real-time physiological indicators, such as inflammatory factor concentrations, blood pressure, heart rate, blood glucose, and blood oxygen saturation—indicators related to chronic inflammatory diseases—as well as dynamic health status data, such as daily symptom scores, weight change trajectories, symptom frequency, and sleep duration and quality—are overlaid as dynamic node attributes onto an existing structured variation interaction network. In practice, each node representing a specific gene variation in the network is assigned an attribute vector. This vector is formed by concatenating the aforementioned real-time and dynamic data after standardization, thereby binding the static gene variation network to the dynamic individual physiological context.
[0033] The network with superimposed dynamic attributes is then fed into a graph convolutional network module for feature propagation. The graph convolutional network takes as input the node feature matrix and adjacency matrix of the structured variation interaction network. The node feature matrix contains the basic genetic variation features of each node and dynamic node attributes (real-time physiological indicators and dynamic health status data). The adjacency matrix is composed of network edge weights. Through the convolutional operations of the graph convolutional network, the feature information of each node and its neighboring nodes is aggregated. For example, for a given node, the contribution of neighboring node features is allocated according to the edge weights in the adjacency matrix; the higher the edge weight of a neighboring node, the greater its weight proportion in the aggregation process. After multiple layers of convolutional propagation, the feature information of each node is updated, and the final output is a propagated node embedding vector. This vector integrates the node's own features with the features of other related nodes in the network, enabling it to capture the potential connections between variations and the interactive relationships between genes and real-time physiological and health data.
[0034] The proportion of interaction effects between variant nodes is calculated and defined as the first interaction effect proportion. Specifically, it is calculated by selecting any two variant nodes and calculating their interaction effect value using their node embedding vectors (interaction effect value = dot product of node embedding vectors / product of the magnitudes of the two vectors). The sum of all interaction effect values between variant nodes in the network is then calculated, and the proportion of a single interaction effect value to the total sum is the first interaction effect proportion between those two variant nodes. A first preset proportion threshold is set, determined based on clinical research data on gene variant interactions in chronic inflammatory diseases. If the first interaction effect proportion between two variant nodes is higher than this threshold, it indicates that their interaction has a significant impact on the risk of chronic inflammatory diseases. A feature subset related to this interaction is extracted from the node embedding vectors. This feature subset contains vector fragments that can characterize the interaction pattern. If the first interaction effect proportion is not higher than this threshold, while extracting the corresponding interaction feature subset, a single variant feature is also extracted. The single variant feature is a feature description of a single gene variant site itself, including variant type, association strength, and corresponding dynamic physiological indicator response values, ensuring that no single variant factor affecting disease risk is overlooked.
[0035] Next, the features extracted in the above steps are integrated. Whether it is a subset containing only interaction features or a combination of interaction features and single variant features, they will be merged to form an initial set of individualized variant features. This set constitutes a preliminary mathematical representation of the patient's genetic background and its interaction with environmental and physiological factors.
[0036] To ensure the reliability of this feature set, quality assessment and iterative optimization are required. Raw sequence data from patient gene databases, historical medical records, family genetic background data, and public chronic inflammatory disease gene feature databases are selected as validation data sources. Each feature in the initial individualized variant feature set is compared with the corresponding information in each validation data source. For example, the variant interaction patterns in the features are matched with confirmed chronic inflammation-related variant interaction patterns in the public database. The proportion of matching features to the total number of features is counted to assess the reliability of the initial set. If inconsistent features are found, errors in the data processing process are checked and corrected. Subsequently, the feature coverage rate of the set is calculated. Feature coverage is the ratio of the number of chronic inflammatory disease-related variant feature types in the initial individualized variant feature set to the total number of known chronic inflammatory disease-related variant feature types. If the calculated feature coverage rate is lower than a preset coverage threshold, the current feature set is deemed to have insufficient coverage of disease-related information, and an iterative optimization process is initiated. Optimization is achieved by adjusting the propagation parameters of the graph convolutional network. These parameters may include the number of graph convolutional layers, kernel size, neighbor node feature aggregation weights, and / or the type of activation function. For example, the number of propagation layers can be increased to capture deeper node interaction features, and the convolutional kernel size can be adjusted to optimize the granularity of feature extraction. After parameter adjustment, the system re-executes all steps from feature propagation to feature integration: dynamic attributes are superimposed onto the network, the adjusted graph convolutional network is run to obtain new node embedding vectors, the proportion of the first interaction effect is recalculated and conditional feature extraction is performed, and finally, a new initial individualized variant feature set is generated and its coverage is evaluated again. This "propagation-extraction-validation-evaluation-adjustment" loop continues until the feature coverage of the generated feature set meets the preset requirements. At this point, the loop terminates, and the set is output as the final enhanced variant feature set.
[0037] Step 3: Based on the enhanced variant feature set, perform time series analysis by combining the patient's historical health data time series records to generate a time-enhanced variant feature sequence for characterizing dynamic changes in risk.
[0038] In one specific embodiment, the process of performing step 3 may specifically include the following steps: Obtain an enhanced set of variant features, combine patient-specific data stratification results with temporal fluctuation records of risk factors for chronic inflammatory diseases, and generate initial time-series data; If the sequence length of the initial time series data exceeds the preset length threshold, the initial time series data will be divided into multiple subsequences, and subsequences with similar fluctuation patterns will be merged. Based on the specific influencing factors of chronic inflammatory diseases, the merged time series data were labeled with features, and an initial variation feature sequence was generated based on the labeling results. The initial variation feature sequence is decomposed using multi-dimensional time series analysis techniques to identify and correct abnormal fluctuation points. Fluctuation anomaly detection is performed on the corrected initial variant feature sequence. If the fluctuation standard deviation exceeds the preset standard threshold, the subsequence segmentation and reconstruction are re-executed until a time-enhanced variant feature sequence that meets the fluctuation requirements is generated.
[0039] Specifically, the enhanced variant feature set contains quantitative indicators reflecting gene-environment interactions. After obtaining the enhanced variant feature set, the patient's personalized data stratification results and the time fluctuation records of chronic inflammatory disease risk factors are retrieved. The personalized data stratification results are divided based on dimensions such as patient age, gender, disease stage, and family genetic predisposition. For example, by age, patients are divided into youth, middle-aged, and elderly levels; by disease stage, they are divided into acute phase, remission phase, and chronic phase levels. The time fluctuation records of chronic inflammatory disease risk factors cover the changes in factors such as inflammatory factor concentration, gene variant expression level, and environmental exposure intensity over time, with the time dimension in weeks and the recording period ranging from 1 to 3 years. Each feature in the enhanced variant feature set is associated and matched with the corresponding personalized data stratification label and the time fluctuation data of risk factors. For example, the feature of "specific gene insertion / deletion variant" is bound to "middle-aged level" and "inflammatory factor concentration fluctuation data in the past 6 months," arranged in chronological order to form initial time-series data. Each time-series data point contains feature value, stratification label, and risk factor fluctuation value at the corresponding time point, establishing a multidimensional correspondence between features and time and individual stratification.
[0040] Initial time-series data undergoes structural optimization before being input into the analysis process. The sequence length of the time-series data, i.e., the total number of time points included, is checked. If this length exceeds a preset threshold, such as more than one hundred consecutive monitoring days, a data segmentation mechanism is triggered. The segmentation operation divides the long sequence into multiple shorter subsequences, based on a fixed time window, for example, every twenty time points constitute one subsequence. Subsequently, these subsequences are evaluated for similarity, calculating the volatility pattern parameters of each subsequence, including volatility amplitude (the difference between the maximum and minimum values of the eigenvalues in the subsequence), volatility frequency (the number of times the eigenvalues peak), and trend slope (the linear regression slope of the subsequence eigenvalues). The similarity of the volatility pattern parameters between any two subsequences is calculated using the cosine similarity formula. If the similarity between two subsequences is higher than a preset similarity threshold, they are determined to have similar volatility patterns and are merged into a longer, consistent sequence segment. This segmentation and merging mechanism aims to reduce the complexity of long-sequence computation while preserving and strengthening recurring, representative risk evolution patterns in the data.
[0041] After structural optimization, the time-series data enters the feature annotation stage. In this stage, data points are labeled based on specific influencing factors of chronic inflammatory diseases. These factors are derived from a medical knowledge base and include disease-related gene variant types, inflammatory marker levels, rapid decline rates of lung function in the short term, and environmental triggers (such as allergens and pollutants). Each influencing factor corresponds to a specific risk level classification standard. For example, a C-reactive protein concentration higher than 10 mg / L is labeled as "high risk," and a single nucleotide variant expression level of a specific gene higher than twice the average level is also labeled as "high risk." The entire time-series data is scanned to identify data points or time periods that meet these factor definitions and label them as high-risk. For example, if the "gene variant expression level" feature value at a certain time point corresponds to a "high risk" level, it is labeled as "1," and if it corresponds to a "low risk" level, it is labeled as "0." Based on these annotation results, an initial variant feature sequence is generated. This sequence not only contains the original numerical features but also incorporates label information representing the risk level.
[0042] The initial variation feature sequence needs to undergo accuracy optimization to eliminate noise interference. Optimization is achieved through multi-dimensional time series analysis techniques, specifically using a seasonal decomposition method to break down the sequence into three components: a long-term trend term, a seasonal periodic term, and a residual term. The long-term trend term reveals the macroscopic direction of change in risk factors, the seasonal periodic term captures regular fluctuations caused by seasonal changes or periodic treatments, and the residual term contains the remaining information after removing the trend and periodicity, representing fluctuations caused by random interference factors. Based on this decomposition, data points that deviate excessively from the normal range in the residual term are identified as anomalous fluctuation points. These anomalous fluctuation points may represent mutations in the patient's health status or significant changes in gene mutations. These anomalous points are corrected based on the contextual information provided by the trend and periodicity terms, for example, by replacing them with the moving average of neighboring points or predicted values based on the overall trend of the sequence. Preferably, the causes of abnormalities are investigated by combining the patient's clinical records and environmental exposure records from the same period. If the abnormality is caused by accidental factors (such as a single acute infection or short-term high-intensity environmental exposure), the characteristic value of the abnormal fluctuation point is corrected by linear interpolation, that is, the average characteristic value of two adjacent time points before and after the abnormal point is used to replace the abnormal value. If the abnormality is caused by disease progression or changes in gene expression, the abnormal point is retained and marked as a "critical change point" to ensure that the correction process does not destroy the true risk change trend.
[0043] The corrected sequences still need to undergo fluctuation anomaly detection to ensure their overall smoothness and stability. The detection metric is the standard deviation of the entire sequence's fluctuation, which reflects the dispersion of data points around their mean. The calculated standard deviation is compared to a preset threshold, determined based on statistical analysis of data from stable-phase patients. If the sequence's standard deviation exceeds this threshold, it indicates the presence of severe fluctuations or structural breaks in the sequence that were not fully processed in the previous stage. In this case, the subsequence segmentation and reconstruction process is restarted, the data is re-divided, pattern matched and merged, and subsequent annotation and decomposition correction steps are performed again. This iterative cycle of "segmentation-merging-annotation-decomposition-detection" continues until the generated sequence's standard deviation of fluctuation is below the preset threshold. At this point, the sequence is determined to meet the fluctuation requirements and is output as a time-enhanced variant feature sequence.
[0044] The initial time-series data reflects the changes in dynamic physiological indicators over time on a static individual genetic risk basis. The processed time-enhanced variant feature sequence, which is a denoised, labeled, and structurally optimized sequence, can more clearly and stably characterize the dynamic evolution trajectory of risk status associated with chronic inflammatory diseases, providing high-quality input for subsequent risk quantification.
[0045] Step 4: Evaluate the risk factors of the time-enhanced variant feature sequence, determine the weight distribution of each risk factor, and generate a preliminary risk score sequence accordingly.
[0046] In one specific embodiment, the process of performing step 4 may specifically include the following steps: Analyze the frequency of occurrence and influence intensity of each risk factor in the time-enhanced variant feature sequence to determine the weight distribution; The proportion of the interaction effect among each risk factor is calculated and defined as the proportion of the second interaction effect. Time-dependent feature extraction technology is used to process the time-enhanced variation feature sequence. Combined with the dynamic weight adjustment rule and the proportion of the second interaction effect, the weights of each risk factor are dynamically optimized. Based on the optimized risk factor weights, an initial risk score sequence for chronic inflammatory diseases is generated; Analyze the volatility of the initial risk score sequence. If its standard deviation exceeds the preset volatility range, adjust the parameters in the dynamic weight adjustment rule and iterate to optimize and obtain an intermediate risk score sequence that meets the requirements. The stability of the intermediate risk score sequence is optimized by using multi-dimensional risk factor analysis techniques to obtain a preliminary risk score sequence.
[0047] Specifically, the basic weight of each risk factor in the time-enhanced variant feature sequence is evaluated based on its frequency of occurrence within the sequence time range and the influence strength represented by the factor's numerical value. First, the time-series data corresponding to each risk factor are extracted from the sequence. Risk factors include gene variants (such as single nucleotide variants and insertion / deletion variants), physiological indicators (such as inflammatory factor concentration and blood pressure), environmental exposure factors (such as pollutant exposure intensity), and lifestyle factors (such as the impact of smoking). The frequency of each risk factor in the time-series is calculated, i.e., the ratio of the number of time points where the corresponding feature value is non-zero to the total number of time points. Simultaneously, the influence strength is calculated. Influence strength refers to the degree of effect of the risk factor on the risk of chronic inflammatory diseases, determined by directly correlating the factor's feature value (quantitative representation data of the factor, such as gene expression level and inflammatory factor concentration) with the disease risk (logistic regression model regression coefficient). That is, the correlation coefficient quantifies the degree of effect. The correlation coefficient is calculated using a logistic regression model based on clinical case data. The model input is the feature value, the output is the disease probability, and the correlation coefficient is the regression coefficient of the corresponding feature in the model. By combining the frequency of occurrence and the intensity of influence, the initial weights of each risk factor are determined by product or weighted summation, forming an initial weight distribution and establishing the correspondence between risk factors and weights.
[0048] After obtaining the initial weight distribution, a time-dependent feature extraction technique is introduced to capture the dynamic synergistic effects between risk factors. A Long Short-Term Memory (LSTM) network model is used to process the time-enhanced variant feature sequences. This network learns long-term dependencies in the sequence through its internal gating mechanism. The network input is the feature matrix of the time-series sequence, and the output is a feature vector representing temporal context information, containing the dynamic correlation information of risk factors at different time points. Simultaneously, the proportion of interaction effects between each risk factor is calculated; this proportion is defined as the second interaction effect proportion to distinguish it from the interaction evaluation in the feature extraction stage. The calculation of the second interaction effect proportion focuses on the temporal synergistic fluctuations between the features of risk factors. To calculate the second interaction effect proportion between each risk factor, any two risk factors are selected, and their feature values at each time point in the time-series sequence are extracted. The interaction effect value is calculated using the mutual information formula. Where p(x,y) is the joint probability distribution of the two factor eigenvalues, and p(x) and p(y) are the marginal probability distributions of the individual factor eigenvalues. The sum of the pairwise interaction effects of all risk factors is calculated, and the proportion of the single interaction effect value to the sum is the proportion of the second interaction effect.
[0049] The proportion of the second interaction effect, along with contextual information extracted from the time-dependent model, is input into a dynamic weight adjustment rule. This rule can be formulated as a gradient-based optimization process, whose objective function aims to maximize the consistency between risk predictions and actual health status changes. Factor pairs with higher interaction effect proportions receive greater weight increases during gradient updates. Through this process, the weights of risk factors are dynamically optimized to better reflect their true importance in a specific time context.
[0050] Based on the optimized weight distribution, an initial risk score sequence for chronic inflammatory diseases is generated. The generation process is implemented through a weighted aggregation function, which multiplies the values of all factors in the time-enhanced variant feature sequence by their dynamically optimized weights at each time point and sums them to output a continuous risk score value that changes over time, thus forming the initial risk score sequence.
[0051] The initial risk score sequence undergoes volatility analysis to verify its smoothness. The standard deviation of the sequence is calculated; if it exceeds a preset volatility range derived from historical stable-period patient data, it indicates the presence of undesirable sharp fluctuations or abrupt changes in the sequence. This triggers adjustments to parameters in the dynamic weight adjustment rule, such as reducing the optimization learning rate or modifying the coefficient of the smoothness constraint term in the objective function. After parameter adjustment, the process from time-dependent feature extraction to risk score generation is re-executed to produce a new initial risk score sequence, which is then subjected to volatility evaluation again. This iterative optimization loop continues until the output sequence's standard deviation meets the preset requirements; at this point, the sequence is labeled as an intermediate risk score sequence.
[0052] The intermediate risk score sequence underwent further stability optimization using multi-dimensional risk factor analysis, which categorizes risk factors by their origin into genetic, environmental, and lifestyle dimensions. Principal component analysis was performed on factors within each dimension to extract principal components and capture the main variation patterns. Subsequently, these principal components from different dimensions were fused, and the intermediate risk score sequence was smoothed or slightly adjusted based on the fused global information to suppress fluctuations caused by noise in a single dimension, thereby improving the overall stability of the sequence. The sequence obtained after this stability optimization was ultimately determined as the preliminary risk score sequence.
[0053] Step 5: Based on the preliminary risk score sequence, perform probability distribution modeling and bias correction to obtain a comprehensive risk prediction value for chronic inflammatory diseases.
[0054] In one specific embodiment, the process of performing step 5 may specifically include the following steps: For the preliminary risk score sequence, a probability distribution model is applied to calculate its risk probability distribution; The risk probability distribution is corrected based on linear correction logic, and the corrected probability distribution is compared with the historical risk probability benchmark distribution. The deviation between the two is reduced through an iterative correction mechanism to obtain the final probability distribution. Calculate the proportion of the dominant part of the interaction in the preliminary risk score sequence. If the proportion is higher than the preset dominant threshold, adjust the weight of relevant factors through multi-dimensional data fusion technology to highlight the impact of the interaction in the final prediction. Based on the final probability distribution and the weighted adjusted results, an initial comprehensive risk prediction value for chronic inflammatory diseases is generated. The initial comprehensive risk prediction value was cross-validated using multi-source data. The deviation between the validated initial comprehensive risk prediction and the historical clinical baseline is calculated. If the deviation exceeds the preset deviation range, the parameters of the probability distribution model are readjusted and the calculation is iterated until a comprehensive risk prediction that meets the deviation requirements is generated.
[0055] Specifically, the preliminary risk score sequence is fed into a probability distribution model, such as a Gaussian mixture model. This model is suitable for characterizing the continuous distribution of risk scores for chronic inflammatory diseases. The model parameters include the mean and variance of the sequence. The mean is the arithmetic mean of the risk scores at all time points, and the variance is obtained by calculating the sum of the squares of the differences between each risk score and the mean, then dividing by the number of time points. The preliminary risk score sequence is then input into the Gaussian distribution model, and the model parameters are solved using the maximum likelihood estimation method to obtain the risk probability distribution. This distribution is presented as a probability density curve, with the horizontal axis representing the risk score and the vertical axis representing the probability of occurrence of the corresponding score, thus establishing a quantitative correspondence between risk scores and their probabilities.
[0056] After obtaining the risk probability distribution, risk value correction based on linear correction logic is performed. The correction process applies a linear transformation function of the form P_corrected = α1 × P_initial + β, where P_initial is the initial probability value, α1 is the scaling factor, β is the offset, and P_corrected is the corrected probability value. The coefficients α1 and β are pre-set based on the model's performance on the validation set, aiming to systematically adjust the numerical range of probabilities and the calibration curve. The corrected probability distribution is then compared with a historical risk probability baseline distribution, derived from a clinically confirmed risk score distribution observed in a large-scale population cohort. The comparison is achieved by calculating a measure of difference between the two distributions, such as calculating the Kullback-Leibler divergence. If the divergence value is greater than a preset tolerance threshold, an iterative correction mechanism is triggered. This mechanism fine-tunes the parameters of the probability distribution model using gradient descent methods, such as adjusting the mean, variance, or mixture weights of the Gaussian mixture model, thereby generating a new risk probability distribution that is closer to the historical baseline. This "comparison-correction" cycle continues until the difference between the two distributions is below a tolerance threshold, at which point the resulting distribution is determined as the final probability distribution.
[0057] While optimizing the probability distribution, the importance of interaction effects in risk composition is assessed in parallel. This assessment is achieved by calculating the proportion of the interaction-dominant component in the preliminary risk score sequence. The interaction-dominant component refers to the risk score contribution value generated by the interaction of two or more risk factors. First, the risk score composition at each time point is decomposed using a multiple linear regression model. The model input consists of the eigenvalues of each risk factor and pairwise interaction terms, and the output is the risk score. The sum of the risk contribution values corresponding to the interaction terms with non-zero regression coefficients is the value of the interaction-dominant component. The proportion of this value to the total risk score is the percentage of the interaction-dominant component. A preset dominance threshold is set, based on clinical research data on interactions in chronic inflammatory diseases. If the percentage is higher than this threshold, it indicates that the interaction has a significant impact on risk. The weights of relevant factors are adjusted using multi-dimensional data fusion technology. This technology integrates raw data from genetic, physiological, environmental, and lifestyle dimensions, and the weights are adjusted using a weighted average method. Specifically, a subset directly related to the interaction is first extracted from the multi-dimensional data (such as gene variation data and corresponding physiological indicator fluctuation data involved in a high proportion of interaction). This subset is then assigned a weight coefficient positively correlated with the proportion of interaction influence (the higher the proportion of interaction influence, the larger the coefficient). Next, the weight coefficient of this subset is multiplied by the corresponding factor weight in the original weight vector to obtain the adjusted weights of the interaction-related factors. The weights of other non-interaction-related factors remain unchanged or are slightly adjusted proportionally. Then, a weighted average method is used to integrate the adjusted weights of the interaction-related factors and the non-interaction factors. The weight allocation of the weighted average is based on the contribution of each dimension of data to the risk of chronic inflammatory diseases (based on clinical data presets). The final output is a new weight vector. The increase in the weight of interaction-related factors is positively correlated with the proportion of interaction influence, and the weight proportion of risk factors involved in the interaction is significantly increased. This approach retains the rationality of the original weights while accurately highlighting the impact of interactions on risk prediction through the correlation of multi-dimensional data.
[0058] Based on the final probability distribution and the weighted adjusted results, an initial comprehensive risk prediction value for chronic inflammatory diseases is generated. For example, the formula for calculating the initial comprehensive risk prediction value is: Where R0 is the initial comprehensive risk prediction value, P(t) is the risk probability corresponding to time point t in the final probability distribution, and W n For, F n (t) represents the characteristic value of the nth risk factor at time point t, where N is the number of risk factors and T is the length of the time series.
[0059] The initial composite risk prediction is then cross-validated using multi-source data, including the patient's real-time physiological monitoring data, recent clinical diagnostic records, family history of genetic diseases, and disease progression records. For example, the validation process compares this data with the actual disease progression recorded in the patient's medical history and the current health status reflected in the real-time physiological monitoring data. If significant inconsistencies are found—for example, a prediction indicating high risk when recent physiological indicators are all within the normal range—the prediction is flagged and may be slightly revised based on multi-source evidence, resulting in a validated initial composite risk prediction.
[0060] Finally, the deviation between this validated predicted value and the historical clinical baseline is calculated. The historical clinical baseline is derived from the patient's past clinical diagnostic gold standard results or typical values of similar patient groups. The deviation is calculated as the absolute difference |V-Bc|, where V is the predicted value and Bc is the historical clinical baseline. If this deviation exceeds a preset deviation range determined based on clinically acceptable error, it indicates that the current parameter settings of the probability distribution model are still unsatisfactory. The parameters of the probability distribution model will be readjusted, for example, by changing the number of components or the covariance structure of the Gaussian mixture model, and the entire process from probability distribution calculation to multi-source validation will be re-executed. This external iterative loop continues until the deviation between the output composite risk prediction value and the historical clinical baseline falls within the preset range. At this point, the value is determined as the final output composite risk prediction value.
[0061] Step 6: Extract key risk information based on the comprehensive risk prediction value, and combine it with hierarchical visualization technology to generate a chronic inflammatory disease risk assessment report that includes personalized health intervention recommendations.
[0062] In one specific embodiment, the process of performing step 6 may specifically include the following steps: Based on the comprehensive risk prediction value, the portion of the risk value that is higher than the preset risk threshold is selected to form a high-risk subset; The high-risk subset is subjected to hierarchical data processing, and a preliminary suggestion list is generated based on personalized suggestion generation rules; By combining a hierarchical data visualization presentation method with a preliminary suggestion list, and through automated report feedback loop optimization technology, a draft risk assessment report is generated. Calculate the information coverage rate of the initial draft risk assessment report for key risk information. If the information coverage rate is lower than the preset coverage threshold, adjust the hierarchical data visualization presentation method and regenerate the assessment report. The regenerated assessment report is verified using multi-dimensional data verification technology to obtain a chronic inflammatory disease risk assessment report.
[0063] Specifically, the comprehensive risk prediction value is compared with a preset risk threshold, which is set based on the epidemiological data of the target disease. For example, a risk value of 0.7 is used as the threshold, and values higher than this value are considered high risk. From the time-series data and risk factor correlation information corresponding to the comprehensive risk prediction value, all data entries with risk values higher than the preset risk threshold are selected to form a high-risk subset. The high-risk subset includes high-risk time nodes, risk scores at the corresponding time points, the top 5 contributing risk factors (such as specific gene mutations, excessive concentrations of inflammatory factors, long-term exposure to highly polluted environments, etc.), and the type and intensity of interactions among risk factors, establishing a direct correspondence between high-risk levels and core influencing factors.
[0064] The high-risk subset then enters the stratified data processing stage, with stratification dimensions including risk level, risk factor type, and time period. For example, risk levels are divided into moderate-high risk (0.7–0.8), very high risk (0.8–0.9), and extremely high risk (above 0.9) based on risk value; risk factor types are categorized into genetic, physiological, environmental, and lifestyle-related factors; and time periods are divided into short-term sudden onset (1–2 time points), medium-term sustained occurrence (3–5 time points), and long-term recurrence (more than 6 time points) based on the continuity of high-risk occurrence. Simultaneously, a preliminary suggestion list is automatically generated based on personalized suggestion generation rules. These rules correspond one-to-one with the stratification results, and the rule base is derived from clinical intervention guidelines and evidence-based medicine data for chronic inflammatory diseases. For example, for high-risk factors caused by genetic factors, it is recommended to regularly monitor gene expression levels and conduct targeted interventions; for long-term and recurring high-risk factors caused by environmental factors, it is recommended to adjust living areas or strengthen protective measures; for moderate-to-high-risk factors caused by lifestyle factors, it is recommended to develop dietary and exercise plans. Each stratified dimension should have at least two specific and actionable recommendations to ensure that the recommendations are accurately matched with the causes of high-risk factors.
[0065] After data stratification and suggestion generation are completed, a draft risk assessment report is generated by combining the stratified data visualization presentation with the preliminary suggestion list. The visualization presentation can be an interactive network diagram, where nodes represent high-risk variants and are colored hierarchically, and edges represent interaction strength; or a heatmap, showing the trajectory of different risk factors fluctuating over time. The preliminary suggestion list is presented alongside the visualization charts as captions or sidebar explanations. This integration process utilizes an automated report feedback loop optimization technology. This technology first generates a draft containing charts and text, then a simulation verification module scores the draft's clarity and logical coherence. If the score is below internal standards, the chart type, color scheme, or text layout is automatically adjusted, and the report is re-rendered, forming a rapid internal loop optimization.
[0066] After the initial draft of the report is generated, the information coverage of key risk information needs to be evaluated. Key risk information includes high-risk time points, core risk factors, interaction types, and risk change trends. The information coverage calculation formula is Ic = (Nk / Nt) × 100%, where Nk is the number of key risk information items fully presented in the initial draft, and Nt is the total number of all key risk information items. If the calculated information coverage is lower than the preset coverage threshold, it indicates that the initial draft of the report fails to fully reflect the overall risk picture, triggering an adjustment to the visualization presentation. The adjustment may involve switching from a single chart to a dashboard combination, or adding previously hidden secondary connections to the network diagram to reveal a more complex interaction network. After the adjustment, the process from data layering to report generation is re-executed to produce a new initial draft of the report and recalculate the coverage. This cycle continues until the coverage target is met.
[0067] The regenerated assessment report is validated using multi-dimensional data verification technology, which covers three dimensions: data consistency, recommendation rationality, and clinical suitability. Data consistency verification compares the high-risk information in the report with the original data of the preliminary risk score sequence and comprehensive risk prediction value to ensure numerical accuracy. Recommendation rationality verification references the latest clinical guidelines for chronic inflammatory diseases to assess the scientific validity and feasibility of the recommendations. Clinical suitability verification considers the patient's individual characteristics (age, disease duration, underlying diseases) to evaluate the relevance of the recommendations. If data inconsistencies are found during verification, corrections are made at the high-risk subset screening or visualization transformation stage. If recommendations are deemed unreasonable, the rule base is retrieved to update the recommendation entries. If clinical suitability is insufficient, the recommendation details are adjusted based on the patient's individual characteristics. The final chronic inflammatory disease risk assessment report is obtained after verification.
[0068] Figure 2 The bar chart comparing the main performance indicators shows the numerical differences between the proposed method and traditional methods in terms of accuracy, precision, recall, and F1 score. It demonstrates that the proposed method, through multi-source data fusion and dynamic interactive analysis, significantly outperforms traditional methods in all core performance indicators, and has more outstanding prediction accuracy. Figure 3This diagram illustrates the comparison of ROC curves for traditional methods, random methods, and our proposed method in disease risk prediction. The ROC curve depicts the relationship between the false positive rate and the true positive rate. As the decision threshold changes, the shape of the curve reflects how the classifier makes decisions at different thresholds. Generally, the closer the ROC curve is to the top left corner, the better the model's performance. AUC, the area under the ROC curve, is a numerical indicator of the model's classification performance. A value closer to 1 indicates better classification performance, while a value closer to 0.5 indicates worse performance. The diagram uses AUC values to illustrate the performance differences between the two methods. Our proposed method has an AUC value of 0.864, significantly higher than the traditional method's 0.765. This demonstrates that our method, through time-enhanced risk tracking, has a much higher ability to distinguish disease risks than traditional methods, resulting in superior prediction reliability. Figure 4 The line graph comparing the detection capabilities at different disease stages shows the accuracy of this method versus traditional methods at each stage (asymptomatic, early, middle, and late). It demonstrates that this method, through dynamic time-series analysis, significantly improves detection capabilities at each stage (especially the early stage), with an early detection accuracy 45% higher than traditional methods, showcasing its significant advantage in early prediction of chronic inflammatory disease risk. In the three comparative analyses above, the traditional method refers to risk assessment methods that rely on single-dimensional data, static analysis logic, or basic statistical models before the maturity of multi-source data fusion technology and deep application of machine learning. Its core characteristic is that it fails to fully capture the complexity, dynamism, and temporal correlations of disease occurrence.
[0069] The above describes a method for predicting the risk of chronic inflammatory diseases in the embodiments of this application. The following describes a system for predicting the risk of chronic inflammatory diseases in the embodiments of this application. Please refer to [link to relevant documentation]. Figure 5 This application provides a schematic diagram of the structure of a chronic inflammatory disease risk prediction system, which includes: Network building module 10 is used to obtain raw sequence data related to chronic inflammatory diseases from the patient gene database, clean and standardize them, and build a structured variant interaction network containing individualized patient characteristics based on the processed data.
[0070] The set extraction module 20 is used to extract an enhanced set of variant features that can reflect gene-environment interactions by integrating real-time physiological indicators and dynamic health status data of patients for structured variant interaction networks and through feature propagation technology.
[0071] The time series analysis module 30 is used to perform time series analysis based on the enhanced variant feature set and the patient's historical health data time series records, and generate a time-enhanced variant feature sequence to characterize the dynamic changes in risk.
[0072] The weight assessment module 40 is used to assess the risk factors of the time-enhanced variation feature sequence, determine the weight distribution of each risk factor, and generate a preliminary risk score sequence accordingly.
[0073] The risk prediction module 50 is used to perform probability distribution modeling and bias correction based on the preliminary risk score sequence to obtain a comprehensive risk prediction value for chronic inflammatory diseases.
[0074] The report generation module 60 is used to extract key risk information based on the comprehensive risk prediction value and, combined with hierarchical visualization technology, generate a chronic inflammatory disease risk assessment report containing personalized health intervention recommendations.
[0075] This application also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of the method for predicting the risk of chronic inflammatory diseases.
[0076] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for predicting the risk of chronic inflammatory diseases, characterized in that, The methods include: Step 1: Obtain raw sequence data related to chronic inflammatory diseases from the patient gene database, clean and standardize them, and construct a structured variant interaction network containing individualized patient characteristics based on the processed data; Step 2: For the structured variant interaction network, integrate the patient's real-time physiological indicators and dynamic health status data, and extract the set of enhanced variant features that can reflect gene-environment interactions through feature propagation technology. Step 3: Based on the enhanced variant feature set, perform time series analysis by combining the patient's historical health data time series records to generate a time series enhanced variant feature sequence for characterizing dynamic changes in risk; Step 4: Evaluate the risk factors of the time-enhanced variation feature sequence, determine the weight distribution of each risk factor, and generate a preliminary risk score sequence accordingly. Step 5: Based on the preliminary risk score sequence, perform probability distribution modeling and bias correction to obtain a comprehensive risk prediction value for chronic inflammatory diseases; Step 6: Extract key risk information based on the comprehensive risk prediction value, and combine it with hierarchical visualization technology to generate a chronic inflammatory disease risk assessment report that includes personalized health intervention recommendations.
2. The method according to claim 1, characterized in that, In step 1, raw sequence data related to chronic inflammatory diseases are obtained from the patient's gene database and then cleaned and standardized, including: Raw sequence data containing multiple variant forms were extracted from the patient's gene database, and the corresponding variant site information was obtained. The variant forms included at least single nucleotide variants and insertion / deletion variants. The distribution of gene mutation sites was deeply analyzed, and a standardized filter was used to remove non-mutation-related repetitive segments and segments unrelated to chronic inflammatory diseases. By combining patients' medical history and family genetic background data, cross-analysis and association strength calculation were performed on the retained variant sites, and a basic dataset of gene variants was constructed accordingly. The integrity of the basic dataset of gene variations is assessed and multi-source validation is performed, and inconsistent variant sites are corrected.
3. The method according to claim 2, characterized in that, In step 1, a structured variation interaction network containing individualized patient characteristics is constructed based on the processed data, including: The initial network structure is constructed by using the mutation sites in the basic gene mutation dataset as nodes; The strength of the interaction is assessed by calculating the co-occurrence frequency of the corresponding variants of each node in the patient's medical record information, and the statistical correlation between variants is quantified by combining the Pearson correlation coefficient. The two are then fused to determine the edge weights of the initial network. The patient's environmental exposure data and lifestyle data are quantified into attribute features and incorporated into the initial network as node attributes or edge attributes to generate an initial variant interaction network with individual characteristics. The connectivity of the initial mutated interactive network is evaluated. If the average node connectivity is lower than the preset connectivity threshold, the weight calculation method is adjusted and the edge weights are updated. The structured mutation interaction network is obtained by optimizing the initial mutation interaction network after adjusting the edge weights based on the fusion of multi-source individual feature data.
4. The method according to claim 1, characterized in that, Step 2 includes: Real-time physiological indicator data and dynamic health status data are superimposed as dynamic node attributes onto the structured variation interaction network. A graph convolutional network is applied to propagate features of the superimposed network to obtain the propagated node embedding vectors. The proportion of interaction effects between mutated nodes is calculated and defined as the first interaction effect proportion. It is determined whether the first interaction effect proportion is higher than the first preset proportion threshold. If it is, a feature subset related to the interaction is extracted from the node embedding vector. If not, a single mutation feature is extracted while extracting the feature subset. The single mutation feature is a feature description of a single gene mutation site itself. Integrate feature subsets and / or single variant features to generate an initial set of individualized variant features; The initial individualized variant feature set is validated using multi-source data. Then, the feature coverage rate is calculated. If the feature coverage rate is lower than the preset coverage rate threshold, the propagation parameters of the graph convolutional network are adjusted and iteratively optimized until an enhanced variant feature set that meets the coverage requirements is generated.
5. The method according to claim 1, characterized in that, Step 3 includes: Obtain an enhanced set of variant features, combine patient-specific data stratification results with temporal fluctuation records of risk factors for chronic inflammatory diseases, and generate initial time-series data; If the sequence length of the initial time series data exceeds the preset length threshold, the initial time series data will be divided into multiple subsequences, and subsequences with similar fluctuation patterns will be merged. Based on the specific influencing factors of chronic inflammatory diseases, the merged time series data were labeled with features, and an initial variation feature sequence was generated based on the labeling results. The initial variation feature sequence is decomposed using multi-dimensional time series analysis techniques to identify and correct abnormal fluctuation points. Fluctuation anomaly detection is performed on the corrected initial variant feature sequence. If the fluctuation standard deviation exceeds the preset standard threshold, the subsequence segmentation and reconstruction are re-executed until a time-enhanced variant feature sequence that meets the fluctuation requirements is generated.
6. The method according to claim 1, characterized in that, Step 4 includes: Analyze the frequency of occurrence and influence intensity of each risk factor in the time-enhanced variant feature sequence to determine the weight distribution; The proportion of the interaction effect among each risk factor is calculated and defined as the proportion of the second interaction effect. Time-dependent feature extraction technology is used to process the time-enhanced variation feature sequence. Combined with the dynamic weight adjustment rule and the proportion of the second interaction effect, the weights of each risk factor are dynamically optimized. Based on the optimized risk factor weights, an initial risk score sequence for chronic inflammatory diseases is generated; Analyze the volatility of the initial risk score sequence. If its standard deviation exceeds the preset volatility range, adjust the parameters in the dynamic weight adjustment rule and iterate to optimize and obtain an intermediate risk score sequence that meets the requirements. The stability of the intermediate risk score sequence is optimized by using multi-dimensional risk factor analysis techniques to obtain a preliminary risk score sequence.
7. The method according to claim 1, characterized in that, Step 5 includes: For the preliminary risk score sequence, a probability distribution model is applied to calculate its risk probability distribution; The risk probability distribution is corrected based on linear correction logic, and the corrected probability distribution is compared with the historical risk probability benchmark distribution. The deviation between the two is reduced through an iterative correction mechanism to obtain the final probability distribution. Calculate the proportion of the dominant part of the interaction in the preliminary risk score sequence. If the proportion is higher than the preset dominant threshold, adjust the weight of relevant factors through multi-dimensional data fusion technology to highlight the impact of the interaction in the final prediction. Based on the final probability distribution and the weighted adjusted results, an initial comprehensive risk prediction value for chronic inflammatory diseases is generated. The initial comprehensive risk prediction value was cross-validated using multi-source data. The deviation between the validated initial comprehensive risk prediction and the historical clinical baseline is calculated. If the deviation exceeds the preset deviation range, the parameters of the probability distribution model are readjusted and the calculation is iterated until a comprehensive risk prediction that meets the deviation requirements is generated.
8. The method according to claim 1, characterized in that, Step 6 includes: Based on the comprehensive risk prediction value, the portion of the risk value that is higher than the preset risk threshold is selected to form a high-risk subset; The high-risk subset is subjected to hierarchical data processing, and a preliminary suggestion list is generated based on personalized suggestion generation rules; By combining a hierarchical data visualization presentation method with a preliminary suggestion list, and through automated report feedback loop optimization technology, a draft risk assessment report is generated. Calculate the information coverage rate of the initial draft risk assessment report for key risk information. If the information coverage rate is lower than the preset coverage threshold, adjust the hierarchical data visualization presentation method and regenerate the assessment report. The regenerated assessment report is verified using multi-dimensional data verification technology to obtain a chronic inflammatory disease risk assessment report.
9. A chronic inflammatory disease risk prediction system for implementing the method of any one of claims 1 to 8, characterized in that, The system includes: The network construction module is used to obtain raw sequence data related to chronic inflammatory diseases from patient gene databases, clean and standardize them, and construct a structured variant interaction network containing individualized patient characteristics based on the processed data. The set extraction module is used to extract an enhanced set of variant features that reflect gene-environment interactions by integrating real-time physiological indicators and dynamic health status data of patients for structured variant interaction networks and through feature propagation technology. The time series analysis module is used to perform time series analysis based on the enhanced variant feature set and the patient's historical health data time series records, and generate a time series enhanced variant feature sequence to characterize the dynamic changes in risk. The weighting assessment module is used to assess the risk factors of the time-enhanced variation feature sequence, determine the weight distribution of each risk factor, and generate a preliminary risk score sequence accordingly. The risk prediction module is used to perform probability distribution modeling and bias correction based on the preliminary risk score sequence to obtain a comprehensive risk prediction value for chronic inflammatory diseases. The report generation module is used to extract key risk information based on comprehensive risk prediction values and, combined with hierarchical visualization technology, generate a chronic inflammatory disease risk assessment report that includes personalized health intervention recommendations.
10. A computer-readable storage medium storing instructions thereon, characterized in that, When the instruction is executed by the processor, it implements a method for predicting the risk of chronic inflammatory diseases as claimed in any one of claims 1 to 8.