Type 2 diabetes mellitus high-risk group identification system based on dynamic risk prediction

By using a dynamic risk prediction system that combines Lasso-Cox risk proportion regression and latent class growth model, the problems of large-scale screening and insufficient model interpretability in existing technologies are solved. This enables accurate identification and management of high-risk groups for type 2 diabetes, improves identification accuracy and interpretability, and supports applications in primary healthcare.

CN121726091APending Publication Date: 2026-03-24PEKING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve rapid, low-cost, large-scale screening of high-risk individuals with type 2 diabetes while ensuring identification effectiveness. Furthermore, existing models lack the ability to model the cumulative effects and trend changes of individual health risks, making it difficult to accurately control the timing of preventive interventions. Moreover, black-box models cannot provide structured explanations of variable contribution and decision-making paths, making them difficult to implement in primary healthcare and chronic disease management systems.

Method used

A dynamic risk prediction-based system is adopted, including a data interface module, a data preprocessing module, a dynamic risk modeling module, a trajectory recognition and risk classification module, and a result visualization module. The system utilizes the Lasso-Cox proportional hazards regression model and the latent class growth model to construct a dynamic risk prediction model. By combining the linear mixed model and the Cox proportional hazards regression model, the system can achieve dynamic modeling and trajectory recognition of individual health risks.

Benefits of technology

It enables the prediction of individual disease incidence probability at any future point in time, clarifies trend types, improves identification accuracy and clinical credibility, supports system integration and hierarchical management, enhances model interpretability, improves acceptance by medical staff, and is practical in primary healthcare scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121726091A_ABST
    Figure CN121726091A_ABST
Patent Text Reader

Abstract

The invention discloses a type 2 diabetes mellitus high-risk group identification system based on dynamic risk prediction, which relates to the technical field of medical health and comprises a data interface module, a data preprocessing module, a dynamic risk modeling module, a track identification and risk typing module and a result visualization module. The data interface module is used for accessing original structured data in a health information big data platform or a medical data warehouse, and data sources comprise resident health examination records, laboratory inspection indexes, outpatient diagnosis information, medicine use and the like. Through modeling capability breakthrough, dynamic prediction and risk trend double fusion, compared with a traditional static scoring method or a single modeling structure, the method adopts a combined modeling and latent category growth model linkage modeling strategy to form a continuous risk function output and trajectory recognition coexisting double expression model, and the risk prediction and trajectory recognition efficiency is improved. And a basis is provided for health management and decision support of high-risk groups of type 2 diabetes mellitus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical and health technology, specifically to a system for identifying high-risk individuals for type 2 diabetes based on dynamic risk prediction. Background Technology

[0002] Currently, the identification of high-risk individuals for type 2 diabetes mainly relies on two basic technical approaches. One is to rapidly assess the current disease risk using risk scoring tables constructed from blood biochemistry tests or non-invasive indicators. The other is to use an individual's health examination data at a specific point in history as a basis to predict their current risk of developing the disease. Both methods have been widely used for initial screening of high-risk individuals, but neither can achieve rapid, low-cost, large-scale screening while ensuring identification effectiveness. Blood biochemistry-based identification methods rely on invasive testing, which is time-consuming and labor-intensive. Especially in high-frequency screening scenarios for large populations, the complexity of operation and resource consumption cannot meet efficiency requirements. While non-invasive prediction using risk scoring tables can significantly reduce the screening burden, its predictive performance is generally weak, with insufficient sensitivity and decreased specificity, making it difficult to form a reliable high-risk individual identification mechanism. The other prediction method requires the use of an individual's historical health data, but this type of data usually has problems such as irregular time windows, inconsistent variable dimensions, and high data missing rates. These problems limit the predictive model in terms of timeliness, consistency, and generalization ability, ultimately affecting its performance and applicability in real screening scenarios.

[0003] Chinese patent CN102063568A discloses "an individual-level diabetes prediction model, comprising the following: (1) recording the subject's own information; (2) converting the acquired information into model data; (3) inputting the model information conversion result into the calculation system; (4) calculating the probability of disease and judging the risk of disease." This invention can effectively identify high-risk groups for diabetes and intervene in their lifestyles as early as possible, which can better prevent or delay the onset of diabetes and has great public health significance and economic value. Therefore, the purpose of this invention is to establish an effective individual diabetes risk prediction model suitable for the Chinese population. For communities, clinics, and even residents, the risk factor information required by the model can be filled in to obtain the individual's diabetes risk for the next n years.

[0004] Existing technologies only address the technical problem that China currently lacks data for establishing individual-level diabetes prediction models, yet there is a great need for a diabetes prediction model suitable for the characteristics of the Chinese population.

[0005] Some pioneering models attempt to introduce machine learning methods (such as random forests and XGBoost) to improve modeling accuracy. These methods may achieve high predictive performance on the training set, but the complex model structure, uninterpretable parameters, and overfitting make the prediction results difficult for clinicians and administrators to understand and adopt. In addition, machine learning algorithms usually rely on large amounts of continuous and structurally consistent data, which face limitations such as data sparsity and insufficient scale in scenarios such as community healthcare and chronic disease follow-up, resulting in insufficient generalizability and adaptability of the models.

[0006] The above technologies have been widely used for identifying high-risk groups in T2DM, or have been promoted on a small scale, but there are still key structural deficiencies.

[0007] First, static calculation models, based solely on historical single-point health data, cannot identify the cumulative effects and trend changes of individual health risks, lacking the ability to model "future risk trends," thus hindering the precise control of preventive intervention timing. Second, while existing black-box models can improve predictive performance, they rely on high-dimensional features and large-sample training, are sensitive to input data quality and sample size, and cannot provide structured explanations such as variable contribution and decision-making paths, limiting their application in primary healthcare and chronic disease management systems. Third, current high-risk identification systems are mostly "model-as-output," lacking the data governance, dynamic prediction, trajectory recognition, and result visualization capabilities required for integration with real medical information platforms, making them difficult to embed in actual public health business processes. Summary of the Invention

[0008] The purpose of this invention is to provide a system for identifying high-risk individuals for type 2 diabetes based on dynamic risk prediction, in order to solve the problems mentioned in the background art.

[0009] To achieve the above objectives, the present invention provides the following technical solution: a system for identifying high-risk individuals for type 2 diabetes based on dynamic risk prediction, comprising a data interface module, a data preprocessing module, a dynamic risk modeling module, a trajectory recognition and risk classification module, and a result visualization module;

[0010] The data interface module is used to access raw structured data from a health information big data platform or medical data warehouse. Data sources include residents' health check-up records, laboratory test indicators, outpatient diagnosis information, and drug usage, etc.

[0011] The data preprocessing module receives standard structured data from the interface module and performs further processing to meet the technical requirements of model construction.

[0012] The dynamic risk modeling module is based on a joint modeling method, which dynamically models an individual's longitudinal health variables and the outcome of type 2 diabetes. The modeling uses a time-dependent Lasso-Cox proportional hazards regression model to select variables.

[0013] The trajectory recognition and risk classification module identifies the risk change patterns of individuals under different follow-up periods. Based on the dynamic risk prediction results of type 2 diabetes, this module uses a latent class growth model to classify and analyze the dynamic change patterns of T2DM incidence risk.

[0014] The results visualization module displays the system's modeling output to users, including individual risk prediction curves, risk trajectory categories, ranking of key variable contributions, and risk level labels. It can also simultaneously display the historical records of the original variables involved in the modeling process for an individual, supporting longitudinal backtracking and manual verification.

[0015] Preferably, the data preprocessing module ensures that the original data can be reliably input according to the data format, field structure and type requirements required by the system through field mapping, data type standardization, unit conversion, timestamp unification and integrity verification, while ensuring the security and stability of the data transmission process.

[0016] Preferably, the data preprocessing module processes data including: missing value imputation, outlier identification and correction, time window compression and reconstruction, and variable derivation and transformation, ultimately forming a continuous, standardized dataset that can be used for modeling.

[0017] Preferably, the specific screening process in the dynamic risk modeling module is as follows:

[0018] Let the first Individuals in time covariates are Its survival time is The deletion instruction is The risk function of the Cox model is then:

[0019]

[0020] in It is the benchmark risk function. The regression coefficients are the covariates.

[0021] To introduce Lasso regularization for variable selection, the objective function is optimized as follows:

[0022]

[0023] in, It is a partial likelihood function. For regularization parameters, express The absolute value of the i-th regression coefficient, This is the vector of variable coefficients obtained through Lasso estimation;

[0024] Considering censoring, the partial likelihood function is defined as:

[0025]

[0026] in, For censored indicator variables, For the first The individual's observation time, In time The group of individuals that are still in the risk range Indicates the first Individuals in time covariate values, This represents the covariates of other individuals within the risk set.

[0027] Preferably, the dynamic risk prediction model constructed by the joint modeling method in the dynamic risk modeling module first constructs a longitudinal sub-model through a linear mixture model to predict the changes of different measurements over time, and then constructs a survival sub-model using a Cox proportional hazards regression model. The results of the longitudinal sub-model are then input into the survival sub-model as parameters to achieve an accurate prediction of the risk of disease. The specific construction process is as follows:

[0028] Linear mixed models are used to describe the changing trends of longitudinal data, that is, to model the dynamic changes of repeated measures variables (such as biomarker levels) of individuals over time. They can simultaneously capture the average trend at the population level and the differences between individuals, through a fixed effects component (…). ) captures the overall trend of the group, the random effects part ( It reflects individual differences and provides personalized predictions:

[0029]

[0030] Where i represents an individual and j represents the observation time. Let i be the observed value of individual i during the j-th measurement. For fixed effects variables, For random effects variables, and For the corresponding coefficients, it is usually assumed that Error term ;

[0031] The Cox proportional hazards model is used to predict an individual's risk of developing the disease. It dynamically incorporates the output of a linear mixture model through a common time variable t. (This allows for real-time updates to risk estimates. The model assumes that the covariates have an exponential relationship with the risk, and the baseline risk function...) This allows for greater flexibility as no specific format needs to be specified.

[0032]

[0033] in This represents the risk of developing the disease in individual i at time t. This represents an individual's baseline risk at time t. It is a covariate. These are the corresponding regression coefficients. This is the output of the linear mixture model, i.e., the prediction result for the longitudinal data. The parameters connecting longitudinal and survival data are obtained by maximizing likelihood estimation, quantifying the dynamic impact of longitudinal features on disease risk.

[0034] Preferably, the specific construction process and trajectory determination method in the trajectory recognition and risk classification module are as follows:

[0035] The essence of LCGM is to divide an individual's longitudinal change trajectory into several latent classes, each representing a common change pattern. The model assumes that individuals in the sample can be divided into several latent classes (such as low-risk, medium-risk growth, and high-risk growth), and that members of each class share similar change trajectories.

[0036]

[0037] in, Let i be the risk value of individual i at time t. Let represent the initial risk level, linear rate of change, and quadratic rate of change for individual i, respectively. This is the error term;

[0038] Latent categories are introduced by class indicator variables. express;

[0039]

[0040] in, Let k be the probability that an individual belongs to the latent class k. As covariates, For the corresponding coefficients;

[0041] By comparing the fitting effects of disease trajectory models under different k values, the sum of squared errors, Akaike information criterion, and Bayesian information of each model are compared to determine the optimal number of potential categories. The stability of the model parameter estimation is tested by the Bootstrap method. Typical trajectory categories are formed based on the optimal number of clusters, and structured classification labels are generated for further hierarchical intervention and path planning.

[0042] Preferably, the front end of the results visualization module supports the export and printing of charts and tables, which facilitates the generation of individual assessment reports, follow-up tracking, or paper retention, and does not involve structured data interfaces or data push functions between systems. Its main goal is to serve human judgment and decision support.

[0043] Compared with the prior art, the beneficial effects of the present invention are:

[0044] 1. This invention achieves a breakthrough in modeling capabilities, integrating dynamic prediction with risk trend analysis:

[0045] Compared with traditional static scoring methods or single modeling structures, this invention adopts a joint modeling and latent category growth model linkage modeling strategy to form a dual expression model with continuous risk function output and trajectory recognition. This mechanism can output the individual's disease probability at any future time point, while clarifying the trend type, realizing the dual-dimensional information expression of "disease risk + development path".

[0046] 2. This invention improves both identification accuracy and clinical reliability:

[0047] By setting a dual high-risk judgment standard of "risk threshold + high-risk trajectory", the misjudgment problem caused by a single prediction model near the boundary value is effectively reduced, and the screening effect and intervention efficiency are improved.

[0048] 3. This invention provides a clear structural result, supporting system integration and hierarchical management:

[0049] Risk trajectory tags are output in a structured coding format, including fields such as trajectory category, risk value, trend direction, and update time. They can be directly embedded into chronic disease management platforms or third-party systems to achieve automatic risk classification and strategy allocation for the population, which helps to promote the construction of individualized intervention pathways.

[0050] 4. This invention improves the acceptance of medical personnel by enhancing explainability:

[0051] The system interface can simultaneously display risk prediction charts, trajectory classifications, historical values ​​of input variables, and model contribution rankings, making the source of predictions clear at a glance and solving the problem of "incomprehensible" results from black-box models.

[0052] 5. This invention utilizes a visual application format to serve grassroots scenarios:

[0053] The print output function makes the system practical in scenarios such as family doctor contract signing, paper-based follow-up visits, notification of the elderly, and training and education. Output content includes individualized risk charts and trajectory analysis summaries, laying a solid foundation for generating medically interpretable reports and improving the model's actual usage and acceptance. Attached Figure Description

[0054] Figure 1 A schematic diagram of the overall structure is provided for an embodiment of the present invention. Detailed Implementation

[0055] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] Please see Figure 1 The present invention provides a technical solution: a system for identifying high-risk groups of type 2 diabetes based on dynamic risk prediction, including a data interface module, a data preprocessing module, a dynamic risk modeling module, a trajectory recognition and risk classification module, and a result visualization module;

[0057] The data interface module is used to access raw structured data from a health information big data platform or medical data warehouse. Data sources include residents' health check-up records, laboratory test indicators, outpatient diagnosis information, and drug usage, etc.

[0058] The data preprocessing module receives standard structured data from the interface module and performs further processing to meet the technical requirements of model construction.

[0059] The dynamic risk modeling module, based on a joint modeling approach, dynamically models an individual's longitudinal health variables and the outcome of type 2 diabetes. The modeling employs a time-dependent Lasso-Cox proportional hazards regression model for variable selection.

[0060] The trajectory recognition and risk classification module identifies the risk change patterns of individuals under different follow-up periods. Based on the dynamic risk prediction results of type 2 diabetes, this module uses the Latent Class Growth Model (LCGM) to classify and analyze the dynamic change patterns of T2DM incidence risk.

[0061] The results visualization module displays the system's modeling output to users (individuals identifying risks, medical personnel, public health managers, etc.), including individual risk prediction curves, risk trajectory categories, ranking of key variable contributions, and risk level labels. It can also simultaneously display the records of the original variables involved in the modeling process for each individual, supporting longitudinal backtracking and manual verification.

[0062] The data preprocessing module ensures that the raw data can be reliably input according to the data format, field structure and type requirements required by the system through field mapping, data type standardization, unit conversion, timestamp unification and integrity verification, while ensuring the security and stability of the data transmission process.

[0063] The data preprocessing module handles the following tasks: missing value imputation (e.g., multiple imputation), outlier identification and correction, time window compression and reconstruction (e.g., aggregation of variables by year or quarter), and variable derivation and transformation (e.g., weight change rate, blood glucose coefficient of variation, etc.), ultimately forming a continuous, standardized dataset that can be used for modeling.

[0064] The specific screening process in the dynamic risk modeling module is as follows:

[0065] Let the first Individuals in time covariates are Its survival time is The deletion instruction is The risk function of the Cox model is then:

[0066]

[0067] in It is the benchmark risk function. The regression coefficients are the covariates.

[0068] To introduce Lasso regularization for variable selection, the objective function is optimized as follows:

[0069]

[0070] in, It is a partial likelihood function. For regularization parameters, express The absolute value of the i-th regression coefficient, This is the vector of variable coefficients obtained through Lasso estimation;

[0071] Considering censoring, the partial likelihood function is defined as:

[0072]

[0073] in, For censored indicator variables, For the first The individual's observation time, In time The group of individuals that are still in the risk range Indicates the first Individuals in time covariate values, Represents the covariates of other individuals within the risk set;

[0074] Internal 10-fold cross-validation selects λ where the model error is minimized; β≠0 indicates statistical significance.

[0075] The time-dependent Lasso-Cox risk proportion regression model ultimately identified 23 potential characteristic factors, and then through joint modeling, 19 characteristic factors were finally determined to enter the dynamic risk prediction model.

[0076] The dynamic risk modeling module uses a joint modeling method to construct a dynamic risk prediction model. First, a longitudinal sub-model is built using a linear mixture model to predict the changes of different measurements over time. Then, a survival sub-model is built using a Cox proportional hazards regression model. The results of the longitudinal sub-model are used as parameters to input into the survival sub-model to achieve an accurate prediction of the risk of disease. The specific construction process is as follows:

[0077] Linear mixed models are used to describe the changing trends of longitudinal data, that is, to model the dynamic changes of repeated measures variables (such as biomarker levels) of individuals over time. They can simultaneously capture the average trend at the population level and the differences between individuals, through a fixed effects component (…). ) captures the overall trend of the group, the random effects part ( It reflects individual differences and provides personalized predictions:

[0078]

[0079] Where i represents an individual and j represents the observation time. Let i be the observed value of individual i during the j-th measurement. For fixed effects variables, For random effects variables, and For the corresponding coefficients, it is usually assumed that Error term ;

[0080] The Cox proportional hazards model is used to predict an individual's risk of developing the disease. It dynamically incorporates the output of a linear mixture model through a common time variable t. (This allows for real-time updates to risk estimates. The model assumes that the covariates have an exponential relationship with the risk, and the baseline risk function...) This allows for greater flexibility as no specific format needs to be specified.

[0081]

[0082] in This represents the risk of developing the disease in individual i at time t. This represents an individual's baseline risk at time t. It is a covariate. These are the corresponding regression coefficients. This is the output of the linear mixture model, i.e., the prediction result for the longitudinal data. The parameters connecting longitudinal and survival data are obtained by maximizing likelihood estimation, quantifying the dynamic impact of longitudinal features on disease risk.

[0083] In the joint model, the dynamic trend of longitudinal data is extracted through a linear mixture model and passed as a time-related covariate to the Cox model, and then transmitted through parameters. Linking individual dynamic characteristics with survival risk, and combining the parameters of the joint model. Parameter estimation is performed by maximizing the following joint likelihood function;

[0084]

[0085] in Let be the likelihood function of the longitudinal data in the linear mixed model. For survival data in the Cox model, likelihood estimation can simultaneously consider the joint distribution of longitudinal data and survival data, capturing the complex dynamic characteristics of disease progression.

[0086] P < 0.05 was used as the inclusion criterion for model feature factors. The C-index, receiver operating characteristic curve, and area under the curve (AUC) were used to evaluate the model's discriminative power. Internal validation was performed using the bootstrap method, with 500 resampling runs on the final model. The mean of the C-index and its 95% CI were calculated. The calibration curve was used to evaluate the model's calibration. The proportional hazards hypothesis was tested using the Schoenfeld residual method, and no violation of this hypothesis was found (P = 0.90).

[0087] The joint modeling method (constructing a longitudinal sub-model through a linear mixture model to predict the changes of different measurements over time, then using a Cox proportional hazards regression model to construct a survival sub-model, and inputting the results of the longitudinal sub-model as parameters into the survival sub-model) constructs a dynamic prediction model to achieve accurate prediction of disease risk, while also making the model have good interpretability and stability. The model can automatically refresh the output results based on new data during individual follow-up and generate a disease risk curve that changes over time.

[0088] The specific construction process and trajectory determination method in the trajectory recognition and risk classification module are as follows:

[0089] The essence of LCGM is to divide an individual's longitudinal change trajectory into several latent classes, each representing a common change pattern. The model assumes that individuals in the sample can be divided into several latent classes (such as low-risk, medium-risk growth, and high-risk growth), and that members of each class share similar change trajectories.

[0090]

[0091] in, Let i be the risk value of individual i at time t. Let represent the initial risk level, linear rate of change, and quadratic rate of change for individual i, respectively. This is the error term;

[0092] Latent categories are introduced by class indicator variables. express;

[0093]

[0094] in, Let k be the probability that an individual belongs to the latent class k. As covariates, For the corresponding coefficients;

[0095] k-value fitting compares the fitting effects of disease trajectory models under different k values, compares the sum of squares due to error (SSE), Akaike information criterion (AIC), and Bayesian information criterion (BIC) of each model to determine the optimal number of potential categories, and uses the Bootstrap method to test the stability of model parameter estimation. Based on the optimal number of clusters, typical trajectory categories are formed, such as "rapid rise", "slow progression", and "low fluctuation", and structured classification labels are generated for further hierarchical intervention and path planning.

[0096] Criteria for identifying high-risk groups: Based on the dynamic risk model number output by the model, the system calculates the probability of an individual developing the disease within a fixed observation window (e.g., 1, 3, 5, 7, 10 years), fits the disease risk trajectory, and defines high-risk groups as those who meet any of the following conditions.

[0097] a) The predicted incidence rate is ≥20% in the next 3 years (this can be judged in conjunction with specific clinical scenarios);

[0098] b) Classified as either "rapidly rising" or "high-risk sustained" trajectory categories.

[0099] The results visualization module supports the export and printing of charts and tables (such as PDF, PNG, CSV) in the front end, which facilitates the generation of individual assessment reports, follow-up tracking, or paper retention. It does not involve structured data interfaces or data push functions between systems. Its main goal is to serve human judgment and decision support.

[0100] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0101] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A system for identifying high-risk individuals for type 2 diabetes based on dynamic risk prediction, characterized in that: It includes a data interface module, a data preprocessing module, a dynamic risk modeling module, a trajectory recognition and risk classification module, and a results visualization module; The data interface module is used to access raw structured data in health information big data platforms or medical data warehouses through API calls, ETL processes, and FHIR standard protocols. The data preprocessing module receives standard structured data from the interface module and performs further processing to meet the technical requirements of model construction. The dynamic risk modeling module is based on a joint modeling method, which dynamically models an individual's longitudinal health variables and the outcome of type 2 diabetes. The modeling uses a time-dependent Lasso-Cox proportional hazards regression model to select variables. The trajectory recognition and risk classification module identifies the risk change patterns of individuals under different follow-up periods. Based on the dynamic risk prediction results of type 2 diabetes, this module uses a latent class growth model to classify and analyze the dynamic change patterns of T2DM incidence risk. The results visualization module displays the system's modeling output to users, including individual risk prediction curves, risk trajectory categories, ranking of key variable contributions, and risk level labels. It can also simultaneously display the historical records of the original variables involved in the modeling process for an individual, supporting longitudinal backtracking and manual verification.

2. The type 2 diabetes high-risk population identification system based on dynamic risk prediction according to claim 1, characterized in that: The data preprocessing module ensures that the raw data can be reliably input according to the data format, field structure and type requirements required by the system through field mapping, data type standardization, unit conversion, timestamp unification and integrity verification.

3. The type 2 diabetes high-risk population identification system based on dynamic risk prediction according to claim 2, characterized in that: After processing, the data preprocessing module ultimately produces a continuous, standardized dataset that can be used for modeling.

4. The type 2 diabetes high-risk population identification system based on dynamic risk prediction according to claim 3, characterized in that: The specific screening process in the dynamic risk modeling module is as follows: Let the first Individuals in time covariates are Its survival time is The deletion instruction is The risk function of the Cox model is then: ; in It is the benchmark risk function. The regression coefficients are the covariates. To introduce Lasso regularization for variable selection, the objective function is optimized as follows: ; in, This is a partial likelihood function. For regularization parameters, express The absolute value of the i-th regression coefficient, This is the vector of variable coefficients obtained through Lasso estimation; Considering censoring, the partial likelihood function is defined as: ; in, For censored indicator variables, For the first The individual's observation time, In time The group of individuals that are still in the risk range Indicates the first Individuals in time covariate values, This represents the covariates of other individuals within the risk set.

5. The type 2 diabetes high-risk population identification system based on dynamic risk prediction according to claim 4, characterized in that: The dynamic risk modeling module uses a joint modeling method to construct a dynamic risk prediction model. First, a longitudinal sub-model is built using a linear mixture model to predict the changes of different measurements over time. Then, a survival sub-model is built using a Cox proportional hazards regression model. The results of the longitudinal sub-model are used as parameters to input into the survival sub-model to achieve an accurate prediction of the risk of disease. The specific construction process is as follows: Linear mixed models are used to describe the changing trends of longitudinal data, that is, to model the dynamic changes of repeated measures variables of individuals over time. They can simultaneously capture the average trend at the group level and the differences between individuals, through a fixed effects component (…). ) captures the overall trend of the group, the random effects part ( It reflects individual differences and provides personalized predictions: ; Where i represents an individual and j represents the observation time. Let i be the observed value of individual i during the j-th measurement. For fixed effects variables, For random effects variables, and These are the corresponding coefficients. It is usually assumed that... Error term ; The Cox proportional hazards model is used to predict an individual's risk of developing the disease. It dynamically incorporates the output of a linear mixture model through a common time variable t. (This allows for real-time updates to risk estimates. The model assumes that the covariates have an exponential relationship with the risk, and the baseline risk function...) This allows for greater flexibility as no specific format needs to be specified. ; in This represents the risk of developing the disease in individual i at time t. This represents an individual's baseline risk at time t. It is a covariate. These are the corresponding regression coefficients. This is the output of the linear mixture model, i.e., the prediction result for the longitudinal data. The parameters connecting longitudinal and survival data are obtained by maximizing likelihood estimation, quantifying the dynamic impact of longitudinal features on disease risk.

6. The type 2 diabetes high-risk population identification system based on dynamic risk prediction according to claim 5, characterized in that: The specific construction process and trajectory determination method in the trajectory recognition and risk classification module are as follows: The essence of LCGM is to divide an individual's longitudinal change trajectory into several latent classes, each representing a common change pattern. It also integrates a dynamic risk prediction model constructed using joint modeling into LCGM for joint application. The model assumes that individuals in the sample can be divided into several latent classes, and that members of each class share similar change trajectories. ; in, Let i be the risk value of individual i at time t. Let represent the initial risk level, linear rate of change, and quadratic rate of change for individual i, respectively. This is the error term; Latent categories are introduced by class indicator variables. express; ; in, Let k be the probability that an individual belongs to the latent class k. As covariates, For the corresponding coefficients; By comparing the fitting effect of disease trajectory models under different k values, the sum of squared errors, Akaike information criterion, and Bayesian information content of each model are compared. Typical trajectory categories are formed based on the optimal number of clusters, and structured classification labels are generated for further hierarchical intervention and path planning.

7. The type 2 diabetes high-risk population identification system based on dynamic risk prediction according to claim 6, characterized in that: The results visualization module supports the export and printing of charts and tables, facilitating the generation of individual assessment reports, follow-up tracking, or paper-based record keeping. It also supports the synchronous display of longitudinal variable records and risk prediction curves, enabling time-stamp-based retrospective verification. The visualization results feature data security and storage capabilities, with the primary goal of supporting human judgment and decision-making.

Citation Information

Patent Citations

  • Individual diabetes mellitus prediction model

    CN102063568A