Diabetes risk monitoring method and system based on artificial intelligence

Through the deep LSTM-Cox model combined with local characteristics and long-term trend characteristics, the problem of difficulty in accurately assessing diabetes risk in the prior art is solved, and a more efficient diabetes risk prediction is achieved.

CN120236759AInactive Publication Date: 2025-07-01THE SECOND HOSPITAL OF HEBEI MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510305935.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to accurately assess diabetes risk in combination with short-term local patterns and long-term trends.

Method used

Using a deep LSTM-Cox model-based approach, users' health data are obtained from multiple data sources, and local and long-term trend characteristics are extracted by removing outliers and processing missing data, and fused them to comprehensive prediction of diabetes risk.

Benefits of technology

This achieves a more accurate assessment of individual diabetes risk, improving the comprehensiveness and accuracy of diabetes risk prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236759A_ABST
    Figure CN120236759A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence and medical health monitoring, in particular to a diabetes risk monitoring method and system based on artificial intelligence. The method comprises the following steps: acquiring original data of a user from a plurality of data sources, wherein the original data comprises physiological parameters, living habit data and family history case information; preprocessing the original data, including removing abnormal values and processing missing data; inputting the preprocessed original data to a deep LSTM-Cox model, and evaluating the diabetes risk of the user; outputting the diabetes basic occurrence probability of the user; and calculating an optimized comprehensive diabetes risk score, and dividing the users into different diabetes risk grades. According to the diabetes risk monitoring method and system based on artificial intelligence, local features and long-term trend features of health data can be extracted from different dimensions, comprehensive prediction of the diabetes risk is performed after the features are fused, and the diabetes risk of an individual can be evaluated more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and medical health monitoring, and particularly to a method and system for monitoring diabetes risk based on artificial intelligence. Background Art

[0002] The risk prediction of diabetes involves various health data, including various time series data such as blood glucose levels, body weight, heart rate, and exercise. These data have a high dimension and there is noise and uncertainty, and effective features need to be extracted to improve the accuracy of prediction. The risk of diabetes is not only related to short-term blood glucose fluctuations, such as postprandial blood glucose fluctuations and blood glucose changes after exercise, but also closely related to long-term health trends, such as long-term blood glucose fluctuations and body weight change trends. How to combine short-term local patterns and long-term trends for accurate diabetes risk assessment is a challenging technical problem. Summary of the Invention

[0003] To solve the problems in the above background art, the present invention provides a method and system for monitoring diabetes risk based on artificial intelligence, which can extract local features and long-term trend features of health data from different dimensions, and fuse these features for comprehensive prediction of diabetes risk, and can more accurately evaluate the diabetes risk of an individual.

[0004] The present invention adopts the following technical solutions:

[0005] A method for monitoring diabetes risk based on artificial intelligence, the method comprising:

[0006] S1: Obtain the original data of a user from multiple data sources, including physiological parameters, lifestyle data, and family history case information; preprocess the original data, including removing outliers and processing missing data;

[0007] S2: Input the preprocessed original data into a deep LSTM-Cox model to evaluate the diabetes risk of the user; extract a key feature set that has a significant impact on diabetes risk and a non-key feature set that has a smaller impact from historical health data, and the deep LSTM-Cox model outputs the basic diabetes occurrence probability of the user by training the relationship between the key feature set of historical health data and the diabetes onset situation;

[0008] S3: According to the predicted basic diabetes occurrence probability of the user, combine with the non-key feature set to calculate an optimized comprehensive diabetes risk score, and divide the user into different diabetes risk levels.

[0009] Further, the method for removing outliers includes:

[0010] Construct a random forest for the original data, randomly sample from the original data, construct multiple random decision trees, each tree randomly selects a feature, and divides the data at a random value of this feature, denoted as sample x; calculate the anomaly score S(x), and the calculation formula of the anomaly score S(x) is:

[0011]

[0012] Among them, E(h(x)) is the average path length of sample x in all decision trees; c(n) is the average path length of the sample. Set the samples with S(x)>0.6 as anomaly points and remove this data from the original data; the methods for processing missing data include: using the K-nearest neighbor algorithm to fill in the missing values according to the values of other similar samples in the original data; determining the similarity between the sample where the missing value is located and other samples, and calculating the Euclidean distance; selecting K other samples that are most similar to the sample with the missing value; filling in the missing value through the feature values of the K neighboring samples. For numerical data, the average value, median or weighted average of K neighbors can be selected to fill in the missing value; for categorical data, the category that appears most frequently among K neighbors is usually used to fill in the missing value.

[0013] Furthermore, the method for extracting the key feature set and the non-key feature set with less influence on the diabetes risk based on historical health data includes: calculating the Pearson correlation coefficient between the feature and the diabetes occurrence probability:

[0014]

[0015] Among them, X i is the feature value of the i-th sample, Y i is the diabetes occurrence probability P(DM) of the i-th sample, are the mean values of sample X and diabetes occurrence probability Y respectively, and the value range of r(X,Y) is [-1,1],

[0016] Extract the features with |r|>0.3 as the key feature set that has a significant impact on the diabetes risk, and the remaining features as the non-key feature set with less influence. Record the Pearson coefficient of each feature in the non-key feature set as r n .

[0017] Furthermore, the deep LSTM-Cox model includes: a dimensionality reduction layer, a CNN-LSTM layer, and a Cox layer;

[0018] The dimensionality reduction layer is used to reduce the dimension of the key feature set and output a time series matrix X∈R T×N , where T is the number of time steps and N is the number of features at each time step;

[0019] The CNN-LSTM layer includes a one-dimensional convolutional neural network and a long short-term memory network. The one-dimensional convolutional neural network extracts short-term local features from the time series matrix X, and the long short-term memory network is used to extract long-term trend features from the time series matrix X. The short-term local features and long-term trend features are normalized to generate the feature matrix X combine ;

[0020] The Cox layer constructs a risk function based on the feature matrix X combine and outputs the basic occurrence probability of diabetes for the user.

[0021] Furthermore, the method for dimensionality reduction of the key feature set by the dimensionality reduction layer includes: normalizing the key feature set so that its mean is 0 and variance is 1, and the calculation formula is:

[0022]

[0023] where μ is the feature mean of the samples in the key feature set, and σ is the standard deviation of the samples in the key feature set;

[0024] Construct a data matrix X for the key feature set m Calculate the covariance matrix C:

[0025]

[0026] where m is the number of samples;

[0027] Calculate the eigenvalues and eigenvectors:

[0028] Cv i = λ i v i ;

[0029] where λ i is the eigenvalue, and v i is the corresponding eigenvector;

[0030] Arrange the eigenvalues in descending order and select the largest k eigenvectors to form the transformation matrix W:

[0031] W = [v1, v2,..., v k ;

[0032] Project the key feature set X into the low-dimensional space using the transformation matrix W, and output the time series matrix X ∈ R T×N .

[0033] Furthermore, the method for the one-dimensional convolutional neural network to extract short-term local features from the time series matrix X includes:

[0034] Perform a convolution operation on the time series matrix X, and the output of the convolutional layer is:

[0035]

[0036] Among them, ω i is the weight of the convolution kernel, x t-i+1 is the data of the time series matrix X, b is the bias term, and y t is the convolution result, and k is the size of the convolution kernel;

[0037] Use the max pooling layer to downsample the output of the convolutional layer:

[0038]

[0039] Stack multiple convolutional layers and pooling layers together to further extract short-term local features;

[0040] Furthermore, the method for the long short-term memory network to extract long-term trend features from the time series matrix X includes:

[0041] The long short-term memory network processes the input time series matrix X through memory units and gating mechanisms;

[0042] Gradually adjust the memory units and hidden layer states of the long short-term memory network to output long-term trend features;

[0043] Furthermore, the method for normalizing the short-term local features and long-term trend features to generate the feature matrix X combine includes:

[0044] Normalize each short-term local feature and long-term trend feature to unify the scales of the short-term local features and long-term trend features. The normalization formula is:

[0045]

[0046] Among them, x is the sample original data of the short-term local features and long-term trend features, μ is the sample mean of the short-term local features and long-term trend features, and σ is the sample standard deviation of the short-term local features and long-term trend features;

[0047] Take the short-term local features as one part and the long-term trend features as another part, and concatenate them together by column; assume there are m short-term local features and n long-term trend features, then the dimension of the new matrix is (m + n) × 1, and the combined feature matrix X combine is:

[0048] X combine =[x1, x2,..., x m , x m+1 ,..., x m+n .

[0049] Further, the Cox layer is based on the feature matrix X combine The method for constructing the risk function includes:

[0050] Group the feature matrix X combine by time steps to obtain the feature matrix X at each time point combine (t);

[0051] Based on the feature matrix X combine Construct the risk function, and the risk function is:

[0052] Risk(t) = h(t) = h0(t)·exp(β1x1 + β2x2 +... + β n x n );

[0053] where h(t) is the risk of an individual at time t, and h0(t) is the baseline risk function; x1, x2,..., x n are the features of the feature matrix X combine (t); β1, β2,..., β n are the regression coefficients, indicating the impact of each feature on the diabetes risk;

[0054] Estimate the regression coefficients β1, β2,..., β of each feature through the feature matrix X combine , and calculate the risk of diabetes for each individual at time t. n

[0055] Further, the method for calculating the optimized comprehensive diabetes risk score according to the predicted basic diabetes occurrence probability of the user and combining the non-critical feature set includes:

[0056] Introduce non-critical feature set weight correction and define the optimized diabetes risk score:

[0057] R DM = P(DM) + Σ n r n ·f(X n );

[0058] where P(DM) is the basic diabetes occurrence probability, and f(X n ) is the normalized value of the non-critical feature impact factor.

[0059] Further, an artificial intelligence-based diabetes risk monitoring system, the system is used to implement the above method, and the system includes:

[0060] A data acquisition module, used to collect the user's health data from various devices and sensors:

[0061] ​A data processing module for preprocessing and feature extraction of the collected health data:

[0062] A risk assessment module for assessing the diabetes risk of the user at each time point:

[0063] An early warning module for sending an early warning notice to the user according to the diabetes risk assessment result.

[0064] An electronic device, comprising:

[0065] A processor;

[0066] A memory for storing instructions executable by the processor;

[0067] Wherein, the processor is configured to call the instructions stored in the memory to execute the method of any one of the above.

[0068] Compared with the prior art, the beneficial effects of the present invention are:

[0069] The diabetes risk monitoring method and system based on artificial intelligence of the present invention can extract local features and long-term trend features of health data from different dimensions, and comprehensively predict the diabetes risk after fusing these features, and can more accurately evaluate the diabetes risk of individuals. The convolutional neural network can capture long-term dependencies and is particularly suitable for extracting long-term trends in time series data such as blood glucose and body weight; the long short-term memory network can extract valuable information from local fluctuation patterns. The combination of the two can improve the comprehensiveness and accuracy of diabetes risk prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings according to these drawings without creative efforts.

[0071] Figure 1 It is a flowchart of a diabetes risk monitoring method and system based on artificial intelligence of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0072] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0073] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0074] Specifically, a method for monitoring diabetes risk based on artificial intelligence, the method includes:

[0075] S1: Obtain the user's raw data from multiple data sources, including physiological parameters, lifestyle data, and family history case information. Physiological parameters include blood glucose, blood pressure, heart rate, weight, blood lipids, etc. Lifestyle data includes diet, exercise, sleep, smoking, drinking, etc. and family history case information. The data sources include smart wearable devices, blood glucose sensors, heart rate monitors, weight sensors, smart sphygmomanometers, etc. Physiological data is collected in real time through smart devices and uploaded to the system;

[0076] Preprocess the raw data, including removing outliers and handling missing data; removing outliers and handling missing data are the basis for obtaining accurate and stable diabetes risk prediction results.

[0077] S2: Input the preprocessed raw data into the deep LSTM-Cox model to evaluate the user's diabetes risk; extract a key feature set that has a significant impact on diabetes risk and a non-key feature set with less impact from historical health data. The deep LSTM-Cox model outputs the basic diabetes occurrence probability of the user by training the relationship between the key feature set of historical health data and the diabetes onset situation. Combining short-term local features and long-term trend features into a new feature matrix is an important step to improve the accuracy of the diabetes risk prediction model. By effectively processing and combining these features, the individual's health status can be more comprehensively reflected, thereby improving the prediction accuracy. Combining the characteristics of time series data, local and long-term health change patterns can be extracted to further optimize the diabetes risk assessment system.

[0078] S3: According to the predicted basic diabetes occurrence probability of the user, combined with the non-key feature set, calculate the optimized comprehensive diabetes risk score, and divide the user into different diabetes risk levels. Through machine learning, risk assessment is carried out on important factors to obtain the basic diabetes occurrence probability; at the same time, secondary factors are integrated, thereby improving the accuracy of diabetes risk prediction, with strong scalability. More influencing factors, such as insulin resistance index, genetic data, etc., can be added to the non-key feature set to improve the accuracy of diabetes risk prediction.

[0079] Specifically, the method for removing outliers includes:

[0080] Construct a random forest for the original data. Randomly sample from the original data to construct multiple random decision trees. Each tree randomly selects a feature and divides the data at a random value of this feature, denoted as sample x. Calculate the anomaly score S(x). The calculation formula for the anomaly score S(x) is as follows:

[0081]

[0082] Among them, E(h(x)) is the average path length of sample x in all decision trees; c(n) is the average path length of the sample. Set the samples with S(x) > 0.6 as anomaly points and remove this data from the original data. The methods for handling missing data include: using the K-nearest neighbor algorithm to fill in the missing values according to the values of other similar samples in the original data; determining the similarity between the sample with the missing value and other samples, and calculating the Euclidean distance; selecting K other samples that are most similar to the sample with the missing value; filling in the missing value through the feature values of the K neighboring samples. For numerical data, the average value, median, or weighted average of the K neighbors can be selected to fill in the missing value; for categorical data, the category that appears most frequently among the K neighbors is usually used to fill in the missing value. Diabetes risk prediction involves various health index data, such as blood sugar, weight, exercise situation, eating habits, etc. These data sets often contain some outliers. If not processed, these outliers may affect the prediction effect of the model and lead to inaccurate prediction results.

[0083] Specifically, the methods for extracting the key feature set that has a significant impact on diabetes risk and the non-key feature set that has a smaller impact from historical health data include: calculating the Pearson correlation coefficient between the feature and the probability of diabetes occurrence:

[0084]

[0085] Among them, X i is the feature value of the i-th sample, and Y i is the probability of diabetes occurrence P(DM) of the i-th sample. are the mean values of sample X and the probability of diabetes occurrence Y respectively. The value range of r(X,Y) is [-1, 1].

[0086] Extract the features with |r| > 0.3 as the key feature set that has a significant impact on diabetes risk, and the remaining features as the non-key feature set that has a smaller impact. Record the Pearson coefficient of each feature in the non-key feature set as r n . In the diabetes prediction model, calculate the Pearson correlation coefficient between the health features and the probability of diabetes occurrence to help analyze which features have a significant impact on the occurrence of diabetes and provide guidance for subsequent feature selection and risk assessment. At the same time, the Pearson coefficient of the non-key feature set is rn For subsequent risk optimization.

[0087] Furthermore, the deep LSTM-Cox model includes: a dimensionality reduction layer, a CNN-LSTM layer, and a Cox layer;

[0088] The dimensionality reduction layer is used to reduce the dimensionality of the key feature set and output a time series matrix X ∈ R T×N , where T is the number of time steps and N is the number of features at each time step; thereby reducing data redundancy and retaining the most important features for diabetes risk prediction.

[0089] The CNN-LSTM layer includes a one-dimensional convolutional neural network and a long short-term memory network. The one-dimensional convolutional neural network extracts short-term local features from the time series matrix X; the long short-term memory network is used to extract long-term trend features from the time series matrix X; the short-term local features and long-term trend features are normalized to generate a feature matrix X combine ; The one-dimensional convolutional neural network CNN is good at processing local features and can efficiently extract important local information from the original data. In diabetes prediction, the one-dimensional convolutional neural network CNN can be used to extract local patterns of health monitoring data, such as blood glucose fluctuations and blood glucose changes after exercise, to help capture short-term fluctuations and trends. The long short-term memory network LSTM is a recurrent neural network used to process sequence data and is good at capturing long-term dependencies in time series, and can effectively model the trend of data changing over time. In diabetes prediction, the long short-term memory network LSTM can be used to analyze long-term health trends, such as blood glucose fluctuation trends and weight change curves, and predict future risks. The short-term local features and long-term trend features are normalized to generate a feature matrix X combine , scale the values of different features to the same range, ensure that the short-term local features and long-term trend features have the same scale, so that each feature is treated fairly during model training, avoid certain features dominating the model learning process, and improve the training efficiency and stability of the model.

[0090] The Cox layer constructs a risk function based on the feature matrix X combine and outputs the basic diabetes occurrence probability of the user. The Cox proportional hazards regression model is a widely used statistical method in survival analysis for studying the relationship between time points and the occurrence of diseases.

[0091] Specifically, the method for the dimensionality reduction layer to reduce the dimensionality of the key feature set includes: normalizing the key feature set so that its mean is 0 and variance is 1, and the calculation formula is:

[0092]

[0093] Among them, μ is the feature mean of the samples in the key feature set, and σ is the standard deviation of the samples in the key feature set;

[0094] Construct a data matrix X for the key feature set m Calculate the covariance matrix C:

[0095]

[0096] Among them, m is the number of samples;

[0097] Calculate the eigenvalues and eigenvectors:

[0098] Cv i = λ i v i ;

[0099] Among them, λ i is the eigenvalue, and v i is the corresponding eigenvector;

[0100] Arrange the eigenvalues in descending order, and select the largest k eigenvectors to form the transformation matrix W:

[0101] W = [v1, v2,..., v k ;

[0102] Use the transformation matrix W to project the key feature set X into a low-dimensional space, and output the time series matrix X ∈ R T×N .

[0103] Specifically, the method for a one-dimensional convolutional neural network to extract short-term local features from the time series matrix X includes:

[0104] Perform a convolution operation on the time series matrix X, and the output of the convolutional layer is:

[0105]

[0106] Among them, ω i is the weight of the convolution kernel, x t-i+1 is the data of the time series matrix X, b is the bias term, y t is the convolution result, and k is the size of the convolution kernel;

[0107] Use the max pooling layer to downsample the output of the convolutional layer:

[0108]

[0109] Use multiple convolutional layers and pooling layers stacked together to further extract short-term local features;

[0110] Specifically, the method for a long short-term memory network to extract long-term trend features from the time series matrix X includes:

[0111] The long short - term memory network processes the input time - series matrix X through memory units and gating mechanisms;

[0112] Gradually adjust the memory units and hidden - layer states of the long short - term memory network to output long - term trend features;

[0113] Specifically, the method of standardizing the short - term local features and long - term trend features to generate the feature matrix X combine includes:

[0114] Standardize each short - term local feature and long - term trend feature to unify the scales of the short - term local features and long - term trend features. The standardization formula is:

[0115]

[0116] where x is the sample raw data of the short - term local features and long - term trend features, μ is the sample mean of the short - term local features and long - term trend features, and σ is the sample standard deviation of the short - term local features and long - term trend features;

[0117] Take the short - term local features as one part and the long - term trend features as another part, and splice them together by column. Suppose there are m short - term local features and n long - term trend features, then the dimension of the new matrix is (m + n)×1, and the combined feature matrix X combine is:

[0118] X combine =[x1,x2,...,x m ,x m+1 ,...,x m+n .

[0119] Specifically, the method for the Cox layer to construct the risk function based on the feature matrix X combine includes:

[0120] Group the feature matrix X combine by time steps to obtain the feature matrix X combine (t) at each time point;

[0121] Construct a risk function based on the feature matrix X combine The risk function is:

[0122] Risk(t)=h(t)=h0(t)·exp(β1x1 + β2x2+...+β n x n );

[0123] where h(t) is the risk of an individual at time t, and h0(t) is the baseline risk function; x1,x2,…,x nis the feature matrix X combine (t); β1, β2,..., β n are the regression coefficients, representing the impact of each feature on the risk of diabetes;

[0124] Through the feature matrix X combine The regression coefficients β1, β2,..., β of each feature are estimated through training n , and the risk of diabetes occurrence for each individual at time t is calculated.

[0125] Specifically, the method for calculating the optimized comprehensive diabetes risk score based on the predicted basic diabetes occurrence probability of the user and in combination with the non-critical feature set includes:

[0126] Introduce the weight correction of the non-critical feature set and define the optimized diabetes risk score:

[0127] R DM = P(DM) + Σ n r n ·f(X n );

[0128] Among them, P(DM) is the basic diabetes occurrence probability, and f(X n ) is the normalized value of the non-critical feature impact factor. According to the diabetes occurrence probability, the risk levels are divided. In this embodiment, it is set that R DM ≤20% is the low risk level, representing that the user maintains a healthy lifestyle and has regular physical examinations; it is set that 20% < R DM ≤50% is the medium-low risk level, representing that the user needs to monitor blood glucose and adjust diet and exercise habits; it is set that 50% < R DM ≤80% is the medium-high risk level, representing that the user needs to manage health under the guidance of a doctor and closely monitor blood glucose; it is set that R DM > 80%, and medical intervention and glucose tolerance testing are recommended.

[0129] Specifically, an artificial intelligence-based diabetes risk monitoring system, the system is used to implement the above method, and the system includes:

[0130] The data acquisition module is used to collect the user's health data from various devices and sensors: automatically collect the user's health data, synchronize it with the cloud platform or local database, perform real-time data transmission and update to ensure the timeliness of the monitoring data, and at the same time support the integration of data from different devices and sensors to form comprehensive health data.

[0131] The data processing module is used to preprocess and extract features from the collected health data: thereby converting the time series data into a format suitable for model input and processing time series features.

[0132] A risk assessment module for assessing the diabetes risk of a user at each time point: assessing the diabetes risk at each time point.

[0133] An early warning module for sending an early warning notice to the user according to the diabetes risk assessment result; timely reminding the user to carry out health management, adjust the lifestyle or seek medical treatment.

[0134] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0135] It should be noted that the terms "first", "second", etc. in the description, claims and drawings of the present application are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein.

[0136] The above further describes the present invention with the aid of specific embodiments. However, it should be understood that the specific descriptions herein should not be construed as limiting the essence and scope of the present invention. Various modifications made by those of ordinary skill in the art to the above embodiments after reading this specification all fall within the scope protected by the present invention.

Claims

1. A diabetes risk monitoring method based on artificial intelligence, characterized in that: The method comprises: S1: Obtaining the user's original data from multiple data sources, including physiological parameters, living habits data, and family history case information; preprocessing the original data, including removing outliers and processing missing data; S2: Input the preprocessed raw data into the deep LSTM-Cox model to evaluate the diabetes risk of the user; extract the key feature set that has a significant impact on the diabetes risk and the non-key feature set with a smaller impact based on the historical health data, and the deep LSTM-Cox model outputs the user's basic diabetes occurrence probability by training the relationship between the key feature set of the historical health data and the incidence of diabetes; S3: Based on the predicted basic diabetes occurrence probability of the user, combined with the non-critical feature set, an optimized comprehensive diabetes risk score is calculated to divide the user into different diabetes risk levels.

2. The method according to claim 1, characterized in that The method for removing outliers includes: constructing a random forest for the original data, randomly sampling the original data, constructing multiple random decision trees, each tree randomly selecting a feature, and dividing the data at the random value of the feature, recorded as sample x; calculating the anomaly score S(x), the calculation formula of the anomaly score S(x) is: Wherein, E(h(x)) is the average path length of sample x in all decision trees; c(n) is the average path length of the sample, and the sample with S(x)>0.6 is set as an outlier, and the data in the original data is removed; the method for processing missing data includes: using the K nearest neighbor algorithm to fill the missing value according to the values ​​of other similar samples in the original data; determining the similarity between the sample where the missing value is located and other samples, and calculating the Euclidean distance; selecting K other samples that are most similar to the missing value sample; filling by the characteristic values ​​of the K neighboring samples, for numerical data, the average value, median or weighted average value of the K neighbors can be selected to fill the missing value; for categorical data, the category with the most occurrences among the K neighbors is usually used to fill the missing value.

3. The method according to claim 1, characterized in that The method of extracting the key feature set that has a significant impact on the risk of diabetes and the non-key feature set with a smaller impact based on historical health data includes: calculating the Pearson correlation coefficient between the feature and the probability of diabetes occurrence: Among them, X i is the eigenvalue of the i-th sample, Y i is the probability of diabetes occurrence P(DM) of the i-th sample, are the means of sample X and diabetes probability Y, respectively. The value range of r(X,Y) is [-1,1]. The features with |r|>0.3 are extracted as the key feature set that has a significant impact on the risk of diabetes, and the remaining features are non-key feature sets with less impact. The Pearson coefficient of each feature in the non-key feature set is recorded as r n .

4. The method according to claim 3, characterized in that The deep LSTM-Cox model includes: a dimensionality reduction layer, a CNN-LSTM layer, and a Cox layer; The dimension reduction layer is used to reduce the dimension of the key feature set and output the time series matrix X∈R T×N , where T is the number of time steps and N is the number of features for each time step; The CNN-LSTM layer includes a single-dimensional convolutional neural network and a long short-term memory network. The single-dimensional convolutional neural network extracts short-term local features from the time series matrix X; the long short-term memory network is used to extract long-term trend features from the time series matrix X; the short-term local features and the long-term trend features are standardized to generate a feature matrix X combine ; The Cox layer is based on the feature matrix X combine Construct a risk function and output the user's basic probability of diabetes occurrence.

5. The method according to claim 4, characterized in that The method for reducing the dimension of the key feature set by the dimension reduction layer includes: standardizing the key feature set so that its mean is 0 and its variance is 1, and the calculation formula is: Among them, μ is the characteristic mean of the samples of the key feature set, σ is the standard deviation of the samples of the key feature set; construct the data matrix X for the key feature set m Calculate the covariance matrix C: Where m is the number of samples; Compute eigenvalues ​​and eigenvectors: Cv i =λ i v i ; Among them, λ i is the eigenvalue, v i The corresponding eigenvector; Arrange the eigenvalues ​​in descending order and select the largest k eigenvectors to form the transformation matrix W: <h2 style=";text-align:left;direction:ltr">W=[v1,v2,...,v<h2 style=";text-align:left;direction:ltr"> k <h2 style=";text-align:left;direction:ltr"> ]; Use the transformation matrix W to project the key feature set X into a low-dimensional space and output the time series matrix X∈R T×N .

6. The method according to claim 4, characterized in that The method for extracting short-term local features from a time series matrix X using a single-dimensional convolutional neural network includes: Perform a convolution operation on the time series matrix X, and the output of the convolution layer is: Among them, ω i is the weight of the convolution kernel, x t-i+1 is the data of the time series matrix X, b is the bias term, y t is the convolution result, k is the size of the convolution kernel; Use a max pooling layer to downsample the output of the convolutional layer: Use multiple convolutional layers and pooling layers stacked together to further extract short-term local features; The method for extracting long-term trend features from a time series matrix X using the long short-term memory network includes: The LSTM network processes the input time series matrix X through memory units and gating mechanisms; gradually adjusting the memory units and hidden layer states of the long short-term memory network to output long-term trend characteristics; The short-term local features and long-term trend features are standardized to generate a feature matrix X combine The methods include: Each short-term local feature and long-term trend feature is standardized to unify the scales of the short-term local features and the long-term trend features. The standardized formula is: Among them, x is the sample original data of short-term local features and long-term trend features, μ is the sample mean of short-term local features and long-term trend features, and σ is the sample standard deviation of short-term local features and long-term trend features; short-term local features are taken as one part and long-term trend features are taken as another part, and they are spliced ​​together by column; if there are m short-term local features and n long-term trend features, the dimension of the new matrix is ​​(m+n)×1, and the combined feature matrix X combine for: X combine =[x1,x2,…,x m , x m+1 ,…,x m+n ]。 7. The method according to claim 4, characterized in that The Cox layer is based on the feature matrix X combine Methods for constructing risk functions include: The feature matrix X combine Group by time step and get the feature matrix X at each time point combine (t); Based on the feature matrix X combine Construct a risk function, the risk function is: Risk(t)=h(t)=h0(t)·exp(β1x1+β2x2+...+β n x n ); Where h(t) is the risk of an individual at time t, h0(t) is the baseline risk function; x1, x2, …, x n is the feature matrix X combine (t) characteristics; β1,β2,...,β n is the regression coefficient, indicating the effect of each characteristic on the risk of diabetes; Through the feature matrix X combine Train to estimate the regression coefficients β1, β2, ..., β for each feature n , and calculate the risk of each individual developing diabetes at time t.

8. The method according to claim 3, characterized in that The method for calculating the optimized comprehensive diabetes risk score based on the predicted diabetes basic occurrence probability of the user and in combination with the non-critical feature set includes: Introduce non-critical feature set weight correction and define the optimized diabetes risk score: R DM =P(DM)+Σ n r n ·f(X n ); Among them, P(DM) is the basic probability of diabetes, f(X n ) is the normalized value of the influencing factor of non-critical features.

9. A diabetes risk monitoring system based on artificial intelligence, characterized in that: The system is used to implement the method described in any one of claims 1 to 8, and the system includes: Data collection module, used to collect user health data from various devices and sensors: The data processing module is used to pre-process and extract features of the collected health data: The risk assessment module is used to assess the user's diabetes risk at each time point: The early warning module is used to send early warning notifications to users based on the diabetes risk assessment results.

10. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 8.