A blood glucose prediction method and device based on sparse feature processing and causal analysis
By processing sparse features and performing causal analysis, the continuity of time series is restored, and a personalized blood glucose prediction model is constructed. This solves the problem of the correlation between sparse features and blood glucose fluctuations, and achieves high-precision, personalized blood glucose prediction and safe diabetes management.
Patent Information
- Application Number
- CN202510055518.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-01-14
AI Technical Summary
In existing technologies, the temporal sparsity of sparse features such as dietary behavior and drug use data makes it difficult for blood glucose prediction models to capture the dynamic relationship between sparse features and blood glucose fluctuations, thus affecting prediction accuracy.
By processing sparse features and performing causal analysis, the continuity of time series is restored, a personalized blood glucose prediction model is constructed, and the influence patterns of sparse features are captured using physiological formula transformation and multi-scale segmentation methods. Causal relationship analysis is performed to determine the optimal time lag and perform time alignment. Blood glucose prediction is then performed in conjunction with high-dimensional feature encoding.
It significantly improves the utilization rate of sparse features, quantifies the causal relationship between covariates and blood glucose fluctuations, provides personalized blood glucose prediction results, supports precision diabetes management, and reduces medical risks.
Smart Images

Figure CN120126752B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of blood glucose prediction, and in particular relates to a blood glucose prediction method and device based on sparse feature processing and causal analysis. Background Technology
[0002] Diabetes mellitus is a widespread metabolic disease worldwide, classified into type 1 and type 2. Type 1 diabetes patients are highly dependent on exogenous insulin supplementation for blood glucose levels. Due to the lack or inadequacy of insulin secretion in diabetic patients, blood glucose levels are easily affected by various factors such as diet, exercise, medication, and physiological stress, resulting in complex and unstable fluctuations. Abnormal fluctuations in blood glucose levels can lead to acute complications such as hyperglycemic crisis and hypoglycemic coma. Long-term fluctuations can also increase the risk of microvascular and macrovascular complications, such as retinopathy, neuropathy, and cardiovascular disease. Therefore, accurately predicting blood glucose level trends is of great significance for the daily management of diabetic patients, insulin dosage adjustment, and prevention of complications.
[0003] For example, Chinese patent document CN117766145A discloses a method for constructing a continuous blood glucose prediction model, a blood glucose prediction method, and a device. This method trains the model by acquiring historical data from blood glucose measurements of different patients over several consecutive days, thereby improving the generalization ability and prediction accuracy of the prediction model. Chinese patent document CN107174258A discloses a blood glucose concentration prediction method. Based on physiological data, spectral data, true blood glucose concentration values, and non-blood glucose concentration data from multiple sample test subjects, a multivariate correction algorithm is used to establish a blood glucose concentration prediction model based on the "M+N" theory. This model effectively improves the prediction accuracy of blood glucose concentration.
[0004] Sparse features are a common but complex category of data features in the field of blood glucose prediction. They primarily include data related to patient eating behavior and medication use, such as meal timing and intake, insulin injection dosage and frequency. These features have a significant impact on changes in blood glucose levels, but their sparsity presents numerous challenges in data analysis and modeling. For example, patients' eating behaviors are typically recorded only a few times a day, and the recording time is often uneven, appearing sparse and scattered compared to continuously collected blood glucose data recorded every 5 minutes. Medication data, especially insulin injection data, also exhibits temporal sparsity; for instance, relevant information is recorded only before meals or when blood glucose is abnormal, failing to cover the entire time frame. Temporal sparsity leads to a significant mismatch between the distribution of these features over time and the continuity of blood glucose changes, posing challenges to data fusion and modeling. This event-driven sparsity further increases the complexity of feature processing, as models struggle to capture the dynamic relationship between sparse features and blood glucose fluctuations.
[0005] Despite their sparsity, these features are undeniably important in blood glucose prediction. Dietary behaviors are the primary drivers of blood glucose fluctuations; for example, carbohydrate intake can rapidly raise blood glucose levels. Pharmacological data, particularly insulin injections, are the core means of regulating blood glucose. The dosage and time interval of injections directly determine the magnitude and speed of the blood glucose decrease. The profound impact of these features on blood glucose fluctuations means they must be effectively incorporated into predictive models. Summary of the Invention
[0006] This invention discloses a blood glucose prediction method and device based on sparse feature processing and causal analysis, which can solve the problem of insufficient accuracy in blood glucose prediction in the prior art.
[0007] A blood glucose prediction method based on sparse feature processing and causal analysis includes the following steps:
[0008] (1) Obtain blood glucose data and multi-source covariate data related to blood glucose, including food intake data and medication data; after checking and completing the blood glucose data and corresponding multi-source data, restore the continuity of the time series.
[0009] (2) Construct a personalized blood glucose prediction model. The blood glucose prediction model utilizes the sparse features of food intake data and drug data to perform physiological formula conversion. Then, through a multi-scale segmentation method in the time dimension, it captures the influence pattern of sparse features at different time scales and distinguishes between short-term and long-term effects. After that, it performs causal relationship analysis with blood glucose data to determine the optimal time lag of the sparse variable's influence on blood glucose. Then, based on the time lag analysis results, it performs time alignment processing on multi-source covariate data and blood glucose data. It learns individual blood glucose characteristics based on the processed data. In the learning process, the model adopts high-dimensional feature encoding, multi-scale learning, extracts individualized key features, and finally decodes to obtain the blood glucose prediction value for future time periods.
[0010] (3) Use the data processed in step (1) to train the blood glucose prediction model. After training, input the food data and drug data into the blood glucose prediction model to obtain the predicted blood glucose value.
[0011] Using this invention, patients can achieve personalized blood glucose management through accurate blood glucose prediction. For example, they can adjust their diet, exercise levels, or insulin dosage based on predicted blood glucose levels, thereby avoiding unnecessary health risks. Simultaneously, the predictive model can also help medical personnel monitor changes in their condition, providing data support for optimizing treatment plans.
[0012] In step (1), the following formula is used to complete the data:
[0013]
[0014] Where, x i and x j These are known adjacent data points, respectively at time t. i and t j place, t k It is missing data x k The point in time.
[0015] In step (2), the physiological formula conversion transforms the values of single points such as food intake data and drug data into curves over a period of time.
[0016] For food intake data, the following formula is used to convert it into a curve over a period of time:
[0017]
[0018] Among them, t s It is the sampling time, t meal It refers to the time of eating, where Carb represents the effective carbohydrates at a given time, C meal It is the total amount of carbohydrates consumed in a meal, t peak This is the time when Carb reaches its maximum value. a1 is the growth rate, and a2 is the decline rate.
[0019] For drug data, the following formula is used to convert it into a curve over a period of time:
[0020]
[0021] Among them, t s It is the sampling time, t d T is the duration of insulin activity, a is the time of exponential decay, and S is the growth coefficient.
[0022] In step (2), the causal relationship analysis is performed on each transformed sparse variable and blood glucose variable. The specific process is as follows:
[0023] First, we need to consider the historical data of the blood glucose variable itself, using the following formula:
[0024]
[0025] Among them, X t This is the current blood glucose level, a i The regression coefficient of the historical blood glucose value after the incident, where p is the size of the historical blood glucose value window. These are the residual values after modeling blood glucose levels;
[0026] Taking into account historical data for blood glucose levels and other variables, the following formula is used:
[0027]
[0028] Among them, Y t b is a candidate causal variable. j is the regression coefficient of the candidate causal variable, and q is the size of the historical value window of the candidate causal variable. It is the residual value after modeling historical data that simultaneously considers blood glucose levels and other variables;
[0029] To determine other variables Y t Does it affect blood sugar X? t If there is a significant causal effect, the F-test statistic is calculated using the following formula:
[0030]
[0031] Where N is the total number of samples in the time series, and the critical value F of the F-distribution is found using the F-statistic. critical If F > F critical This indicates that Y t For X t There is a significant causal relationship.
[0032] The formula for determining the optimal time lag of the effect of sparse variables on blood glucose is as follows:
[0033]
[0034] Choose the historical value window q that minimizes the sum of squared residuals. When q = τ, the causal relationship is most significant. τ is the optimal lag, reflecting the time of the main influence of other variables on blood glucose.
[0035] The multi-source covariate data and blood glucose data were time-aligned using the following formula:
[0036]
[0037] in, Is with X t Aligned covariate values, Y t-τ This is the covariate value lagged by τ; through this adjustment, the covariate Y is guaranteed to... t Historical information Y t-τ For the current blood glucose X t It makes a direct contribution to the prediction.
[0038] A blood glucose prediction device based on sparse feature processing and causal analysis includes a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they implement the above-mentioned blood glucose prediction method.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] 1. This invention transforms the original single-value sparse features into function curves by processing the sparse covariate features, which significantly improves the utilization rate of sparse features and provides more complete data input for blood glucose prediction.
[0041] 2. By introducing causal analysis, this invention quantifies the causal relationship between covariates and blood glucose fluctuations, determines the optimal time lag, and performs time series alignment, thus ensuring the scientific validity and reliability of the model.
[0042] 3. Through the processing of sparse features and multi-feature encoding, the model of this invention can provide personalized blood glucose prediction results and support precise diabetes management.
[0043] 4. Through multi-indicator evaluation, this invention can achieve a smaller error in prediction and a higher proportion of areas with no medical risk, thus better supporting clinical decision-making and ensuring safety. Attached Figure Description
[0044] Figure 1 This is a flowchart of a blood glucose prediction method based on sparse feature processing and causal analysis according to the present invention.
[0045] Figure 2 This is a schematic diagram of food intake data conversion provided in an embodiment of the present invention.
[0046] Figure 3 This is a schematic diagram of drug data conversion provided in an embodiment of the present invention.
[0047] Figure 4 The image shows the model prediction results provided in the embodiments of the present invention.
[0048] Figure 5 This is a comparison diagram between the model of this invention and other models.
[0049] Figure 6 This is a comparison chart of the model of this invention with other models in the evaluation of medical indicators. Detailed Implementation
[0050] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the embodiments described below are intended to facilitate the understanding of the present invention and do not constitute any limitation thereof.
[0051] like Figure 1 As shown, a blood glucose prediction method based on sparse feature processing and causal analysis is mainly implemented through three modules: data preprocessing module 101, model building module 102, and blood glucose prediction performance evaluation and comparison module 103.
[0052] The data preprocessing module 101 is used to perform data inspection and completion on blood glucose data. The inspection involves identifying missing blood glucose values and setting them to zero. The completion involves interpolating data features to restore the continuity of the time series.
[0053] Model building module 102 is used to train and construct a personalized blood glucose prediction model. The model employs sparse feature processing for sparse variables such as food intake, basal metabolism, and medication. A systematic sparse feature processing method is designed, along with a multi-scale segmentation method using the time dimension to capture the influence patterns of sparse features at different time scales, distinguishing between short-term and long-term effects. Combined with blood glucose features, causal relationship analysis and optimal time lag learning are performed to determine the optimal time lag of the sparse variables' influence on blood glucose. Based on the time lag analysis results, multi-source covariate data and blood glucose data are time-aligned to ensure the consistency of all features over time and eliminate time mismatch interference. High-dimensional feature encoding is performed on the aligned food intake, basal metabolism, medication, and blood glucose data to extract individualized key features. A multi-scale learning framework is used to decode the multi-scale features into predicted blood glucose values for future time periods, and a deep decoding structure is combined to generate the final prediction result.
[0054] The blood glucose prediction performance evaluation and comparison module 103 is used to predict blood glucose values over different time spans and compare the prediction results of different models with actual data, as well as the significance of the predicted values in the medical field. It can calculate a series of indicators to evaluate the model's predictive accuracy. These evaluation indicators include Mean Absolute Error (MAE) and Root Mean Square Error (RMSE), used to measure the accuracy of predicted blood glucose values, and a medical indicator assessment (Clarke), used to measure the risk of the model's predictions of blood glucose values in the medical field.
[0055] Data completion uses the following formula:
[0056]
[0057] Where, x i and x j These are known adjacent data points, respectively at time t. i and t j place, t k It is missing data x k The point in time.
[0058] Blood glucose data often has missing parts, while complete data is crucial for model training and prediction. The data preprocessing module 101 improves the continuity and completeness of the data, thereby enhancing the model's prediction accuracy and reliability.
[0059] Model building module 102 is a model for feature learning and prediction of individual blood glucose. Its core idea is to utilize sparse features, transform them into physiological formulas, and capture the influence patterns of sparse features at different time scales through multi-scale segmentation in the time dimension, distinguishing between short-term and long-term effects. Then, causal relationship analysis is performed with blood glucose data to determine the optimal influence lag. Based on the lag analysis results, multi-source covariate data and blood glucose data are time-aligned to ensure consistency of all features over time and eliminate time mismatch interference. The model learns individual blood glucose characteristics based on the processed data. During the learning process, the model uses high-dimensional feature encoding, multi-scale learning, and extracts individualized key features, finally decoding to obtain predicted blood glucose values for future time periods.
[0060] The model prediction module 102 includes physiological formula conversion, causal relationship analysis, optimal time lag determination, and time step alignment. The physiological formula conversion transforms single-point values (e.g., food intake and medication) into curves over a time period. The following formula is used for food intake:
[0061]
[0062] Among them, t s It is the sampling time, t meal It refers to the time of eating, where Carb represents the effective carbohydrates at a given time, C meal It is the total amount of carbohydrates consumed in a meal, t peak This is the time when Carb reaches its maximum value. a1 is the growth rate, and a2 is the decline rate.
[0063] The specific conversion method is as follows: Figure 2 As shown, the dashed line represents the original feeding data, with a value of -1 indicating no feeding event and a value other than -1 representing the total amount of carbohydrates consumed at that time point. The solid line represents the transformed feeding data. The value is 0 for the first three time points after a feeding event, then begins to rise, reaching an extreme value at the 12th time point, and then begins to decline until returning to 0 at the 48th time point. It is particularly important to note that the effects of different feeding events are cumulative.
[0064] The following formula is used for drugs:
[0065]
[0066] Among them, t s It is the sampling time, t d T is the duration of insulin activity, a is the time of exponential decay, and S is the growth coefficient.
[0067] Because the effects of insulin are time-dependent, it takes a long time for insulin to affect blood glucose levels: typically, it peaks after 1 hour and then gradually declines over 6 hours. The percentage of insulin activity remaining after a postprandial insulin injection, or the percentage of active insulin, can be modeled using an exponential decay curve, representing the effect of insulin from the perspective of active insulin.
[0068] The specific conversion method is as follows: Figure 3 As shown, the dashed line represents the original drug injection data, with a value of -1 indicating no drug injection and a value other than -1 indicating the total number of drug injections at that time point. The solid line represents the converted drug injection data. The data peaks at the 12th time point after a drug injection and returns to 0 at the 72nd time point. It is particularly important to note that the effects of each drug injection event are cumulative.
[0069] Causal relationship analysis is crucial in blood glucose prediction, as it establishes the interaction mechanisms between blood glucose levels and various sparse variables. By analyzing causal relationships, we can not only quantify the degree of influence of these variables on blood glucose but also determine their time lag, thereby optimizing the input structure of the prediction model. This analysis helps eliminate irrelevant or weakly correlated variables, avoids interference from data noise on the model, and captures the long-term and short-term influence patterns of important variables on blood glucose. First, we need to consider the historical data of the blood glucose variable itself, using the following formula:
[0070]
[0071] Among them, X t This is the current blood glucose level, a i The regression coefficient of the historical blood glucose value after the incident, where p is the size of the historical blood glucose value window. This is the residual value after modeling blood glucose levels. Considering historical data for both blood glucose levels and other variables, the following formula is used:
[0072]
[0073] Among them, Y t These are candidate causal variables, such as medication, diet, etc., b j is the regression coefficient of the candidate causal variable, and q is the size of the historical value window of the candidate causal variable. It is the residual value after modeling historical data considering both blood glucose levels and other variables. This is to determine the other variables Y. t Does it affect blood sugar X? t If there is a significant causal effect, the F-test statistic is calculated using the following formula:
[0074]
[0075] Where N is the total number of samples in the time series, and the critical value F of the F-distribution is found using the F-statistic. critical If F > F critical This indicates that Y t For X t There is a significant causal relationship.
[0076] The optimal time lag for influence is determined using the following formula:
[0077]
[0078] Choose the historical value window q that minimizes the sum of squared residuals. When q = τ, the causal relationship is most significant. τ is the optimal lag, reflecting the time of the main influence of other variables on blood glucose.
[0079] Time alignment processing is characterized by employing the following formula:
[0080]
[0081] in, Is with X t Aligned covariate values, Y t-τ This is the covariate value lagged by τ. Through this adjustment, the covariate Y is guaranteed to... t Historical information Y t-τ For the current blood glucose X t It makes a direct contribution to the prediction.
[0082] In model optimization, multi-feature encoding is used to fuse data on blood glucose, food intake, basal metabolism, and medication, extracting key features as input. Subsequently, multi-scale learning is employed to capture short-term fluctuations and long-term trends in blood glucose changes, fusing information from these different time scales. Finally, in the multi-feature decoding stage, the multi-scale features are transformed into predictions of future blood glucose values. A comparison of the model's predictions with the actual data, based on data from the first six time points, is as follows. Figure 4 As shown.
[0083] In the blood glucose prediction performance evaluation module, various evaluation metrics can be used to objectively assess the model's prediction accuracy. Among these, Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) are two of the most commonly used metrics. MAE represents the average absolute error between the predicted and actual values, reflecting the degree of deviation in the model's prediction results and directly reflecting the magnitude of the model's average error. RMSE, calculated by taking the square root of the average of squared errors, assigns higher weight to larger prediction errors, highlighting the impact of outliers.
[0084] Furthermore, the significance of Clarke's medical indicator assessment lies in its closer relevance to practical applications. For diabetes management, the goal of predictive models is not only numerical accuracy but also ensuring that prediction errors do not lead to erroneous clinical decisions. Clarke analysis provides a direct assessment of the clinical reliability of predictive models, ensuring their practical value in real-world scenarios.
[0085] The mean absolute error described in the blood glucose prediction performance evaluation comparison module is calculated using the following formula:
[0086]
[0087] Where n is the number of predicted values, y i ′ is the predicted value, y i This is the actual value.
[0088] The root mean square error is expressed by the following formula:
[0089]
[0090] Where n is the number of predicted values, y i ′ is the predicted value, y i This is the actual value.
[0091] The medical indicator assessment uses the following formula:
[0092]
[0093] Among them, G ref This represents the actual blood glucose level, G. pred This is for predicting blood glucose levels. Zone A indicates no clinical risk, Zone B indicates acceptable deviation, Zone C indicates high-risk prediction, Zone D indicates dangerous prediction, and Zone E indicates extremely dangerous prediction that is highly likely to cause medical accidents.
[0094] Figure 5 The comparison chart shows that our model outperforms other models in terms of both mean absolute error (MAE) and root mean square error (RMSE) across tasks with varying prediction time spans. This indicates that our model possesses high prediction accuracy and stability in blood glucose prediction tasks, and can better reflect the true trend of blood glucose fluctuations.
[0095] Figure 6 This chart compares the model of this invention with other models in the evaluation of medical indicators. Analysis of the proportion of clinically risk-free areas shows that the proportion of clinically risk-free areas in this model is higher than that of other models, indicating that it can effectively avoid high-risk blood glucose predictions and verifying the reliability and safety of this model in medical scenarios.
[0096] The embodiments described above provide a detailed explanation of the technical solutions and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A blood glucose prediction method based on sparse feature processing and causal analysis, characterized in that, The method comprises the following steps: (1) obtaining blood glucose data and multi-source covariate data related to blood glucose, the multi-source covariate data including food data and drug data; after checking and data completion of the blood glucose data and the corresponding multi-source data, continuity of time series is restored; (2) constructing a personalized blood glucose prediction model, the blood glucose prediction model using sparse features of the food data and the drug data for physiological formula conversion, and then capturing influence patterns of the sparse features on different time scales through a time-dimension multi-scale segmentation method to distinguish short-term and long-term effects; then performing causal relationship analysis on the blood glucose data to determine the optimal influence time lag of the sparse variables on the blood glucose; and then performing time alignment processing on the multi-source covariate data and the blood glucose data based on the time lag analysis result; learning individual blood glucose features according to the processed data, in the learning process, the model uses high-dimensional feature coding, multi-scale learning, extracts individualized key features, and finally decodes to obtain blood glucose prediction values in a future time period; the causal relationship analysis is performed on each converted sparse variable and the blood glucose variable, and the specific process is as follows: firstly, historical data of the blood glucose variable itself is considered, and the following formula is used: wherein X t is the blood glucose value at the current time, a i is the regression coefficient of the accident blood glucose history value, p is the size of the blood glucose history value window, is the residual value after modeling the blood glucose value; meanwhile, historical data of the blood glucose value and other variables are considered, and the following formula is used: where Y t is the candidate causal variable, b j is the regression coefficient of the candidate causal variable, and q is the window size of the historical values of the candidate causal variable, is the residual value after modeling the historical data of blood glucose values and other variables simultaneously. To determine whether other variables Y t have a significant causal effect on blood glucose X t F-test statistic is calculated using the following formula: Wherein, N is the total sample number of time series, the critical value F of F distribution is found according to F statistics critical If F>F critical , it is explained that Y t There is a significant causal relationship for X t (3) training the blood glucose prediction model using the data processed in step (1), and after the training is completed, inputting the food data and the drug data into the blood glucose prediction model to obtain predicted blood glucose values. 2.The blood glucose prediction method based on sparse feature processing and causal analysis according to claim 1, wherein, in step (1), the following formula is used for data completion: where x i and x j are known adjacent data points at times t i and t j respectively, and t k is the time point for which the missing data x k is desired. 3.The blood glucose prediction method based on sparse feature processing and causal analysis according to claim 1, characterized in that, in step (2), the physiological formula conversion converts the single-point values of the food data and the drug data into a time curve. 4.The blood glucose prediction method based on sparse feature processing and causal analysis according to claim 3, characterized in that, for the food data, the following formula is used to convert into a time curve: where t s is the time of the sample, t meal is the time of the meal, Carb represents the available carbohydrates at a given time, C meal is the total amount of carbohydrates ingested in a meal, t peak is the time at which Carb reaches its maximum value, at which time a1 is the rate of increase, and a2 is the rate of decrease. 5.The blood glucose prediction method based on sparse feature processing and causal analysis according to claim 3, characterized in that, for the drug data, the following formula is used to convert into a time curve: where t s is the time of sampling, t d is the time of insulin activity persistence, T is the time of exponential decay, a is the growth coefficient, and S is the proportionality coefficient. 6.The blood glucose prediction method based on sparse feature processing and causal analysis according to claim 1, wherein, in step (2), the optimal influence time lag of the sparse variables on the blood glucose is determined, and the formula is as follows: selecting a historical value window q that makes the residual sum of squares minimum, when q = τ, the causal relationship is most significant, and τ is the optimal influence time lag, reflecting the main influence time of other variables on the blood glucose.
7. The sparse feature processing and causal analysis based blood glucose prediction method according to claim 6, characterized in that, the time alignment processing of the multi-source covariate data and the blood glucose data uses the following formula: wherein, is the covariate value aligned with X t the covariate value Y t-τ is the covariate value lagged by τ; by this adjustment, the historical information Y t of the covariate Y t-τ directly contributes to the prediction of the current blood glucose X t .
8. A blood glucose prediction device based on sparse feature processing and causal analysis, characterized by, a memory and one or more processors, the memory storing executable code, the one or more processors executing the executable code to implement the blood glucose prediction method in any one of claims 1-7.
Citation Information
Patent Citations
Blood glucose concentration prediction method
CN107174258A
Continuous blood glucose prediction model construction method, blood glucose prediction method and device
CN117766145A
Hypoglycemia early warning method and system based on sensing data and physiological information fusion
CN111631730A
Diabetes blood sugar prediction method and device, electronic equipment and storage medium
CN117012388A