A method and system for blood glucose fluctuation pattern extraction and insulin management based on continuous glucose monitoring data
By performing missing value completion, peak detection, gamma distribution modeling, and independent component analysis on continuous blood glucose monitoring data, the problem of rare event identification and interpretability in existing blood glucose fluctuation analysis technologies has been solved, enabling personalized insulin management recommendations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- KUNMING UNIV OF SCI & TECH
- Filing Date
- 2026-03-17
- Publication Date
- 2026-06-16
AI Technical Summary
Existing methods for analyzing blood glucose fluctuations struggle to identify and quantify rare, individualized fluctuation events. Black-box prediction models lack interpretability, have limited cross-object comparable coding spaces, and are unable to output directly corresponding individualized management recommendations.
By acquiring continuous blood glucose monitoring data and combining it with information related to diet, insulin administration, and exercise, missing value completion and noise reduction are performed, local peak values are detected, interpeak intervals are calculated, a gamma distribution model is established, self-information values are calculated, ternary state encoding is performed, an object-time encoding matrix is constructed, and independent component analysis is conducted to determine blood glucose fluctuation patterns and their intensity, and insulin management prompts are generated.
It enables sensitive detection of rare events in asymmetric dynamic processes, establishes a comparable coding space, enhances the interpretability of the model and the targeting of clinical interventions, and provides personalized insulin management recommendations.
Smart Images

Figure CN122224497A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital health and medical data analysis technology, and more specifically, to a method and system for extracting blood glucose fluctuation patterns and managing insulin based on continuous blood glucose monitoring data. Background Technology
[0002] Continuous glucose monitoring (CGM) technology, capable of recording interstitial fluid glucose levels with minute-level temporal resolution, has become a crucial source of fundamental data for diabetes management. However, commonly used clinical indicators for evaluating blood glucose fluctuations, such as standard deviation, coefficient of variation, MAGE, and TIR, are mostly based on assumptions of symmetrical distribution or amplitude statistics, making it difficult to characterize the asymmetric dynamic processes commonly seen in individuals with type 1 diabetes, such as rapid decline followed by slow recovery and short fluctuations followed by long-tailed hyperglycemia. On the other hand, while deep learning-based prediction methods can predict events, they often suffer from model black boxes, difficulty in interpretation, and difficulty in attributing causes to specific behaviors and treatments. The concept of self-information encoding proposed in neuroscience suggests that low-probability deviations / surprise events within time intervals can carry higher information content and can reveal potential independent structures through blind source separation.
[0003] In implementing the embodiments of the present invention, the prior art has at least the following problems or defects: existing blood glucose fluctuation analysis methods are difficult to effectively identify and quantify individualized rare fluctuation events, and lack the ability to characterize the asymmetric features of blood glucose fluctuations; black box prediction models have insufficient interpretability and cannot directly attribute fluctuation patterns to specific interventions such as diet, insulin, or exercise; existing technologies have failed to establish a cross-object comparable coding space, which is not conducive to population classification and follow-up assessment; and there is a lack of blood glucose fluctuation decoupling and decomposition methods based on probability modeling and blind source separation, making it difficult to output individualized management recommendations that directly correspond to specific clinical interventions. Summary of the Invention
[0004] This invention provides a method and system for extracting blood glucose fluctuation patterns and managing insulin based on continuous blood glucose monitoring data.
[0005] In a first aspect of the present invention, a method for extracting blood glucose fluctuation patterns and managing insulin based on continuous blood glucose monitoring data is provided, comprising: S1. Obtain the continuous blood glucose monitoring sequence of the target object within a preset monitoring period, and obtain covariate data aligned with the time of the continuous blood glucose monitoring sequence. The covariate data includes at least one of dietary event information, insulin administration information, and exercise-related physiological signals. S2. Perform missing value completion and noise reduction on the continuous blood glucose monitoring sequence to obtain a preprocessed sequence; S3. Detect local peak moments in the preprocessed sequence, and calculate the interval between adjacent peak moments based on the local peak moments to obtain the inter-peak interval sequence; S4. Estimate the gamma distribution parameters for each target object based on the interpeak interval sequence to obtain an individualized gamma distribution model; S5. Calculate the self-information value of each interpeak interval according to the gamma distribution model to obtain the self-information sequence; S6. Based on the distribution of the self-information sequence, determine the first threshold and the second threshold, and discretize the self-information value into a ternary state code to obtain the ternary code sequence of the target object; S7. Align the ternary encoding sequences of multiple target objects along a unified time axis to construct an object-time encoding matrix; S8. Perform independent component analysis on the object-time encoding matrix to obtain the independent component time activation matrix and the mixture matrix; S9. Perform correlation analysis between the time activation matrix of the independent components and the covariate data to determine the blood glucose fluctuation pattern and its intensity for each independent component. S10. Generate insulin management prompts based on the blood glucose fluctuation pattern and its intensity.
[0006] In a second aspect of the present invention, a blood glucose fluctuation pattern extraction and insulin management system based on continuous blood glucose monitoring data is provided, comprising: The data acquisition module is used to acquire continuous blood glucose monitoring sequences and covariate data of the target object; The preprocessing module is used to perform missing value completion and noise reduction on the continuous blood glucose monitoring sequence to generate a preprocessed sequence; The peak and interval extraction module is used to detect local peak moments and calculate the inter-peak interval sequence in the preprocessed sequence; The gamma modeling and self-information calculation module is used to estimate the gamma distribution parameters and calculate the self-information value based on the interpeak interval sequence. The ternary coding and matrix construction module is used to discretize self-information values into ternary state codes and construct object-time coding matrices. The independent component analysis module is used to perform independent component decomposition on the object-time encoding matrix to obtain the independent component time activation matrix and the mixture matrix. The pattern interpretation module is used to perform correlation analysis between the independent component time activation matrix and the covariate data to determine the blood glucose fluctuation pattern and its intensity. The management output module is used to generate insulin management prompts based on the blood glucose fluctuation pattern and its intensity. Each module is configured to perform any of the methods described above for extracting blood glucose fluctuation patterns and managing insulin based on continuous blood glucose monitoring data.
[0007] The embodiments of the present invention have at least the following beneficial effects: 1. By using gamma distribution to perform individualized probability modeling of interpeak intervals and calculating self-information values, blood glucose fluctuation events are mapped to rarity in a probabilistic sense. This solves the problem that existing technologies are difficult to identify and quantify low-probability surprise events in asymmetric dynamic processes. It enables rare fluctuation patterns that are difficult to capture by traditional indicators, such as nighttime drift, decreased mobility, and dawn phenomenon, to be sensitively detected and highlighted. At the same time, by using ternary state coding, a comparable coding space is established between different objects, which facilitates population typing and long-term follow-up assessment.
[0008] 2. Independent component analysis was used to perform blind source separation on the object-time coding matrix, decoupling the mixed blood glucose fluctuations into several statistically independent components. This solved the problems of insufficient interpretability and difficulty in attributing fluctuation patterns to specific physiological mechanisms in existing black-box prediction models. Through correlation analysis with covariates such as diet, insulin administration, and exercise physiological signals, the direct association between postprandial fluctuations, nighttime baseline drift, exercise-induced decline, and specific behavioral treatment links was realized, enhancing clinical interpretability and intervention targeting.
[0009] 3. Based on the temporal activation intensity of independent components and their association with covariates, personalized management auxiliary information is generated, including hypoglycemia risk alerts, nighttime basal rate matching assessments, postprandial dosing timing assessments, and exercise-related hypoglycemia risk alerts. This solves the problem of existing technologies lacking management outputs that directly correspond to specific clinical interventions. It enables medical staff to make refined insulin management decisions based on traceable pattern intensity, such as adjusting preprandial dosing timing, optimizing nighttime basal rate protocols, and formulating carbohydrate replenishment strategies before and after exercise, thereby improving clinical usability and implementation. Attached Figure Description
[0010] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. Several embodiments of the invention are illustrated in the drawings by way of example and not limitation, wherein: Figure 1 This is a flowchart illustrating a method for extracting blood glucose fluctuation patterns and managing insulin based on continuous blood glucose monitoring data, provided in an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the effect of blood glucose fluctuation event modeling and self-information ternary encoding based on continuous blood glucose monitoring data according to an embodiment of the present invention; Figure 3This is a schematic diagram illustrating the effect of independent component analysis based on an object-time coding matrix and outputting blood glucose fluctuation patterns according to an embodiment of the present invention. Figure 4 This is a schematic diagram of the structure of a blood glucose fluctuation pattern extraction and insulin management system based on continuous blood glucose monitoring data, provided in an embodiment of the present invention. Detailed Implementation
[0011] The principles and spirit of the invention will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are provided merely to enable those skilled in the art to better understand and implement the invention, and are not intended to limit the scope of the invention in any way. Rather, these embodiments are provided to make the invention more thorough and complete, and to fully convey the scope of the invention to those skilled in the art.
[0012] The following is for reference. Figure 1 , Figure 1 This is a schematic flowchart illustrating a method for extracting blood glucose fluctuation patterns and managing insulin based on continuous blood glucose monitoring data, according to an embodiment of the present invention. Figure 1 As shown, a method for extracting blood glucose fluctuation patterns and managing insulin based on continuous blood glucose monitoring data includes: S1. Obtain the continuous blood glucose monitoring sequence of the target object within a preset monitoring period, and obtain covariate data aligned with the time of the continuous blood glucose monitoring sequence. The covariate data includes at least one of dietary event information, insulin administration information, and exercise-related physiological signals. S2. Perform missing value completion and noise reduction on the continuous blood glucose monitoring sequence to obtain a preprocessed sequence; S3. Detect local peak moments in the preprocessed sequence, and calculate the interval between adjacent peak moments based on the local peak moments to obtain the inter-peak interval sequence; S4. Estimate the gamma distribution parameters for each target object based on the interpeak interval sequence to obtain an individualized gamma distribution model; S5. Calculate the self-information value of each interpeak interval according to the gamma distribution model to obtain the self-information sequence; S6. Based on the distribution of the self-information sequence, determine the first threshold and the second threshold, and discretize the self-information value into a ternary state code to obtain the ternary code sequence of the target object; S7. Align the ternary encoding sequences of multiple target objects along a unified time axis to construct an object-time encoding matrix; S8. Perform independent component analysis on the object-time encoding matrix to obtain the independent component time activation matrix and the mixture matrix; S9. Perform correlation analysis between the time activation matrix of the independent components and the covariate data to determine the blood glucose fluctuation pattern and its intensity for each independent component. S10. Generate insulin management prompts based on the blood glucose fluctuation pattern and its intensity. The insulin management prompts include at least one of the following: hypoglycemia risk prompts, nighttime basal rate matching assessment prompts, postprandial dosing timing assessment prompts, or exercise-related hypoglycemia risk prompts.
[0013] This method first acquires continuous blood glucose monitoring sequences of the target subject within a preset monitoring period. These sequences are collected by a subcutaneous glucose sensor at a sampling period of 5 minutes. Simultaneously, time-aligned covariate data are acquired, including dietary event information, insulin administration information, and exercise-related physiological signals. Dietary event records include meal timestamps and carbohydrate intake. Insulin administration information includes basal infusion rate and dosing delay. Exercise signals include heart rate and heart rate variability. The monitoring sequences undergo missing value completion and noise reduction. When observations exist, exponential smoothing is used for updates with a smoothing coefficient ranging from 0.2 to 0.4. When observations are missing, Kalman filtering is used for prediction updates. The state is set as either a one-dimensional blood glucose value or a two-dimensional rate of change. The state transition matrix describes the natural evolution of blood glucose. The covariance of process noise and observation noise is calibrated according to sensor accuracy. Finally, the Kalman estimation and smoothing results are fused to obtain the preprocessed sequence.
[0014] Local peaks are detected in the preprocessed sequence. A minimum peak spacing of 20 to 60 minutes is set to suppress spurious peaks. The time difference between adjacent peaks is calculated to obtain the interpeak interval sequence. Based on this sequence, the shape and scale parameters of the gamma distribution are estimated using maximum likelihood estimation with boundary constraints. The shape parameter controls skewness, and the scale parameter controls broadening. The boundaries are set to 0.1 to 20 minutes and 0.1 to 200 minutes, respectively. The probability density function value for each interval is calculated, and the negative natural logarithm is taken as the self-information value; a larger value indicates a rarer event. First and second thresholds are determined based on the 33rd and 67th percentiles of the self-information sequence. The self-information is discretized and mapped to a ternary state code: less than or equal to the second threshold is the ground state (0), greater than the second threshold with a short interval is a positive surprise (+1), and a long interval is a negative surprise (-1). The multi-object encoded sequences are aligned along a unified time axis to construct an object-time encoding matrix, with rows corresponding to objects and columns corresponding to time coordinates.
[0015] Independent component analysis (ICA) was performed on the matrix using either Infomax or FastICA algorithms. The number of components was set to 5 to 7, or determined by the reconstruction error. The minimum number of components was used when the error was less than 5% to 10% of the matrix energy, resulting in an independent component time activation matrix and a mixture matrix. Spearman rank correlation analysis was performed on the activation matrix and covariates. A correlation with postprandial carbohydrates indicated a postprandial fluctuation pattern; a correlation with nighttime basal rate indicated a nighttime drift pattern; a correlation with heart rate variability indicated an exercise-induced decline pattern; and a regular peak in the morning with a weak correlation to insulin indicated a dawn phenomenon. Management prompts were generated based on pattern intensity. Frequent negative surprises associated with delayed dosing or exercise triggered a hypoglycemia risk warning; strong nighttime drift components associated with basal rate variability triggered a nighttime basal rate matching assessment; strong postprandial components associated with delayed dosing triggered a dosing timing assessment; and strong exercise components associated with heart rate variability triggered an exercise-induced hypoglycemia risk warning.
[0016] In some embodiments, the missing value completion and denoising process in step S2 includes: when there are continuous blood glucose monitoring observations, exponential smoothing is used to update and obtain a smoothed estimate; when the observations are missing, Kalman filtering is used for state prediction and updating, and the Kalman filter estimate and the exponential smoothing result are fused to obtain the preprocessed sequence.
[0017] Specifically, the smoothing coefficient λ in the exponential smoothing update controls the weighting of the current observation and the historical smoothed value. The value of λ ranges from 0.1 to 0.6, preferably from 0.2 to 0.4. When λ is 0.3, it means that the current blood glucose observation has a weight of 30% and the smoothed estimate from the previous time point has a weight of 70%. The larger the λ is, the more sensitive it is to new observations, and the smaller the λ is, the more dependent it is on historical trends. This coefficient can be adaptively adjusted according to the individual blood glucose fluctuation characteristics or fixed to the empirical value of the population.
[0018] The state vector of a Kalman filter can be set to one dimension containing only blood glucose concentration, or extended to two dimensions containing blood glucose concentration and its rate of change. The state transition matrix describes the natural evolution of blood glucose without external intervention. For a one-dimensional state, it can be set as an identity matrix to represent a random walk. For a two-dimensional state, it can be set as an upper triangular matrix containing the sampling period parameter to represent the integral evolution of blood glucose value with the rate of change. The process noise covariance matrix reflects the inherent uncertainty of blood glucose dynamics, and its magnitude is set according to the individual blood glucose coefficient of variation. The observation noise covariance matrix reflects the measurement error of the continuous blood glucose monitoring sensor, and is usually calculated based on the average absolute relative difference specified by the manufacturer, such as 8% to 10%.
[0019] The Kalman filter's prediction step uses the state transition matrix and the previous best estimate to deduce the current prior estimate. The update step, when observations are available, uses Kalman gain weighting to fuse the prior estimate with the actual observations to obtain the posterior estimate; when observations are missing, the update is skipped, and the prior estimate is retained as the best estimate. The fusion step performs a weighted average of the Kalman filter estimate and the exponential smoothing result. The weights are adaptively determined based on the real-time estimation variances of the two methods, with the method with smaller variance receiving a higher weight, or a fixed weight (e.g., 50% each) can be used. The fused sequence undergoes a final smoothing process to eliminate possible splicing traces.
[0020] In some embodiments, the detection of local peak moments in step S3 includes: setting the minimum peak spacing to any value between 10 and 60 minutes, suppressing false peaks caused by noise based on the minimum peak spacing, and obtaining a set of effective local peak moments.
[0021] The minimum peak interval can be set from 10 to 60 minutes, corresponding to 2 to 12 sampling intervals when the continuous blood glucose monitoring sampling cycle is 5 minutes. The selection of this parameter requires a trade-off between the sensitivity and specificity of peak detection. Setting it too short will lead to over-segmentation caused by noise, while setting it too long will miss real dense fluctuation events. 10 minutes corresponds to the minimum resolvable interval for people with rapid postprandial fluctuations, and 60 minutes corresponds to the maximum necessary interval for people with slow basal fluctuations. The preferred value is 20 to 30 minutes to cover most physiological scenarios.
[0022] Local peak detection employs either the sliding window extremum method or the difference zero-crossing method. The sliding window extremum method sets the window half-width to half the minimum peak interval. Points within the window that simultaneously satisfy a value higher than both their left and right neighbors and a peak significance threshold are identified as candidate peaks. The peak significance threshold is set as the difference in blood glucose levels between a candidate point and the minimum values of the left and right windows; a difference exceeding 10 mg / dL or 0.6 mmol / L is considered significant. The difference zero-crossing method calculates the first-order difference of the blood glucose sequence, with zeros from positive to negative corresponding to peak positions. Valid peaks are then selected based on the minimum peak interval constraint. After detection, a set of valid local peak times is obtained, which is then arranged chronologically and the interval between adjacent peaks is calculated for subsequent gamma distribution modeling.
[0023] In some embodiments, estimating the gamma distribution parameters in step S4 includes: fitting the gamma distribution to the interpeak interval sequence of each target object using maximum likelihood estimation to obtain shape parameters and scale parameters; the probability density function of the gamma distribution is: in, For shape parameters, For scale parameters, This is a gamma function.
[0024] Specifically, the optimization objective of maximum likelihood estimation is to maximize the log-likelihood function, which is the sum of the logarithmic probability densities of all interpeak interval samples under the gamma distribution. The optimization algorithm employs a quasi-Newton method with boundary constraints, such as L-BFGS-B or the Newton-Raphson iterative method. The boundary constraints are set as follows: the shape parameter is greater than 0.1 and less than 20, and the scale parameter is greater than 0.1 minutes and less than 200 minutes. These boundaries are set based on the range of blood glucose fluctuations observed in clinical practice to avoid numerical optimization from getting stuck in extreme values. The initial values are set as the results of moment estimation, that is, the shape parameter is calculated by back-calculating the sample mean and variance by dividing the sample mean by the square of the sample mean, and the scale parameter is calculated by dividing the sample variance by the sample mean.
[0025] The iterative convergence condition is set to a parameter change of less than 10 to the power of -6 or a function value change of less than 10 to the power of -8, and the maximum number of iterations is set to 100. For cases with insufficient sample size, such as fewer than 10 intervals, Bayesian regularization estimation is used, introducing a conjugate prior distribution of the shape and scale parameters, such as a gamma prior. The mode of the posterior distribution is used instead of the maximum likelihood estimation to prevent overfitting. After the shape and scale parameters are determined, the probability density function of the gamma distribution can be used to calculate the probability density of any interval value. This density value reflects the prevalence of that interval in the individual's oscillation rhythm; a higher density indicates that the interval better matches the individual's typical oscillation pattern, and a lower density indicates that it is rarer.
[0026] In some embodiments, the formula for calculating the self-information value in step S5 is: in, For the first Interpeak intervals Let g be the probability density function of the gamma distribution. This is the self-information value.
[0027] The calculation of the self-information value is performed independently for each interpeak interval. The input parameters are the value of that interval, the shape parameter and scale parameter of the individual gamma distribution. First, the probability density function value of that interval under the gamma distribution is calculated. This calculation involves power functions, exponential functions, and gamma functions. The gamma function is calculated through numerical approximation, such as the Lancaster approximation or the Stirling formula. To avoid numerical underflow caused by excessively small probability density function values, when the calculated result is less than 10 to the power of -12, the value is truncated to 10 to the power of -12 as a lower limit to ensure the numerical stability of logarithmic operations. Then, the negative natural logarithm of this truncated probability density function value is taken as the self-information value. The base of the logarithm is the natural logarithm, i.e., base e. Alternatively, a base-2 logarithm can be used; the difference between the two is only a constant multiple and does not affect the subsequent threshold segmentation. The self-information value ranges from 0 to 27.6, where 0 represents a certain event with a probability of 1, and 27.6 corresponds to a lower limit probability of 10 to the power of -12. Actual self-information values are typically distributed between 0 and 10, with values greater than 5 considered rare events with high information content. The statistical properties of the self-information sequence include the mean, median, standard deviation, and quantiles. These properties are used to determine the threshold for subsequent ternary encoding. The 33rd and 67th percentiles serve as the dividing points between low-to-medium information content and medium-to-high information content, respectively.
[0028] Preferably, self-information calculation can employ an adaptive lower bound strategy, dynamically adjusting the probability lower bound based on individual sample size and distribution fit quality. When the sample size is large and the goodness of fit is high, a stricter lower bound, such as 10 to the power of -15, is used to preserve the discriminative power of extremely rare events. When the sample size is small or the fit is poor, a more lenient lower bound, such as 10 to the power of -9, is used to avoid excessive amplification of noise. Self-information values can be further standardized by subtracting the sequence mean and dividing by the standard deviation to obtain standardized self-information in z-score form, making self-information comparable across different individuals or time windows, facilitating cross-object analysis. For stratified gamma modeling, self-information calculation requires first determining the time window or blood glucose state stratum to which each interpeak interval belongs, then calling the corresponding stratum's gamma distribution parameters for calculation, ensuring that the self-information reflects the relative rarity under specific physiological conditions. The self-information sequence can be extended by incorporating the directional information of the interpeak intervals, considering not only the interval length but also the trend of blood glucose changes before and after the interval. For example, a peak value after a rapid rise corresponds to a high self-information positive marker, and a trough value after a rapid fall corresponds to a high self-information negative marker, enhancing sensitivity to the direction of fluctuations. The quality of the self-information calculation results can be assessed through the continuity of the sequence and the proportion of outliers. If isolated outliers exceeding five standard deviations of the mean appear in the self-information sequence, it indicates a possible peak detection error or sensor malfunction, triggering a data quality warning. Finally, the self-information sequence maintains the same length and time index as the original interpeak interval sequence to ensure time alignment with subsequent covariate analyses.
[0029] In some embodiments, in step S6, the first threshold and the second threshold are the 33rd percentile and the 67th percentile of the self-information sequence, respectively; the mapping rule of the ternary state coding is: when hour, , represents the ground state; when and When the interval is less than the preset short interval threshold, , indicating a state of surprise, corresponding to unusually short interval events; when and When the interval exceeds the preset long interval threshold, , indicates a negative surprise state, corresponding to an abnormally long interval event; in, The second threshold, This represents the peak interval.
[0030] Specifically, the first threshold is taken from the 33rd percentile of the information sequence, and the second threshold is taken from the 67th percentile of the information sequence. The percentile calculation adopts the linear interpolation method. After the information sequence is arranged in ascending order, the 33rd percentile is located at the 33% position of the sequence length, and the 67th percentile is located at the 67% position. The mapping rule of the ternary state coding is a three-layer judgment structure: First, the self-information value is compared with the second threshold. When the self-information value is less than or equal to the second threshold, regardless of the length of the interpeak interval, it is mapped to the ground state code 0, indicating that the fluctuation event belongs to a common pattern in the individual and does not carry special physiological information. When the self-information value is greater than the second threshold, the corresponding interpeak interval is further compared with the preset short and long interval thresholds. The preset short interval threshold and the preset long interval threshold are taken as the 33rd percentile and 67th percentile of the interpeak interval sequence, respectively, or set to a fixed number of minutes based on the clinical expert experience, such as a short interval threshold of 15 minutes and a long interval threshold of 120 minutes. When the interpeak interval is less than the preset short interval threshold, it is mapped to a positive surprise state code +1, indicating abnormally frequent blood glucose fluctuations, i.e., short interval events, which usually correspond to rapid postprandial fluctuations or hormone counter-regulation. When the interpeak interval is greater than the preset long interval threshold, it is mapped to a negative surprise state code -1, indicating abnormally sparse blood glucose fluctuations, i.e., long interval events, which usually correspond to the nighttime basal state or delayed recovery after hypoglycemia. For cases where the self-information value is greater than the second threshold but the interpeak interval is between the short interval threshold and the long interval threshold, the judgment is made according to the degree of deviation of the interval direction from the median, or the default is to be classified into the ground state to avoid misjudgment.
[0031] In some embodiments, the independent component analysis in step S8 uses the Infomax independent component analysis algorithm, and the number of independent components is set to 5 to 7; or by gradually increasing the number of independent components and calculating the reconstruction error, the minimum number of independent components is taken when the reconstruction error is lower than a preset threshold.
[0032] It should be noted that this step aligns the ternary coding sequences of multiple target objects along a unified time axis to construct an object-time coding matrix. The rows of this matrix correspond to different target objects, while the columns correspond to a unified time coordinate position. Matrix elements are ternary state coding values, i.e., -1, 0, or +1. Unified time axis alignment maps monitoring data from different objects to the same time reference system, resolving temporal misalignment issues caused by different monitoring start times, sampling frequencies, or monitoring period lengths. The object-time coding matrix serves as the input data structure for subsequent independent component analysis. It requires the number of rows to equal the number of target objects, the number of columns to equal the number of discrete time points on the unified time axis, and that matrix elements be complete without missing elements or filled with specific values to fill missing positions. This matrix construction is a crucial step in achieving cross-object blood glucose fluctuation pattern extraction. By unifying the coding space, it makes fluctuation events from different individuals comparable, providing a standardized data organization format for blind source separation algorithms.
[0033] Specifically, a fixed sampling time axis strategy is adopted for unified time axis alignment, setting a unified time resolution of 5 minutes or 10 minutes to cover the entire time range from the earliest monitoring start time to the latest monitoring end time. For each target object's ternary coding sequence, the coded value is mapped to the nearest unified time coordinate based on its original timestamp. The mapping adopts either the nearest neighbor principle or the linear interpolation principle. The nearest neighbor principle selects the nearest unified time point for assignment, while the linear interpolation principle performs interpolation calculations between adjacent coded values before assignment. For time locations where no fluctuation events occurred, the ground state code 0 is filled to indicate that there was no significant fluctuation at that time. For time intervals where the object did not participate in monitoring, a mask is used or specific missing values such as 9 are filled to distinguish it from the valid codes. After the matrix is constructed, a data integrity check is performed, requiring that objects with more than 70% of valid data be included in the analysis. Objects with less than this threshold are removed or require supplementary monitoring data. The number of columns in the matrix, i.e., the length of the time dimension, is set according to the analysis purpose. When focusing on intraday rhythms in longitudinal analysis, 24 hours (144 5-minute points) are used; when focusing on weekly patterns in longitudinal analysis, 168 hours (2016 5-minute points) are used. The data type of the matrix elements uses signed 8-bit integers to save storage. -1, 0, and +1 correspond to the three states of negative surprise, ground state, and positive surprise, respectively.
[0034] In some embodiments, the correlation analysis in step S9 employs the Spearman rank correlation analysis method; the covariate data includes at least one of postprandial carbohydrate intake, carbohydrate-to-protein ratio, basal insulin infusion rate, insulin administration delay, heart rate variability, or exercise intensity index.
[0035] Specifically, the Infomax algorithm comprises four stages: data preprocessing, whitening, separation matrix estimation, and convergence determination. In the data preprocessing stage, the object-time encoding matrix is centered and optionally dimensionality reduced. Centering ensures each column has a mean of 0. Dimensionality reduction uses principal component analysis to retain principal components that explain more than 95% of the variance. The dimensionality after dimensionality reduction is typically set to 10 to 30 to reduce computation and remove noise. In the whitening stage, the centered data is transformed into a covariance matrix with an identity matrix. The whitening matrix is calculated using the eigenvector matrix and the inverse square root matrix of the eigenvalues from principal component analysis. The whitened data is uncorrelated across dimensions and has unit variance.
[0036] In the separation matrix estimation stage, iterative optimization using the natural gradient ascent method is employed. The initial separation matrix is set as an identity matrix or a random orthogonal matrix. The iterative update rule is the difference between the separation matrix multiplied by the identity matrix and the learning rate multiplied by the negative entropy gradient. The negative entropy gradient is approximated using a nonlinear function, either the sigmoid or tanh function, with the sigmoid function parameter set to 1.0. The learning rate is initially set to 0.1 and adaptively adjusted based on convergence. During the convergence judgment stage, the change norm of the separation matrix or the growth rate of the objective function is monitored. The process stops when the change is less than 10 to the power of -6 for five consecutive iterations or when the maximum number of iterations (1000) is reached. The initial number of independent components is set to 5 to 7, covering major physiological fluctuation patterns such as postprandial fluctuations, nocturnal drift, dawn phenomenon, motion response, and unexplained noise; alternatively, it can be determined by the reconstruction error. The reconstruction error is calculated as the relative error of the Frobenius norm between the original matrix and the estimated product, with a preset threshold of 5% to 10%. The number of components is gradually increased from 1 until the error falls below the threshold, and the minimum value satisfying the condition is taken to avoid overfitting.
[0037] In some embodiments, the method further includes: calculating an asymmetric index based on the preprocessed sequence, the asymmetric index being used to quantify the skewness characteristics of blood glucose fluctuations, and classifying the target object according to the asymmetric index to generate differentiated insulin management prompts; the calculation formula for the asymmetric index is: in, These are the 10th, 50th, and 90th percentiles of blood glucose values or interpeak intervals within the target time window, respectively.
[0038] Correlation analysis was performed on the time activation matrix of independent components and the covariate data to determine the blood glucose fluctuation pattern and its intensity for each independent component. The correlation analysis was performed using Spearman rank correlation analysis. The covariate data included at least one of the following: postprandial carbohydrate intake, carbohydrate to protein ratio, basal insulin infusion rate, insulin administration delay, heart rate variability, or exercise intensity index.
[0039] Spearman rank correlation is a nonparametric correlation analysis method that measures the degree of monotonic association between two variables by calculating the rank correlation coefficient. It does not assume that the data follows a normal distribution and is robust to outliers. Each row of the independent component time activation matrix represents the time activation sequence of an independent component, and the covariate data must be aligned with the activation sequence time to form a corresponding time series. Blood glucose fluctuation patterns are explanatory labels for the physiological significance of independent components, such as postprandial fluctuation patterns and nocturnal drift patterns. The pattern strength is quantified by the absolute value of the correlation coefficient; the larger the absolute value, the stronger the association, and the more representative the independent component is of the physiological pattern. This analysis realizes the mapping from statistically independent components to physiologically interpretable patterns, which is a key step in connecting data-driven discovery with clinical knowledge.
[0040] Specifically, the calculation steps for Spearman rank correlation analysis are as follows: First, convert the time-activated sequences of independent components and the time series of covariates into ranks, i.e., the positions after sorting by numerical value, and take the average rank for the same value; then calculate the Pearson correlation coefficient between the ranks of the two sequences, which is the product of the covariance of the rank difference and the standard deviation of each rank. The result ranges from -1 to +1, where +1 indicates perfect positive correlation, -1 indicates perfect negative correlation, and 0 indicates no correlation. A t-test is used for significance testing. The test statistic is the square root of the correlation coefficient multiplied by the sample size minus 2, divided by the square root of the correlation coefficient squared (1 minus 2). The degrees of freedom are the sample size minus 2, and the significance level is set to 0.05. The statistical significance of the correlation is determined by consulting the t-distribution table or calculating the p-value.
[0041] Postprandial carbohydrate intake was recorded as the number of grams of carbohydrates per meal, time-aligned to the 0-4 hour postprandial window; the carbohydrate-to-protein ratio was the ratio of grams of carbohydrates to grams of protein, reflecting the macronutrient composition of the meal; the basal insulin infusion rate was the number of basal insulin units infused per hour by the insulin pump, or the equivalent hourly rate converted from the daily dose of long-acting insulin; insulin dosing delay was the difference between the actual dosing time and the planned dosing time, with positive values indicating delay and negative values indicating advance; heart rate variability was expressed as the standard deviation of adjacent RR intervals or high-frequency power components, reflecting the state of autonomic nervous system regulation; exercise intensity indicators were expressed as metabolic equivalents or the percentage of heart rate to maximum heart rate, collected in real time by wearable devices. Correlation coefficients were calculated for each independent component and all covariates to form a correlation matrix. Rows corresponded to independent components, and columns to covariate types. A pattern label was assigned based on the largest correlation coefficient in each row and its corresponding covariate type. For example, if the correlation coefficient with postprandial carbohydrates was the largest and most significant, it was marked as a postprandial fluctuation pattern. The absolute value of the correlation coefficient served as a quantitative indicator of pattern strength.
[0042] like Figure 2 As shown, for a given 24-hour CGM sequence, peak detection is first performed to obtain the set of peak times, and then the interval sequence between adjacent peaks is calculated. .right Parameters are obtained by fitting the gamma distribution. , Then calculate self-information .according to quantile threshold Events are filtered, and rare events are encoded into ternary sequences according to the interval length. .Depend on Figure 2 As can be seen from (c)-(d), this encoding can highlight abnormally short / abnormally long interval events, realize the expression of rare events in blood glucose fluctuations, and provide input for subsequent ICA decomposition and pattern interpretation.
[0043] like Figure 3 As shown, Figure 3 The independent components and their interpretations obtained after performing independent component analysis (ICA) on the object-time encoding matrix are shown, wherein... Figure 3 The upper part shows a statistical illustration of the explanatory strength / contribution of each independent component (e.g., expressed as the percentage of explained variance). Figure 3 The lower part shows a schematic diagram of the time activation sequences of several major independent components.
[0044] In one embodiment, the object-temporal ternary encoding matrix constructed in steps S6-S7 is used as the ICA input. After whitening and decomposition, an independent component set IC1…ICN is obtained, along with the corresponding temporal activation sequence. For ease of demonstration, each independent component is sorted from highest to lowest contribution. Figure 3 The upper bar chart shows the contribution percentages of IC1-IC5 (e.g., 44%, 28%, 16%, 9%, 3%), used to characterize the degree to which different fluctuation patterns explain the overall coding matrix.
[0045] Furthermore, Figure 3 Examples of time-activated sequences of some independent components and their physiological implications are shown: (1) Post-meal pattern (IC1): A significant activation peak appears near the meal event, which is characterized by high-frequency / high-amplitude activation within the post-meal time window; (2) Nighttime pattern (IC2): Activation changes that are persistent or slowly fluctuating during the nighttime period, used to characterize nighttime baseline drift or fluctuations related to basal insulin matching; (3) Exercise pattern (IC3): A short-term enhanced activation or characteristic waveform appears near the exercise event, reflecting the exercise-induced decrease / fluctuation of blood glucose.
[0046] It should be noted that the aforementioned "meal / exercise" events can be obtained from subject self-reports, wearable devices, or automatic labeling based on time window rules; the semantic labels of independent components such as "post-meal / nighttime / exercise" can be determined through correlation analysis or regression analysis with covariates (e.g., Spearman rank correlation or multiple regression), thereby achieving interpretable attribution of ICA components. Figure 3 As can be seen, the method of the present invention can further decouple the encoded representation of blood glucose fluctuations into multiple independent modes, and output the contribution and time activation characteristics of each mode, providing a basis for the generation of subsequent risk warnings and management assistance information.
[0047] like Figure 4 As shown in some embodiments, a blood glucose fluctuation pattern extraction and insulin management system based on continuous glucose monitoring data includes: Data acquisition module 201 is used to acquire continuous blood glucose monitoring sequences and covariate data of the target object; Preprocessing module 202 is used to perform missing value completion and noise reduction on the continuous blood glucose monitoring sequence to generate a preprocessed sequence; Peak and interval extraction module 203 is used to detect local peak moments and calculate the inter-peak interval sequence in the preprocessed sequence; Gamma modeling and self-information calculation module 204 is used to estimate gamma distribution parameters and calculate self-information values based on the interpeak interval sequence; Ternary coding and matrix construction module 205 is used to discretize self-information values into ternary state codes and construct an object-time coding matrix; Independent component analysis module 206 is used to perform independent component decomposition on the object-time encoding matrix to obtain independent component time activation matrix and mixture matrix; Pattern interpretation module 207 is used to perform correlation analysis between the independent component time activation matrix and covariate data to determine the blood glucose fluctuation pattern and its intensity. Management output module 208 is used to generate insulin management prompt information based on the blood glucose fluctuation pattern and its intensity; Each module is configured to perform any of the methods described above for extracting blood glucose fluctuation patterns and managing insulin based on continuous blood glucose monitoring data.
[0048] It is understandable that this extraction of blood glucose fluctuation patterns based on continuous glucose monitoring data is consistent with the modules and references recorded in the insulin management system. Figure 1The steps described correspond to those in the glucose fluctuation pattern extraction and insulin management method based on continuous glucose monitoring data. Therefore, the operations, features, and beneficial effects described above for the glucose fluctuation pattern extraction and insulin management method based on continuous glucose monitoring data are also applicable to the glucose fluctuation pattern extraction and insulin management system based on continuous glucose monitoring data and its included modules, and will not be repeated here.
[0049] Furthermore, the storage medium in the embodiments of this application stores program instructions capable of implementing all the above methods. These program instructions can be stored in the storage medium in the form of a software product, including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, or terminal devices such as computers, servers, mobile phones, and tablets.
[0050] The above description is merely an explanation of some preferred embodiments of the present invention and the technical principles employed. Those skilled in the art should understand that the scope of the invention as described in the embodiments of the present invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present invention.
Claims
1. A method for extracting blood glucose fluctuation patterns and managing insulin based on continuous blood glucose monitoring data, characterized in that, Includes the following steps: S1. Obtain the continuous blood glucose monitoring sequence of the target object within a preset monitoring period, and obtain covariate data aligned with the time of the continuous blood glucose monitoring sequence. The covariate data includes at least one of dietary event information, insulin administration information, and exercise-related physiological signals. S2. Perform missing value completion and noise reduction on the continuous blood glucose monitoring sequence to obtain a preprocessed sequence; S3. Detect local peak moments in the preprocessed sequence, and calculate the interval between adjacent peak moments based on the local peak moments to obtain the inter-peak interval sequence; S4. Estimate the gamma distribution parameters for each target object based on the interpeak interval sequence to obtain an individualized gamma distribution model; S5. Calculate the self-information value of each interpeak interval according to the gamma distribution model to obtain the self-information sequence; S6. Based on the distribution of the self-information sequence, determine the first threshold and the second threshold, and discretize the self-information value into a ternary state code to obtain the ternary code sequence of the target object; S7. Align the ternary encoding sequences of multiple target objects along a unified time axis to construct an object-time encoding matrix; S8. Perform independent component analysis on the object-time encoding matrix to obtain the independent component time activation matrix and the mixture matrix; S9. Perform correlation analysis between the time activation matrix of the independent components and the covariate data to determine the blood glucose fluctuation pattern and its intensity for each independent component. S10. Generate insulin management prompts based on the blood glucose fluctuation pattern and its intensity.
2. The method according to claim 1, characterized in that, The missing value completion and denoising process in step S2 includes: when there are continuous blood glucose monitoring observations, exponential smoothing is used to update and obtain a smoothed estimate; when the observations are missing, Kalman filtering is used for state prediction and updating, and the Kalman filter estimate is fused with the exponential smoothing result to obtain the preprocessed sequence.
3. The method according to claim 1, characterized in that, The detection of local peak moments in step S3 includes: setting the minimum peak spacing to any value between 10 and 60 minutes, suppressing false peaks caused by noise based on the minimum peak spacing, and obtaining a set of effective local peak moments.
4. The method according to claim 1, characterized in that, Step S4, estimating the gamma distribution parameters, includes: fitting the gamma distribution to the interpeak interval sequence of each target object using maximum likelihood estimation to obtain the shape and scale parameters; the probability density function of the gamma distribution is: in, For shape parameters, For scale parameters, This is a gamma function.
5. The method according to claim 1, characterized in that, The formula for calculating the self-information value in step S5 is as follows: in, For the first Interpeak intervals Let g be the probability density function of the gamma distribution. This is the self-information value.
6. The method according to claim 1, characterized in that, In step S6, the first threshold and the second threshold are the 33rd and 67th percentiles of the self-information sequence, respectively; the mapping rule for the ternary state encoding is: when hour, , represents the ground state; when and When the interval is less than the preset short interval threshold, , indicating a state of surprise, corresponding to unusually short interval events; when and When the interval exceeds the preset long interval threshold, , indicates a negative surprise state, corresponding to an abnormally long interval event; in, The second threshold, This represents the peak interval.
7. The method according to claim 1, characterized in that, The independent component analysis in step S8 uses the Infomax independent component analysis algorithm and sets the number of independent components to 5 to 7; or by gradually increasing the number of independent components and calculating the reconstruction error, the minimum number of independent components is taken when the reconstruction error is lower than a preset threshold.
8. The method according to claim 1, characterized in that, The correlation analysis in step S9 uses the Spearman rank correlation analysis method; the covariate data includes at least one of postprandial carbohydrate intake, carbohydrate to protein ratio, basal insulin infusion rate, insulin administration delay, heart rate variability, or exercise intensity index.
9. The method according to claim 1, characterized in that, Also includes: Based on the preprocessed sequence, an asymmetric index is calculated. This asymmetric index is used to quantify the skewness characteristics of blood glucose fluctuations, and the target population is categorized according to the asymmetric index to generate differentiated insulin management prompts. The calculation formula for the asymmetric index is as follows: in, These are the 10th, 50th, and 90th percentiles of blood glucose values or interpeak intervals within the target time window, respectively.
10. A blood glucose fluctuation pattern extraction and insulin management system based on continuous blood glucose monitoring data, characterized in that, include: The data acquisition module is used to acquire continuous blood glucose monitoring sequences and covariate data of the target object; The preprocessing module is used to perform missing value completion and noise reduction on the continuous blood glucose monitoring sequence to generate a preprocessed sequence; The peak and interval extraction module is used to detect local peak moments and calculate the inter-peak interval sequence in the preprocessed sequence; The gamma modeling and self-information calculation module is used to estimate the gamma distribution parameters and calculate the self-information value based on the interpeak interval sequence. The ternary coding and matrix construction module is used to discretize self-information values into ternary state codes and construct object-time coding matrices. The independent component analysis module is used to perform independent component decomposition on the object-time encoding matrix to obtain the independent component time activation matrix and the mixture matrix. The pattern interpretation module is used to perform correlation analysis between the independent component time activation matrix and the covariate data to determine the blood glucose fluctuation pattern and its intensity. The management output module is used to generate insulin management prompts based on the blood glucose fluctuation pattern and its intensity. Each module is configured to perform the method as described in any one of claims 1 to 9.