Noninvasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion

By using a method that combines reinforcement learning with multimodal dynamic weight fusion, the problems of individual differences and environmental changes in non-invasive blood glucose testing are solved, achieving high-precision and stable blood glucose testing, which is suitable for smart wearable devices.

CN120913857AActive Publication Date: 2025-11-07NORTHEASTERN UNIV CHINA

Patent Information

Application Number
CN202511352107.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-11-07
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Existing non-invasive blood glucose testing technologies face challenges such as individual differences, environmental changes, and insufficient long-term stability, resulting in poor testing accuracy and adaptability, making it difficult to meet the needs of continuous monitoring.

Method used

We employ a method based on reinforcement learning and multimodal dynamic weight fusion. By aligning ECG, PPG and other sensor signals through hardware timestamps, we use a deep reinforcement learning network for adaptive weight allocation and signal processing to construct a predictive and early warning dual-branch structure for blood glucose detection.

Benefits of technology

It improves the accuracy and stability of non-invasive blood glucose testing, can dynamically adapt to individual differences and environmental changes, enhances the real-time performance and safety of testing, and meets the requirements of clinical-grade continuous monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913857A_ABST
    Figure CN120913857A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical health monitoring, and relates to a noninvasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion, and the method comprises the following steps: carrying out adaptive dynamic weight distribution on the weight of an ECG signal and a PPG signal based on a reinforcement learning network, optimizing a signal processing parameter, and determining whether to trigger threshold adjustment; performing weighted feature fusion on the ECG signal and the PPG signal to obtain a fusion feature vector; the fusion feature vector is input into a prediction and early warning double-branch output structure, the prediction and early warning double-branch output structure comprises a regression branch and a classification branch, the regression branch is used for continuously predicting the blood glucose value, and finally a single blood glucose concentration value is output; and the classification branch is used for early warning level judgment, and finally outputting classification alarms for continuous prediction results so as to complete noninvasive blood glucose detection. The method has the beneficial effects that the fluctuation influence of user movement and temperature on noninvasive blood glucose data acquisition is reduced, and the individual difference suitability and long-term stability of noninvasive blood glucose detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical health monitoring, and particularly relates to a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion. BACKGROUND

[0002] Blood glucose monitoring is the core of diabetes management. Traditional invasive blood glucose detection (such as finger blood sampling) has limitations such as pain, complicated operation, and inability to monitor in real time, and cannot meet the needs of users for continuous and convenient blood glucose monitoring. With the popularity of smart wearable devices, non-invasive blood glucose detection technology based on biological signals has become a research hotspot in the field of medical health due to its non-invasiveness, wearability and real-time nature.

[0003] The implementation of non-invasive blood glucose detection technology faces complex real-world challenges. For example, PPG (Photoplethysmogram, based on photoplethysmogram) signals are easily affected by motion artifacts, changes in skin keratin layer thickness, and other disturbances, and ECG (Electrocardiogram, electrocardiogram) signals are affected by significant differences in heart rate variability and waveform morphology between individuals, making it difficult for a single modality signal to stably capture blood glucose-related feature patterns. At the same time, dynamic changes in user exercise intensity, environmental temperature fluctuations and other factors can further cause nonlinear drift of physiological signals, making it difficult for traditional detection methods to adapt to the spatiotemporal changes in signal characteristics in real time. More importantly, in long-term wear scenarios, sensor aging, physiological changes in human skin light transmittance over time, and other time-varying factors can continuously degrade model prediction accuracy. If the problems of robust fusion of multi-modal signals, dynamic environmental adaptation, and individual difference calibration cannot be effectively solved, the blood glucose detection results will deviate significantly, making it difficult to meet the accuracy requirements of clinical-level continuous monitoring, and even potentially delaying the timing of health interventions due to data misjudgment, posing potential risks to user health management.

[0004] Current non-invasive blood glucose detection methods mainly focus on the following three methods: single-modal data-based blood glucose prediction methods, static fusion-based multi-modal methods, and minimally invasive sensor-based detection methods.

[0005] The main steps of the single-modal data-based blood glucose prediction algorithm are: (1) physiological signals are collected by a single sensor (such as a PPG or ECG sensor); (2) the signals are preprocessed and features are extracted; (3) the extracted single-modal features are input into a regression model and the model is trained; (4) the trained model is used to predict the blood glucose value of the real-time collected signals.

[0006] The main steps of the multi-modal data method based on static fusion are: (1) synchronously collecting PPG, ECG and motion sensor data; (2) using fixed weight distribution or static fusion algorithm to fuse multi-modal features; (3) using the fused features to construct a blood glucose prediction model and completing training; and (4) outputting the final blood glucose detection result through the model to the fused real-time data.

[0007] The main steps of the sensor technology based on minimally invasive method are: (1) directly contacting interstitial fluid through an implantable or patch sensor; (2) wirelessly transmitting to a terminal device; and (3) combining a simple calibration model to output a blood glucose value.

[0008] The blood glucose prediction algorithm based on single modal has limitations: (1) single PPG or ECG signal is easily disturbed by motion; and (2) individual differences have a great impact on data, resulting in insufficient blood glucose prediction accuracy.

[0009] The multi-modal data based on static fusion has defects: (1) the traditional fixed weight and Choquet integral multi-modal fusion method cannot dynamically adapt to changes in user's motion, temperature fluctuations and other environmental changes.

[0010] The sensor based on minimally invasive method has poor long-term stability: sensor aging and physiological parameter time variation cause the model performance to decay over time. SUMMARY

[0011] Technical problems to be solved In view of the above-mentioned shortcomings and deficiencies of the prior art, the present application provides a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion, which solves the technical problems of fluctuation influence of user's motion and temperature on non-invasive blood glucose data acquisition, and poor individual difference adaptability and insufficient long-term stability of existing non-invasive blood glucose detection.

[0012] Technical scheme In order to achieve the above-mentioned purpose, the main technical scheme adopted by the present application includes: The present application provides a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion, comprising: Aligning the acquired ECG signal, PPG signal, three-axis accelerometer signal and skin temperature sensor signal through hardware time stamp, and constructing a state space; Taking the state space as the input of the reinforcement learning network, based on the reinforcement learning network, performing adaptive dynamic weight distribution on the weight of the ECG signal and the PPG signal, optimizing the signal processing parameters, and confirming whether to trigger threshold adjustment; Constructing a feature vector of the ECG signal and the PPG signal to perform weighted feature fusion on the ECG signal and the PPG signal, and obtaining a fused feature vector; The fusion feature vector is input into a prediction and early warning double-branch output structure, wherein the prediction and early warning double-branch output structure comprises a regression branch and a classification branch, the regression branch is used for continuous prediction of the blood glucose value, and finally outputs a single blood glucose concentration value; the classification branch is used for early warning level determination, and finally outputs a graded alarm for the continuous prediction result, so as to complete the non-invasive blood glucose detection.

[0013] Optionally, the method further comprises pre-processing the acquired PPG signal and ECG signal, which specifically comprises: eliminating baseline drift using a 0.5Hz-8Hz band-pass filter; The PPG signal is pre-processed according to the following steps: adopting an LMS adaptive filter to eliminate PPG signal motion interference, and the filter output is wherein is the kth filtered weight result, L is the filter order, r(n-k) is a delayed sampling value of the reference interference signal, and y(n) is the estimated motion interference of the filter output; the error is calculated as wherein is the original PPG signal, is the denoised signal, and the weight is updated according to wherein is the k+1th filtered weight result, μ is a step factor, is a constant for preventing zero, r(n) is a reference motion interference signal synchronized with the PPG signal, e(n) is low-pass filtered to suppress residual high-frequency noise; Based on the physiological characteristics and signal-to-noise ratio of the PPG waveform, a dynamic window width is set; only high signal-to-noise ratio data in the systolic phase is intercepted, and the diastolic phase and noise interference period are discarded; the systolic phase starting point corresponds to the maximum value of the second derivative wherein, the time point at which the output represents reaches the maximum value, argmax represents parameter maximization; the window width is dynamically adjusted according to the output of reinforcement learning, the data in the intercepted window is smoothed and interpolated to ensure waveform continuity; The ECG signal is pre-processed according to the compensation calculation formula to realize drift compensation, wherein is the reference ECG signal without motion interference, is the original ECG signal with motion interference, a y is the acceleration value in the Y-axis direction, a z is the acceleration value in the Z-axis direction, wherein j1 and j2 are compensation coefficients, the initial values are set to 0.15 and 0.20 respectively, and are dynamically adjusted according to the output results of reinforcement learning subsequently.

[0014] Optionally, the state space set is: ; in, The ratio of AC to DC components of the PPG signal. For pulse conduction time, The standard deviation of the RR interval of ECG. The ECG low-frequency / high-frequency power ratio, For the time-domain variance of acceleration, For acceleration frequency domain, For skin temperature gradient, This represents the average historical blood glucose prediction error.

[0015] Optionally, adaptive dynamic weight allocation is performed on the weights of the ECG and PPG signals based on a reinforcement learning network, including: The modal non-invasive blood glucose model's regulatory actions are used for adaptive dynamic weight allocation of signal feature weights, optimization of signal processing parameters, and confirmation of whether to trigger threshold adjustment. The action space set is defined as follows: ; in, The feature fusion weights satisfy the constraints. , The order of PPG adaptive filtering (16th to 64th order). The ECG motion artifact compensation coefficient is (0.05mV / g–0.2mV / g). Dynamic window width contraction period ratio Calibration triggers decision; The goal of dynamic adjustment of the multimodal noninvasive blood glucose model is to optimize prediction accuracy, signal quality, and the rationality of the prediction strategy. This optimization is transformed into maximizing reinforcement learning rewards, and the agent's hybrid reward function is defined as follows: ; In the formula Let be the reward value at time t. These are respectively: prediction accuracy reward, signal processing quality reward, and prediction strategy rationality reward; Among them, the blood glucose prediction accuracy bonus is: , To predict blood glucose levels using a regression model, For invasive input measurement reference values, These are the weighting coefficients. It is the error function; Signal quality reward : , where R ppgThe PPG signal is optimized for reward, AC is an alternating current component, and DC is a direct current component; Recg is an ECG signal optimized reward, SDNN t is the standard deviation of normal sinus rhythm intervals; wherein the standard value is the AC / DC average of the user in a resting state for 30 consecutive cardiac cycles.

[0016] Optionally, the reinforcement learning network further comprises: By analyzing the action space set elements of the multi-modal non-invasive blood glucose model, the feature weights of the PPG and ECG signals satisfy ω ECG +ω PPG =1, 0≤ω ECG ≤1, 0≤ω PPG ≤1, the PPG filter order should be an integer in the interval [16, 64], and the PPG filter order lsm-order is mapped to the interval [16, 64] using a Sigmoid function: , wherein is a Sigmoid function; the ECG compensation coefficient k interval is [0.05, 0.2], and the ECG compensation coefficient k is mapped to the interval [0.05, 0.2] by linear transformation: ; the dynamic window width is cut off at a proportion of (0~1); The deep reinforcement learning network structure is composed of a set S representing the environment state, a set A representing the agent's exploration action, and a reward r obtained by the agent; in the process of continuous exploration of the agent, the Q value function is used to represent the mathematical expectation of the long-term cumulative reward that the agent can obtain by following the policy to perform subsequent actions from the action; ; In the formula: is the state of state s at time t, the Q value obtained by policy decision to take action a at time t; ; is the expectation under policy ; is the discount factor when calculating the return at the subsequent state k; is the return at the subsequent state k; The V value function is used to represent the mathematical expectation of the long-term cumulative reward that the agent can obtain by following the policy to perform subsequent actions from the state; ; In the formula: is the state of state s at time t, the Q value obtained by policy The V value obtained by decision making; The parameters of the policy network Actor and the evaluation network Critic are set, and the deep reinforcement learning network uses an Actor network and a Critic network Each network has its corresponding target network parameters ; The sampling policy gradient of the Actor network is updated by the following formula: ; In the formula, is the probability distribution of the output action, is the logarithmic gradient of the policy function, is the advantage function, and the calculation formula is , is the state , the state distribution under the policy π, represents the action is sampled from the probability distribution of the policy ; The parameters of the Actor network are updated by the following formula: ; In the formula, is the learning rate under the policy ; The minimum loss function parameters of the Critic network are optimized by the following formula: ; The Critic network parameters are updated by the following formula: ; In the formula, is the learning rate; In the learning process, the experience vector at each time is stored to form an experience pool M by using the experience replay mechanism; during training, experience samples with a size of mini-batch are randomly extracted from the experience pool M for network parameter updating each time; PPG, ECG features, three-axis accelerometer and temperature sensor data and blood glucose prediction network prediction values are collected for training; the deep reinforcement learning strategy is continuously improved by iteratively training the Actor and Critic networks until the strategy converges or a predetermined number of training steps is reached to update the network parameters, and the network parameters after training are fixed for solving the problem of multi-modal non-invasive blood glucose model.

[0017] Optionally, the feature vectors of the ECG signal and the PPG signal are constructed to perform feature fusion on the ECG signal and the PPG signal weighted, to obtain a fusion feature vector, including: An ECG signal feature is acquired, the ECG signal feature including an RR interval, an ST slope, and an LF / HF energy ratio, and an ECG signal feature vector is constructed ; A PPG signal feature is acquired, the PPG signal feature including a main wave height, a wave shape area ratio, and a PTT, and a PPG signal feature vector is constructed ; The PPG and ECG signals are subjected to feature fusion weighted to obtain a fusion feature vector F .

[0018] Optionally, the regression branch uses a blood glucose value regression model to perform continuous prediction of the blood glucose value, and specifically includes: The fusion feature vector is input, and a time sequence sliding window is processed, with a window length of 10 continuous sampling points After two convolutional layers and one pooling layer, a blood glucose prediction value is finally generated through full connection The prediction formula is as follows: , is a weight of the full connection layer, is a bias term of the full connection layer, is a hidden state; Wherein, when the regression is trained, the loss function is defined as a mean square error : Wherein is a true blood glucose value, is a predicted blood glucose value, and N is the total number of training samples.

[0019] Optionally, the classification branch combines a cost-sensitive SVM classifier and a classification cost matrix to perform blood glucose classification alarm, and specifically includes: A classification cost matrix is established according to the clinical risk priority and the severity of the consequences of misjudgment; A cost-sensitive SVM classifier is constructed in combination with the classification cost matrix; A three-level decision system is constructed based on the absolute value of the blood glucose concentration as a basis and in combination with the time continuity of the prediction result; The blood glucose prediction value output by the blood glucose value regression model is input into the cost-sensitive SVM classifier after multi-dimensional feature expansion, and then a three-level risk probability distribution is output; The three-level risk probability distribution is input into the three-level decision system to perform classification alarm on the continuous prediction result.

[0020] ​​​Optionally, a classification cost matrix is established according to the clinical risk priority and the severity of the misjudgment consequences, including: With the principle of the highest cost of hypoglycemia missed report, the second cost of hyperglycemia false report, and the lowest cost of correct prediction in the normal range, the specific establishment process of the matrix is as follows: a 3*3 dimensional matrix is defined to correspond to three levels of early warning categories; the diagonal elements are correct prediction with the lowest cost; in the non-diagonal elements, the misjudgment cost related to hypoglycemia is higher than that related to hyperglycemia, and the missed report cost of hyperglycemia is higher than that of misjudgment of normal as hyperglycemia; in addition, the matrix is dynamically adjusted according to real-time expansion features, and each cost weight is real-time corrected in combination with exercise intensity and prediction reliability.

[0021] Optionally, a cost-sensitive SVM classifier is constructed in combination with the classification cost matrix, including: The blood glucose prediction value output by the blood glucose value regression model is supplemented with blood glucose change trend, circadian rhythm factor, exercise correlation index and risk confidence to form a 5-dimensional feature vector, wherein the risk confidence is calculated based on the output of the blood glucose value regression model. A 3*3 matrix is set according to the clinical risk priority, wherein the cost of hypoglycemia missed report is the highest, the cost of hyperglycemia false report is the second, and the cost of normal prediction is the lowest, and the weights are dynamically adjusted in combination with exercise intensity and risk confidence, wherein the exercise intensity is determined according to the variance and main frequency in the three-axis accelerometer signal. During training, the classification cost matrix is used as the loss function weight to optimize the parameters, and during inference, the comprehensive cost is calculated: classification cost matrix* posterior probability, the warning level with the minimum cost is selected, the three-level risk probability distribution is output, and the graded alarm of continuous prediction is triggered.

[0022] Advantages The beneficial effects of the present application are: the present application is a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion, a dynamic weight distribution mechanism replaces the traditional fixed weight or rule-driven adaptive method, the weight proportion of PPG (photoplethysmogram) and ECG (electrocardiogram) is adaptively adjusted according to real-time environmental parameters such as exercise state and temperature change, and the MARD value in dynamic scenes is improved. Reinforcement learning (PPO algorithm) driven adaptive adjustment balances prediction accuracy and weight stability through a reward function, and the following three principles are considered in the design process: clinical accuracy penalty: measure the error between predicted blood glucose value and invasive measurement reference value based on L2 norm, ensure that the result meets the medical standard; weight stability constraint: constrain the difference between current and previous weights through L2 norm, avoid model instability caused by sharp weight fluctuations. Multi-dimensional signal preprocessing and motion interference suppression provide environmental feature basis for subsequent signal preprocessing and weight distribution. A scene-based noise reduction strategy is used to solve the problem that a single filtering method cannot adapt to dynamic motion interference. The robustness and environmental adaptability of the signal are improved. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 A flowchart of a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion is provided for an embodiment of the application. DETAILED DESCRIPTION

[0024] In order to better explain the present application, in order to facilitate understanding, the present application will be described in detail by specific embodiments in combination with the accompanying drawings.

[0025] The embodiment of the application specifically relates to a non-invasive blood glucose detection method based on photoelectric plethysmogram (PPG), electrocardiogram (ECG) and environmental sensor signals and multi-modal dynamic weight fusion, and is suitable for real-time monitoring of blood glucose in an intelligent wearable device.

[0026] In order to better understand the above technical solutions, the exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a clearer, more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0027] In a first aspect, with reference to Figure 1 The embodiment provides a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion, comprising: S1, aligning the acquired ECG signal, PPG signal, three-axis accelerometer signal and skin temperature sensor signal by hardware timestamp, and constructing a state space.

[0028] S2, taking the state space as the input of the reinforcement learning network, based on the reinforcement learning network, performing adaptive dynamic weight distribution on the weights of the ECG signal and the PPG signal, optimizing the signal processing parameters, and confirming whether the threshold adjustment is triggered.

[0029] S3, constructing a feature vector of the ECG signal and the PPG signal to perform weighted feature fusion on the ECG signal and the PPG signal, and obtaining a fused feature vector.

[0030] S4, inputting the fused feature vector into a prediction and early warning double-branch output structure, wherein the prediction and early warning double-branch output structure comprises a regression branch and a classification branch, the regression branch is used for continuous prediction of blood glucose value, and finally outputs a single blood glucose concentration value; the classification branch is used for early warning level determination, and finally outputs a graded alarm to the continuous prediction result, so as to complete the non-invasive blood glucose detection.

[0031] The embodiment provides a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion. The multi-modal physiological signals (ECG, PPG, accelerometer, temperature) are aligned through hardware time stamping, ensuring the consistency of the time sequence, laying a data foundation for multi-modal fusion, and avoiding feature distortion caused by time offset. Reinforcement learning is introduced to realize dynamic weight distribution and parameter optimization, so that the model can adapt to the signal characteristics (such as motion interference and individual differences) of different individuals and different physiological states, and solve the problem of poor generalization ability of the traditional fixed weight model. The feature weighted fusion fully excavates the complementarity of ECG and PPG signals (ECG reflects the electrical activity of the heart, and PPG reflects the hemodynamic changes), and improves the feature representation capability. The prediction and early warning double-branch structure takes into account the continuous prediction (regression branch) and graded alarm (classification branch) of blood glucose value, which not only meets the quantitative detection demand, but also realizes the timely intervention of clinical risk through graded early warning, and improves the practicability and safety of non-invasive blood glucose detection.

[0032] Optionally, the method further comprises preprocessing the acquired PPG signal and ECG signal, which specifically comprises: eliminating baseline drift using a band-pass filter of 0.5Hz-8Hz; The PPG signal is preprocessed according to the following steps: adopting an LMS adaptive filter to eliminate PPG signal motion interference, and the filter output is , wherein is the kth filtered weight result, L is the filter order, r(n-k) is the delayed sampling value of the reference interference signal, and y(n) is the estimated motion interference of the filter output; the error is calculated as , wherein is the original PPG signal, is the denoised signal, and the weight is updated according to , wherein is the k+1th filtered weight result, μ is a step factor, is a constant to prevent zero, r(n) is a reference motion interference signal synchronized with the PPG signal, e(n) is low-pass filtered to suppress residual high-frequency noise; Based on the physiological characteristics and signal-to-noise ratio of the PPG waveform, a dynamic window width is set; only high signal-to-noise ratio data in the systolic period is intercepted, and the diastolic period and noise interference period are discarded; the systolic period starting point corresponds to the maximum value of the second derivative , wherein, the time point at which the output represents reaches the maximum value, and argmax represents parameter maximization; the window width is dynamically adjusted according to the output of the reinforcement learning, and the data in the intercepted window is smoothed and interpolated to ensure the continuity of the waveform; According to the compensation calculation formula The ECG signal is preprocessed to realize drift compensation, wherein is a reference ECG signal without motion interference, is an original ECG signal with motion interference, a y is an acceleration value in the Y-axis direction, a z is an acceleration value in the Z-axis direction, wherein j1 and j2 are compensation coefficients, and the initial values are set to 0.15 and 0.20 respectively, and are dynamically adjusted according to the output results of reinforcement learning subsequently.

[0033] The band-pass filter (0.5Hz-8Hz) effectively eliminates baseline drift and preliminarily purifies the signal to provide a stable foundation for subsequent processing. The LMS adaptive filter specifically eliminates the motion interference of the PPG signal, significantly improves the signal-to-noise ratio of the PPG signal through dynamic weight update and error feedback, and solves the core problem of the decrease in non-invasive detection accuracy in the motion scene. Based on the dynamic window width of the PPG physiological characteristics (only the high signal-to-noise ratio data in the systolic period is retained), the diastolic period noise interference is reduced, and the window width is dynamically adjusted and smoothed by interpolation through reinforcement learning, ensuring the waveform continuity and improving the reliability of feature extraction. The motion drift compensation formula of the ECG signal specifically eliminates the motion artifact, dynamically adjusts the compensation coefficient, solves the distortion problem of the ECG signal in the body movement, and provides protection for the accurate extraction of the heart rate related features (such as RR interval).

[0034] Optionally, the state space set is: ; wherein, is the ratio of alternating current / direct current components of the PPG signal, is the pulse transit time, is the standard deviation of the RR interval of the ECG, is the low / high frequency power ratio of the ECG, is the acceleration time domain variance, is the main frequency of the acceleration frequency domain, is the skin temperature gradient, is the mean of the historical blood glucose prediction error.

[0035] Through multi-dimensional state quantization, the reinforcement learning agent is provided with comprehensive decision basis, so that it can accurately judge the current physiological and signal state, and improve the rationality of dynamic adjustment.

[0036] Optionally, the reinforcement learning network is used for self-adaptive dynamic weight distribution of the weights of the ECG signal and the PPG signal, including: The adjustment action of the multi-modal non-invasive blood glucose model is used for self-adaptive dynamic weight distribution of the feature weights of the signal, optimization of the signal processing parameters, and confirmation of whether the threshold adjustment is triggered, and the action space set is defined as: ; in, The feature fusion weights satisfy the constraints. , The order of PPG adaptive filtering (16th to 64th order). The ECG motion artifact compensation coefficient is (0.05mV / g–0.2mV / g). Dynamic window width contraction period ratio Calibration triggers decision; The goal of dynamic adjustment of the multimodal noninvasive blood glucose model is to optimize prediction accuracy, signal quality, and the rationality of the prediction strategy. This optimization is transformed into maximizing reinforcement learning rewards, and the agent's hybrid reward function is defined as follows: ; In the formula Let be the reward value at time t. These are respectively: prediction accuracy reward, signal processing quality reward, and prediction strategy rationality reward; Among them, the blood glucose prediction accuracy bonus is: , To predict blood glucose levels using a regression model, For invasive input measurement reference values, These are the weighting coefficients. It is the error function; Signal quality reward : , where R ppg Optimize rewards for PPG signals. AC represents the alternating current component, and DC represents the direct current component; Recg is the ECG signal optimization reward. SDNN t The standard deviation of the normal sinus interval is given; the standard value is the mean AC / DC ratio over 30 consecutive cardiac cycles in the user's resting state.

[0037] The action space encompasses key parameters such as feature weights, filter order, and compensation coefficients, enabling multi-dimensional collaborative optimization and overcoming the limitations of traditional single-parameter adjustments. A hybrid reward function (prediction accuracy + signal quality + strategy rationality) guides the model to improve prediction accuracy while simultaneously considering signal processing stability and long-term strategy effectiveness (e.g., avoiding frequent calibration), balancing short-term performance with long-term robustness. The blood glucose prediction accuracy reward (based on the error between predicted and reference values) is directly related to the core detection target, while the signal quality reward (AC / DC optimization for PPG and SDNN optimization for ECG) ensures data quality from the source. The combination of these two elements forms a closed-loop optimization, significantly improving the model's anti-interference capability.

[0038] Optionally, the reinforcement learning network further comprises: By analyzing the action space set elements of the multi-modal non-invasive blood glucose model, the feature weights of the PPG and ECG signals satisfy ω ECG +ω PPG =1, 0≤ω ECG ≤1, 0≤ω PPG ≤1, the PPG filter order should be an integer in the interval [16, 64], and the PPG filter order lsm-order is mapped to the interval [16, 64] using a Sigmoid function: , wherein is a Sigmoid function; the ECG compensation coefficient k is in the interval [0.05, 0.2], and the ECG compensation coefficient k is mapped to the interval [0.05, 0.2] by linear transformation: ; the clipping ratio of the dynamic window width is (0~1); The deep reinforcement learning network structure is composed of a set S representing the environment state, a set A representing the action explored by the agent, and a reward r obtained by the agent; in the process of continuous exploration of the agent, the Q value function is used to represent the mathematical expectation of the long-term cumulative reward that the agent can obtain by following the policy to perform subsequent actions from the action; ; In the formula: is the state of the state s at time t , the Q value obtained by the policy decision to take action a at time t ; is the expectation under the policy ; is the discount factor when the return is calculated at the subsequent state k is the return at the subsequent state k The V value function is used to represent the mathematical expectation of the long-term cumulative reward that the agent can obtain by following the policy to perform subsequent actions from a certain state; ; In the formula: is the state of the state s at time t , the V value obtained by the policy decision; The parameters of the policy network Actor and the evaluation network Critic of the deep reinforcement learning network structure are set; the deep reinforcement learning network uses an Actor network and a Critic network , and each network has its corresponding target network parameters ; Update the sampling policy gradient of the Actor network using the following formula: ; In the formula, This represents the probability distribution of the output action. Let be the logarithmic gradient of the policy function. The dominant function is calculated using the following formula: , For state It follows the state distribution under policy π. Indicates action From strategy Sampling from the probability distribution; The parameters of the Actor network are updated using the following formula: ; In the formula, For strategy The learning rate is below; The following formula is used to optimize the parameters of the loss function of the Critic network: ; Update the Critic network parameters using the following formula: ; In the formula, The learning rate; During the learning process, an experience playback mechanism is employed to store the experience vector at each moment. The experience samples are stored to form an experience pool M; during training, a mini-batch of experience samples of size M are randomly drawn from the experience pool M each time for network parameter updates. The deep reinforcement learning strategy was trained by collecting PPG and ECG features, triaxial accelerometer and temperature sensor data, and predictions from the blood glucose prediction network. The deep reinforcement learning strategy was continuously improved by iteratively training the Actor and Critic networks until the strategy converged or the predetermined number of training steps was reached to update the network parameters. The network parameters after training were fixed and used to solve the multimodal non-invasive blood glucose model problem.

[0039] The action parameter mapping (Sigmoid function mapping filter order, linear transformation mapping compensation coefficient) ensures that the parameters are within a physiological reasonable range (such as filter order 16-64), avoids invalid exploration, and improves learning efficiency. The Actor-Critic double network structure combines the target network and the experience replay mechanism to reduce sample correlation, stabilize the training process, accelerate policy convergence, and solve the common training fluctuation problem in reinforcement learning. The policy gradient and loss function optimization formula (gradient update of the Actor network, MSE loss of the Critic network) realizes the fine adjustment of the parameters, so that the agent can efficiently learn the optimal strategy, and the final output stable parameters can be directly used for real-time inference of the multi-modal model, improving the response speed of the detection system.

[0040] Optionally, a feature vector of the ECG signal and the PPG signal is constructed to perform weighted feature fusion on the ECG signal and the PPG signal to obtain a fusion feature vector, including: An ECG signal feature is acquired, the ECG signal feature including an RR interval, an ST slope, and an LF / HF energy ratio, and an ECG signal feature vector is constructed . ; A PPG signal feature is acquired, the PPG signal feature including a main wave height, a waveform area ratio, and a PTT, and a PPG signal feature vector is constructed . ; The PPG and the ECG signal are subjected to weighted feature fusion to obtain a fusion feature vector F . .

[0041] The core physiological features of the ECG (RR interval, ST slope, etc.) and the PPG (main wave height, PTT, etc.) are extracted, and these features have clear physiological correlations with the blood glucose concentration (for example, the PTT changes with the increase of the blood glucose), which ensures the effectiveness of the features. The weighted fusion dynamically allocates the weights through reinforcement learning, so that the model can adaptively highlight the effective features and suppress the noise features in different scenarios (such as increasing the PPG weight when the ECG signal quality is poor), solve the problems of “information redundancy” and “feature conflict” of multi-modal data, and improve the discriminability of the fusion features.

[0042] Optionally, the regression branch uses a blood glucose value regression model to perform continuous prediction of the blood glucose value, which specifically includes: The fusion feature vector is input, and a time sequence sliding window is used for processing, the window length being 10 continuous sampling points , after two convolutional layers and one pooling layer, a full connection is finally generated to generate a blood glucose prediction value The prediction formula is as follows: , is the weight of the full connection layer, bias term for the fully connected layer, is the hidden state; where the loss function is defined as mean square error during regression training : where is the true blood glucose value, is the predicted blood glucose value, and N is the total number of training samples.

[0043] The time series sliding window captures the dynamic trend of blood glucose values, solving the problem of accidental sampling at a single time point. The convolutional layer and the pooling layer are combined to effectively extract local correlation information (such as short-term blood glucose fluctuation patterns) in the time series features. The fully connected layer outputs continuous blood glucose values, taking into account the depth of feature extraction and the continuity of prediction. The mean square error (MSE) loss function directly optimizes the deviation between the predicted value and the true value, ensuring the accuracy of the regression and providing accurate blood glucose quantitative reference for clinical practice.

[0044] Optionally, the classification branch combines the cost-sensitive SVM classifier and the classification cost matrix to perform blood glucose classification warning, which specifically includes: establishing a classification cost matrix according to the clinical risk priority and the severity of the consequences of misjudgment; constructing a cost-sensitive SVM classifier combined with the classification cost matrix; establishing a three-level decision system based on the absolute value of blood glucose concentration and the time continuity of the prediction results; inputting the blood glucose prediction value output by the blood glucose value regression model into the cost-sensitive SVM classifier after multi-dimensional feature expansion, and then outputting a three-level risk probability distribution; inputting the three-level risk probability distribution into the three-level decision system to perform classification warning on the continuous prediction results.

[0045] The cost-sensitive SVM combines the clinical cost matrix to quantify the medical risk (such as the serious consequences of low blood glucose false negatives) as a classification cost, solving the problem of traditional classifiers being insensitive to “high-risk misjudgment”. The three-level decision system combines the absolute value of blood glucose and the time continuity to avoid false alarms caused by single prediction errors, improving the stability of the warning. Multi-dimensional feature expansion (blood glucose trend, circadian rhythm, etc.) provides more abundant basis for classification, and the output three-level risk probability distribution can directly guide clinical intervention (such as low blood glucose emergency treatment and high blood glucose routine monitoring), improving the clinical practical value of non-invasive detection.

[0046] Optionally, a classification cost matrix is established according to the clinical risk priority and the severity of the consequences of misjudgment, including: The specific establishment process of the matrix is: defining a 3*3 dimensional matrix corresponding to three levels of early warning categories; the diagonal elements are correct prediction with the lowest cost; in the non-diagonal elements, the misjudgment cost related to hypoglycemia is higher than that related to hyperglycemia, and the hypoglycemia missed report cost is higher than the misjudgment cost of normal misjudgment as hyperglycemia; in addition, the matrix will be dynamically adjusted according to real-time expansion features, and the cost weight is corrected in real time in combination with the exercise intensity and the prediction reliability.

[0047] According to the principle of "the highest cost of hypoglycemia missed report and the second cost of hyperglycemia false alarm", the clinical risk priority (hypoglycemia may cause coma and other emergencies) is met, and the classification decision is more suitable for the actual needs of medical treatment. The dynamic adjustment mechanism (combined with exercise intensity and prediction reliability) makes the cost matrix adapt to complex scenarios (such as large blood glucose fluctuations during exercise, appropriately increasing the early warning sensitivity), avoiding the limitations of fixed matrix in dynamic physiological state, and improving the flexibility and accuracy of early warning.

[0048] Optionally, a cost-sensitive SVM classifier is constructed in combination with the classification cost matrix, comprising: The blood glucose prediction value output by the blood glucose value regression model is supplemented with blood glucose change trend, circadian rhythm factor, exercise correlation index and risk confidence to form a 5-dimensional feature vector, wherein the risk confidence is calculated based on the output of the blood glucose value regression model; According to the clinical risk priority, a 3*3 matrix is set, wherein the cost of hypoglycemia missed report is the highest, the cost of hyperglycemia false alarm is the second, and the cost of normal prediction is the lowest, and the weights are dynamically adjusted in combination with the exercise intensity and the risk confidence, wherein the exercise intensity is determined according to the variance and dominant frequency in the three-axis accelerometer signal; During training, the classification cost matrix is used as the loss function weight to optimize the parameters, and during inference, the total cost is calculated by: classification cost matrix * posterior probability, the minimum cost warning level is selected, the three-level risk probability distribution is output, and the graded alarm of continuous prediction is triggered.

[0049] The 5-dimensional feature vector breaks through the limitation of single blood glucose value, captures the multi-dimensional rules of blood glucose change, and improves the adaptability of the classifier to complex physiological states. The introduction of exercise intensity (based on accelerometer signal) and risk confidence enables the model to distinguish between "true hyperglycemia" and "pseudo-hyperglycemia signal caused by exercise", and reduces misjudgment under the influence of exercise. According to the decision basis of classification cost matrix * posterior probability, the minimum cost warning level is selected, the clinical safety of the decision is ensured from the mathematical level, and the serious medical risks (such as hypoglycemia missed report) are minimized.

[0050] It is apparent that those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, the present application should be construed to encompass all such modifications and variations as fall within the scope of the present claims and their equivalents.

[0051] Although the embodiments of the present application have been shown and described above, it is to be understood that the above-described embodiments are merely exemplary and are not to be taken in a limiting sense, but the scope of the present application is not to be understood to be limited to the above-described embodiments but various changes could be made by those skilled in the art without departing from the scope of the present application.

Claims

1. A non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion, characterized in that, The method comprises the following steps: Aligning the acquired ECG signal, PPG signal, three-axis accelerometer signal and skin temperature sensor signal by hardware timestamp, and constructing a state space; Taking the state space as the input of the reinforcement learning network, and performing adaptive dynamic weight distribution on the weights of the ECG signal and PPG signal based on the reinforcement learning network, optimizing the signal processing parameters, and confirming whether to trigger threshold adjustment; Constructing a feature vector of the ECG signal and PPG signal to perform weighted feature fusion on the ECG signal and PPG signal, and obtaining a fusion feature vector; Inputting the fusion feature vector into a prediction and early warning double-branch output structure, wherein the prediction and early warning double-branch output structure comprises a regression branch and a classification branch, the regression branch is used for continuous prediction of blood glucose value, and finally outputs a single blood glucose concentration value; The classification branch is used for early warning level determination, and finally outputs a graded alarm to the continuous prediction result, so as to complete the non-invasive blood glucose detection.

2. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 1, characterized in that, The method further comprises pre-processing the acquired PPG signal and ECG signal, which specifically comprises: Using a 0.5Hz-8Hz band-pass filter to eliminate baseline drift; The PPG signal is preprocessed according to the following steps: the LMS adaptive filter is used to eliminate the motion interference of the PPG signal, and the filter output is wherein is the kth filtered weight result, L is the filter order, r(n-k) is the delayed sampling value of the reference interference signal, and y(n) is the estimated motion interference of the filter output; the error is calculated by wherein is the original PPG signal, is the denoised signal, and the weight is updated according to wherein is the k+1th filtered weight result, μ is the step factor, is the anti-zero constant, r(n) is the reference motion interference signal synchronized with the PPG signal, and e(n) is low-pass filtered to suppress the residual high-frequency noise; Based on the physiological characteristics and signal-to-noise ratio of the PPG waveform, a dynamic window width is set; only high signal-to-noise ratio data is intercepted in the systolic phase, and the diastolic phase and noise interference period are discarded; the systolic phase starting point corresponds to the maximum value of the second derivative wherein, The output represents The time point at which the maximum value is reached, argmax represents the parameter maximization; the output is dynamically adjusted according to the reinforcement learning, the window width is adjusted, the data in the intercepted window is smoothed and interpolated, and the waveform continuity is ensured; According to the compensation calculation formula The ECG signal is preprocessed to realize drift compensation, wherein is a reference ECG signal without motion interference, is an original ECG signal with motion interference, a y is an acceleration value in the Y-axis direction, a z is an acceleration value in the Z-axis direction, wherein j1 and j2 are compensation coefficients, the initial values of which are set to 0.15 and 0.20 respectively, and are dynamically adjusted according to the output results of reinforcement learning subsequently.

3. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 2, characterized in that, The state space set is: ; wherein, is the PPG signal AC / DC component ratio, is the pulse transit time, is the ECG RR interval standard deviation, is the ECG low / high frequency power ratio, is the acceleration time domain variance, is the acceleration frequency domain dominant frequency, is the skin temperature gradient, is the historical blood glucose prediction error mean.

4. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 3, characterized in that, The adaptive dynamic weight distribution of the reinforcement learning network on the weights of the ECG signal and PPG signal comprises: The adjustment action of the multi-modal non-invasive blood glucose model is used for adaptive dynamic weight distribution of the feature weights of the signal, optimization of the signal processing parameters, and confirmation of whether to trigger threshold adjustment, and the action space set is defined as: ; wherein, is a feature fusion weight, satisfying the constraint , is a PPG adaptive filter order (16-64 orders), is an ECG motion artifact compensation coefficient (0.05 mV / g - 0.2 mV / g), is a dynamic window width shrinkage period ratio, is a calibration trigger decision; The target of the dynamic adjustment of the multi-modal non-invasive blood glucose model is the optimization of prediction accuracy, signal quality and prediction strategy rationality, the optimization of prediction accuracy, signal quality and prediction strategy rationality of the multi-modal non-invasive blood glucose model are converted into the maximization of the reinforcement learning reward, and the mixed reward function of the agent is defined as: ; In the formula is the reward value at time t, respectively, are the prediction accuracy reward, the signal processing quality reward, and the prediction strategy rationality reward. Wherein, the blood glucose prediction accuracy reward: , is the blood glucose value predicted by the regression model, is the reference value of the invasive input measurement, is the weight coefficient, is the error function; Signal quality reward : where R ppg is a PPG signal optimization reward, AC is an alternating current component and DC is a direct current component; Recg is an ECG signal optimization reward, SDNN t is a standard deviation of normal sinus intervals; wherein the standard value is an average of AC / DC of a user in a resting state for 30 consecutive cardiac cycles.

5. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 4, characterized in that, The reinforcement learning network further comprises: By analyzing the action space set elements of the multi-modal non-invasive blood glucose model, the feature weights of the PPG and ECG signals satisfy ω ECG +ω PPG =1, 0≤ω ECG ≤1, 0≤ω PPG ≤1, the PPG filter order should be an integer in the interval [16, 64], and the PPG filter order lsm-order is mapped to the interval [16, 64] using a Sigmoid function: , wherein is a Sigmoid function; the ECG compensation coefficient k is in the interval [0.05, 0.2], and the ECG compensation coefficient k is mapped to the interval [0.05, 0.2] by linear transformation: ; the dynamic window width is cut off at a proportion of (0~1). The deep reinforcement learning network structure is composed of a set S representing the environment state, a set A representing the exploration action of the agent, and a reward r obtained by the agent; in the continuous exploration process of the agent, a Q value function is used to represent the mathematical expectation of the long-term cumulative reward that can be obtained by the agent after following the strategy to perform subsequent actions from the action; ; wherein: s(t) is the state at time t the policy a(t) is the action at time t the Q-value resulting from the policy the expected reward r(k) is the reward at the subsequent state k r(k) is the reward at the subsequent state k A V value function is used to represent the mathematical expectation of the long-term cumulative reward that can be obtained by the agent after following the strategy to perform subsequent actions from the state; ; In the formula: is the state s at time t Next, the V value obtained by the policy decision; a policy network Actor and a critic network Critic setting the parameters of the deep reinforcement learning network structure; the deep reinforcement learning network uses an Actor network and a Critic network each with its corresponding target network parameters ; The sampling policy gradient of the Actor network is updated by using the following formula: ; where, is the probability distribution over actions, is the log gradient of the policy function, is the advantage function, computed as , is the state subject to the state distribution under policy π, denotes an action sampled from the probability distribution of actions under policy π. The parameters of the Actor network are updated by using the following formula: ; In the formula, The strategy under the learning rate; The minimum loss function parameters of the Critic network are optimized by using the following formula: ; The Critic network parameters are updated by using the following formula: ; In the formula, is the learning rate; During the learning process, an experience playback mechanism is employed to store the experience vector at each moment. The experience samples are stored to form an experience pool M; during training, a mini-batch of experience samples of size M are randomly drawn from the experience pool M each time for network parameter updates. The PPG, ECG features, three-axis accelerometer and temperature sensor data and the prediction value of the blood glucose prediction network are collected for training; the deep reinforcement learning strategy is continuously improved in the way of iterative training of the Actor and Critic networks until the strategy converges or a predetermined number of training steps is reached to update the network parameters, and the network parameters after training are fixed for solving the multi-modal non-invasive blood glucose model problem.

6. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 5, characterized in that, The feature vectors of the ECG signal and the PPG signal are constructed to perform feature fusion on the ECG signal and the PGG signal by weighting, to obtain a fusion feature vector, including: Obtaining ECG signal features, the ECG signal features including RR interval, ST slope and LF / HF energy ratio, and constructing an ECG signal feature vector For ; Obtaining PPG signal features, the PPG signal features including a main wave height, a wave shape area ratio and a PTT, and constructing a PPG signal feature vector For ; The PPG and ECG signals are weighted to obtain a fusion feature vector F To .

7. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 6, characterized in that, The regression branch uses a blood glucose value regression model to perform continuous prediction of the blood glucose value, which specifically includes: Input fusion feature vector, use time sequence sliding window processing, window length is 10 continuous sampling points After two convolutional layers and a pooling layer, the final full connection generates blood glucose prediction value The prediction formula is as follows: , are weights of the fully connected layer, are bias terms of the fully connected layer, is a hidden state; where regression training defines the loss function as mean square error : , where is the true blood glucose value, is the predicted blood glucose value, and N is the total number of training samples.

8. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 7, characterized in that, The classification branch combines a cost-sensitive SVM classifier and a classification cost matrix to perform blood glucose classification alarm, which specifically includes: A classification cost matrix is established according to the clinical risk priority and the severity of the consequences of misjudgment. A cost-sensitive SVM classifier is constructed in combination with the classification cost matrix. A three-level decision system is constructed based on the absolute value of the blood glucose concentration as a boundary and the time continuity of the prediction result. The blood glucose prediction value output by the blood glucose value regression model is input into the cost-sensitive SVM classifier after multi-dimensional feature expansion, and then a three-level risk probability distribution is output. The three-level risk probability distribution is input into the three-level decision system to perform classification alarm on the continuous prediction result.

9. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 8, characterized in that, A classification cost matrix is established according to the clinical risk priority and the severity of the consequences of misjudgment, including: The specific establishment process of the matrix is as follows: a 3x3-dimensional matrix corresponding to three levels of warning categories is defined; the diagonal elements are correct predictions with the lowest cost; in the non-diagonal elements, the misjudgment cost related to hypoglycemia is higher than that related to hyperglycemia, and the hypoglycemia miss report cost is higher than the misjudgment cost of misjudging normal as hyperglycemia; in addition, the matrix is dynamically adjusted according to real-time expanded features, and the weight of each cost is real-time corrected in combination with the exercise intensity and the prediction reliability.

10. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 9, characterized in that, A cost-sensitive SVM classifier is constructed in combination with the classification cost matrix, including: The blood glucose prediction value output by the blood glucose value regression model is supplemented with blood glucose trend, circadian rhythm factor, exercise correlation index and risk confidence to form a 5-dimensional feature vector, wherein the risk confidence is calculated based on the output of the blood glucose value regression model; A 3x3 matrix is set according to the clinical risk priority, wherein the cost of hypoglycemia miss report is the highest, the cost of hyperglycemia false alarm is the second, and the cost of normal prediction is the lowest, and the weight is dynamically adjusted in combination with the exercise intensity and the risk confidence, wherein the exercise intensity is determined according to the variance and dominant frequency in the three-axis accelerometer signal; During training, the classification cost matrix is used as the loss function weight to optimize the parameters, and during inference, the total cost is calculated by: classification cost matrix x posterior probability, the warning level with the minimum cost is selected, and a three-level risk probability distribution is output, which is used to trigger the classification alarm of the continuous prediction.

Citation Information

Patent Citations

  • Noninvasive blood glucose detection method and device, electronic equipment and medium

    CN118078275A

  • Systems and methods for non-invasive blood pressure measurement

    US20170156606A1

  • Systems and methods for computationally efficient non-invasive blood quality measurement

    US20200237303A1

  • Multi-sensor upper arm band for physiological measurements and algorithms to predict glycemic events

    US20240245307A1

Cited By

  • Sinking compensation method and system for CGM sensor

    CN121242569A

  • Noninvasive blood glucose calibration detection method and system based on multi-mode signal fusion

    CN121301861A

  • Non-invasive blood glucose calibration detection method and system based on multi-mode signal fusion

    CN121301861B