Non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion

By using a method that combines reinforcement learning with multimodal dynamic weights, the problems of individual differences and environmental changes in non-invasive blood glucose testing are solved, achieving high-precision and stable blood glucose testing that is suitable for smart wearable devices.

CN120913857BActive Publication Date: 2025-12-26NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511352107.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-12-26
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Existing non-invasive blood glucose testing technologies face challenges such as individual differences, environmental changes, and insufficient long-term stability, resulting in low testing accuracy and difficulty in meeting the needs of continuous monitoring.

Method used

We employ a method based on reinforcement learning and multimodal dynamic weight fusion. By aligning ECG, PPG and other sensor signals with hardware timestamps, we construct a state space. We then use a reinforcement learning network for adaptive weight allocation and signal processing, and combine regression and classification branches for blood glucose prediction and early warning.

Benefits of technology

It improves the accuracy and stability of non-invasive blood glucose testing, can dynamically adapt to individual differences and environmental changes, and enhances the real-time performance and safety of the test.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913857B_ABST
    Figure CN120913857B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of medical health monitoring, and relates to a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion. The method comprises: based on a reinforcement learning network, performing adaptive dynamic weight distribution on the weights of ECG signals and PPG signals, optimizing signal processing parameters, and confirming whether a threshold value is triggered for adjustment; performing weighted feature fusion on the ECG signals and PPG signals to obtain a fusion feature vector; inputting the fusion feature vector into a prediction and early warning double-branch output structure, wherein the prediction and early warning double-branch output structure comprises a regression branch and a classification branch, the regression branch is used for continuous prediction of blood glucose values, and finally outputs a single blood glucose concentration value; the classification branch is used for early warning level determination, and finally outputs a graded alarm for the continuous prediction result, so as to complete non-invasive blood glucose detection. The beneficial effects are that the fluctuation influence of user motion and temperature on non-invasive blood glucose data acquisition is reduced, and the individual difference adaptability and long-term stability of non-invasive blood glucose detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical health monitoring, and particularly relates to a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion. BACKGROUND

[0002] Blood glucose monitoring is the core of diabetes management. Traditional invasive blood glucose detection (such as finger blood sampling) has limitations such as pain, complicated operation, and inability to monitor in real time, and cannot meet the needs of users for continuous and convenient blood glucose monitoring. With the popularity of smart wearable devices, non-invasive blood glucose detection technology based on biological signals has become a research hotspot in the field of medical health due to its non-invasiveness, wearability and real-time nature.

[0003] The implementation of non-invasive blood glucose detection technology faces complex real-world challenges. For example, PPG (Photoplethysmogram, based on photoplethysmogram) signals are easily affected by motion artifacts, changes in skin keratin layer thickness, and other disturbances, and ECG (Electrocardiogram, electrocardiogram) signals are affected by significant differences in heart rate variability and waveform morphology between individuals, making it difficult for a single modality signal to stably capture blood glucose-related feature patterns. At the same time, dynamic changes in user exercise intensity, environmental temperature fluctuations and other factors can further cause nonlinear drift of physiological signals, making it difficult for traditional detection methods to adapt to the spatiotemporal changes in signal characteristics in real time. More importantly, in long-term wear scenarios, sensor aging, physiological changes in human skin light transmittance over time, and other time-varying factors can continuously degrade model prediction accuracy. If the problems of robust fusion of multi-modal signals, dynamic environmental adaptation, and individual difference calibration cannot be effectively solved, the blood glucose detection results will deviate significantly, making it difficult to meet the accuracy requirements of clinical-level continuous monitoring, and even potentially delaying the timing of health interventions due to data misjudgment, posing potential risks to user health management.

[0004] Current non-invasive blood glucose detection methods mainly focus on the following three methods: single-modal data-based blood glucose prediction methods, static fusion-based multi-modal methods, and minimally invasive sensor-based detection methods.

[0005] The main steps of the single-modal data-based blood glucose prediction algorithm are: (1) physiological signals are collected by a single sensor (such as a PPG or ECG sensor); (2) the signals are preprocessed and features are extracted; (3) the extracted single-modal features are input into a regression model and the model is trained; (4) the trained model is used to predict the blood glucose value of the real-time collected signals.

[0006] The main steps of the multi-modal data method based on static fusion are: (1) synchronously collecting PPG, ECG and motion sensor data; (2) using fixed weight distribution or static fusion algorithm to fuse multi-modal features; (3) using the fused features to construct a blood glucose prediction model and completing training; and (4) outputting the final blood glucose detection result through the model to the fused real-time data.

[0007] The main steps of the sensor technology based on minimally invasive method are: (1) directly contacting interstitial fluid through an implantable or patch sensor; (2) wirelessly transmitting to a terminal device; and (3) combining a simple calibration model to output a blood glucose value.

[0008] The blood glucose prediction algorithm based on single modal has limitations: (1) single PPG or ECG signal is easily disturbed by motion; and (2) individual differences have a great impact on data, resulting in insufficient blood glucose prediction accuracy.

[0009] The multi-modal data based on static fusion has defects: (1) the traditional fixed weight and Choquet integral multi-modal fusion method cannot dynamically adapt to changes in user's motion, temperature fluctuations and the like.

[0010] The sensor based on minimally invasive method has poor long-term stability: sensor aging and physiological parameter time variation cause the model performance to decay over time. SUMMARY

[0011] Technical problems to be solved

[0012] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present application provides a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion, which solves the technical problems of fluctuation influence of user's motion and temperature on non-invasive blood glucose data acquisition, and poor individual difference adaptability and insufficient long-term stability of existing non-invasive blood glucose detection.

[0013] Technical scheme

[0014] In order to achieve the above-mentioned purpose, the main technical scheme adopted by the present application includes:

[0015] The present application provides a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion, comprising:

[0016] The acquired ECG signal, PPG signal, three-axis accelerometer signal and skin temperature sensor signal are aligned by a hardware time stamp, and a state space is constructed;

[0017] The state space is used as the input of the reinforcement learning network, the reinforcement learning network is used to perform adaptive dynamic weight distribution on the weight of the ECG signal and the PPG signal, optimize the signal processing parameters, and confirm whether the threshold value is adjusted;

[0018] constructing a feature vector of the ECG signal and the PPG signal to perform feature fusion on the ECG signal and the PPG signal by weighting, to obtain a fusion feature vector;

[0019] inputting the fusion feature vector into a prediction and early warning double-branch output structure, wherein the prediction and early warning double-branch output structure comprises a regression branch and a classification branch, the regression branch is used for continuous prediction of the blood glucose value, and finally outputs a single blood glucose concentration value; the classification branch is used for early warning level determination, and finally outputs a graded alarm on the continuous prediction result, so as to complete the non-invasive blood glucose detection.

[0020] Optionally, the method further comprises pre-processing the acquired PPG signal and ECG signal, which specifically comprises:

[0021] using a 0.5 Hz-8 Hz band-pass filter to eliminate baseline drift;

[0022] the PPG signal is pre-processed according to the following steps: an LMS adaptive filter is used to eliminate motion interference of the PPG signal, and the filter output is wherein is the kth filtered weight result, L is the filter order, r(n-k) is a delayed sampling value of a reference interference signal, and y(n) is an estimated motion interference of the filter output; an error is calculated by wherein is the original PPG signal, is the denoised signal, and the weight is updated according to wherein is the k+1th filtered weight result, μ is a step factor, is a constant for preventing zero, and r(n) is a reference motion interference signal synchronized with the PPG signal; e(n) is low-pass filtered to suppress residual high-frequency noise;

[0023] based on the physiological characteristics and signal-to-noise ratio of the PPG waveform, a dynamic window width is set; only high signal-to-noise ratio data in the systolic phase is intercepted, and the diastolic phase and noise interference period are discarded; the systolic phase starting point corresponds to the maximum value of the second derivative wherein, the time point at which the output represents reaches the maximum value, and argmax represents parameter maximization; the window width is dynamically adjusted according to the output of the reinforcement learning, and the data in the intercepted window is smoothed and interpolated to ensure the continuity of the waveform;

[0024] the ECG signal is pre-processed according to a compensation calculation formula to realize drift compensation, wherein is a reference ECG signal without motion interference, is an original ECG signal with motion interference, and ay is the acceleration value in Y-axis direction, a z is the acceleration value in Z-axis direction, wherein j1, j2 are compensation coefficients, the initial values are set to 0.15 and 0.20 respectively, and subsequent dynamic adjustment is made according to the output results of reinforcement learning.

[0025] Optionally, the state space set is:

[0026] ;

[0027] wherein, is the AC / DC component ratio of the PPG signal, is the pulse transit time, is the standard deviation of the RR interval of the ECG, is the low / high frequency power ratio of the ECG, is the acceleration time domain variance, is the acceleration frequency domain main frequency, is the skin temperature gradient, is the mean of historical blood glucose prediction error.

[0028] Optionally, the adaptive dynamic weight distribution based on the reinforcement learning network is performed on the weight of the ECG signal and the PPG signal, including:

[0029] The adjustment action of the multi-modal non-invasive blood glucose model is used for adaptive dynamic weight distribution of the feature weight of the signal, optimization of the signal processing parameter, and confirmation of whether to trigger threshold adjustment, and the action space set is defined as:

[0030] ;

[0031] wherein, is the feature fusion weight, satisfying the constraint , is the PPG adaptive filter order (16-64), is the ECG motion artifact compensation coefficient (0.05 mV / g-0.2 mV / g), is the dynamic window width contraction period ratio, is the calibration trigger decision;

[0032] The target of the dynamic adjustment of the multi-modal non-invasive blood glucose model is the optimization of the prediction accuracy, the signal quality and the rationality of the prediction strategy, the optimization of the multi-modal non-invasive blood glucose prediction accuracy, the signal quality and the rationality of the prediction strategy are converted into the maximization of the reinforcement learning reward, and the mixed reward function of the intelligent agent is defined as:

[0033] ;

[0034] wherein is the reward value at time t, are prediction accuracy reward, signal processing quality reward, and prediction strategy rationality reward, respectively;

[0035] wherein the blood glucose prediction accuracy reward is: , is the blood glucose value predicted by the regression model, is the reference value measured by the invasive input, is the weight coefficient, is the error function;

[0036] signal quality reward : wherein R ppg is the PPG signal optimization reward, , AC is the alternating component, DC is the direct current component; Recg is the ECG signal optimization reward, , SDNN t is the standard deviation of normal sinus rhythm intervals; wherein the standard value is the average of AC / DC of the user in a resting state for 30 consecutive cardiac cycles.

[0037] Optionally, the reinforcement learning network further comprises:

[0038] By analyzing the action space set elements of the multi-modal non-invasive blood glucose model, the feature weights of the PPG and ECG signals satisfy ω ECG +ω PPG =1, 0≤ω ECG ≤1, 0≤ω PPG ≤1, the PPG filter order should be an integer in the interval [16, 64], and the Sigmoid function is used to map the PPG filter order lsm-order to the interval [16, 64]: wherein is the Sigmoid function; the ECG compensation coefficient k interval is [0.05, 0.2], and the ECG compensation coefficient k is mapped to the interval [0.05, 0.2] by linear transformation: ; the dynamic window width is cut off at a proportion of (0~1);

[0039] The deep reinforcement learning network structure is composed of a set S representing the environment state, a set A representing the agent's exploration action, and a reward r obtained by the agent; in the process of continuous exploration of the agent, the Q value function is used to represent the mathematical expectation of the long-term cumulative reward that the agent can obtain by following the strategy to perform subsequent actions after starting from this action;

[0040] ;

[0041] wherein: is the state The policy Decides to take action a at time t The Q-value obtained by the policy The policy The expected value under the policy The discount factor when calculating the return at subsequent state k The return at subsequent state k

[0042] The V-value function represents the mathematical expectation of the long-term cumulative reward that can be obtained by following the policy to perform subsequent actions starting from a certain state of the agent

[0043] ;

[0044] In the formula: The state of state s at time t The policy The V-value obtained by the policy

[0045] Set the parameters of the policy network Actor and the evaluation network Critic of the deep reinforcement learning network structure; the deep reinforcement learning network uses an Actor network and a Critic network Each network has its corresponding target network parameters ;

[0046] Update the sampling policy gradient of the Actor network using the following formula:

[0047] ;

[0048] In the formula, The probability distribution of the output action, The logarithmic gradient of the policy function, The advantage function, calculated as , The state obeys the state distribution under the policy π, Indicates the action is sampled from the probability distribution of the policy ;

[0049] Update the parameters of the Actor network using the following formula:

[0050] ;

[0051] In the formula, The learning rate under the policy ;

[0052] The minimization loss function parameters of the Critic network are optimized by using the following formula:

[0053] ;

[0054] The Critic network parameters are updated by using the following formula:

[0055] ;

[0056] In the formula, is a learning rate;

[0057] In the learning process, an experience replay mechanism is used to store the experience vector of each moment to form an experience pool M; during training, experience samples with a size of mini-batch are randomly extracted from the experience pool M each time for network parameter updating;

[0058] PPG, ECG features, three-axis accelerometer and temperature sensor data and blood glucose prediction network prediction values are collected for training; the iterative training of the Actor and Critic networks is used to continuously improve the deep reinforcement learning strategy until the strategy converges or a predetermined number of training steps is reached to update the network parameters, and the network parameters after training are fixed for solving the multi-modal non-invasive blood glucose model problem.

[0059] Optionally, a feature vector of the ECG signal and the PPG signal is constructed to perform weighted feature fusion on the ECG signal and the PPG signal to obtain a fusion feature vector, including:

[0060] ECG signal features are obtained, the ECG signal features include RR intervals, ST slopes and LF / HF energy ratios, and an ECG signal feature vector is constructed. ;

[0061] PPG signal features are obtained, the PPG signal features include main wave heights, wave area ratios and PTTs, and a PPG signal feature vector is constructed. ;

[0062] The PPG and ECG signals are subjected to weighted feature fusion to obtain a fusion feature vector F .

[0063] Optionally, the regression branch uses a blood glucose value regression model to continuously predict blood glucose values, which specifically includes:

[0064] The fusion feature vector is input, and a time sequence sliding window is used for processing, the window length is 10 continuous sampling points , after two convolutional layers and one pooling layer, a final full connection generates a blood glucose prediction value.​ The prediction formula is as follows:

[0065] , is the weight of the full connection layer, is the bias term of the full connection layer, is the hidden state;

[0066] where the loss function is defined as mean square error during regression training : where is the true blood glucose value, is the predicted blood glucose value, and N is the total number of training samples.

[0067] Optionally, the classification branch combines the cost-sensitive SVM classifier and the classification cost matrix to perform blood glucose grading alarm, which specifically includes:

[0068] A classification cost matrix is established according to the clinical risk priority and the severity of the consequences of misjudgment.

[0069] A cost-sensitive SVM classifier is constructed in combination with the classification cost matrix.

[0070] A three-level decision system is constructed based on the absolute value of blood glucose concentration as a boundary and the time continuity of the prediction result.

[0071] The blood glucose prediction value output by the blood glucose value regression model is input into the cost-sensitive SVM classifier after multi-dimensional feature expansion, and then a three-level risk probability distribution is output.

[0072] The three-level risk probability distribution is input into the three-level decision system to perform grading alarm on the continuous prediction result.

[0073] Optionally, a classification cost matrix is established according to the clinical risk priority and the severity of the consequences of misjudgment, including:

[0074] The principle is that the cost of hypoglycemia false negative is the highest, the cost of hyperglycemia false positive is the second, and the cost of correct prediction in the normal range is the lowest. The specific establishment process of the matrix is as follows: a 3x3 dimensional matrix corresponding to three levels of warning categories is defined; the diagonal elements are correct predictions with the lowest cost; in the non-diagonal elements, the misjudgment cost related to hypoglycemia is higher than that related to hyperglycemia, and the cost of hypoglycemia false negative is higher than that of misjudgment of normal as hyperglycemia; in addition, the matrix will be dynamically adjusted according to real-time expansion features, and the weight of each cost will be real-time corrected in combination with exercise intensity and prediction reliability.

[0075] Optionally, a cost-sensitive SVM classifier is constructed in combination with the classification cost matrix, including:

[0076] The blood glucose prediction value output by the blood glucose value regression model is supplemented with a blood glucose change trend, a circadian rhythm factor, a motion correlation index and a risk confidence to form a 5-dimensional feature vector, wherein the risk confidence is calculated based on the output of the blood glucose value regression model.

[0077] A 3*3 matrix is set according to clinical risk priority, wherein the cost of low blood glucose missed alarm is the highest, the cost of high blood glucose false alarm is the second, and the cost of normal prediction is the lowest, and the weights are dynamically adjusted in combination with the motion intensity and the risk confidence, wherein the motion intensity is determined according to the variance and the main frequency in the three-axis accelerometer signal;

[0078] During training, the classification cost matrix is used as the loss function weight to optimize the parameters, and during inference, the classification cost matrix*posterior probability is calculated, the warning level with the minimum cost is selected, the three-level risk probability distribution is output, and the graded alarm of continuous prediction is triggered.

[0079] Beneficial effects

[0080] The beneficial effects of the present application are: the present application is a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion, a dynamic weight distribution mechanism replaces the traditional fixed weight or rule-driven adaptive method, the weight proportion of PPG (photoelectric plethysmogram) and ECG (electrocardiogram) is adaptively adjusted according to the real-time environmental parameters such as motion state and temperature change, and the MARD value in a dynamic scene is improved. Reinforcement learning (PPO algorithm) driven adaptive adjustment balances the prediction accuracy and weight stability through a reward function, and the following three principles are considered during the design process: clinical accuracy penalty: the error between the predicted blood glucose value and the invasive measurement reference value is measured based on L2 norm to ensure that the result meets the medical standard; weight stability constraint: the difference between the current and previous weights is constrained by L2 norm to avoid model instability caused by drastic weight fluctuations. Multi-dimensional signal preprocessing and motion interference suppression provide environmental characteristics for subsequent signal preprocessing and weight distribution. A scene-specific noise reduction strategy is used to solve the problem that a single filtering method cannot adapt to dynamic motion interference. The robustness and environmental adaptability of the signal are improved. BRIEF DESCRIPTION OF DRAWINGS

[0081] Figure 1 A flowchart of a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0082] In order to better explain the present application and facilitate understanding, the present application is described in detail below through specific embodiments in combination with the drawings.

[0083] The embodiment of the present application relates to a non-invasive blood glucose detection method based on a multi-modal dynamic weight fusion of a photoelectric volume pulse wave (PPG), an electrocardiogram (ECG) and an environmental sensor signal, and is suitable for blood glucose real-time monitoring in an intelligent wearable device.

[0084] In order to better understand the above technical solutions, the exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided so that the present application can be more clearly, thoroughly understood, and the scope of the present application can be completely conveyed to those skilled in the art.

[0085] In a first aspect, with reference to Figure 1 The embodiment provides a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion, comprising:

[0086] S1, aligning the acquired ECG signal, PPG signal, three-axis accelerometer signal and skin temperature sensor signal by hardware timestamp, and constructing a state space.

[0087] S2, taking the state space as the input of the reinforcement learning network, and based on the reinforcement learning network, performing adaptive dynamic weight distribution on the weight of the ECG signal and the PPG signal, optimizing the signal processing parameters, and confirming whether to trigger threshold adjustment.

[0088] S3, constructing a feature vector of the ECG signal and the PPG signal to perform weighted feature fusion on the ECG signal and the PPG signal, and obtaining a fusion feature vector.

[0089] S4, inputting the fusion feature vector into a prediction and early warning double-branch output structure, wherein the prediction and early warning double-branch output structure comprises a regression branch and a classification branch, the regression branch is used for continuous prediction of blood glucose value, and finally outputs a single blood glucose concentration value; the classification branch is used for early warning level determination, and finally outputs a graded alarm to the continuous prediction result, so as to complete the non-invasive blood glucose detection.

[0090] The embodiment provides a non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion. The multi-modal physiological signals (ECG, PPG, accelerometer, temperature) are aligned through hardware time stamping, ensuring time consistency, laying a data foundation for multi-modal fusion, and avoiding feature distortion caused by time offset. Reinforcement learning is introduced to realize dynamic weight distribution and parameter optimization, so that the model can adapt to signal characteristics (such as motion interference and individual differences) under different individuals and different physiological states, solving the problem of poor generalization ability of traditional fixed weight models. Feature weighted fusion fully excavates the complementarity of ECG and PPG signals (ECG reflects cardiac electrical activity, and PPG reflects hemodynamic changes), and improves the feature representation capability. The prediction and early warning double-branch structure takes into account continuous prediction (regression branch) and graded warning (classification branch) of blood glucose value, which not only meets the quantitative detection demand, but also realizes timely intervention of clinical risk through graded warning, improving the practicability and safety of non-invasive blood glucose detection.

[0091] Optionally, the method further comprises preprocessing the acquired PPG signal and ECG signal, specifically comprising:

[0092] using a 0.5Hz-8Hz band-pass filter to eliminate baseline drift;

[0093] The PPG signal is preprocessed according to the following steps: an LMS adaptive filter is used to eliminate motion interference of the PPG signal, and the filter output is , wherein is the kth filtered weight result, L is the filter order, r(n-k) is the delayed sampling value of the reference interference signal, and y(n) is the estimated motion interference of the filter output; the error is calculated as , wherein is the original PPG signal, is the denoised signal, and the weight is updated according to , wherein is the k+1th filtered weight result, μ is the step factor, is the anti-zero constant, r(n) is the reference motion interference signal synchronized with the PPG signal, e(n) is low-pass filtered to suppress residual high-frequency noise;

[0094] Based on the physiological characteristics and signal-to-noise ratio of the PPG waveform, a dynamic window width is set; only high signal-to-noise ratio data in the systolic period is intercepted, and the diastolic period and noise interference period are discarded; the systolic period starting point corresponds to the maximum value of the second derivative , wherein, the time point at which the output represents reaches the maximum value, and argmax represents parameter maximization; the window width is dynamically adjusted according to the output of the reinforcement learning, and the data in the intercepted window is smoothed and interpolated to ensure the continuity of the waveform;

[0095] According to the compensation calculation formula The ECG signal is pre-processed to realize drift compensation, wherein is a reference ECG signal without motion interference, is an original ECG signal with motion interference, a y is the acceleration value in the Y-axis direction, a z is the acceleration value in the Z-axis direction, wherein j1 and j2 are compensation coefficients, and the initial values are set to 0.15 and 0.20 respectively, and are dynamically adjusted according to the output results of reinforcement learning subsequently.

[0096] The baseline drift is effectively eliminated by the band-pass filter (0.5Hz-8Hz), and the signal is preliminarily purified to provide a stable basis for subsequent processing. The LMS adaptive filter specifically eliminates the motion interference of the PPG signal, significantly improves the signal-to-noise ratio of the PPG signal through dynamic weight update and error feedback, and solves the core problem of the decrease in the accuracy of non-invasive detection in the motion scene. Based on the dynamic window width of the PPG physiological characteristics (only the high signal-to-noise ratio data in the systolic period is retained), the diastolic period noise interference is reduced, and the window width is dynamically adjusted and smoothed by interpolation through reinforcement learning, ensuring the waveform continuity and improving the reliability of feature extraction. The motion drift compensation formula of the ECG signal specifically eliminates the motion artifacts, dynamically adjusts the compensation coefficients, solves the distortion problem of the ECG signal in the body movement, and provides protection for the accurate extraction of the heart rate related features (such as RR interval).

[0097] Optionally, the state space set is:

[0098] ;

[0099] wherein, is the ratio of alternating current / direct current components of the PPG signal, is the pulse transit time, is the standard deviation of the RR interval of the ECG, is the low / high frequency power ratio of the ECG, is the acceleration time domain variance, is the main frequency of the acceleration frequency domain, is the skin temperature gradient, is the mean value of the historical blood glucose prediction error.

[0100] Through multi-dimensional state quantization, the reinforcement learning agent is provided with comprehensive decision basis, so that it can accurately judge the current physiological and signal state, and improve the rationality of dynamic adjustment.

[0101] Optionally, the reinforcement learning network is used for self-adaptive dynamic weight distribution of the weights of the ECG signal and the PPG signal, including:

[0102] The adjusting action of the multi-modal non-invasive blood glucose model is used for adaptive dynamic weight distribution of feature weight of signals, optimization of signal processing parameters, and confirmation of whether to trigger threshold adjustment, and defines the action space set as:

[0103] ;

[0104] wherein, is a feature fusion weight, satisfying the constraint , is a PPG adaptive filter order (16-64 orders), is an ECG motion artifact compensation coefficient (0.05 mV / g-0.2 mV / g), is a dynamic window width shrinkage period ratio, is a calibration trigger decision;

[0105] The target of dynamic adjustment of the multi-modal non-invasive blood glucose model is the optimization of prediction accuracy, signal quality and prediction strategy rationality, and the optimization of multi-modal non-invasive blood glucose prediction accuracy, signal quality and prediction strategy rationality is converted into the maximization of reinforcement learning reward, and a mixed reward function of the agent is defined as:

[0106] ;

[0107] wherein is a reward value at time t, are respectively a prediction accuracy reward, a signal processing quality reward and a prediction strategy rationality reward;

[0108] wherein, the blood glucose prediction accuracy reward is: , is a regression model predicted blood glucose value, is an invasive input measurement reference value, is a weight coefficient, is an error function;

[0109] The signal quality reward is : wherein R ppg is a PPG signal optimization reward, , AC is an alternating current component, and DC is a direct current component; Recg is an ECG signal optimization reward, , SDNN t is a standard deviation of normal sinus intervals; wherein the standard value is an AC / DC average value of 30 consecutive cardiac cycles in a user's resting state.

[0110] The action space encompasses key parameters such as feature weights, filter order, and compensation coefficients, enabling multi-dimensional collaborative optimization and overcoming the limitations of traditional single-parameter adjustments. A hybrid reward function (prediction accuracy + signal quality + strategy rationality) guides the model to improve prediction accuracy while simultaneously considering signal processing stability and long-term strategy effectiveness (e.g., avoiding frequent calibration), balancing short-term performance with long-term robustness. The blood glucose prediction accuracy reward (based on the error between predicted and reference values) is directly related to the core detection target, while the signal quality reward (AC / DC optimization for PPG and SDNN optimization for ECG) ensures data quality from the source. The combination of these two elements forms a closed-loop optimization, significantly improving the model's anti-interference capability.

[0111] Optionally, the reinforcement learning network further includes:

[0112] By analyzing the action space set elements of the multimodal noninvasive blood glucose model, the feature weights of PPG and ECG signals satisfy ω ECG +ω PPG =1, 0≤ω ECG ≤1, 0≤ω PPG ≤1, the PPG filter order should be an integer in the interval [16, 64]. Use the Sigmoid function to map the PPG filter order lsm-order to the interval [16, 64]. ,in The function is the Sigmoid function; the ECG compensation coefficient k is in the interval [0.05, 0.2]. A linear transformation is used to map the ECG compensation coefficient k to the interval [0.05, 0.2]. The dynamic window width is truncated at a ratio of (0~1).

[0113] The deep reinforcement learning network structure consists of a set S representing the environmental state, a set A representing the agent's exploration actions, and the reward r obtained by the agent. During the agent's continuous exploration, the Q-value function is used to represent the mathematical expectation of the long-term cumulative reward that the agent can obtain by following the policy and executing subsequent actions from the beginning of the action.

[0114] ;

[0115] In the formula: Let s be the state at time t. Next, through strategy Decision to take action a at time t The Q value obtained below; For strategy The following expectations; This is the discount factor used when calculating the payoff in subsequent state k; This represents the reward for the subsequent state k.

[0116] The V-value function is used to represent the mathematical expectation of the long-term cumulative reward that can be obtained by following the policy to perform subsequent actions starting from a certain state of the agent;

[0117] ;

[0118] In the formula: is the state of state s at time t , the V-value obtained by the policy decision-making;

[0119] The parameters of the policy network Actor and the evaluation network Critic of the deep reinforcement learning network structure are set; the deep reinforcement learning network uses an Actor network and a Critic network , each of which has its corresponding target network parameters ;

[0120] The sampling policy gradient of the Actor network is updated by using the following formula:

[0121] ;

[0122] In the formula, is the probability distribution of the output action, is the logarithmic gradient of the policy function, is the advantage function, and the calculation formula is , is the state subject to the state distribution under the policy π, indicates the action is sampled from the probability distribution of the policy ;

[0123] The parameters of the Actor network are updated by using the following formula:

[0124] ;

[0125] In the formula, is the learning rate under the policy ;

[0126] The minimum loss function parameters of the Critic network are optimized by using the following formula:

[0127] ;

[0128] The Critic network parameters are updated by using the following formula:

[0129] ;

[0130] In the formula, learning rate;

[0131] In the learning process, an experience replay mechanism is adopted to store the experience vector of each moment to form an experience pool M; during training, experience samples with a size of mini-batch are randomly extracted from the experience pool M each time for network parameter updating;

[0132] PPG, ECG features, three-axis accelerometer and temperature sensor data and blood glucose prediction network prediction values are collected for training; the iterative training of the Actor and Critic networks is used to continuously improve the deep reinforcement learning strategy until the strategy converges or a predetermined number of training steps is reached to update the network parameters, and the network parameters after training are fixed for solving the multi-modal non-invasive blood glucose model problem.

[0133] Action parameter mapping (Sigmoid function mapping filter order, linear transformation mapping compensation coefficient) ensures that the parameters are within a physiological reasonable range (such as filter order 16-64), avoids invalid exploration, and improves learning efficiency. The Actor-Critic double network structure combined with the target network and the experience replay mechanism reduces sample correlation, stabilizes the training process, speeds up policy convergence, and solves the common training fluctuation problem in reinforcement learning. The policy gradient and loss function optimization formula (gradient update of the Actor network, MSE loss of the Critic network) realizes fine adjustment of the parameters, so that the agent can efficiently learn the optimal strategy, and the stable parameters finally output can be directly used for real-time inference of the multi-modal model, improving the response speed of the detection system.

[0134] Optionally, a feature vector of the ECG signal and the PPG signal is constructed to perform weighted feature fusion on the ECG signal and the PPG signal to obtain a fusion feature vector, including:

[0135] An ECG signal feature is obtained, the ECG signal feature including an RR interval, an ST slope, and an LF / HF energy ratio, and an ECG signal feature vector is constructed ;

[0136] A PPG signal feature is obtained, the PPG signal feature including a main wave height, a waveform area ratio, and a PTT, and a PPG signal feature vector is constructed ;

[0137] The PPG and ECG signals are subjected to weighted feature fusion to obtain a fusion feature vector F .

[0138] ​​​The core physiological characteristics of ECG (RR interval, ST slope, etc.) and PPG (main wave height, PTT, etc.) are extracted, which have clear physiological correlation with blood glucose concentration (such as PTT changes with the increase of blood glucose), to ensure the effectiveness of the characteristics. Weighted fusion dynamically allocates weights through reinforcement learning, so that the model can adaptively highlight effective features and suppress noise features in different scenarios (such as improving PPG weight when ECG signal quality is poor), solve the problems of "information redundancy" and "feature conflict" of multi-modal data, and improve the discriminability of fused features.

[0139] Optionally, the regression branch uses a blood glucose value regression model for continuous prediction of blood glucose value, which specifically includes:

[0140] The input fusion feature vector is processed using a time series sliding window, and the window length is 10 consecutive sampling points After two convolutional layers and one pooling layer, the final full connection generates a blood glucose prediction value The prediction formula is as follows:

[0141] , is the weight of the full connection layer, is the bias term of the full connection layer, is the hidden state;

[0142] The loss function is defined as mean square error during regression training : , wherein is the true blood glucose value, is the predicted blood glucose value, and N is the total number of training samples.

[0143] The time series sliding window captures the dynamic trend of blood glucose value, solving the problem of accidental sampling at a single time point. The convolutional layer and the pooling layer are combined to effectively extract local correlation information in the time series features (such as short-term blood glucose fluctuation pattern), and the full connection layer outputs continuous blood glucose value, taking into account the depth of feature extraction and the continuity of prediction. The mean square error (MSE) loss function directly optimizes the deviation between the predicted value and the true value, ensuring the regression accuracy and providing accurate blood glucose quantitative reference for clinical practice.

[0144] Optionally, the classification branch combines a cost-sensitive SVM classifier and a classification cost matrix for blood glucose classification and warning, which specifically includes:

[0145] A classification cost matrix is established according to the clinical risk priority and the severity of the consequences of misjudgment;

[0146] A cost-sensitive SVM classifier is constructed in combination with the classification cost matrix;

[0147] Based on the absolute value of blood glucose concentration, a three-level decision system is constructed in combination with the time continuity of the prediction result;

[0148] The blood glucose prediction value output by the blood glucose value regression model is input into a cost-sensitive SVM classifier after multi-dimensional feature expansion, and then a three-level risk probability distribution is output;

[0149] The three-level risk probability distribution is input into a three-level decision system to grade and alarm the continuous prediction results.

[0150] The cost-sensitive SVM combines with the clinical cost matrix to quantify the medical risk (such as the serious consequences of hypoglycemia missed report) as a classification cost, solving the problem that the traditional classifier is not sensitive to “high risk misjudgment”. The three-level decision system combines the absolute value of blood glucose and the continuity of time to avoid false alarms caused by single prediction errors and improve the stability of early warning. Multi-dimensional feature expansion (blood glucose trend, circadian rhythm, etc.) provides more abundant basis for classification, and the output three-level risk probability distribution can directly guide clinical intervention (such as emergency treatment of hypoglycemia and routine monitoring of hyperglycemia), improving the clinical practical value of non-invasive detection.

[0151] Optionally, a classification cost matrix is established according to the clinical risk priority and the severity of misjudgment consequences, including:

[0152] According to the principle that the cost of hypoglycemia missed report is the highest, the cost of hyperglycemia false report is the second, and the cost of correct prediction in the normal range is the lowest, the specific establishment process of the matrix is as follows: define a 3x3 dimensional matrix corresponding to the three-level warning categories; the diagonal elements are correct predictions with the lowest cost; in the non-diagonal elements, the misjudgment cost related to hypoglycemia is higher than that related to hyperglycemia, and the cost of hypoglycemia missed report is higher than that of misjudgment of normal as hyperglycemia; in addition, the matrix will be dynamically adjusted according to the real-time expansion features, and the cost weights will be real-time corrected in combination with the exercise intensity and the prediction reliability.

[0153] According to the principle that “the cost of hypoglycemia missed report is the highest, and the cost of hyperglycemia false report is the second”, the clinical risk priority (hypoglycemia may cause coma and other emergencies) is met, so that the classification decision is more in line with the actual needs of medical treatment. The dynamic adjustment mechanism (combined with exercise intensity and prediction reliability) makes the cost matrix adapt to complex scenarios (such as large blood glucose fluctuations during exercise, appropriately increasing the warning sensitivity), avoiding the limitations of fixed matrix in dynamic physiological state, and improving the flexibility and accuracy of early warning.

[0154] Optionally, a cost-sensitive SVM classifier is constructed in combination with the classification cost matrix, including:

[0155] The blood glucose prediction value output by the blood glucose value regression model is supplemented with blood glucose change trend, circadian rhythm factor, exercise correlation index and risk confidence to form a 5-dimensional feature vector, wherein the risk confidence is calculated based on the output of the blood glucose value regression model;

[0156] A 3x3 matrix is set according to clinical risk priority, in which the cost of low blood sugar missed alarm is the highest, the cost of high blood sugar false alarm is the second, and the cost of normal prediction is the lowest, and the weights are dynamically adjusted in combination with exercise intensity and risk confidence, wherein the exercise intensity is determined according to the variance and main frequency in the three-axis accelerometer signal;

[0157] During training, the classification cost matrix is used as a loss function weight to optimize parameters, and during inference, the total cost is calculated by combining the classification cost matrix and the posterior probability, the warning level with the minimum cost is selected, and the three-level risk probability distribution is output, which is used to trigger the hierarchical alarm of continuous prediction.

[0158] The 5-dimensional feature vector breaks through the limitation of a single blood glucose value, captures the multi-dimensional rules of blood glucose changes, and improves the adaptability of the classifier to complex physiological states. The introduction of exercise intensity (based on accelerometer signal) and risk confidence enables the model to distinguish between "true high blood sugar" and "motion-induced false high blood sugar signal", and reduces misjudgment under the interference of exercise. The classification cost matrix x posterior probability is used as the basis for decision-making, the warning level with the minimum cost is selected, and the clinical safety of the decision-making is ensured from the mathematical level, and the serious medical risks (such as low blood sugar missed alarm) are minimized.

[0159] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application should also include these modifications and variations.

[0160] Although the embodiments of the present application have been shown and described above, it should be understood that the above-described embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above-described embodiments within the scope of the present application.

Claims

1. A non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion, characterized in that, The method comprises the following steps: Aligning the acquired ECG signal, PPG signal, three-axis accelerometer signal and skin temperature sensor signal by hardware timestamp, and constructing a state space; The state space set is: ; wherein, is the PPG signal AC / DC component ratio, is the pulse transit time, is the ECG RR interval standard deviation, is the ECG low / high frequency power ratio, is the acceleration time domain variance, is the acceleration frequency domain dominant frequency, is the skin temperature gradient, is the historical blood glucose prediction error mean; Taking the state space as the input of the reinforcement learning network, the reinforcement learning network is used to perform adaptive dynamic weight distribution on the weights of the ECG signal and the PPG signal, optimize the signal processing parameters, and confirm whether to trigger threshold adjustment; Constructing a feature vector of the ECG signal and the PPG signal to perform weighted feature fusion on the ECG signal and the PPG signal, and obtaining a fused feature vector; The fused feature vector is input into a prediction and early warning double-branch output structure, wherein the prediction and early warning double-branch output structure comprises a regression branch and a classification branch, the regression branch is used for continuous prediction of blood glucose value, and finally outputs a single blood glucose concentration value; the classification branch is used for early warning level determination, and finally outputs a graded alarm to the continuous prediction result, so as to complete non-invasive blood glucose detection.

2. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 1, characterized in that, The method further comprises pre-processing the acquired PPG signal and ECG signal, which specifically comprises: Using a 0.5Hz-8Hz band-pass filter to eliminate baseline drift; The PPG signal is preprocessed according to the following steps: the LMS adaptive filter is used to eliminate the motion interference of the PPG signal, and the filter output is wherein is the kth filtered weight result, L is the filter order, r(n-k) is the delayed sampling value of the reference interference signal, and y(n) is the estimated motion interference of the filter output; the error is calculated by wherein is the original PPG signal, is the denoised signal, and the weight is updated according to wherein is the k+1th filtered weight result, μ is the step factor, ϵ is the anti-zero constant, r(n) is the reference motion interference signal synchronized with the PPG signal, and e(n) is low-pass filtered to suppress the residual high-frequency noise. Based on the physiological characteristics and signal-to-noise ratio of the PPG waveform, a dynamic window width is set; only high signal-to-noise ratio data is intercepted in the systolic phase, and the diastolic phase and noise interference period are discarded; the systolic phase starting point corresponds to the maximum value of the second derivative wherein, The output represents The time point at which the maximum value is reached, argmax represents the parameter maximization; the output is dynamically adjusted according to the reinforcement learning, the window width is adjusted, the data in the intercepted window is smoothed and interpolated, and the waveform continuity is ensured; According to the compensation calculation formula The ECG signal is preprocessed to realize drift compensation, wherein is the reference ECG signal without motion interference, is the original ECG signal with motion interference, a y is the acceleration value in the Y-axis direction, a z is the acceleration value in the Z-axis direction, wherein j1 and j2 are compensation coefficients, the initial values of which are set to 0.15 and 0.20 respectively, and are dynamically adjusted according to the output results of reinforcement learning subsequently.

3. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 2, characterized in that, The reinforcement learning network is used to perform adaptive dynamic weight distribution on the weights of the ECG signal and the PPG signal, which comprises: The adjustment action of the multi-modal non-invasive blood glucose model is used to perform adaptive dynamic weight distribution on the feature weights of the signal, optimize the signal processing parameters, and confirm whether to trigger threshold adjustment, and the set of action spaces is defined as: ; wherein, is a feature fusion weight, satisfying the constraint , is a PPG adaptive filter order (16-64 orders), is an ECG motion artifact compensation coefficient (0.05 mV / g - 0.2 mV / g), is a dynamic window width shrinkage period ratio, is a calibration trigger decision; The target of the dynamic adjustment of the multi-modal non-invasive blood glucose model is the optimization of prediction accuracy, signal quality and prediction strategy rationality, the optimization of prediction accuracy, signal quality and prediction strategy rationality of the multi-modal non-invasive blood glucose model are converted into the maximization of the reinforcement learning reward, and the mixed reward function of the agent is defined as: ; In the formula is the reward value at time t, respectively, are the prediction accuracy reward, the signal processing quality reward, and the prediction strategy rationality reward. Wherein, the blood glucose prediction accuracy reward: , is the blood glucose value predicted by the regression model, is the reference value of the invasive input measurement, is the weight coefficient, is the error function; Signal quality reward : where R ppg is a PPG signal optimization reward, AC is an alternating current component and DC is a direct current component; Recg is an ECG signal optimization reward, SDNN t is a standard deviation of normal sinus intervals; wherein the standard value is an average of AC / DC of 30 consecutive heart cycles in a resting state of the user.

4. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 3, characterized in that, The reinforcement learning network further comprises: By analyzing the action space set elements of the multi-modal non-invasive blood glucose model, the feature weights of the PPG and ECG signals satisfy ω ECG +ω PPG =1, 0≤ω ECG ≤1, 0≤ω PPG ≤1, the PPG filter order should be an integer in the interval [16, 64], and the PPG filter order lsm-order is mapped to the interval [16, 64] using a Sigmoid function: , wherein is a Sigmoid function; the ECG compensation coefficient k is in the interval [0.05, 0.2], and the ECG compensation coefficient k is mapped to the interval [0.05, 0.2] by linear transformation: ; the dynamic window width is cut off at a proportion of (0~1). The deep reinforcement learning network structure is composed of a set of state representations S, a set of agent exploration actions A, and a reward r obtained by the agent; in the process of continuous exploration of the agent, a Q value function is used to represent the mathematical expectation of the long-term cumulative reward that can be obtained by the agent after following the strategy to perform subsequent actions from the action; ; wherein: s(t) is the state at time t the policy a(t) is the action taken at time t the Q-value resulting from the policy the expected reward r is the discount factor when computing the reward at the subsequent state k r(k) is the reward at the subsequent state k A V value function is used to represent the mathematical expectation of the long-term cumulative reward that can be obtained by the agent after following the strategy to perform subsequent actions from the state; ; In the formula: is the state s at time t Next, the V value obtained by the policy decision; a policy network Actor and a critic network Critic setting the parameters of the deep reinforcement learning network structure; the deep reinforcement learning network uses an Actor network and a Critic network each with its corresponding target network parameters ; The sampling policy gradient of the Actor network is updated by using the following formula: ; where, is the probability distribution over actions, is the log gradient of the policy function, is the advantage function, computed as , is the state is the state distribution under policy denotes an action is sampled from the probability distribution over actions The parameters of the Actor network are updated by using the following formula: ; In the formula, The strategy under the learning rate; The minimum loss function parameters of the Critic network are optimized by using the following formula: ; The Critic network parameters are updated by using the following formula: ; In the formula, is the learning rate; During the learning process, an experience playback mechanism is employed to store the experience vector at each moment. The experience samples are stored to form an experience pool M; during training, a mini-batch of experience samples of size M are randomly drawn from the experience pool M each time for network parameter updates. The PPG, ECG features, three-axis accelerometer and temperature sensor data and the predicted values of the blood glucose prediction network are collected for training; the iterative training of the Actor and Critic networks is used to continuously improve the deep reinforcement learning strategy until the strategy converges or a predetermined number of training steps is reached to update the network parameters, and the network parameters after training are fixed, which are used for solving the multi-modal non-invasive blood glucose model problem.

5. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 4, characterized in that, The feature vectors of the ECG signal and the PPG signal are constructed to perform feature fusion on the ECG signal and the PGG signal by weighting, to obtain a fusion feature vector, including: An ECG signal feature is acquired, the ECG signal feature including an RR interval, an ST slope, and an LF / HF energy ratio, and an ECG signal feature vector is constructed as ; Obtaining PPG signal features, the PPG signal features including a main wave height, a wave shape area ratio, and a PTT, and constructing a PPG signal feature vector as ; The feature fusion of the PPG and ECG signals is weighted to obtain a fusion feature vector F as .

6. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 5, characterized in that, The regression branch uses a blood glucose value regression model to perform continuous prediction of the blood glucose value, which specifically includes: Input fusion feature vector, use time sequence sliding window processing, window length is 10 continuous sampling points After two convolutional layers and a pooling layer, the final full connection generates blood glucose prediction value The prediction formula is as follows: , are weights of the fully connected layer, are bias terms of the fully connected layer, is a hidden state; where regression training defines the loss function as mean square error : , where is the true blood glucose value, is the predicted blood glucose value, and N is the total number of training samples.

7. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 6, characterized in that, The classification branch combines a cost-sensitive SVM classifier and a classification cost matrix to perform blood glucose classification alarm, which specifically includes: A classification cost matrix is established according to the clinical risk priority and the severity of the consequences of misjudgment. A cost-sensitive SVM classifier is constructed in combination with the classification cost matrix. A three-level decision system is constructed based on the absolute value of the blood glucose concentration as a boundary and the time continuity of the prediction result. The blood glucose prediction value output by the blood glucose value regression model is input into the cost-sensitive SVM classifier after multi-dimensional feature expansion, and then a three-level risk probability distribution is output. The three-level risk probability distribution is input into the three-level decision system to perform classification alarm on the continuous prediction result.

8. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 7, characterized in that, A classification cost matrix is established according to the clinical risk priority and the severity of the consequences of misjudgment, including: The specific establishment process of the matrix is as follows: a 3x3 dimensional matrix corresponding to three levels of warning categories is defined; the diagonal elements are correct predictions with the lowest cost; in the non-diagonal elements, the misjudgment cost related to hypoglycemia is higher than that related to hyperglycemia, and the hypoglycemia false negative cost is higher than that of misjudgment of normal as hyperglycemia; in addition, the matrix is dynamically adjusted according to real-time expansion features, and the weight of each cost is real-time corrected in combination with the exercise intensity and the prediction reliability.

9. The non-invasive blood glucose detection method based on reinforcement learning and multi-modal dynamic weight fusion according to claim 8, characterized in that, A cost-sensitive SVM classifier is constructed in combination with the classification cost matrix, including: The blood glucose prediction value output by the blood glucose value regression model is supplemented with blood glucose trend, circadian rhythm factor, exercise correlation index and risk confidence to form a 5-dimensional feature vector, wherein the risk confidence is calculated based on the output of the blood glucose value regression model; A 3x3 matrix is set according to the clinical risk priority, wherein the cost of hypoglycemia false negative is the highest, the cost of hyperglycemia false positive is the second, and the cost of normal prediction is the lowest, and the weight is dynamically adjusted in combination with the exercise intensity and the risk confidence, wherein the exercise intensity is determined according to the variance and dominant frequency in the three-axis accelerometer signal; During training, the classification cost matrix is used as the loss function weight to optimize the parameters, and during inference, the total cost is calculated by: classification cost matrix x posterior probability, the warning level with the minimum cost is selected, and a three-level risk probability distribution is output, which is used to trigger the classification alarm of the continuous prediction.

Citation Information

Patent Citations

  • Noninvasive blood glucose detection method and device, electronic equipment and medium

    CN118078275A

  • Systems and methods for non-invasive blood pressure measurement

    US20170156606A1