Unbalanced data-oriented blood glucose regression modeling method based on bioelectrical impedance

By dynamically enhancing sparse region samples using a multi-band impedance sensor and the SMOGN-Boundary algorithm, and optimizing model hyperparameters using the Sparrow Search algorithm, the problem of unbalanced data in non-invasive blood glucose testing is solved, achieving high-precision blood glucose prediction.

CN120977599APending Publication Date: 2025-11-18GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510985059.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing non-invasive blood glucose testing technologies suffer from poor generalization performance due to imbalanced data, especially when there are many healthy samples but few hyperglycemic samples. As a result, existing technologies are unable to effectively predict the accuracy of high and low blood glucose levels.

Method used

A multi-band impedance sensor is used to acquire bioelectrical signals. The SMOGN-Boundary algorithm is combined to enhance the samples in sparse regions. A balanced training set is constructed through dynamic mixed sampling and sample weight adjustment. The sparrow search algorithm is used to optimize the hyperparameters of the extreme random tree model and improve the generalization ability of the model.

Benefits of technology

It effectively expands the sample size for both high and low blood sugar levels, maintains the continuity of data distribution characteristics and boundaries, and improves the model's learning ability and prediction accuracy in abnormal regions, especially exhibiting higher fitting accuracy and stability in both high and low blood sugar levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977599A_ABST
    Figure CN120977599A_ABST
Patent Text Reader

Abstract

The invention discloses an unbalanced data-oriented bioelectrical impedance-based blood glucose regression modeling method, which comprises the following steps of: 1, acquiring the amplitude, phase and impedance characteristics of the forearm of a testee by using a multi-band electrical impedance sensor, performing invasive blood glucose detection at the same time, and constructing an annotation data set between a bioelectrical signal and a standard blood glucose value; 2, dynamically enhancing the diversity and coverage range of sparse region samples by using an SMOGN-Boundy algorithm, and expanding the sample size and balancing data distribution; 3, on the basis of a sample reweighting strategy of reciprocal square root weighting, sample weights are embedded into a loss function of the model; 4, training an extreme random tree blood glucose prediction model based on the loss function, and introducing a sparrow search algorithm to optimize model hyper-parameters to obtain optimal hyper-parameters of the model; and step 5, performing model training based on the optimal hyper-parameter to obtain a final extreme random tree blood glucose prediction model. According to the method, the generalization ability of the blood glucose prediction model of the unbalanced data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of blood glucose prediction model, and particularly relates to a blood glucose regression modeling method based on bioelectrical impedance for unbalanced data. BACKGROUND

[0002] At present, blood glucose detection mainly relies on invasive detection means, and needs to detect glucose level through blood sampling, which brings pain and infection risk to patients. Bioelectrical impedance spectroscopy as a non-invasive detection means has been used to study human tissue composition and physiological state, but the correlation between its signal and blood glucose is complex, and in actual data collection, there are usually more healthy population samples and relatively few high blood glucose samples, resulting in an unbalanced problem in the data set, and the existing technology is difficult to achieve good generalization performance on such unbalanced regression data set. SUMMARY

[0003] In order to solve the problem that the existing non-invasive blood glucose detection technology cannot achieve good generalization performance due to unbalanced data, the present application provides a blood glucose regression modeling method based on bioelectrical impedance for unbalanced data.

[0004] In order to achieve the above purpose, the present application provides the following technical scheme:

[0005] Step 1, using a multi-frequency band impedance sensor, the amplitude, phase and impedance characteristics of the tester's forearm are obtained, and at the same time, invasive blood glucose detection is performed to construct a labeled data set between bioelectric signals and blood glucose .

[0006] Step 2, using SMOGN-Boundary algorithm to dynamically enhance the diversity and coverage of sparse area samples, expand sample size and balance data distribution, including steps 2.1 to 2.5:

[0007] Step 2.1, based on the density distribution characteristics of the target variable (blood glucose value), the correlation function is calculated , the calculation formula is:

[0008] (1)

[0009] Wherein, is the sample density of the target value , and is the maximum density of the target variable .

[0010] Step 2.2, target-driven sample partition, dividing key area DR and normal area DN, and the partition threshold tR is set to satisfy:

[0011] (2)

[0012] wherein, is the target value is the sample density, is the target variable is the maximum density.

[0013] Step 2.3, boundary judgment based on target variable difference mean, by calculating the target variable difference mean of the sparse area sample and its neighbor, if the difference exceeds the preset threshold, it is marked as boundary sample, the calculation formula is as follows:

[0014] (3)

[0015] wherein is the near neighbor set of the sample , τ is the dynamic threshold. This mechanism can effectively capture the critical samples sensitive to model decision.

[0016] Step 2.4, dynamic hybrid sampling, conditional oversampling is implemented for DR area and boundary samples, linear interpolation is used to generate new samples, if the sample is far away from the neighbor, adaptive Gaussian noise is injected to enhance global diversity.

[0017] Step 2.5, the newly generated samples are fused with the original data to construct the balanced training set .

[0018] Step 3, based on the square root reciprocal weighted sample reweighting strategy, the sample weight is embedded into the loss function of the model, including the following steps 3.1 to step 3.2:

[0019] Step 3.1, assuming the target variable blood glucose value appears in the data set , the weight of the calculation formula is:

[0020] (4)

[0021] Step 3.2, embed the sample weight into the loss function of the model, the loss function is adjusted to weighted mean square error (WMSE):

[0022] (5)

[0023] Based on the loss function, the extreme random tree model is trained, and the sparrow search algorithm is introduced to optimize the model hyperparameters, including steps 4.1 to 4.6:

[0024] Step 4.1, set the hyperparameters and range of the extreme random tree blood glucose prediction model, including the number of base learners , the maximum tree depth , the minimum sample division number and the minimum leaf node sample number mapping the hyperparameters of the extreme random tree blood glucose prediction model to the positions of the sparrows .

[0025] Step 4.2, define the sparrow individual position set the total number of sparrows and the maximum number of iterations initialize the sparrow population .

[0026] Step 4.3, input the sparrow individual position into the extreme random tree for model training, and calculate its WMSE on the validation set as the fitness value:

[0027] (6)

[0028] where, is the fitness value of the sparrow individual at position .

[0029] Step 4.4, update the individual position and population iteration according to the behavior mechanism of the sparrow algorithm, including:

[0030] the top individuals in the population as foragers for global search, and the position update formula is as follows

[0031] (7)

[0032] where, and represent the positions of the sparrow individual at the th and the th iteration, is the maximum number of iterations, is a random number, is a warning value, is a warning threshold, is a random variable following standard normal distribution, is a full 1 vector;

[0033] for the individuals ranked after as followers, local search is performed around the foragers, and the position update formula is:

[0034] (8)

[0035] where represents the current global worst individual position;

[0036] If some sparrows think that the current environment is dangerous (such as falling into a local optimum), an alert update mechanism is adopted to quickly jump out of the current search area:

[0037] (9)

[0038] wherein x represents the current global fitness optimal individual position, is a disturbance coefficient. Step 4.5, re-evaluating the WMSE according to the new position.

[0039] Step 4.6, repeating S4.3 to S4.5 until the maximum number of iterations is reached, and finally outputting the optimal individual position

[0040] , which is used to train the final blood glucose prediction model. Compared with the prior art, the present application has the following advantages:

[0041] The present application uses the SMOGN-Boundary algorithm to synthesize and enhance the sample sparse boundary region (such as the high and low blood glucose segment) in the original blood glucose data set, can generate new samples according to the distribution of the target variable, not only effectively expands the sample number in the extreme interval, but also maintains the distribution characteristics and boundary continuity of the data, balances the data distribution, and thus improves the learning ability and generalization ability of the model in the abnormal area.

[0042] The present application calculates the square root reciprocal of the frequency of occurrence of the target value to which the sample belongs as a weight, and embeds it into the loss function of the regression model, realizes the weighted mean square error (WMSE) optimization goal, which significantly improves the focus of the model in the training process, makes the model pay more attention to predicting low-frequency key samples, especially in the high and low blood glucose segment, has higher fitting precision.

[0043] The present application uses the sparrow search algorithm to intelligently search for model hyperparameters, uses the sparrow search algorithm which has the characteristics of fast global search and local development ability, can effectively jump out of the local optimum trap, so that the model has better generalization ability and improves the prediction stability in practical application.

[0044] BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0046] Figure 1 ​It is the overall framework diagram of a blood glucose regression modeling method based on bioelectrical impedance for unbalanced data.

[0047] Figure 2 It is a flow chart of optimizing an extreme random tree model by using a sparrow search algorithm in the application.

[0048] Figure 3 It is a Clarke error grid map of the blood glucose value predicted by the trained model and the standard blood glucose value. DETAILED DESCRIPTION

[0049] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.

[0050] In order to illustrate the technical solutions of the application, specific embodiments are used for illustration.

[0051] As shown in the overall framework diagram of a blood glucose regression modeling method based on bioelectrical impedance for unbalanced data, the overall framework diagram comprises S101, S102, S103 and S104. Figure 1

[0052] In the S101, a bioelectrical impedance detection hardware platform is built by using a resonance method to detect bioelectrical impedance, an inductive element is externally added to form a resonance circuit with human tissues, the system collects bioimpedance spectrum with an excitation signal output frequency band in the range of 1MHz~70MHz, and the frequency is swept at an interval of 0.5MHz for detection, so as to accurately obtain the measurement data of bioelectrical impedance, and the standard blood glucose value is collected by a method of invasive collection to construct a data set .

[0053] The S102 comprises S2.1 to S2.5.

[0054] S1, a density judgment stage, a correlation function is calculated based on the density distribution characteristics of the target variable (blood glucose value) , and the calculation formula is as follows:

[0055] (1)

[0056] wherein, is the sample density of the target value , and is the maximum density of the target variable .

[0057] ​S2.2, target-driven sample partitioning, dividing the critical region DR and the normal region DN, and setting the partition threshold tR to satisfy:

[0058] (2)

[0059] S2.3, boundary judgment based on target variable difference mean, by calculating the target variable difference mean of the sparse region sample and its neighbor, if the difference exceeds the preset threshold, it is marked as a boundary sample, the calculation formula is as follows:

[0060] (3)

[0061] wherein is the neighbor set of sample , and τ is a dynamic threshold. This mechanism can effectively capture the critical samples sensitive to model decision.

[0062] S2.4, dynamic hybrid sampling, conditional oversampling is implemented for DR region and boundary samples, linear interpolation is used to generate new samples, if the sample is far away from the neighbor, adaptive Gaussian noise is injected to enhance global diversity.

[0063] S2.5, merging the newly generated samples with the original data to construct an equalized training set .

[0064] The S103 includes steps S3.1 to S3.2:

[0065] S3.1, assuming that the target variable blood glucose value occurs in the data set with a frequency , the weight is calculated as follows:

[0066] (4)

[0067] S3.2, embedding the sample weight into the loss function of the model, and adjusting the loss function to weighted mean square error (WMSE):

[0068] (5)

[0069] wherein, denotes the total number of samples in the data set.

[0070] The S104 sparrow search algorithm optimizes the model hyperparameters, as shown in Figure 2 , the sparrow search algorithm is used to optimize the extreme random tree hyperparameters, including the following steps S201-S208:

[0071] S201, setting the hyperparameters and range of the extreme random tree blood glucose prediction model, including the number of base learners Maximum tree depth Minimum sample split Minimum leaf node sample Mapping extreme random tree blood glucose prediction model hyperparameters to sparrow individual positions .

[0072] S202, define sparrow individual positions Set the total number of sparrow individuals And the maximum number of iterations Initialize the sparrow population .

[0073] S203, input the sparrow individual positions Into the extreme random tree for model training, and calculate its WMSE on the validation set as the fitness value:

[0074] (6)

[0075] Where, is the fitness value of the sparrow individual at position .

[0076] According to the behavior mechanism of the sparrow algorithm, individual position updating and population iteration updating are carried out, including S204 to S206:

[0077] S204, the top Individuals in the population are foragers for global search, and the position updating formula is as follows:

[0078] (7)

[0079] Where, And Respectively represent the positions of the sparrow individual At the And the Iteration, is the maximum number of iterations, is a random number, is a warning value, is a warning threshold, is a random variable subject to standard normal distribution, is a full 1 vector;

[0080] S205, for the individuals ranked After that, as followers, local search is carried out around the foragers, and the position updating formula is:

[0081] (8)

[0082] Where The current global fitness worst individual position is represented.

[0083] S205, for the individual after ranking As a follower, a local search is performed around the forager, and the position update formula is:

[0084] S206, if some sparrows think that the current environment is dangerous (such as falling into a local optimum), an alert update mechanism is adopted to quickly jump out of the current search area, and the position update formula is:

[0085] (9)

[0086] S207, repeat S203 to S206 until the maximum number of iterations is reached.

[0087] S208, output the optimal individual position .

[0088] Based on the optimal individual position Model training is performed to obtain the S105 extreme random tree blood glucose prediction model.

[0089] The optimized model is used for regression prediction of non-invasive feature samples, and the output blood glucose value is as follows: the mean absolute relative deviation (MAPE) is 9.3%, the root mean square error (RMSE) is 1.12mmol / L, and the Clark error grid analysis result is as shown in Figure 3 Most of the data points are in area A, indicating that the model prediction result has high accuracy and practicality in the clinically acceptable range.

[0090] The biological impedance-based blood glucose regression modeling method for unbalanced data provided by the present application overcomes the problems in the prior art such as the scarcity of extreme blood glucose value samples, the model training biasing towards mainstream samples, and the significant decline in prediction performance in the boundary area, realizes systematic improvement of the blood glucose prediction model in terms of sample enhancement, loss function design and hyperparameter optimization, and effectively improves the prediction accuracy of the model.

[0091] The above only describes the preferred embodiments of the present application, and it should be noted that for those skilled in the art, without departing from the technical principles of the present application, several improvements and modifications can be made, and these improvements and modifications should also be considered as the protection scope of the present application.

Claims

1. A bioelectrical impedance-based method for blood glucose regression modeling for imbalanced data, characterized in that, Includes the following steps: Step 1: Using a multi-band impedance sensor, the amplitude, phase, and impedance characteristics of the test subject's forearm are acquired, while invasive blood glucose testing is performed simultaneously to construct a bioelectrical signal. Compared with standard blood glucose value Annotated datasets between ; Step 2: Use the SMOGN-Boundary algorithm to dynamically enhance the diversity and coverage of sparse region samples, expand the sample size and balance the data distribution. Step 3: Based on the sample reweighting strategy of square root reciprocal weighting, the sample weights are embedded into the model's loss function; Step 4: Train the extreme random tree model based on the loss function, and introduce the sparrow search algorithm to optimize the model hyperparameters to obtain the optimal hyperparameters of the model; Step 5: Train the model based on the optimal hyperparameters to obtain the final extreme random tree blood glucose prediction model.

2. The bioelectrical impedance-based glucose regression modeling method for imbalanced data according to claim 1, characterized in that, In step 2, the SMOGN-Boundary algorithm includes the following steps: S2.1, Density Judgment Stage: Based on the density distribution characteristics of the target variable (standard blood glucose level), calculate the correlation function. The calculation formula is as follows: (1) in, For target value The sample density, For target variable Maximum density; S2.2, Target-driven sample partitioning, dividing the critical region DR and the regular region DN, with the partitioning threshold tR set to satisfy: (2) S2.3 Boundary determination is based on the mean difference of the target variable. By calculating the mean difference of the target variable between a sparse region sample and its neighbors, if the difference exceeds a preset threshold, it is marked as a boundary sample. The calculation formula is as follows: (3) in For the sample The nearest neighbor set, where τ is the dynamic threshold; S2.4, Dynamic Hybrid Sampling: Conditional oversampling is implemented for DR region and boundary samples, and new samples are generated by linear interpolation. If the sample is far from its neighbors, adaptive Gaussian noise is injected to enhance global diversity. S2.5 merges the newly generated samples with the original data to construct a balanced training set. .

3. The bioelectrical impedance-based glucose regression modeling method for imbalanced data according to claim 1, characterized in that, In step 3, the sample reweighting strategy based on the inverse square root weighting includes the following steps: S3.1, Assume the target variable is blood glucose level. Frequency of occurrence in the dataset Then the weight The calculation formula is: (4) S3.2, embed the sample weights into the model's loss function, and adjust the loss function to the weighted mean squared error (WMSE): (5) in, This indicates the total number of samples in the dataset.

4. The bioelectrical impedance-based glucose regression modeling method for imbalanced data according to claim 1, characterized in that, In step 4, the sparrow search algorithm optimizes the model hyperparameters, including the following steps: S4.1, Define the hyperparameters and range of the extreme random tree blood glucose prediction model, including the number of base learners. Maximum tree depth Minimum number of sample splits Minimum number of leaf node samples The hyperparameters of the extreme random tree blood glucose prediction model are mapped to the locations of individual sparrows. ; S4.2, Define the position of individual sparrows Set the total number of individual sparrows. and maximum number of iterations Initialize the sparrow population ; S4.3, Position the individual sparrows The input is fed into an extreme random tree for training the blood glucose prediction model, and its WMSE is calculated as the fitness value on the validation set: (6) in, It is the location of individual sparrows. fitness value; S4.4, based on the behavior mechanism of the sparrow algorithm, performs individual position updates and population iterative updates, including: Top in the population Individuals, acting as foragers, perform a global search, and the position update formula is as follows: (7) in, and Each represents an individual sparrow. In the and the The position of the next iteration. It is the maximum number of iterations. It is a random number. It is a warning value. It is a warning threshold. It is a random variable that follows a standard normal distribution. It is a vector consisting entirely of 1s; After ranking Individuals acting as followers conduct local searches around the forager, and their position update formula is: (8) in This indicates the position of the individual with the worst global fitness at present; If some sparrows perceive the current environment as dangerous, they will use an alert update mechanism to quickly jump out of the current search area. (9) in This indicates the position of the individual with the best global fitness at present. The disturbance coefficient; S4.5, Reassess WMSE based on the new location; S4.6, repeat S4.3 to S4.5 until the maximum number of generations is reached, and finally output the optimal individual position. This parameter is used to train the final blood glucose prediction model.