Diabetes nursing data collecting and processing system based on machine learning
By adopting machine learning technology in the diabetes care data processing system, combining data preprocessing, stability analysis and circle-like degree analysis, and selecting appropriate noise reduction algorithms, the limitations of traditional noise reduction methods in processing unstable blood glucose data are solved, and more efficient noise reduction effect and data accuracy are achieved.
Patent Information
- Application Number
- CN202411992748.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional denoising methods are difficult to maximize the denoising effect when processing blood sugar concentration data, especially when blood sugar changes are unstable, and a single algorithm is difficult to fully utilize the denoising effect.
Using a machine learning-based diabetes care data acquisition and processing system, through data preprocessing, stability analysis, circle-like degree analysis and noise reduction modules, appropriate noise reduction algorithms (such as normalized minimum mean square error algorithm or adaptive minimum mean square error algorithm) are selected for noise reduction.
It improves the accuracy and stability of blood sugar concentration data, enhances the accuracy of blood sugar monitoring data, and provides more reliable data support for disease prevention and management.
Smart Images

Figure CN119943240A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical data processing, and in particular to a diabetes care data acquisition and processing system based on machine learning. Background Art
[0002] The management of diabetes depends on accurate blood sugar monitoring. With the advancement of technology, smart blood sugar monitoring devices have become increasingly popular and have become an important tool for the management and treatment of diabetic patients. Since the equipment generates a large amount of data during the data collection process, and the data may be affected by equipment failure, measurement errors or the environment, it is very important to remove the noise before analyzing the collected data.
[0003] Diabetes care data is mainly collected through continuous glucose monitoring systems (CGMs), which can continuously record patients' blood sugar levels. However, the collected data is often interfered by noise, such as data anomalies caused by equipment failure, environmental interference, or patient behavior. In order to ensure the accuracy of the data, it is usually necessary to denoise the original data, and the least mean square error (LMS) algorithm is widely used in this scenario. Its improved algorithms such as the normalized least mean square error (NLMS) algorithm and the adaptive least mean square error (ALMS) algorithm can better optimize the characteristics of the data. The NLMS algorithm can quickly adapt to data with large blood sugar fluctuations, while the ALMS algorithm is more suitable for situations where blood sugar changes are relatively stable. When faced with unstable changes in blood sugar concentration due to emotional fluctuations, lack of sleep, or pathological reasons (such as infection, fever), a single algorithm is often difficult to fully exert the denoising effect. The traditional method does not have a suitable screening or processing method, and cannot maximize the denoising effect of the two. Summary of the invention
[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a diabetes care data acquisition and processing system based on machine learning, and the technical solutions adopted are as follows: An embodiment of the present invention provides a diabetes care data collection and processing system based on machine learning, the system comprising: a data preprocessing module, for collecting blood glucose concentration data, and preprocessing the blood glucose concentration data to obtain a blood glucose concentration data sequence; A data screening module, used to calculate the slope of every two adjacent data in the blood glucose concentration data sequence, and screen the data in the blood glucose concentration data sequence according to the slope to obtain stable data and data to be analyzed; The stability analysis module is used to segment the data to be analyzed, and calculate the degree of change of the data segment according to the standard deviation of the data segment in the data to be analyzed and every two adjacent data; classify each data segment according to the degree of change to obtain the fluctuating data segment; A circularity analysis module is used to construct a convex hull based on the data in the fluctuating data segment and obtain the minimum circumscribed circle of the convex hull; and calculate the circularity of the fluctuating data segment based on the area of the convex hull, the edge points and the area of the minimum circumscribed circle; The denoising module is used to select a suitable denoising algorithm for the fluctuating data segment based on the degree of circularity, and to perform denoising on other data segments in the data to be analyzed except the fluctuating data segment.
[0005] Preferably, preprocessing the blood glucose concentration data to obtain a blood glucose concentration data sequence includes: The preprocessing of the collected blood glucose concentration data included filling in missing values, smoothing extreme values, and logarithmic transformation.
[0006] Preferably, screening the data in the blood glucose concentration data sequence according to the slope to obtain stable data and data to be analyzed includes: Search for continuous data segments with equal slopes in the blood glucose concentration data sequence. If the number of data in the data segment is greater than or equal to the preset number, the data segment is stable data; data other than stable data in the blood glucose concentration data sequence is data to be analyzed.
[0007] Preferably, the calculation formula for the degree of change of the data segment is: , Among them, S represents the degree of change of a data segment in the data to be analyzed; α and β represent weight coefficients respectively; represents the standard deviation of the data segment; e represents a natural constant; n represents the number of data in the data segment; and They respectively represent the data corresponding to the n-i+1th moment and the nith moment in the data segment; Norm represents the normalization operation.
[0008] Preferably, each data segment is classified according to the degree of change to obtain the fluctuating data segment, including: If the degree of change of a data segment in the data to be analyzed is greater than a set threshold, the data segment is a fluctuating data segment; if the degree of change of a data segment in the data to be analyzed is less than or equal to the set threshold, the data segment is a non-fluctuating data segment.
[0009] Preferably, the calculation formula for the circularity of the fluctuation data segment is: , Among them, F represents the circularity of the fluctuating data segment; m represents the number of data in the fluctuating data segment; , , and Respectively represent the time corresponding to the i-th data, i+1-th data, m-th data and 1st data in the fluctuation data segment; , , and They represent the i-th data, i+1-th data, m-th data and 1st data in the fluctuation data segment respectively; e represents a natural constant; represents the area of the minimum circumscribed circle of the convex hull corresponding to the fluctuating data segment; n represents the number of edge points in the fluctuating data segment constituting the convex hull; m represents the number of data in the fluctuating data segment; and They respectively represent the c+1th point and the cth point in the edge points of the fluctuating data segment constituting the convex hull; Norm represents the normalization operation; represents the average value of the difference between two adjacent edge points in the fluctuation data segment constituting the convex hull; γ, δ, and ε represent weight coefficients respectively.
[0010] Preferably, selecting a suitable noise reduction algorithm for the fluctuation data segment based on the degree of circularity to perform noise reduction includes: If the circularity of the fluctuating data segment is greater than the screening threshold, the normalized minimum mean square error algorithm is selected for denoising; if the circularity of the fluctuating data segment is less than or equal to the screening threshold, the adaptive minimum mean square error algorithm is selected for denoising.
[0011] Preferably, the noise reduction is performed on the data segments other than the fluctuation data segments in the data to be analyzed, including: The adaptive minimum mean square error algorithm is selected to reduce the noise of other data segments in the analyzed data except the fluctuating data segment.
[0012] The embodiments of the present invention have at least the following beneficial effects: the present invention pre-processes blood glucose concentration data to obtain a blood glucose concentration data sequence, and the pre-processing can facilitate the subsequent data processing; at the same time, the data in the blood glucose concentration data sequence is subjected to stability analysis according to the slope between two adjacent data in the blood glucose concentration data sequence to obtain stable data and data to be analyzed; then the data to be analyzed is segmented, and the degree of change of the data segment is calculated, and a data segment with large volatility is obtained according to the degree of change, which is recorded as a fluctuating data segment, and then a convex hull in the fluctuating data segment is constructed and the minimum circumscribed circle of the convex hull is obtained, and the circularity of the fluctuating data segment is obtained, and the analysis is performed using geometric features to obtain the circularity of the fluctuating data segment, and a suitable denoising algorithm is selected for the fluctuating data segment according to the circularity to perform denoising, and other data segments in the data to be analyzed except the fluctuating data segment are denoised, and this method can flexibly adjust the denoising strategy according to the distribution characteristics of the blood glucose concentration data, select a suitable denoising algorithm, improve the accuracy and stability of data processing, and thus enhance the accuracy of the blood glucose concentration monitoring data, and provide more reliable data support for disease prevention and management. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0014] Figure 1 A system block diagram of a diabetes care data acquisition and processing system based on machine learning provided by an embodiment of the present invention; Figure 2 A schematic diagram of the minimum circumscribed circle of a diabetes care data acquisition and processing system based on machine learning provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0015] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the specific implementation, structure, features and effects of a diabetes care data acquisition and processing system based on machine learning provided by the embodiment of the present invention proposed by the present invention in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics in one or more embodiments may be combined in any suitable form.
[0016] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0017] The following is a detailed description of a specific scheme of a diabetes care data acquisition and processing system based on machine learning provided by the present invention in conjunction with the accompanying drawings. Example
[0018] The main application scenarios of the present invention are: The present invention aims at the problem of the inadequacy of the screening mechanism of the traditional denoising method when the blood glucose concentration data changes unstably. The core idea is to first judge the stability of the data by combining and analyzing the scene characteristics, and then use the convex hull algorithm to analyze the characteristics of the circularity of the data for the unstable data, so as to determine an optimal screening and processing mechanism. Through this method, it is possible to accurately and flexibly select an appropriate denoising algorithm according to the distribution characteristics of the data to improve the accuracy and stability of the blood glucose concentration data.
[0019] See also Figure 1 , which shows a system block diagram of a diabetes care data acquisition and processing system based on machine learning provided by an embodiment of the present invention, the system includes the following modules: The data preprocessing module is used to collect blood glucose concentration data and preprocess the blood glucose concentration data to obtain a blood glucose concentration data sequence.
[0020] With the help of a continuous glucose monitoring system (CGM), the patient's real-time blood glucose concentration data is dynamically collected, and combined with the patient's care scenarios (such as diet, exercise, and medication use), key moments that may cause blood glucose fluctuations (such as after meals, after high-intensity exercise, or after insulin injection) are marked. This scenario-based annotation can provide a more accurate basis for subsequent data processing. The blood glucose concentration data has corresponding labels, as shown in Table 1, the various labels of the blood glucose concentration data at each moment, including diet type, diet amount, exercise intensity, exercise time, drug type, drug dosage, causes of blood glucose changes, and blood glucose measurement equipment.
[0021] Table 1
[0022] Furthermore, the collected blood glucose concentration data need to be preprocessed, mainly including filling in missing values in different situations (such as sudden breakpoints), smoothing extreme values and logarithmic transformation.
[0023] Missing values are repaired using a scenario-based dynamic interpolation method to avoid data distortion caused by simple mean interpolation. Extreme values are retained based on the actual time of data collection. For example, extreme values at some critical moments that are prone to cause sudden changes in blood sugar are retained, and at other times, the data can be smoothed to a normal range using a smoothing algorithm (STL time series decomposition). 3) Logarithmic transformation of the data. Logarithmic transformation can compress extreme values in the data to reduce the impact of local extreme values in the data, ensure that the fluctuation range matches the analysis requirements of the time series model, and retain the key impact of special data in the nursing scenario (such as a surge in blood sugar after high-intensity exercise) on the prediction. In this way, the preprocessed blood glucose concentration data can be obtained, and the preprocessed blood glucose concentration data constitutes a blood glucose concentration data sequence.
[0024] When blood glucose concentration data fluctuates dramatically, NLMS can quickly adapt to the dynamic changes of the input signal, but it may overfit the noise when the changes are small, resulting in unnecessary processing. When blood glucose values fluctuate less at the basal level (such as during sleep), ALMS can process more stably, but it also has some significant defects. The algorithm responds slowly to sudden data changes and cannot effectively process drastically fluctuating data.
[0025] After processing the collected diabetes care data, that is, the blood glucose concentration data, if the data distribution shows a fixed increase or fixed decrease, it is not suitable for the convex hull algorithm, because the convex hull algorithm ultimately needs to construct a polygon. If the data has this distribution law, only one straight line cannot construct a convex polygon. Filter out the data that meets the convex hull construction rules, and then analyze the stability of the data. When the data changes stably, the adaptive minimum mean square error algorithm can be used for denoising. For unstable data, further analyze the circularity of the data. When the circularity of the data is large, the normalized minimum mean square error algorithm is used to better adapt to the large fluctuations in data changes; and when the circularity of the data is small, the adaptive minimum mean square error algorithm is used to fine-tune the denoising effect of the data. For data that cannot construct a convex hull, that is, when the data rises or falls fixedly, it can be considered that the distribution of this part of the data is stable and does not need to be denoised.
[0026] A brief overview of the normalized minimum mean square error algorithm for denoising is as follows (known technology): Before starting processing, initialize the filter weights and learning rate The weights are usually initialized to zero or small random values. The learning rate controls the step size of each weight update.
[0027] For each input data point Normalize and calculate the current predicted value Then calculate the error (ie actual value Difference from predicted value): , Through this error, the LMS algorithm automatically updates the weight vector To reduce errors, gradually optimize the filter and remove noise: , in, is the learning rate, which controls the step size of each update, is the current input data, It's an error.
[0028] Repeat step 2 and iteratively update the weights until the error When the filter converges to the minimum value or reaches the predetermined stop condition, the output of the filter It is getting closer to the original data , the noise is gradually filtered out.
[0029] A brief overview of the adaptive minimum mean square error algorithm for denoising is as follows (known technology): Unlike the normalized minimum mean square error, there is no need to normalize the input data. Instead, the step size needs to be adaptively adjusted based on the error calculated in step 2). When the error is large, the step size is increased to converge quickly, and when the error is small, the step size is reduced to smooth the data.
[0030] The steps to construct the minimum convex polygon (convex hull algorithm) are as follows (known technology): Assume that the preprocessed data set is , where each point Indicates the coordinates corresponding to a blood glucose concentration data.
[0031] Select a starting point from the point set, select the point with the smallest ordinate (if the ordinates are the same, select the point with the smallest abscissa), recorded as .
[0032] by As the base point, calculate each other point and The polar angle formed , the calculation formula of the polar angle is: , Formula logic and explanation: coordinates They are The corresponding coordinates. Sort the points from small to large polar angles to ensure that Start from the beginning and visit all the points in a clockwise direction.
[0033] The data screening module is used to calculate the slope of every two adjacent data in the blood glucose concentration data sequence, and screen the data in the blood glucose concentration data sequence according to the slope to obtain stable data and data to be analyzed.
[0034] For blood sugar concentration data, there may be a fixed rise or fall in blood sugar concentration data. This kind of data distribution cannot construct a convex hull, so it is necessary to analyze the data distribution and remove this part of the data before constructing the convex hull. Then, the data stability of the filtered data is determined. For data with high stability, the adaptive minimum mean square error method can be selected to make denoising more accurate and faster. For unstable data, the convex hull algorithm is used to further determine the data distribution.
[0035] By calculating the rate of change of blood glucose concentration data in the blood glucose concentration data sequence, it is detected whether the data has a fixed upward or downward trend. Specifically, the slope of each two adjacent data in the blood glucose concentration data sequence is calculated by dividing the difference between the latter data and the previous data by the time difference between the two adjacent data. In this way, the slope corresponding to each data can be obtained, and the specific calculation formula is: , in, Indicates the slope corresponding to the ath blood glucose concentration data in the blood glucose concentration data sequence; Indicates the ath blood glucose concentration data in the blood glucose concentration data sequence; represents the a-1th blood glucose concentration data in the blood glucose concentration data sequence; ∆t represents the time interval between the data at two adjacent moments.
[0036] If the blood glucose concentration data shows a fixed upward or downward trend within a period of time (from the starting point with equal change rate to the end point with equal change rate), that is, the slope in this time period is unchanged, and the number of data in this time period is greater than or equal to 3, then this part of the data is considered to lack volatility (a fixed rise or fall can only construct a line segment using the convex hull algorithm and has no geometric features) and cannot be used to construct a convex hull.
[0037] Further, according to the slope corresponding to each blood glucose concentration data, the data in the blood glucose concentration data sequence is screened to obtain stable data and data to be analyzed, and continuous data segments with equal slopes corresponding to the data are found in the blood glucose concentration data sequence. If the number of data in the data segment is greater than or equal to the preset number, the data segment belongs to stable data; the data in the blood glucose concentration data sequence other than the stable data is the data to be analyzed. In the present invention, the preset number is 3, and the implementer can adjust it according to the actual situation.
[0038] The stability analysis module is used to segment the data to be analyzed, and calculate the degree of change of the data segment based on the standard deviation of the data segment in the data to be analyzed and every two adjacent data; and classify each data segment using the degree of change to obtain the fluctuating data segment.
[0039] Blood sugar value is a dynamically changing time series data, and the fluctuation is more obvious in certain periods (such as after meals and after exercise). When constructing the convex hull, using data within a period of time can capture the overall change characteristics of these stages, not just single-point or short-term random fluctuations. Moreover, a period of time covers more data points, which can reflect local fluctuations.
[0040] Furthermore, the data to be analyzed is segmented. In the present invention, the data is segmented in units of a preset time length. The preset time length of the present invention is 1 hour. The implementer can segment according to the needs of the actual situation, and can be segmented evenly or according to certain specific moments that need to be analyzed. In the blood glucose concentration data sequence, since stable data is obtained, the data to be analyzed in the sequence is discontinuous in time series. For one segment of the data to be analyzed, if the time length of the segment of the data to be analyzed is greater than the preset time length, when it is segmented, the time length of the segmented data segments is equal to the preset time length except for the last data segment, which may be less than the preset time length. If the time length of a segment of the data to be analyzed is less than or equal to the preset time length, the segment of the data to be analyzed is a data segment alone, thereby completing the segmentation of the data to be analyzed and obtaining the corresponding data segment.
[0041] Furthermore, the stability of the data segments in the data to be analyzed is analyzed. Specifically, according to the standard deviation of the data segments in the data to be analyzed and the degree of change of the data segments calculated between every two adjacent data, the calculation formula for the degree of change of the data segments in the data to be analyzed is: , Among them, S represents the degree of change of a data segment in the data to be analyzed; α and β represent weight coefficients respectively; represents the standard deviation of the data segment; e represents a natural constant; n represents the number of data in the data segment; and They respectively represent the data corresponding to the n-i+1th moment and the nith moment in the data segment; Norm represents the normalization operation.
[0042] The value of weight coefficients α and β is 0.5, and implementers can adjust them according to actual conditions. It can reflect the volatility of the data, and combined with the average difference of the subsequent adjacent data, it can reflect the stability of the data in the data segment.
[0043] Furthermore, each data segment is classified according to the degree of change to obtain a fluctuating data segment. Specifically, if the degree of change of a data segment in the data to be analyzed is greater than a set threshold, it can be considered that the data in the data segment is not stable enough and the degree of change is high, and it is necessary to judge the degree of circularity. This data segment is a fluctuating data segment. Therefore, each data segment in the data to be analyzed can be classified into fluctuating data segments and non-fluctuating data segments.
[0044] The circularity analysis module is used to construct a convex hull based on the data in the fluctuating data segment and obtain the minimum circumscribed circle of the convex hull; and calculate the circularity of the fluctuating data segment based on the area of the convex hull, the edge points and the area of the minimum circumscribed circle.
[0045] The circularity of the data is intuitively evaluated through geometric methods, and the data in the fluctuating data segment is mapped to a two-dimensional coordinate system, with time as the horizontal coordinate in minutes and blood glucose concentration as the vertical coordinate in mg / dl. By obtaining the minimum convex hull of the data in the fluctuating data segment and the minimum circumscribed circle of the convex hull, the convex hull can obtain the boundary and geometric characteristics of the data change, and the constructed circle can quantify the circularity of the data points in the convex hull. When the circularity of the data is large, the area of the convex hull will be close to the circle formed, and the original data points used to form the convex hull will be fewer. Vice versa. Through this geometric analysis, the system can judge the amplitude of blood glucose fluctuations according to the degree of change of the data and select appropriate processing methods for subsequent denoising.
[0046] The convex hull algorithm is used to construct the convex hull of the data in the fluctuating data segment, and the minimum circumscribed circle of the convex hull is obtained. After the convex hull and the minimum circumscribed circle are constructed, the circularity of the data is analyzed according to the calculated convex hull geometric features. Assume that the data in the fluctuating data segment is composed of time and blood glucose concentration data values Constructing point set .
[0047] The circularity of the fluctuating data segment is calculated based on the area of the convex hull, the edge points, and the area of the minimum circumscribed circle. The calculation formula for the circularity of the fluctuating data segment is as follows: , Among them, F represents the circularity of the fluctuating data segment; m represents the number of data in the fluctuating data segment; , , and Respectively represent the time corresponding to the i-th data, i+1-th data, m-th data and 1st data in the fluctuation data segment; , , and They represent the i-th data, i+1-th data, m-th data and 1st data in the fluctuation data segment respectively; e represents a natural constant; represents the area of the minimum circumscribed circle of the convex hull corresponding to the fluctuating data segment; n represents the number of edge points in the fluctuating data segment constituting the convex hull; m represents the number of data in the fluctuating data segment; and They respectively represent the c+1th point and the cth point in the edge points of the fluctuating data segment constituting the convex hull; Norm represents the normalization operation; represents the average value of the difference between two adjacent edge points in the fluctuation data segment constituting the convex hull; γ, δ, and ε represent weight coefficients respectively.
[0048] What is calculated is the area of the convex hull of the fluctuating data segment, the ratio of the convex hull area to the area of the minimum circumscribed circle. The closer this value is to 1, the greater the circularity of the data. It represents the ratio of the number of points (edge points) used to construct the convex hull to the number of points in the fluctuation data segment. It means that the ratio is inversely normalized. For the fluctuating data segment with a large degree of circularity, the point set used to construct the convex hull is small. Therefore, it is necessary to perform inverse normalization so that the value is closer to 1, which can better represent the data distribution with a large degree of circularity. It indicates the degree of change in the distance between two adjacent edge points in the vertical direction. The closer the value is to 1, the greater the circularity of the fluctuation data segment. The edge point mentioned here is the data point that is the edge point when the convex hull is constructed in the fluctuation data segment.
[0049] According to the fused feature formula, the geometric features of the convex hull can fully represent the circularity of the fluctuating data segment, providing an effective basis for the subsequent selection of denoising methods.
[0050] The denoising module is used to select a suitable denoising algorithm for the fluctuating data segment based on the degree of circularity, and to perform denoising on other data segments in the data to be analyzed except the fluctuating data segment.
[0051] A screening threshold is set to divide the circularity of each fluctuation data segment into two parts: a large circularity and a small circularity. If the circularity of the fluctuation data segment is greater than the screening threshold, it is considered that the circularity of the blood glucose concentration data in this time period is large. If the circularity of the fluctuation data segment is less than or equal to the screening threshold, it is considered that the circularity of the blood glucose concentration data in this time period is small. The screening threshold is taken as 0.5, and the implementer can adjust it according to actual conditions.
[0052] After calculating the circularity of the fluctuating data segment through the convex hull algorithm, the fluctuating data segment is divided into two parts: a part with a large circularity and a part with a small circularity. When the circularity of the fluctuating data segment is large, it means that the blood sugar value may be significantly affected by external factors (such as diet, exercise or measurement error). At this time, the normalized minimum mean square error algorithm needs to be used to respond to the drastic changes in the data more quickly. For data with a small circularity of the fluctuating data segment, the adaptive minimum mean square error can be selected to further smooth the data through fine adjustment to avoid excessive denoising from interfering with trend analysis.
[0053] The specific process of the present invention for adaptively selecting different denoising algorithms according to the circularity of the fluctuation data segment is as follows: When the circularity of the fluctuating data segment is large, that is, the circularity of the fluctuating data segment is greater than 0.5, the normalized minimum mean square error algorithm is selected to reduce the drastic changes in the data, so as to achieve a better denoising effect. When the circularity of the fluctuating data segment is small, that is, the circularity of the fluctuating data segment is less than or equal to 0.5, the adaptive minimum mean square error algorithm is selected to adjust the step size factor more finely to improve the accuracy of denoising. For the data in other data segments (non-fluctuating data segments) except the fluctuating data segments in the data to be analyzed, the adaptive minimum mean square error algorithm is also selected for denoising. In summary, in the blood glucose concentration data sequence, for the fluctuating data segments with a large circularity, the normalized minimum mean square error algorithm is selected for denoising; for the fluctuating data segments with a small circularity and the non-fluctuating data segments in the data to be analyzed, the adaptive minimum mean square error algorithm is selected for denoising, and for stable data, no denoising is required.
[0054] The collection and denoising of diabetes care data is the basis of management and analysis. The denoised data can more accurately reflect the true trend of blood sugar changes and provide more reliable support for diabetes management. During the data collection process, by combining the patient's diet, exercise and drug use information, the key time points of blood sugar fluctuations are recorded to lay the foundation for denoising and subsequent analysis. By processing and analyzing the collected data, the fluctuation pattern of blood sugar data can be deeply revealed, and potential health risks can be discovered using time series prediction models. In addition, through the processing of the present invention, the utilization depth of diabetes care data collection can be comprehensively improved, facilitating more refined and scientific disease management.
[0055] In view of the limitations of traditional denoising methods when dealing with unstable data, an optimization scheme combining geometric features to analyze the circularity of data is proposed. Different denoising algorithms are selected by analyzing the circularity of blood glucose concentration data. The circularity of blood glucose concentration data is analyzed by using the convex hull algorithm to determine the change trend and distribution characteristics of the data. Based on this, the most suitable denoising method is selected, including the normalized minimum mean square error algorithm or the adaptive minimum mean square error algorithm. This method can flexibly adjust the denoising strategy according to the distribution characteristics of blood glucose concentration data, improve the accuracy and stability of data processing, and thus enhance the accuracy of blood glucose monitoring data, providing more reliable data support for disease prevention and management.
[0056] It should be noted that the sequence of the above embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The above is a description of a specific embodiment of this specification. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0057] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.
[0058] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A diabetes care data acquisition and processing system based on machine learning, characterized in that: The system includes: A data preprocessing module is used to collect blood glucose concentration data and preprocess the blood glucose concentration data to obtain a blood glucose concentration data sequence; A data screening module, used to calculate the slope of every two adjacent data in the blood glucose concentration data sequence, and screen the data in the blood glucose concentration data sequence according to the slope to obtain stable data and data to be analyzed; The stability analysis module is used to segment the data to be analyzed, and calculate the degree of change of the data segment according to the standard deviation of the data segment in the data to be analyzed and every two adjacent data; classify each data segment according to the degree of change to obtain the fluctuating data segment; A circularity analysis module is used to construct a convex hull based on the data in the fluctuating data segment and obtain the minimum circumscribed circle of the convex hull; and calculate the circularity of the fluctuating data segment based on the area of the convex hull, the edge points and the area of the minimum circumscribed circle; The denoising module is used to select a suitable denoising algorithm for the fluctuating data segment based on the degree of circularity, and to perform denoising on other data segments in the data to be analyzed except the fluctuating data segment.
2. A diabetes care data acquisition and processing system based on machine learning according to claim 1, characterized in that: The preprocessing of the blood glucose concentration data to obtain a blood glucose concentration data sequence includes: The preprocessing of the collected blood glucose concentration data included filling in missing values, smoothing extreme values, and logarithmic transformation.
3. A diabetes care data acquisition and processing system based on machine learning according to claim 1, characterized in that: The step of screening the data in the blood glucose concentration data sequence according to the slope to obtain stable data and data to be analyzed includes: Search for continuous data segments with equal slopes in the blood glucose concentration data sequence. If the number of data in the data segment is greater than or equal to the preset number, the data segment is stable data; data other than stable data in the blood glucose concentration data sequence is data to be analyzed.
4. A diabetes care data acquisition and processing system based on machine learning according to claim 1, characterized in that: The calculation formula for the degree of change of the data segment is: , Among them, S represents the degree of change of a data segment in the data to be analyzed; α and β represent weight coefficients respectively; represents the standard deviation of the data segment; e represents a natural constant; n represents the number of data in the data segment; and They respectively represent the data corresponding to the n-i+1th moment and the nith moment in the data segment; Norm represents the normalization operation.
5. The diabetes care data acquisition and processing system based on machine learning according to claim 1, characterized in that: The method of classifying each data segment by using the degree of change to obtain the fluctuating data segment includes: If the degree of change of a data segment in the data to be analyzed is greater than a set threshold, the data segment is a fluctuating data segment; if the degree of change of a data segment in the data to be analyzed is less than or equal to the set threshold, the data segment is a non-fluctuating data segment.
6. A diabetes care data acquisition and processing system based on machine learning according to claim 1, characterized in that: The calculation formula of the circularity of the fluctuation data segment is: , Among them, F represents the circularity of the fluctuating data segment; m represents the number of data in the fluctuating data segment; , , and Respectively represent the time corresponding to the i-th data, i+1-th data, m-th data and 1st data in the fluctuation data segment; , , and They represent the i-th data, i+1-th data, m-th data and 1st data in the fluctuation data segment respectively; e represents a natural constant; represents the area of the minimum circumscribed circle of the convex hull corresponding to the fluctuating data segment; n represents the number of edge points in the fluctuating data segment constituting the convex hull; m represents the number of data in the fluctuating data segment; and They respectively represent the c+1th point and the cth point in the edge points of the fluctuating data segment constituting the convex hull; Norm represents the normalization operation; represents the average value of the difference between two adjacent edge points in the fluctuation data segment constituting the convex hull; γ, δ, and ε represent weight coefficients respectively.
7. The diabetes care data acquisition and processing system based on machine learning according to claim 1, characterized in that: The selecting a suitable noise reduction algorithm for the fluctuation data segment based on the degree of circularity to perform noise reduction includes: If the circularity of the fluctuating data segment is greater than the screening threshold, the normalized minimum mean square error algorithm is selected for denoising; if the circularity of the fluctuating data segment is less than or equal to the screening threshold, the adaptive minimum mean square error algorithm is selected for denoising.
8. The diabetes care data acquisition and processing system based on machine learning according to claim 1, characterized in that: The denoising of other data segments in the data to be analyzed except the fluctuating data segment comprises: The adaptive minimum mean square error algorithm is selected to reduce the noise of other data segments in the analyzed data except the fluctuating data segment.
Citation Information
Cited By
Patient diet digital visual management method and system based on multi-source data
CN120148766A