A CKD full-process information management system based on active learning
Through the CKD full-process information management system with active learning, the problem of information collection accuracy of CKD patients is solved, and accurate monitoring and credible judgment of patients' diet, blood pressure and blood sugar information is achieved, which improves the reliability of information monitoring.
Patent Information
- Application Number
- CN202510309642.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-17
AI Technical Summary
In the prior art, the accuracy of information collection of CKD patients, especially diet, blood pressure and blood sugar information, is difficult to guarantee, resulting in poor monitoring effects.
A CKD full-process information management system based on active learning is adopted to build an abnormality verification model through data collection, diet clustering, information analysis and abnormality verification model to achieve accurate monitoring and credible judgment of patient diet, blood pressure and blood sugar information.
Accurate monitoring of CKD patients' information can be achieved, typical dietary patterns can be effectively extracted, health trends can be predicted, and abnormal verification models can be constructed to judge the credibility of information, thereby greatly improving the reliability of information monitoring.
Smart Images

Figure CN119833077B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, it relates to a CKD full-process information management system based on active learning. Background Art
[0002] CKD (Chronic Kidney Disease) is a common and difficult-to-cure chronic disease in modern times, and potential patients need to be monitored for a long time. Patients with diabetes and hypertension are usually the main potential populations. The monitoring methods include hospital physical examinations and information collection. However, hospital physical examinations are time-consuming and costly, so information collection monitoring has gradually become the mainstream. Information collection mainly covers blood pressure, blood sugar, and diet information. However, for hospitals, only reliable data can be used as an effective monitoring basis. Therefore, the accuracy of information collection is particularly important. However, some patients may provide inaccurate diet data due to subjective or objective reasons, resulting in the problem of invalid collected information. Summary of the Invention
[0003] The present invention provides a CKD full-process information management system based on active learning to solve the technical problems proposed in the background art.
[0004] The present invention provides a CKD full-process information management system based on active learning, including:
[0005] A data collection module, configured to collect the diet information, blood pressure information, and blood sugar information of a number of patients within a preset confidence time period;
[0006] Among them, the diet information includes: the timestamp of the ingested food, the ingredients of the ingested food, the cooking method of the ingredients, and the intake of the ingredients; the blood pressure information and blood sugar information respectively include: the blood pressure values and blood sugar values at the 0th, 1st, and 2nd moments after the timestamp of the ingested food;
[0007] A diet clustering module, configured to cluster the diet information of patients to obtain the typical diet of each patient within a preset confidence time period;
[0008] An information analysis module, configured to obtain the blood pressure information and blood sugar information corresponding to the typical diet of patients, and perform trend analysis to obtain the characteristic trends of patients;
[0009] A model construction module, configured to construct an anomaly verification model based on the typical diet information and characteristic trends; the anomaly verification model is used to perform anomaly verification on the diet information, blood pressure information, and blood sugar information of a target patient within a second preset time period to obtain a confidence judgment of the target patient.
[0010] Further, clustering the diet information of patients includes:
[0011] Step 21, establish the initial encoding of dietary information, including:
[0012] Use the timestamp of the ingested food in the dietary information as the first element of the initial encoding;
[0013] Query the ingredients of the ingested food in the dietary information to obtain characteristic parameters; the characteristic parameters include the carbohydrate value, protein value, fat value, and sodium value per unit intake of the ingredients of the ingested food, as well as the glycemic index value and glycemic load value;
[0014] Use each parameter in the characteristic parameters as the second to seventh elements of the initial encoding;
[0015] Use the cooking method of the ingredients in the dietary information encoded by real numbers as the eighth element of the initial encoding;
[0016] Use the intake of the ingredients in the dietary information as the ninth element of the initial encoding;
[0017] Step 22, normalize the initial encoding to obtain the characteristic encoding;
[0018] Step 23, initialize N coding cluster centers, and calculate the similarity between each characteristic encoding and each coding cluster center. The similarity calculation formula is as follows:
[0019] ;
[0020] Among them, represents the similarity between the th characteristic encoding and the th coding cluster center, , represents index, represents the th element of the th characteristic encoding, represents the th element of the th coding cluster center, represents the exponential weight, represents the absolute value of , represents the Manhattan distance between the
[0021] Step 24, if the similarity between the th characteristic encoding and the th coding cluster center is the smallest, then assign the
[0022] Step 25: Based on the N clusters obtained in Step 24, encode the average value of the feature encodings in each cluster as the updated encoding cluster center corresponding to the cluster.
[0023] Step 26: Repeat Step 25 for a preset number of times to obtain N final clusters.
[0024] Step 27: Set a cluster merging threshold. If the similarity between the updated encoding cluster centers of any two final clusters is less than the preset threshold, then merge the corresponding final clusters to obtain merged clusters, and calculate the updated encoding clusters of the merged clusters.
[0025] Step 28: Repeat Step 27 until the similarity between the updated encoding clusters of any two remaining final clusters or merged clusters is greater than the preset threshold. Then, take the final cluster or merged cluster with the largest number of feature encodings as the typical cluster, and the dietary information corresponding to the feature encodings in the typical cluster as the typical diet.
[0026] Furthermore, perform trend analysis, including:
[0027] Construct a data set with the blood pressure information and blood glucose information corresponding to each typical diet.
[0028] Sort the data set in chronological order to obtain a feature sorting.
[0029] Calculate the feature trend of each typical diet based on the feature sorting. The calculation formula for the feature trend is as follows:
[0030] ;
[0031] ;
[0032] , ;
[0033] , ;
[0034] Where, represents the feature trend of the typical diet, represents the trend of blood pressure, represents the trend of blood glucose, represents the weight of the trend of blood pressure, represents the weight of the trend of blood glucose, represents the interaction factor between the trend of blood pressure and the trend of blood glucose, represents the i-th data set, represents the standard blood pressure value of the i-th data set, represents the standard blood glucose value of the i-th data set, Denotes the blood pressure value at time 0 of the i-th data set, Denotes the blood pressure value at time 1 of the i-th data set, Denotes the blood pressure value at time 2 of the i-th data set, Denotes the blood glucose value at time 0 of the i-th data set, Denotes the blood glucose value at time 1 of the i-th data set, Denotes the blood glucose value at time 2 of the i-th data set, Denotes Activation function.
[0035] Furthermore, construct an anomaly verification model, including;
[0036] Construct the training samples of the anomaly verification model, and construct the sample labels of the anomaly verification model;
[0037] The training samples include: Obtain the physical examination data of each patient at the start time and end time of the preset confidence period, analyze the physical examination data based on experts to obtain the health scores of the physical examination data of each patient at the start time and end time of the preset confidence period, and calculate the characteristic differences of the health scores of the physical examination data of each patient at the start time and end time; Based on experts' confidence division of the characteristic differences, the confidence division includes: excellent, good, poor, and very poor;
[0038] Match excellent with the Interval of the characteristic trend value, match good with the Interval of the characteristic trend value, match poor with the Interval of the characteristic trend value, match very poor with the Interval;
[0039] If the characteristic trend of the patient in the preset confidence period matches the confidence division, then judge the diet information, blood pressure information, and blood glucose information of the corresponding patient as confidence data, and use the characteristic encoding of the diet information of the corresponding patient as the training sample, and use the characteristic trend of the corresponding patient as the sample label;
[0040] If the characteristic trend of the patient in the preset confidence period does not match the confidence division, then judge the diet information, blood pressure information, and blood glucose information of the corresponding patient as non-confidence data, and discard the diet information and characteristic trend of the corresponding patient;
[0041] Based on the training samples and sample labels of the patient in the preset confidence period, train to obtain an anomaly verification model, and the anomaly verification model includes: an input layer, a hidden layer, and a classifier.
[0042] Furthermore, the input layer, hidden layer, and classifier include:
[0043] An input layer, which is used to stack a number of feature encodings of a target patient in a second preset time period in chronological order into a feature matrix, and input the feature matrix into a hidden layer;
[0044] A hidden layer, which is used to perform a non-linear mapping on the feature matrix to obtain the hidden state of the feature matrix;
[0045] A classifier, which is used to input the hidden state, and the classification space of the classifier represents the predicted feature trend of the target patient in the second preset time period.
[0046] Furthermore, the calculation formula of the hidden layer includes:
[0047] ;
[0048] ;
[0049] ;
[0050] ;
[0051] Among them, represents the output of the reset gate of the th hidden layer, represents the output of the update gate of the th hidden layer, represents the candidate hidden state of the th hidden layer, represents the hidden state of the th hidden layer, represents activation function, represents activation function, , and represent the first weight matrix, the second weight matrix and the bias matrix of the reset gate of the th hidden layer, , and represent the first weight matrix, the second weight matrix and the bias matrix of the update gate of the th hidden layer, , and represent the first weight matrix, the second weight matrix and the bias matrix of the th candidate hidden state, represents the feature encoding of the th row of the feature matrix input to the th hidden layer, represents the hidden state output by the th hidden layer, Denotes element-wise multiplication.
[0052] Furthermore, the anomaly verification model updates the weight matrix and bias matrix of the hidden layer through backpropagation based on the mean squared error loss function. The calculation formula of the mean squared error loss function is as follows:
[0053] ;
[0054] Where, Denotes the mean squared error loss value of the th patient, Denotes the value of the characteristic trend corresponding to the sample label of the th patient, Denotes the predicted characteristic trend value of the th patient.
[0055] Furthermore, anomaly verification includes:
[0056] Obtaining the characteristic trend of the target patient within the second preset time period, and the predicted characteristic trend of the target patient within the second preset time period by the anomaly verification model;
[0057] If the difference between the characteristic trend and the predicted characteristic trend is less than the preset confidence threshold, it indicates that the diet information, blood pressure information, and blood glucose information of the target patient within the second preset time period are credible;
[0058] If the difference between the characteristic trend and the predicted characteristic trend is greater than the preset confidence threshold, it indicates that the diet information, blood pressure information, and blood glucose information of the target patient within the second preset time period are not credible.
[0059] The beneficial effects of the present invention are as follows: Through the active learning mechanism and multi-dimensional data fusion, the system realizes the precise monitoring and analysis of information such as the diet, blood pressure, and blood glucose of CKD patients, can effectively extract typical diet patterns, predict health trends, and construct an anomaly verification model to judge the credibility of the information of the target patient, thereby greatly improving the reliability of information monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 Is a module diagram of a CKD full-process information management system based on active learning of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0061] Reference will now be made to example embodiments to discuss the subject matter described herein. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Changes can be made to the functions and arrangements of the elements discussed without departing from the scope of protection of the content of this specification. Each example can omit, substitute, or add various processes or components as needed. In addition, the features described relative to some examples can also be combined in other examples.
[0062] As Figure 1 shown, an active learning-based CKD full-process information management system includes:
[0063] A data acquisition module, configured to collect dietary information, blood pressure information, and blood glucose information of a plurality of patients within a preset confidence time period;
[0064] Among them, the dietary information includes: the timestamp of the ingested food, the ingredients of the ingested food, the cooking method of the ingredients, and the intake of the ingredients; the blood pressure information and the blood glucose information respectively include: the blood pressure values and blood glucose values at the 0th, 1st, and 2nd moments after the timestamp of the ingested food;
[0065] A dietary clustering module, configured to cluster the dietary information of the patients to obtain the typical diet of each patient within the preset confidence time period;
[0066] An information analysis module, configured to obtain the blood pressure information and blood glucose information corresponding to the typical diet of the patients, and perform trend analysis to obtain the characteristic trends of the patients;
[0067] A model construction module, configured to construct an anomaly verification model based on the typical dietary information and the characteristic trends; the anomaly verification model is used to perform anomaly verification on the dietary information, blood pressure information, and blood glucose information of the target patient within the second preset time period to obtain the confidence judgment of the target patient.
[0068] In an embodiment of the present invention, clustering the dietary information of the patients includes:
[0069] Step 21, establishing an initial encoding of the dietary information, including:
[0070] Taking the timestamp of the ingested food in the dietary information as the first element of the initial encoding;
[0071] Querying the ingredients of the ingested food in the dietary information to obtain characteristic parameters; the characteristic parameters include the carbohydrate value, protein value, fat value, and sodium value per unit intake of the ingredients of the ingested food, as well as the glycemic index value and the glycemic load value;
[0072] Taking each parameter in the characteristic parameters as the second to seventh elements of the initial encoding;
[0073] The cooking method of the ingredients in the diet information is used as the eighth element of the initial coding through real number coding;
[0074] The intake of the ingredients in the diet information is used as the ninth element of the initial coding;
[0075] Step 22: Normalize the initial coding to obtain the feature coding;
[0076] Step 23: Initialize N coding cluster centers, and calculate the similarity between each feature coding and each coding cluster center. The similarity calculation formula is as follows:
[0077] ;
[0078] Among them, represents the similarity between the th feature coding and the th coding cluster center, , represents the index of , represents the a-th element of the th feature coding, represents the a-th element of the th coding cluster center, represents the exponential weight, represents the absolute value of, represents the linear weight, represents the th feature coding and the th coding cluster center's Manhattan distance;
[0079] Step 24: If the similarity between the th feature coding and the th coding cluster center is the smallest, then assign the i-th feature coding to the j-th coding cluster center;
[0080] Step 25: Based on the results of Step 24, obtain N clusters, and use the average coding of the feature codings in each cluster as the updated coding cluster center corresponding to the cluster;
[0081] Step 26: Repeat Step 25 for a preset number of times to obtain N final clusters;
[0082] Step 27: Set a cluster merging threshold. If the similarity between the updated coding cluster centers of any two final clusters is less than the preset threshold, then merge the corresponding final clusters to obtain merged clusters, and calculate the updated coding clusters of the merged clusters;
[0083] Step 28: Repeat Step 27 until the similarity of the updated encoded clusters between any two remaining final clusters or merged clusters is greater than the preset threshold. Then, take the final cluster or merged cluster with the largest number of feature encodings as the typical cluster, and the dietary information corresponding to the feature encodings in the typical cluster as the typical diet.
[0084] It should be noted that the ingredients can be queried through the Mint Health APP. For example, through the Mint Health APP, query the content of various nutrient elements, glycemic index, and glycemic load value of bread per unit weight (100g per serving).
[0085] Specifically, this solution conducts clustering analysis by initializing as many coding clustering centers as possible. Each coding clustering center represents an artificially fabricated feature code. The construction of the feature code is based on statistical analysis of each element of the feature codes obtained during the historical time period to determine the range of possible values for each element. Subsequently, according to the range of possible values, a number of different feature codes are generated using a randomized or systematic method, and these feature codes are initialized as clustering centers. The initialization process of each clustering center fully considers the distribution characteristics of the feature space to ensure that all potential dietary information patterns can be reasonably covered and represented. After initialization, the similarity between the dietary information feature codes of the patient and these clustering centers is calculated, and each feature code is classified into the most similar clustering center according to the similarity. The similarity calculation formula used here comprehensively considers factors such as element differences, Manhattan distance, and weight factors between feature codes, and can accurately measure the difference between the feature code and the clustering center. The K-means clustering algorithm is used in the clustering process, which specifically includes the following steps: calculating the similarity between the feature code and the clustering center, assigning the feature code to the nearest clustering center, updating the clustering center to the average code of the cluster it belongs to, and continuously iterating until the preset convergence condition is met or the specified number of iterations is reached. After clustering, in order to further optimize the clustering results, subsequent merging operations are performed on the clusters by setting a merging threshold. Specifically, if the similarity between any two clustering centers is lower than the preset merging threshold, they are merged into a new cluster, and this process is repeated until the similarity between all clustering centers is higher than the merging threshold. Finally, the cluster with the largest number of feature codes is selected from all clusters as the typical cluster, and the dietary information corresponding to all feature codes in this typical cluster is extracted as the typical diet of the patient. The typical diet obtained in this way can more realistically reflect the eating habits of the patient during the preset time period because this typical diet is a representative pattern extracted from a large amount of dietary information and can comprehensively reflect the patient's dietary preferences and habit characteristics. Based on the typical diet, the impact on characteristic parameters such as the patient's blood pressure and blood sugar can be further analyzed, so as to infer and judge the characteristic trend of the patient and provide a basis for subsequent health management and intervention measures. This method not only has high accuracy and representativeness, but also can reduce the interference of noise data, laying a foundation for realizing scientific and precise patient monitoring.
[0086] In one embodiment of the present invention, trend analysis is performed, including:
[0087] Construct a data set with the blood pressure information and blood sugar information corresponding to each typical diet;
[0088] Sort the data set in chronological order to obtain a feature sorting;
[0089] Calculate the characteristic trend of each typical diet based on feature ranking. The calculation formula for the characteristic trend is as follows:
[0090] ;
[0091] ;
[0092] , ;
[0093] , ;
[0094] where, represents the characteristic trend of the typical diet, represents the trend of blood pressure, represents the trend of blood glucose, represents the weight of the trend of blood pressure, represents the weight of the trend of blood glucose, represents the interaction factor between the trend of blood pressure and the trend of blood glucose, represents the i-th data set, represents the standard blood pressure value of the i-th data set, represents the standard blood glucose value of the i-th data set, represents the blood pressure value at time 0 of the i-th data set, represents the blood pressure value at time 1 of the i-th data set, represents the blood pressure value at time 2 of the i-th data set, represents the blood glucose value at time 0 of the i-th data set, represents the blood glucose value at time 1 of the i-th data set, represents the blood glucose value at time 2 of the i-th data set, represents the activation function.
[0095] Specifically, by analyzing the blood pressure information and blood glucose information based on the typical diet, the health characteristic trend of the patient within a specific time period is inferred. This process combines the nutritional components of the ingredients in the patient's diet, conducts systematic fusion analysis and calculation, aims to comprehensively reveal the long-term impact of eating habits on the patient's physical health, and quantifies it as a characteristic trend value. The characteristic trend is used to reflect the physical health trend of the patient within the preset confidence time period, and can effectively describe the dynamic impact of eating habits on key health indicators such as the patient's blood pressure and blood glucose. Through scientific calculation, the system combines the diet characteristics with the blood pressure and blood glucose characteristics to form a comprehensive trend model. For example, for patients with a preference for foods rich in high fat and high carbohydrates in their diet (such as fried chicken, beer, etc.), after the system analyzes their blood pressure and blood glucose information, it may be concluded that their characteristic trend value falls within within the range, indicating that such a diet structure may have an adverse impact on the patient's health status. For patients with a scientific and nutritionally balanced diet every day (such as a high-fiber, low-fat diet with scientific proportioning), the healthy change trends of their blood pressure and blood sugar may be more positive, and the system calculates its characteristic trend value as , thus reflecting the positive effect of good eating habits on health. According to the analysis results of typical diet patterns, the change trends of the patient's future health characteristics are inferred. Therefore, if the predicted trend does not match the actual trend, it indicates that there is untrue content in the patient's diet information and is determined to be abnormal. The calculation of characteristic trends combines the blood pressure and blood sugar information of typical diets and scientific nutritional analysis, converting complex health data into intuitive trend values, which can truly reflect the relationship between the patient's health status and diet.
[0096] In an embodiment of the present invention, an anomaly verification model is constructed, including;
[0097] Constructing training samples of the anomaly verification model and sample labels of the anomaly verification model;
[0098] The training samples include: obtaining the physical examination data of each patient at the start and end times of a preset confidence period, obtaining the health scores of the physical examination data of each patient at the start and end times of the preset confidence period based on expert analysis of the physical examination data, and calculating the characteristic difference of the health scores of the physical examination data of each patient at the start and end times; based on expert confidence division of the characteristic differences, the confidence division includes: excellent, good, poor, and very poor;
[0099] Matching excellent to the interval of the characteristic trend value, matching good to the interval of the characteristic trend value, matching poor to the interval of the characteristic trend value, matching very poor to the interval;
[0100] If the characteristic trend of the patient within the preset confidence period matches the confidence division, it is determined that the diet information, blood pressure information, and blood sugar information of the corresponding patient are confidence data, and the characteristic code of the diet information of the corresponding patient is used as the training sample, and the characteristic trend of the corresponding patient is used as the sample label;
[0101] If the characteristic trend of the patient within the preset confidence period does not match the confidence division, it is determined that the diet information, blood pressure information, and blood sugar information of the corresponding patient are non-confidence data, and the diet information and characteristic trend of the corresponding patient are discarded;
[0102] Based on the training samples and sample labels of the patient in the preset confidence time period, an anomaly verification model is trained. The anomaly verification model includes: an input layer, a hidden layer, and a classifier.
[0103] Specifically, in order to train an effective anomaly verification model, the real data of patients within the confidence time period needs to be used. However, due to subjective or objective reasons, the data provided by patients may be untrue. For example, the omission of diet information, incorrect recording, or intentional concealment of certain details. If this untrue data is used to train the anomaly verification model, it may cause the model to deviate and reduce its accuracy. Therefore, the present invention proposes a method to distinguish real data from untrue data through the matching relationship between the health changes and feature trends of patients, so as to ensure that the data used for model training has high credibility. First, within the confidence time period, the data collection of patients not only includes diet information, blood pressure information, and blood glucose information, but also a comprehensive physical health examination of patients is carried out at the start time and end time of the confidence time period. The physical examination data includes the changes in key health indicators such as the patient's blood pressure and blood glucose. These changes are objective and accurate and can be used as a reference for the patient's health status. Subsequently, by introducing a grading system annotated by experts, the health changes of patients are graded. The grading system divides patients into several levels according to the degree of health changes within the confidence time period, such as "excellent", "good", "poor", and "very poor". These grades can quantitatively describe the quality of the patient's health changes and provide a standard for subsequent feature trend matching. Based on the grading of the patient's health changes, the system matches the feature trend values extracted within the confidence time period with the grading results. The feature trend value is a quantitative index used to reflect the comprehensive impact of the patient's diet information on health indicators such as blood pressure and blood glucose within the confidence time period. Assuming that the feature trend value falls within a reasonable range (for example, consistent with the health change grading result), it can be judged that the calculation results of these feature trends are correct, and the corresponding diet information, blood pressure information, and blood glucose information are recognized as real data. On the contrary, if the feature trend value does not match the grading result of the health change, it indicates that there may be an error in this feature trend, and further speculation is that the patient's diet information, blood pressure information, and blood glucose information may be untrue. In this case, the present invention marks these untrue data as invalid data and discards them to prevent these data from interfering with the training of the anomaly verification model. Through the above process, a high-quality and credible data set can be obtained for the training of the anomaly verification model. Based on these real data, the anomaly verification model can learn accurate feature trend patterns, helping the system to effectively judge whether new data is abnormal in subsequent operations. For example, when the model is actually applied, the feature trend value of the target patient is input and compared with the anomaly verification model. If the feature trend value meets the model prediction range, the data is considered credible; otherwise, it is considered uncredible. In summary, the present invention screens out the real data of patients within the confidence time period through the matching relationship between health change grading and feature trend values, and uses these data to train the anomaly verification model.This method not only ensures the authenticity and high quality of the model training data, but also effectively improves the accuracy and robustness of the anomaly verification model, thus providing scientific support for the subsequent detection and processing of abnormal data in the system.
[0104] In an embodiment of the present invention, the input layer, the hidden layer, and the classifier include:
[0105] The input layer is used to stack a number of feature encodings of the target patient in the second preset time period in chronological order into a feature matrix, and input the feature matrix into the hidden layer;
[0106] The hidden layer is used to perform a non-linear mapping on the feature matrix to obtain the hidden state of the feature matrix;
[0107] The classifier is used to input the hidden state, and the classification space of the classifier represents the predicted feature trend of the target patient in the second preset time period.
[0108] Specifically, the internal structure of the anomaly verification model is based on the design of a neural network, which can extract deep features from the input data and efficiently judge the credibility of the data. The main function of the input layer is to receive the feature data extracted from the patient within the confidence time period. The feature data includes: the feature encoding is transformed from the patient's diet information (intake, nutritional components of ingredients, cooking methods, etc.). The feature trend value is a trend index obtained based on the time series analysis of blood pressure and blood sugar data. The input data is organized into a feature matrix, where each row represents the feature data at a time point, and the columns represent the specific dimensions of the features.
[0109] In an embodiment of the present invention, the calculation formula of the hidden layer includes:
[0110] ;
[0111] ;
[0112] ;
[0113] ;
[0114] Among them, represents the output of the reset gate of the th hidden layer, represents the output of the update gate of the th hidden layer, represents the candidate hidden state of the th hidden layer, represents the hidden state of the th hidden layer, represents the activation function, represent activation function, , and represent the first weight matrix, the second weight matrix, and the bias matrix of the reset gate of the th hidden layer, , and represent the first weight matrix, the second weight matrix, and the bias matrix of the update gate of the th hidden layer, , and represent the first weight matrix, the second weight matrix, and the bias matrix of the candidate hidden state of the th candidate hidden state, represent the th row feature encoding of the feature matrix of the input of the th hidden layer, represent the hidden state output by the th hidden layer, represent element-wise multiplication.
[0115] Specifically, by introducing a dynamic adjustment mechanism, the anomaly verification model can flexibly balance the weights between historical information and current input, thereby achieving precise processing of complex time series data. The model designs multiple layers of computing units, including a reset mechanism, an update mechanism, and a candidate state generation mechanism, to deeply interact with and extract the historical relevance and current features of the input data, and finally generates a hidden state that can reflect the dynamic changes of features. This design enhances the model's ability to express non-linear features and improves the robustness and sensitivity to abnormal data through layer-by-layer dynamic adjustment, ensuring accurate judgment of abnormal feature data.
[0116] In an embodiment of the present invention, the anomaly verification model updates the weight matrix and bias matrix of the hidden layer through backpropagation based on the mean squared error loss function. The calculation formula of the mean squared error loss function is as follows:
[0117] ;
[0118] where, represents the mean squared error loss value of the th patient, represents the value of the feature trend corresponding to the sample label of the th patient, represents the predicted feature trend value of the th patient.
[0119] It should be noted that the backpropagation update of the mean squared error loss function is a prior art, so it will not be elaborated here.
[0120] In an embodiment of the present invention, anomaly verification is performed, including:
[0121] Obtain the characteristic trend of the target patient within the second preset time period, and the predicted characteristic trend of the target patient within the second preset time period by the anomaly verification model;
[0122] If the difference between the characteristic trend and the predicted characteristic trend is less than the preset confidence threshold, it indicates that the diet information, blood pressure information, and blood glucose information of the target patient within the second preset time period are credible;
[0123] If the difference between the characteristic trend and the predicted characteristic trend is greater than the preset confidence threshold, it indicates that the diet information, blood pressure information, and blood glucose information of the target patient within the second preset time period are not credible.
[0124] Specifically, the core of anomaly verification lies in judging the credibility of patient data by comparing the characteristic trend of the target patient with the predicted characteristic trend of the anomaly verification model. Specifically, anomaly verification includes the following steps. Obtain the characteristic trend: The characteristic trend is calculated based on the diet information, blood pressure information, and blood glucose information of the target patient within the second preset time period. This trend reflects the characteristics of the patient's health status changing over time during this period, such as the amplitude of blood glucose increase or decrease, the blood pressure fluctuation range, etc. Model predicted characteristic trend: The anomaly verification model is constructed by training high-quality data in the confidence time period and can predict the characteristic trend of the patient within the second preset time period based on the historical characteristic data and patterns of the target patient. The predicted characteristic trend represents the reasonable expected value of the model for the patient's health data. Compare the difference: Compare the actual characteristic trend with the predicted characteristic trend of the model and calculate the difference between the two. This difference represents the degree of deviation between the actual health change of the target patient and the predicted health change of the model.
[0125] Perform credibility judgment. The difference is less than the preset confidence threshold: If the difference is small, it indicates that the characteristic trend of the target patient is relatively close to the prediction of the model, indicating that the patient's diet information, blood pressure information, and blood glucose information are consistent with their health changes, and these data can be considered credible. The difference is greater than the preset confidence threshold: If the difference is large, it indicates that there is a large deviation between the characteristic trend of the target patient and the model prediction, which may be due to inaccurate recording of diet information, abnormal blood pressure or blood glucose information, etc. The system determines that these data are unreliable data. In this way, the present invention can automatically screen the data of the target patient without relying on manual intervention to ensure that the data entering the subsequent analysis and application has high reliability.
[0126] The above has described the embodiments of this embodiment, but this embodiment is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make many forms, all of which fall within the protection scope of this embodiment.
Claims
1. A whole-process information management system for chronic kidney disease based on active learning, characterized in that Including: A data acquisition module, configured to collect the dietary information, blood pressure information, and blood glucose information of a plurality of patients within a preset confidence time period; Among them, the dietary information includes: the timestamp of the ingested food, the ingredients of the ingested food, the cooking method of the ingredients, and the intake of the ingredients; the blood pressure information and the blood glucose information respectively include: the blood pressure values and blood glucose values at the 0th, 1st, and 2nd moments after the timestamp of the ingested food; A dietary clustering module, configured to cluster the dietary information of the patients to obtain the typical diet of each patient within the preset confidence time period; An information analysis module, configured to obtain the blood pressure information and blood glucose information corresponding to the typical diet of the patients, and perform trend analysis to obtain the characteristic trends of the patients; A model construction module, configured to construct an anomaly verification model based on the typical dietary information and the characteristic trends; the anomaly verification model is used to perform anomaly verification on the dietary information, blood pressure information, and blood glucose information of the target patient within a second preset time period to obtain a confidence judgment of the target patient; the performing of the trend analysis includes: Constructing a data set from the blood pressure information and blood glucose information corresponding to each typical diet; Sorting the data set in chronological order to obtain a characteristic sorting; Calculating the characteristic trend of each typical diet based on the characteristic sorting, and the calculation formula of the characteristic trend is as follows: ; ; , ; , ; Among them, represents the characteristic trend of a typical diet, represents the trend of blood pressure, represents the trend of blood glucose, represents the weight of the trend of blood pressure, represents the weight of the trend of blood glucose, represents the interaction factor between the trend of blood pressure and the trend of blood glucose, represents the i-th data set, represents the standard blood pressure value of the i-th data set, represents the standard blood glucose value of the i-th data set, represents the blood pressure value at time 0 of the i-th data set, represents the blood pressure value at time 1 of the i-th data set, represents the blood pressure value at time 2 of the i-th data set, represents the blood glucose value at time 0 of the i-th data set, represents the blood glucose value at time 1 of the i-th data set, represents the blood glucose value at time 2 of the i-th data set, represents the activation function; The performing of the anomaly verification includes: Obtaining the characteristic trend of the target patient within the second preset time period, and the predicted characteristic trend of the target patient within the second preset time period by the anomaly verification model; If the difference between the characteristic trend and the predicted characteristic trend is less than the preset confidence threshold, it indicates that the dietary information, blood pressure information, and blood glucose information of the target patient within the second preset time period are credible; If the difference between the characteristic trend and the predicted characteristic trend is greater than the preset confidence threshold, it indicates that the dietary information, blood pressure information, and blood glucose information of the target patient within the second preset time period are not credible.
2. The full-process information management system for chronic kidney disease based on active learning according to claim 1, wherein Clustering the dietary information of the patients includes: Step 21, establishing an initial encoding of the dietary information, including: Taking the timestamp of the ingested food in the dietary information as the first element of the initial encoding; Querying the ingredients of the ingested food in the dietary information to obtain characteristic parameters; the characteristic parameters include the carbohydrate value, protein value, fat value, and sodium value per unit intake of the ingredients of the ingested food, as well as the glycemic index value and the glycemic load value; Taking each parameter in the characteristic parameters as the second to seventh elements of the initial encoding; Encoding the cooking method of the ingredients in the dietary information as the eighth element of the initial encoding by real numbers; Taking the intake of the ingredients in the dietary information as the ninth element of the initial encoding; Step 22, normalizing the initial encoding to obtain a characteristic encoding; Step 23, initializing N encoding clustering centers, and calculating the similarity between each characteristic encoding and each encoding clustering center, and the similarity calculation formula is as follows: ; Among them, represents the similarity between the th feature code and the th coding cluster center. , represents the index. represents the th element of the th feature code, and represents the th element of the th coding cluster center. represents the absolute value of represents the linear weight. represents the Manhattan distance between the th feature code and the th coding cluster center; Step 24, if the similarity between the th feature code and the th coding cluster center is the smallest, then assign the A-th feature code to the B-th coding cluster center; Step 25, obtaining N clusters based on Step 24, and taking the average encoding of the characteristic encodings in each cluster as the updated encoding clustering center of the corresponding cluster; Step 26, repeating Step 25 for a preset number of times to obtain N final clusters; Step 27: Set the clustering merging threshold. If the similarity between the updated coding cluster centers of any two final clusters is less than the preset threshold, then merge the corresponding final clusters to obtain merged clusters, and calculate the updated coding clusters of the merged clusters. Step 28: Repeat Step 27 until the similarity of the updated coding clusters between any two remaining final clusters or merged clusters is greater than the preset threshold. Then, take the final cluster or merged cluster with the largest number of feature encodings as the typical cluster, and the dietary information corresponding to the feature encodings in the typical cluster as the typical diet.
3. The chronic kidney disease full-process information management system based on active learning according to claim 2, characterized in that, Construct an anomaly verification model, including: Construct the training samples of the anomaly verification model and the sample labels of the anomaly verification model. The training samples include: Obtain the physical examination data of each patient at the start time and end time of the preset confidence period. Based on the analysis of the physical examination data by experts, obtain the health scores of the physical examination data of each patient at the start time and end time of the preset confidence period, and calculate the feature differences of the health scores of the physical examination data of each patient at the start time and end time. Based on the experts' confidence division of the feature differences, the confidence division includes: excellent, good, poor, and very poor. The interval that matches well with the characteristic trend value The interval that matches moderately well with the characteristic trend value The interval that matches poorly with the characteristic trend value The interval that matches very poorly with the characteristic trend value interval; If the feature trend of the patient in the preset confidence period matches the confidence division, then determine that the corresponding patient's dietary information, blood pressure information, and blood glucose information are confidence data, and use the feature encoding of the corresponding patient's dietary information as the training sample, and use the corresponding patient's feature trend as the sample label. If the feature trend of the patient in the preset confidence period does not match the confidence division, then determine that the corresponding patient's dietary information, blood pressure information, and blood glucose information are non-confidence data, and discard the corresponding patient's dietary information and feature trend. Based on the training samples and sample labels of the patients in the preset confidence period, train to obtain an anomaly verification model. The anomaly verification model includes: an input layer, a hidden layer, and a classifier.
4. The full-process information management system for chronic kidney disease based on active learning according to claim 3, characterized in that, The input layer, hidden layer, and classifier include: The input layer is used to stack a number of feature encodings of the target patient in the second preset period in chronological order into a feature matrix, and input the feature matrix into the hidden layer. The hidden layer is used to perform a non-linear mapping on the feature matrix to obtain the hidden state of the feature matrix. The classifier is used to input the hidden state. The classification space of the classifier represents the predicted feature trend of the target patient in the second preset period.
5. The full-process information management system for chronic kidney disease based on active learning according to claim 4, wherein The calculation formula of the hidden layer includes: ; ; ; ; Among them, represents the output of the reset gate of the th hidden layer, represents the output of the update gate of the th hidden layer, represents the candidate hidden state of the th hidden layer, represents the hidden state of the th hidden layer, represents activation function, represents activation function, , and represent the first weight matrix, second weight matrix and bias matrix of the reset gate of the th hidden layer, , and represent the first weight matrix, second weight matrix and bias matrix of the update gate of the th hidden layer, , and represent the first weight matrix, second weight matrix and bias matrix of the th candidate hidden state, represents the feature encoding of the th row of the feature matrix of the input of the th hidden layer, represents the hidden state output by the th hidden layer, represents element-wise multiplication.
6. The full-process information management system for chronic kidney disease based on active learning according to claim 5, wherein, The anomaly verification model updates the weight matrix and bias matrix of the hidden layer through backpropagation based on the mean squared error loss function. The calculation formula of the mean squared error loss function is as follows: ; Among them, represents the mean squared error loss value of the th patient, represents the value of the feature trend corresponding to the sample label of the th patient, represents the predicted feature trend value of the th patient.
Citation Information
Patent Citations
Big data-based nephropathy patient diet condition analysis system and method
CN117594195A