System and method for inferring diabetes based on error value detection and correction algorithm
Through a system based on error value detection and correction algorithms, combined with causal inference and machine learning, the problems of data inconsistency and error values in medical data are solved, the accuracy and comprehensiveness of early diagnosis of diabetes are achieved, and it adapts to the medical needs of different regions and populations.
Patent Information
- Application Number
- CN202510474684.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-09-16
AI Technical Summary
When existing technologies use medical data to diagnose diabetes, there are problems with inconsistent data sources, many errors, and a lack of comprehensive consideration of the patient's living conditions and family genetic information, resulting in low accuracy in early diagnosis and delayed treatment.
A system based on error value detection and correction algorithm is used, including data collection and integration, initial error value screening, deep error value correction, feature extraction and association, and diabetes inference module. The causal inference algorithm combining structural equation model and Markov random field is used to correct abnormal data, and data correction is performed through Bayesian network, integrating machine learning algorithm for inference.
It improves the accuracy of early diagnosis of diabetes, reduces the risk of missed diagnosis and misdiagnosis, provides a more sufficient and comprehensive basis for diagnosis, adapts to the medical needs of different regions and populations, and realizes precise medical decision-making.
Smart Images

Figure CN120656681A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical data processing, and in particular to a system and method for inferring diabetes based on an error value detection and correction algorithm. Background Art
[0002] In today's medical practice, with the continuous development of medical information technology, patients accumulate a vast amount of historical examination and test records during their medical treatment. This data contains a wealth of health information and can theoretically provide an important basis for disease diagnosis and treatment. However, relying solely on these historical examination and test records to infer disease has many limitations.
[0003] First, data comes from a wide range of sources, including hospital information systems (HIS), laboratory information management systems (LIS), and picture archiving and communication systems (PACS). The data collection standards of different sources are inconsistent, resulting in uneven data quality. For example, different hospitals or laboratories may have different testing equipment and testing methods, making the test results of the same indicator lack comparability between different institutions. In addition, there may be potential error factors in the data, such as problems in sample collection, transportation, and storage, or errors in testing instruments, which may lead to erroneous values in the data.
[0004] Secondly, traditional data analysis methods often find it difficult to effectively mine deep clues related to the disease when faced with these complex and mixed data. Take diabetes as an example. This is a chronic disease with a complex pathogenesis and hidden early symptoms. Conventional means find it difficult to accurately capture key signals from messy data. Moreover, existing methods usually only focus on the patient's examination and test records, and lack comprehensive consideration of key information such as the patient's daily life status and family genetic profile. For example, the patient's eating habits, exercise status, sleep quality, psychological state, etc. may all have an impact on the occurrence and development of diabetes, but this information is often ignored or not fully utilized.
[0005] As a result, the current accuracy rate for early diabetes diagnosis is low, resulting in many patients not receiving timely diagnosis and treatment in the early stages of the disease, delaying treatment and increasing medical costs and health risks. This not only causes pain and financial burden to individual patients, but also places significant pressure on the entire medical system and society.
[0006] In summary, existing technologies have obvious shortcomings in using medical data for diabetes diagnosis. There is an urgent need for a new technology that can integrate multi-source information, effectively handle data quality issues, and accurately infer the risk of diabetes.
[0007] In response to the above technical problems, the present invention proposes a system and method for inferring diabetes based on an error value detection and correction algorithm. Summary of the Invention
[0008] The purpose of the present invention is to address the shortcomings of the existing technology and provide a system and method for inferring diabetes based on error value detection and correction algorithm, so as to achieve efficient use of multi-source information and accurately infer the risk of diabetes.
[0009] In order to achieve the above objectives, the present invention adopts the following technical solutions:
[0010] A system for inferring diabetes based on an error value detection and correction algorithm includes:
[0011] Data collection and integration module, used to collect multi-source data corresponding to diabetic patients and unify and organize the multi-source data into a unified format;
[0012] Error value screening module, used to preliminarily screen abnormal data in multi-source data;
[0013] The deep error value correction module is used to correct the abnormal data that has been initially screened based on the causal inference algorithm to obtain the corrected data;
[0014] Feature extraction and association module, used to extract potential features related to diabetes from the corrected data and establish an association model between the features;
[0015] The diabetes inference module is used to input the associated feature data into a pre-trained inference model for processing. The inference model outputs the probability of diabetes and related diagnostic recommendations.
[0016] Furthermore, the error value initial screening module initially screens abnormal data from multi-source data using a simple algorithm based on statistical rules and medical common sense.
[0017] Furthermore, the depth error value correction module corrects the abnormal data preliminarily screened based on the causal inference algorithm in the following manner:
[0018] A causal inference algorithm was formed by combining the structural equation model with the Markov random field. A causal relationship network between multidimensional factors related to diabetic patients and various examination and test indicators was constructed through the causal inference algorithm to find the factors leading to abnormal data. The abnormal data were then corrected using the Bayesian network combined with the probability distribution of historical data.
[0019] Furthermore, the pre-trained inference model in the diabetes inference module is specifically:
[0020] The associated feature data is divided into training set, validation set and test set, and the training set is input into the machine learning model for training to obtain a trained inference model.
[0021] Furthermore, the inputting of the training set into the machine learning model for training also includes: using cross-validation technology to prevent overfitting, and accelerating the convergence speed of the machine learning model through an adaptive learning rate adjustment strategy.
[0022] Furthermore, it also includes a result output and verification module, which is used to present the output inference results in the form of a report and optimize the inference model based on the feedback information.
[0023] Accordingly, a method for inferring diabetes based on an error value detection and correction algorithm is also provided, including:
[0024] S1. Collect multi-source data corresponding to diabetic patients and unify and organize the multi-source data into a unified format;
[0025] S2. Preliminary screening of abnormal data from multi-source data;
[0026] S3. Correct the abnormal data found in the preliminary screening based on a causal inference algorithm to obtain corrected data;
[0027] S4. Extract potential features related to diabetes from the corrected data and establish an association model between the features;
[0028] S5. The associated feature data is input into a pre-trained inference model for processing, and the inference model outputs the probability of diabetes and related diagnostic recommendations.
[0029] Furthermore, in step S3, the abnormal data preliminarily screened are corrected based on the causal inference algorithm as follows:
[0030] A causal inference algorithm was formed by combining the structural equation model with the Markov random field. A causal relationship network between multidimensional factors related to diabetic patients and various examination and test indicators was constructed through the causal inference algorithm to find the factors leading to abnormal data. The abnormal data were then corrected using the Bayesian network combined with the probability distribution of historical data.
[0031] Furthermore, the inference model pre-trained in step S5 is specifically: dividing the associated feature data into a training set, a validation set and a test set, inputting the training set into the machine learning model for training, and obtaining a trained inference model.
[0032] Furthermore, after step S5, the following steps are further included:
[0033] S6. Present the output inference results in the form of a report and optimize the inference model based on the feedback information.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] 1. By integrating multi-source information, this invention breaks the limitation of traditional reliance on historical examinations and test records, comprehensively considers the patient's health status, greatly enriches the data dimension, and makes the basis for inferring diabetes more sufficient and comprehensive, effectively improving the accuracy of early diagnosis, reducing the risk of missed diagnosis and misdiagnosis, and gaining precious time for timely treatment of patients.
[0036] 2. The advanced causal inference algorithm used in the deep error value correction module can not only accurately identify the root cause of erroneous data, but also provide a credibility reference for the correction results with the help of uncertainty quantification models, ensuring high-quality data input, laying a solid foundation for subsequent accurate feature extraction and disease inference, and reducing deviations caused by erroneous data.
[0037] 3. The adaptive weighted ensemble learning algorithm of the diabetes inference module fully leverages the advantages of multiple machine learning algorithms, flexibly adjusts the weights of each algorithm based on different data characteristics, and is highly adaptable to complex and changeable medical data scenarios. It further improves the accuracy and reliability of diabetes risk inference and facilitates precision medical decision-making.
[0038] 4. The feedback learning mechanism constructed by the result output and verification module realizes the dynamic optimization of the system, continuously collects clinical feedback information, improves the model and algorithm parameters in a targeted manner, and continuously improves the system performance, enabling it to keep pace with the pace of medical development, adapt to the medical needs of different regions and different populations, and provide long-term and powerful technical support for diabetes prevention and control work. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 1 is a diagram of a system structure for inferring diabetes based on an error value detection and correction algorithm provided in an embodiment. DETAILED DESCRIPTION
[0040] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.
[0041] The purpose of the present invention is to address the deficiencies of the prior art and provide a system and method for inferring diabetes based on an error value detection and correction algorithm.
[0042] Example 1
[0043] This embodiment provides a system for inferring diabetes based on an error value detection and correction algorithm, such as Figure 1 Shown, including:
[0044] The data collection and integration module 11 is used to collect multi-source data corresponding to diabetic patients and unify and organize the multi-source data into a unified format;
[0045] The error value preliminary screening module 12 is used to preliminarily screen abnormal data in multi-source data;
[0046] The deep error value correction module 13 is used to correct the abnormal data initially screened based on the causal inference algorithm to obtain corrected data;
[0047] A feature extraction and association module 14 is used to extract potential features related to diabetes from the corrected data and establish an association model between the features;
[0048] The diabetes inference module 15 is used to input the associated feature data into a pre-trained inference model for processing, and the inference model outputs the probability of diabetes and related diagnostic recommendations;
[0049] The result output and verification module 16 is used to present the output inference results in the form of a report and optimize the inference model based on the feedback information.
[0050] In the data collection and integration module 11, multi-source data corresponding to diabetic patients are collected, and the multi-source data are formatted and organized.
[0051] Connecting with hospital information systems (HIS), laboratory information management systems (LIS), picture archiving and communication systems (PACS), wearable medical devices, and patient self-health management apps enables real-time synchronous data updates to ensure data integrity.
[0052] This embodiment obtains multi-source patient information data from the above channels. The data not only covers traditional examination and testing records such as blood routine, biochemical indicators, urine routine, and imaging examination results, but also includes the patient's daily exercise data, diet records, sleep monitoring data, detailed information on family genetic disease history, etc.
[0053] Initially clean and standardize the aforementioned patient information data, including data of varying formats, semantics, and sources, to ensure uniform and complete data formats. This includes standardizing time formats and test indicator units. Specifically, ETL (Extract, Transform, Load) tools are used to extract data from systems such as HIS, LIS, and PACS. During the conversion process, custom scripts or data cleaning software are used to remove noise from text data, correct spelling errors, and normalize and standardize numerical data.
[0054] Finally, the processed data is loaded into a local database using SQL statements or database loading tools, providing a comprehensive data foundation for subsequent processing. This module establishes secure and stable data transmission channels with various data sources, utilizing standardized data interface protocols such as HL7 to ensure accurate and timely data collection. Furthermore, a database management system (DBMS) is used for data storage, categorizing and storing data to facilitate rapid subsequent retrieval and access.
[0055] In the error value preliminary screening module 12, abnormal data in the multi-source data is preliminarily screened.
[0056] Using a simple algorithm based on statistical rules and medical common sense, the module quickly scans the dataset for data points that significantly deviate from the normal range or exhibit logical anomalies. For example, hemoglobin values exceeding the normal physiological upper limit by several times, blood glucose levels fluctuating significantly over a short period of time without a record of reasonable medical intervention, and so on. This preliminary identification of possible erroneous values significantly reduces the amount of data to be processed subsequently. The module also includes a built-in library of normal reference ranges for common medical indicators, which is regularly updated based on the latest medical research findings and clinical practice experience to ensure the scientific validity of the screening criteria. Furthermore, multi-threaded parallel processing technology is used to improve data screening efficiency to cope with the demands of processing massive amounts of data.
[0057] In the depth error value correction module 13, the abnormal data preliminarily screened are corrected based on the causal inference algorithm to obtain corrected data.
[0058] For abnormal data from the initial screening, a basic causal framework is built using the structural equation model (SEM) to clarify the direct influence paths between various factors. The Markov random field is then introduced to model the high-order interactions between complex factors, capturing the nonlinear correlations hidden behind the data, and thus obtaining a causal inference algorithm. Based on the causal relationship algorithm, a causal relationship network is constructed covering multi-dimensional factors such as basic patient information (age, gender, family medical history, etc.), lifestyle (diet, exercise, smoking and drinking habits, etc.), previous medication history, and psychological state (stress level, anxiety and depression, etc.) and various examination and test indicators.
[0059] Specifically, the implementation steps of combining the structural equation model (SEM) and the Markov random field (MRF) to form a causal inference algorithm are as follows:
[0060] 1. Build a basic causal framework (SEM part).
[0061] Identify variables and their relationships: First, based on medical knowledge and expert experience, identify various factors related to diabetes as variables, such as age, gender, family medical history, lifestyle (diet, exercise, smoking and drinking habits), previous medication history, psychological state (stress level, anxiety and depression), and various examination and test indicators. Then, based on existing medical research results and clinical practice experience, clarify the direct influence path between these variables, that is, the causal relationship.
[0062] Build a model using a SEM package: Use a SEM package such as AMOS, LISREL, or lavaan in the R language to construct an initial causal framework. When building a model, you need to define regression equations between variables, showing how one variable is affected by the others. For example, blood sugar levels may be affected by factors such as diet, physical activity, and psychological stress.
[0063] Estimating model parameters: Using methods such as maximum likelihood estimation, the collected data is used to estimate the parameters in the SEM model, such as regression coefficients, which reflect the strength of the causal relationship between variables.
[0064] Model evaluation and modification: The constructed SEM model is evaluated through goodness-of-fit tests to check whether the model can well explain the covariance structure in the data. If the model fit is poor, the model structure needs to be modified, such as adding or removing certain paths and re-estimating parameters, until a reasonable and well-fitting causal framework is obtained.
[0065] 2. Introduce Markov random fields to model high-order interactions (MRF part).
[0066] Defining potential functions: Building on the foundational causal framework established by SEM, we consider the complex, higher-order interactions between variables. We use MRFs to model these interactions, and define potential functions to describe the correlation or dependency between variables. Potential functions can be designed based on practical medical knowledge and data characteristics. For example, to determine the combined effect of a specific combination of lifestyle factors on blood sugar levels, we can define a corresponding potential function to capture this synergistic effect.
[0067] Constructing an MRF model: Use the MRF module in Python libraries such as pgmpy to construct a Markov random field model. Variables in the SEM are treated as nodes in the MRF. Undirected edges and their weights are defined between nodes based on a potential function to represent the dependencies between the variables. For example, in an MRF, a node representing blood sugar levels might be connected to multiple nodes, such as those representing dietary habits and exercise levels. The weights on these edges, determined by the corresponding potential function, reflect the strength of the interactions between them.
[0068] MRF model parameter estimation: Using methods such as the expectation maximization (EM) algorithm, the parameters of the MRF model, including the parameters of the potential function, are estimated based on the collected data. Through an iterative optimization process, the model is able to better fit the complex dependencies between variables in the data.
[0069] Joint reasoning and causal network construction: This algorithm combines SEM and MRF to form a unified causal inference algorithm. During reasoning, it comprehensively considers the causal paths defined in SEM and the high-order interactions captured in MRF to construct a comprehensive causal network. This network enables more accurate analysis and inference of the various factors and interactions that lead to anomalous data, providing a basis for correcting anomalous data.
[0070] 3. Integrate SEM and MRF to form a causal inference algorithm.
[0071] Model fusion: This approach combines the strengths of SEM and MRF. SEM provides clear causal directions and paths between variables, while MRF can capture complex undirected dependencies between variables. During the fusion process, the causal relationships in SEM can be used as prior knowledge to guide the design of potential functions and parameter estimation in MRF. Alternatively, the dependencies in MRF can be used to supplement higher-order interactions that may be missed in SEM.
[0072] Causal inference and correction of abnormal data: Utilizing the fused causal network, we conduct in-depth analysis of the initially screened abnormal data. By identifying the combination of factors that lead to the abnormal data and using the Bayesian network combined with the probability distribution of historical data based on the causal relationships and dependencies within the network, we can correct the abnormal data. For example, if a patient's blood sugar test value is abnormally high, the causal network reveals that this may be due to a combination of factors, such as a recent high-sugar diet, lack of exercise, and a high-stress work environment. Based on the normal ranges and interrelationships of these factors, we can then make reasonable corrections to the abnormal blood sugar value.
[0073] As can be seen from the above, in this embodiment, SEM constructs a linear causal framework to clarify direct effects, while MRF captures higher-order nonlinear interactions, compensating for SEM's limitations in understanding complex relationships. When these two methods are combined, SEM first determines the underlying path, and then MRF models higher-order dependencies between residuals or latent variables. Through iterative optimization or joint parameter estimation, local linear relationships are integrated with global nonlinear patterns, forming a hybrid causal inference algorithm that balances interpretability with the ability to model complex systems.
[0074] This embodiment constructs a causal network by first building an initial network architecture based on medical literature and expert knowledge. Then, through machine learning training on large-scale real-world medical data, it continuously optimizes the connection weights and causal relationship strengths between network nodes. Furthermore, it introduces an uncertainty quantification model to assess the credibility of each correction result, providing a reference for subsequent decision-making.
[0075] For example, when a patient's blood sugar value is found to be abnormal data, SEM can be used to trace possible influencing factors, such as drug side effects interfering with biochemical indicators, the patient's special physical condition at the time of sample collection affecting the test results, and changes in physiological indicators caused by emotional fluctuations. Taking the interference of drug side effects as an example, Markov random field analysis is used to analyze the comprehensive impact of the drug on blood sugar under the combined action of factors such as the patient's age, gender, other medications taken at the same time, and current psychological state, so as to more accurately locate the root cause of the error and correct it.
[0076] In the specific implementation, we first use structural equation model software (such as AMOS, LISREL, etc.) to build an initial causal framework, determine the direct causal relationship between variables, and estimate model parameters through methods such as maximum likelihood estimation; then, use Markov random field libraries (such as Python libraries such as pomegranate) to model high-order interaction relationships, and use Bayesian networks combined with the probability distribution of a large amount of historical data to infer and calculate the probability of each factor affecting outliers, and finally achieve accurate correction of erroneous values and restore data authenticity.
[0077] In this embodiment, the correction of abnormal data by Bayesian reasoning is specifically as follows:
[0078] In a Bayesian network, this can be achieved by calculating the posterior probability of the data point in the model. If the posterior probability of a data point is much lower than the probability of a normal data point, it can be considered as abnormal data. The detected abnormal data is then corrected using the Bayesian network. In this embodiment, the observation value of the abnormal data point is input into the Bayesian network model to calculate the posterior probability of each variable taking a normal value under the observation value. For example, for a data point with an abnormally high blood sugar test value, the probability of the blood sugar level taking a normal value can be calculated given the values of other related variables (such as eating habits, exercise volume, etc.). According to the posterior probability distribution, the most likely normal value is found to replace the abnormal data point. This can be achieved by selecting the normal value with the highest probability or sampling according to the probability distribution.
[0079] In this embodiment, the Bayesian network constructs a probabilistic model by integrating the linear causal framework of SEM with the high-order interaction relationship of MRF. First, the network structure is initialized based on medical literature and expert knowledge, SEM is used to determine the direct causal path, and MRF captures complex interactions. Subsequently, the Bayesian network is trained with large-scale medical data to learn the conditional probability distribution between nodes. When abnormal blood sugar is detected, the network infers the probability of each factor contributing to the abnormal value based on historical probability patterns such as the patient's age, gender, drug combination, and psychological state. For example, if the probability of drug side effects is significant, the blood sugar value is corrected and the credibility is annotated to achieve accurate data restoration.
[0080] After correcting outliers in this embodiment, the above correction process needs to be iterated until the data quality meets the requirements of subsequent analysis to ensure the authenticity and reliability of the input data. During the iteration process, data quality assessment indicators such as error value ratio and data consistency index are set. When the indicator reaches a preset threshold (such as the error value ratio is less than 5%), the iteration is stopped. At the same time, the correction log of each iteration is recorded to facilitate subsequent tracing and analysis.
[0081] In the feature extraction and association module 14, potential features related to diabetes are extracted from the corrected data, and an association model between the features is established.
[0082] From the revised high-quality data set, potential features related to diabetes are extracted based on medical knowledge, including but not limited to long-term high fasting blood sugar trends, slow rising trajectories of glycated hemoglobin, abnormal fluctuations in insulin secretion, persistent presence of trace albumin in urine, characteristic imaging manifestations of fundus vascular lesions, long-term abnormal fluctuations in body mass index (BMI), and changes in the accumulation level of advanced glycation end products in the skin. A correlation model is established between the features to explore their coordinated changes in the course of diabetes, providing powerful input for subsequent model training. This module integrates multiple medical image analysis algorithms to automatically identify imaging features such as fundus vascular lesions; uses time series analysis methods to capture the dynamic trends of indicators such as blood sugar and glycated hemoglobin; and uses data mining technology to discover hidden association patterns between different features, such as the correlation between the use of certain specific drugs and abnormal fluctuations in insulin secretion, and the connection between long-term high stress and poor blood sugar control.
[0083] During the feature extraction process, this embodiment combines the experience of medical experts and data analysis results to establish a feature importance scoring system, giving priority to features with high scores; when constructing a feature association model, a graph database (such as Neo4j, etc.) is used to store and manage the relationship between features, and the coordinated change rules of features are intuitively displayed through the attribute settings of nodes and edges.
[0084] In the diabetes inference module 15, the associated feature data is input into a pre-trained inference model for processing, and the inference model outputs the probability of diabetes and related diagnostic recommendations.
[0085] The associated feature data is divided into training, validation, and test sets. These feature data include, but are not limited to, long-term fasting blood glucose trends, glycated hemoglobin changes, insulin secretion fluctuations, the presence of microalbumin in urine, fundus vascular disease characteristics, and BMI fluctuations. Multiple rounds of training are then used to optimize the machine learning model parameters, adapting the model to the data characteristics of different patients. Finally, a trained inference model is obtained, which is used to infer the target patient data and output the probability of diabetes and related diagnostic recommendations. When dividing the data set, stratified sampling and other methods are used to ensure that the sample distribution of each data set is similar to the overall population. During the model training process, GPU acceleration is used to shorten training time. During the inference phase, the input data is preprocessed to meet the model requirements and ensure the accuracy of the inference results.
[0086] The inference model is trained as follows:
[0087] Select a machine learning algorithm: Use an ensemble learning strategy to combine the strengths of multiple machine learning algorithms, such as decision trees, support vector machines (SVMs), and neural networks. In practice, use machine learning libraries such as Scikit-learn to build an integrated machine learning model.
[0088] In the diabetes inference module, an ensemble learning strategy constructs an inference model by fusing decision trees, support vector machines (SVMs), and neural networks. First, the bagging method is used to train multiple decision trees in parallel, forming a random forest as a base learner to capture linear relationships and feature interactions in the data. Simultaneously, the SVM is used to process nonlinear relationships and extract complex patterns. Subsequently, the prediction results of the random forest and SVM are used as input to train a neural network as a meta-learner, integrating the predictive advantages of different algorithms. This combination retains the interpretability of decision trees while leveraging the nonlinear processing capabilities of SVMs and the powerful fitting capabilities of neural networks, improving the model's generalization and predictive accuracy.
[0089] Model training uses the training set to train the machine learning model. During the training process, techniques such as cross-validation and random forest are used to prevent overfitting, and an adaptive learning rate adjustment strategy is used to accelerate model convergence.
[0090] Dynamic weight adjustment: At the beginning of training, each learning algorithm is given the same initial weight. As training data continues to be input, the weights are dynamically adjusted based on each algorithm's performance on different sample subsets. For example, decision trees excel when processing data subsets with clear logical rules, and their weights are increased accordingly. Neural networks excel at capturing complex nonlinear relationships, and their weights are increased when encountering such data, ensuring the final model's high accuracy and adaptability for diabetes inference.
[0091] Model evaluation and optimization: Use validation sets to evaluate the model and adjust algorithm parameters to optimize model performance. Regularly evaluate and update the model to adapt to medical developments and new data features.
[0092] Evaluation and updating can also be implemented using Scikit-learn. The dataset is divided into training set, validation set, and test set. The model is trained on the training set and the performance of the model is evaluated on the validation set. The above evaluation indicators are used to adjust the model parameters, select a better model structure, or use a different algorithm based on the evaluation results. The optimized model is finally evaluated on the test set to ensure the generalization ability of the model.
[0093] In the diabetes inference module, model evaluation is performed on a validation set. For classification tasks, model performance is measured using metrics such as accuracy, precision, recall, F1 score, and ROC-AUC. For regression tasks, metrics such as mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE) are used. During evaluation, the degree of match between the model's predictions and the true values is calculated to ensure the model's generalization ability to unseen data. Based on the evaluation results, algorithm parameters are adjusted to optimize model performance and improve prediction accuracy and adaptability.
[0094] In the diabetes inference module, the selection and optimization of model parameters are crucial. First, the optimal combination is searched within the parameter space through methods such as grid search, random search, or Bayesian optimization. The performance of different parameters is evaluated using a validation set, and the best performing parameters are selected. Subsequently, optimization methods such as gradient descent and genetic algorithms are used to adjust the parameters based on the loss function to minimize prediction error. Through iterative training, hyperparameters such as the learning rate are dynamically adjusted to balance model convergence speed and stability. Finally, the parameters are fine-tuned, incorporating the knowledge of medical experts, to ensure that the model not only conforms to clinical principles but also exhibits high accuracy and generalizability.
[0095] This embodiment uses machine learning libraries such as Scikit-learn to build an integrated learning model. Through a custom evaluation index function, the performance of each algorithm on different sample subsets is monitored in real time. The weights are dynamically adjusted according to the performance to achieve adaptive optimization of the model and obtain a trained inference model.
[0096] Diabetes is specifically inferred as:
[0097] The associated feature data in the test set is input into the trained inference model, and the inference model outputs the probability of diabetes and related diagnostic recommendations based on the feature data.
[0098] In the result output and verification module 16, the output inference results are presented in the form of a report and the inference model is optimized based on the feedback information.
[0099] The inference results are presented to medical staff in the form of an intuitive and easy-to-understand report. The report is compared with the patient's subsequent clinical diagnosis results (if any) and newly added examination and test data, and the feedback learning mechanism is used to continuously optimize the performance of the entire system, so that the accuracy of the inference is continuously improved. The report generated by this module is in the form of pictures and texts. In addition to providing the probability of diabetes and related diagnostic recommendations, it also details the changing trends of key feature data, the comparison before and after correction of outliers, and other information, so that medical staff can deeply understand the inference process. In addition, a quality control system for feedback data is established to clean and verify the collected feedback information. The feedback data is used to retrain the inference model, adjust the algorithm parameters, optimize the causal network and feature extraction strategy, and continuously improve the performance of the system. A feedback data management system is established to classify, store and statistically analyze the feedback data. Based on the analysis results, the various modules of the system are adjusted in a targeted manner to achieve continuous evolution of the system.
[0100] The feedback data collection process in this embodiment is as follows:
[0101] 1. Data source:
[0102] A. Clinical feedback from medical staff: collect the doctors’ approval of the inference results and provide supplementary diagnostic opinions.
[0103] B. Patient's subsequent diagnostic results: such as pathology reports, laboratory tests, etc., to confirm the actual condition of the disease.
[0104] C. New inspection and testing data: such as blood sugar, glycosylated hemoglobin and other dynamic monitoring data.
[0105] 2. Data cleaning:
[0106] A. Missing Value Processing: Use mean filling, median filling, or interpolation to fill in missing data. B. Outlier Correction: Detect outliers using statistical methods (such as Z-score) or machine learning methods and correct or eliminate them based on medical expert advice.
[0107] C. Data standardization: Normalize data of different dimensions to facilitate model processing.
[0108] 3. Data Verification
[0109] A. Logical verification: Check whether the data conforms to medical common sense (such as blood sugar value range, drug side effect logic).
[0110] B. Statistical verification: Analyze data distribution and correlation to ensure consistency with the overall data.
[0111] C. Expert review: Medical experts conduct a final review of the cleaned data to ensure clinical rationality.
[0112] Compared with the prior art, this embodiment has the following beneficial effects:
[0113] 1. By integrating multi-source information, it breaks the limitation of traditional reliance on historical examination and test records, comprehensively considers the patient's health status, greatly enriches the data dimension, and makes the basis for inferring diabetes more sufficient and comprehensive, effectively improving the accuracy of early diagnosis, reducing the risk of missed diagnosis and misdiagnosis, and gaining precious time for timely treatment of patients.
[0114] 2. The advanced causal inference algorithm used in the deep error value correction module can not only accurately identify the root cause of erroneous data, but also provide a credibility reference for the correction results with the help of uncertainty quantification models, ensuring high-quality data input, laying a solid foundation for subsequent accurate feature extraction and disease inference, and reducing deviations caused by erroneous data.
[0115] 3. The adaptive weighted ensemble learning algorithm of the diabetes inference module fully leverages the advantages of multiple machine learning algorithms, flexibly adjusts the weights of each algorithm based on different data characteristics, and is highly adaptable to complex and changeable medical data scenarios. It further improves the accuracy and reliability of diabetes risk inference and facilitates precision medical decision-making.
[0116] 4. The feedback learning mechanism constructed by the result output and verification module realizes the dynamic optimization of the system, continuously collects clinical feedback information, improves the model and algorithm parameters in a targeted manner, and continuously improves the system performance, enabling it to keep pace with the pace of medical development, adapt to the medical needs of different regions and different populations, and provide long-term and powerful technical support for diabetes prevention and control work.
[0117] Accordingly, this embodiment also provides a method for inferring diabetes based on an error value detection and correction algorithm, including:
[0118] S1. Collect multi-source data corresponding to diabetic patients and unify and organize the multi-source data into a unified format;
[0119] S2. Preliminary screening of abnormal data from multi-source data;
[0120] S3. Correct the abnormal data found in the preliminary screening based on a causal inference algorithm to obtain corrected data;
[0121] S4. Extract potential features related to diabetes from the corrected data and establish an association model between the features;
[0122] S5. The associated feature data is input into a pre-trained inference model for processing, and the inference model outputs the probability of diabetes and related diagnostic recommendations.
[0123] Furthermore, in step S3, the abnormal data preliminarily screened are corrected based on the causal inference algorithm as follows:
[0124] A causal inference algorithm was formed by combining the structural equation model with the Markov random field. A causal relationship network between multidimensional factors related to diabetic patients and various examination and test indicators was constructed through the causal inference algorithm to find the factors leading to abnormal data. The abnormal data were then corrected using the Bayesian network combined with the probability distribution of historical data.
[0125] Furthermore, before step S5, the method further includes: dividing the associated feature data into a training set, a validation set, and a test set, inputting the training set into a machine learning model for training, and obtaining a trained inference model.
[0126] Furthermore, after step S5, the following steps are further included:
[0127] S6. Present the output inference results in the form of a report and optimize the inference model based on the feedback information.
[0128] Example 2
[0129] The system for inferring diabetes based on the error value detection and correction algorithm provided in this embodiment is different from that in the first embodiment in that:
[0130] This embodiment applies the system to the endocrinology department of a large general hospital, which sees a large number of patients with complex medical histories every day and has accumulated a wealth of historical examination and testing records.
[0131] The data collection and integration module collects relevant data from over 100,000 patients over the past five years from various internal hospital information systems. It also connects with wearable devices worn by patients (such as smart bracelets that monitor exercise and sleep data) and self-health management apps used by patients (which record information such as diet and mood), transmitting this multi-source data to a local database. During this process, when connecting to the HIS system, it uses the HL7 protocol to obtain patient medical records, including medical history and diagnosis records, in real time. It also extracts blood routine tests, biochemical indicators, and other test data from the LIS system, cleans and converts them using ETL tools, and then stores them in the database. It also interacts with the PACS system to download imaging examination results and perform image preprocessing to ensure that image quality meets the requirements of subsequent analysis.
[0132] The initial error value screening module quickly identified approximately 20% of data with significant anomalies, flagging them and handing them over to the in-depth error value correction module. This module, based on a built-in library of the latest medical indicator reference ranges and utilizing multi-threaded parallel processing technology, completes the initial screening of massive amounts of data within an hour, significantly improving work efficiency.
[0133] The Deep Error Correction module constructs a causal network, incorporating information such as a patient's lifestyle, medication history, and psychological state to correct approximately 80% of these errors. For example, if a patient's blood sugar level is falsely elevated due to recent steroid use and a high-stress work environment, this error can be corrected. When constructing the causal network, an initial architecture is built based on the knowledge of medical experts. Then, using five years of real-world medical data accumulated by the hospital, machine learning training optimizes network parameters, enabling the network to accurately reflect the causal relationships between various factors.
[0134] The feature extraction and association module extracts diabetes-related features from the corrected data and constructs a feature association model. Using a variety of medical image analysis algorithms, it automatically identifies fundus vascular lesions with an accuracy rate exceeding 85%. Through time series analysis, it accurately captures the dynamic trends of indicators such as blood glucose and glycosylated hemoglobin. Using data mining techniques, it discovered over ten new diabetes-related feature association patterns, providing rich information for model training.
[0135] The diabetes inference module inputs feature data into an ensemble learning model. After training and optimization, it inferred information on 1,000 suspected diabetes patients. Compared with traditional diagnostic methods, the accuracy of early diabetes diagnosis increased by approximately 25%, reaching approximately 70%. During model training, cross-validation techniques were used to prevent overfitting, and GPU-accelerated computing was used to reduce training time by 30%, ensuring efficient model training and high-precision inference.
[0136] The result output and verification module feeds inference results back to doctors, combines them with subsequent clinical diagnosis, and continuously optimizes the system. After three months of iteration, the diagnostic accuracy rate has further increased to 75%. The module generates a richly illustrated report that details the changing trends of key feature data and a comparison of outlier values before and after correction. It has been widely praised by doctors and provides strong support for clinical diagnosis.
[0137] Example 3
[0138] The system for inferring diabetes based on the error value detection and correction algorithm provided in this embodiment is different from that in the first embodiment in that:
[0139] This embodiment applies the system to community medical centers to conduct diabetes risk screening based on the physical examination data of middle-aged and elderly people.
[0140] The data collection and integration module collects historical health checkup data from over 5,000 middle-aged and elderly residents in the community. Residents are also encouraged to use simple wearable devices (such as pedometers) and accompanying health tracking apps (to record diet, daily physical sensations, etc.) to collect this additional data. Because community medical centers have a relatively limited data source, primarily relying on physical examination systems, this module optimizes the data collection process to ensure data integrity and accuracy, and also utilizes simple data cleaning tools to perform preliminary processing of the physical examination data.
[0141] Through error value processing, feature extraction, and inference, the system conducts a stratified assessment of community residents' diabetes risk and generates personalized health recommendations, such as recommending further specialist examinations for high-risk individuals and providing dietary and exercise guidance for low-risk individuals. Through feature extraction and association modules, it identifies potential diabetes characteristics closely related to lifestyle habits among community residents, such as the correlation between a long-term high-salt diet and insulin resistance, providing a basis for personalized health recommendations.
[0142] By regularly interviewing residents about their health status and collecting feedback to optimize the system, the accuracy of early diabetes screening in the community has gradually increased, effectively reducing the rate of missed diagnoses and providing strong support for community diabetes prevention and control. During the follow-up process, a strict feedback data quality control system was established to verify and cleanse the information provided by residents, ensuring the reliability of the feedback data and providing strong support for system optimization.
[0143] The implementation cases in the above different scenarios show that the system and method for inferring diabetes based on error value detection and correction algorithm of this embodiment can fully tap the value of multi-source information, significantly improve the accuracy of early diagnosis of diabetes, and provide key technical support for precision medicine and disease prevention and control.
[0144] Note that the above are only preferred embodiments of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments and may include many other equivalent embodiments without departing from the concept of the present invention. The scope of the present invention is determined by the scope of the appended claims.
Claims
1. A system for inferring diabetes based on an error value detection and correction algorithm, characterized in that: include: Data collection and integration module, used to collect multi-source data corresponding to diabetic patients and unify and organize the multi-source data into a unified format; Error value screening module, used to preliminarily screen abnormal data in multi-source data; The deep error value correction module is used to correct the abnormal data that has been initially screened based on the causal inference algorithm to obtain the corrected data; Feature extraction and association module, used to extract potential features related to diabetes from the corrected data and establish an association model between the features; The diabetes inference module is used to input the associated feature data into a pre-trained inference model for processing. The inference model outputs the probability of diabetes and related diagnostic recommendations.
2. The system for inferring diabetes based on an error value detection and correction algorithm according to claim 1, characterized in that: The error value initial screening module initially screens abnormal data in multi-source data using a simple algorithm based on statistical rules and medical common sense.
3. The system for inferring diabetes based on an error value detection and correction algorithm according to claim 1, characterized in that: The deep error value correction module corrects the abnormal data initially screened based on the causal inference algorithm in the following manner: A causal inference algorithm was formed by combining the structural equation model with the Markov random field. A causal relationship network between multidimensional factors related to diabetic patients and various examination and test indicators was constructed through the causal inference algorithm to find the factors leading to abnormal data. The abnormal data were then corrected using the Bayesian network combined with the probability distribution of historical data.
4. The system for inferring diabetes based on an error value detection and correction algorithm according to claim 3, characterized in that: The pre-trained inference model in the diabetes inference module is specifically: The associated feature data is divided into training set, validation set and test set, and the training set is input into the machine learning model for training to obtain a trained inference model.
5. The system for inferring diabetes based on an error value detection and correction algorithm according to claim 4, characterized in that: The inputting of the training set into the machine learning model for training also includes: using cross-validation technology to prevent overfitting, and accelerating the convergence speed of the machine learning model through an adaptive learning rate adjustment strategy.
6. The system for inferring diabetes based on an error value detection and correction algorithm according to claim 5, characterized in that: It also includes a result output and verification module, which is used to present the output inference results in the form of a report and optimize the inference model based on the feedback information.
7. A method for inferring diabetes based on an error value detection and correction algorithm, characterized in that: include: S1. Collect multi-source data corresponding to diabetic patients and unify and organize the multi-source data into a unified format; S2. Preliminary screening of abnormal data from multi-source data; S3. Correct the abnormal data found in the preliminary screening based on a causal inference algorithm to obtain corrected data; S4. Extract potential features related to diabetes from the corrected data and establish an association model between the features; S5. The associated feature data is input into a pre-trained inference model for processing, and the inference model outputs the probability of diabetes and related diagnostic recommendations.
8. The method for inferring diabetes based on an error value detection and correction algorithm according to claim 7, characterized in that: In step S3, the abnormal data initially screened are corrected based on the causal inference algorithm as follows: A causal inference algorithm was formed by combining the structural equation model with the Markov random field. A causal relationship network between multidimensional factors related to diabetic patients and various examination and test indicators was constructed through the causal inference algorithm to find the factors leading to abnormal data. The abnormal data were then corrected using the Bayesian network combined with the probability distribution of historical data.
9. The method for inferring diabetes based on an error value detection and correction algorithm according to claim 8, characterized in that: The inference model pre-trained in step S5 is specifically: dividing the associated feature data into a training set, a validation set and a test set, inputting the training set into the machine learning model for training, and obtaining a trained inference model.
10. The method for inferring diabetes based on an error value detection and correction algorithm according to claim 9, characterized in that: After step S5, the following steps are further included: S6. Present the output inference results in the form of a report and optimize the inference model based on the feedback information.
Citation Information
Cited By
Diabetic health data management method and system
CN121011360A