A diabetes prediction algorithm based on KS model

Through the diabetes prediction algorithm based on the KS model, using data cleaning, standardization and cluster analysis technology, the problem of low accuracy of diabetes prediction in the existing technology is solved, and early screening and accurate prediction of diabetic patients are achieved.

CN113921134BActive Publication Date: 2025-05-09SHANGHAI JUNYILING LIFE SCIENCE RESEARCH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111021699.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-01
Publication Date
2025-05-09
Estimated Expiration
2041-09-01

AI Technical Summary

Technical Problem

The prior art is difficult to improve the accuracy of diabetes prediction, especially in the early detection of risks and timely preventive measures.

Method used

The diabetes prediction algorithm based on the KS model was used to obtain the Pima Indian diabetes data set from the UCI machine learning repository, clean and standardize, and filter the relevant attributes using quartile analysis and Person correlation coefficient analysis. Then, unsupervised learning is used using the KMEANS++ algorithm to obtain the correct clustering data and input it into the SVM for prediction.

Benefits of technology

It improves the accuracy of diabetes prediction, can accurately classify positive and negative diabetes patients, and is used for clinical screening and early warning of early diabetes, which may involve initial diagnosis of other diseases in the future.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113921134B_ABST
    Figure CN113921134B_ABST
Patent Text Reader

Abstract

The present invention discloses a diabetes prediction algorithm based on the KS model. The data set adopts the Pima India diabetes data set. The data set is cleaned and standardized to make the data quality higher. The Person correlation coefficient is used to select the data set attributes, remove unimportant attributes, speed up the processing speed of the later model, and improve the performance of the model. The KS model is trained using the data set, and the data is classified and predicted, and the model is evaluated using the correlation index. Experiments show that the model proposed by the present invention has better performance than others, indicating that the results of this study are very promising. The research results are mainly used in the clinical screening and early warning of early diabetes. In the future, it may involve the early diagnosis of other diseases and is widely used in the field of bioinformatics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence bioinformatics technology, and specifically relates to a diabetes prediction algorithm based on a KS model. Background Art

[0002] With the rapid development of social economy, people's lifestyles are also undergoing tremendous changes. Lifestyle is crucial to people's health. If you don't pay attention to it, it will lead to various chronic diseases in the long run. Among them, diabetes is called the "second killer" of modern diseases. If it is not well controlled, it will cause complications of the eyes, cardiovascular and cerebrovascular, kidneys and other diseases, leading to blindness, amputation, heart disease, stroke, renal failure and other consequences. The harm to the human body is second only to cancer. However, with the continuous expansion of the application scope of artificial intelligence, machine learning technology is used to predict chronic diseases, especially diabetes. If risks can be discovered in the early stage and relevant measures such as prevention and blocking can be taken in time, it will be of great benefit to improving people's quality of life and the healthy life expectancy of residents, and it will also effectively reduce the burden of diabetes treatment.

[0003] The difficulty in using machine learning technology to predict diabetes lies in how to improve the accuracy of the prediction so that healthy people can understand their risk of disease in a timely manner, while also allowing patients to receive guidance from professional doctors in the early stages of the disease. Summary of the invention

[0004] The purpose of the present invention is to provide a diabetes prediction algorithm based on the KS model, which can improve the accuracy of diabetes prediction.

[0005] The technical solution adopted by the present invention is a diabetes prediction algorithm based on the KS model, which is implemented in the following steps:

[0006] Step 1: Obtain the Pima Indian Diabetes Dataset from the UCI Machine Learning Repository.

[0007] Step 2: Clean the acquired Pima India diabetes dataset to obtain a normalized dataset;

[0008] Step 3: Use the quartile analysis method to detect and delete extreme values ​​and outliers in the normalized data set;

[0009] Step 4: Use Person correlation coefficient analysis to select four attributes with high correlation with diabetes prediction;

[0010] Step 5: Standardize the corresponding data of the four attributes respectively;

[0011] Step 6: Input the standardized data set into the KMEANS++ algorithm for unsupervised learning to obtain the correct clustered data;

[0012] Step 7: Input the data into SVM for prediction and evaluate its performance.

[0013] The Pima India Diabetes Dataset has nine attribute values, eight of which are related to diabetes diagnosis and one label attribute. In the label attribute, '0' represents healthy people and '1' represents patient people.

[0014] The present invention is also characterized in that:

[0015] The specific process of step 2 is: delete all content except each attribute data and fill in the missing values ​​in the data set.

[0016] The specific process of filling missing values ​​in the data set is: the average value of each attribute data replaces the missing data.

[0017] The specific process of step 4 is as follows:

[0018] Take the data X in each attribute i ,

[0019] Calculate the Person correlation coefficient using the formula:

[0020]

[0021] in, represents the standard score of the data in the attribute, represents the average value of the data in the attribute, σ X Represents the standard deviation of the data in the attribute;

[0022] The Person correlation coefficients obtained for each attribute are selected and sorted by size, four Person correlation coefficients with larger values ​​are selected, and four attributes are obtained according to the Person correlation coefficients.

[0023] The specific process of step 5 is as follows:

[0024] The data in the four attributes are standardized respectively, and the formula is:

[0025]

[0026] Among them, X represents the data in the attribute, represents the mean value of the data in the attribute, σ represents the standard deviation of the data in the attribute, and X * Represents the data obtained after standardization.

[0027] The beneficial effects of the present invention are:

[0028] The present invention discloses a diabetes prediction algorithm based on the KS model. The diabetes data set is first processed, optimized and other operations are performed to ensure the reliability of the data. The Person correlation coefficient is used to remove irrelevant attributes of the data set. Then, the KMEANS++ algorithm is used to perform unsupervised clustering on the data set, and the correct clustering results are used as the input of the SVM algorithm. Finally, classification is performed in the SVM algorithm to improve the efficiency and accuracy of classification. The new diabetes data set is input into the model, and positive and negative diabetes patients can be classified more accurately. The research results are mainly used in the clinical screening and early warning of early diabetes. In the future, it may involve the early diagnosis of other diseases and is widely used in the field of bioinformatics. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a flow chart of a diabetes prediction algorithm based on the KS model of the present invention;

[0030] Figure 2 is a flow chart of preprocessing data sets in the present invention;

[0031] Figure 3 It is a ROC comparison result diagram obtained when the data set is processed by the method of the present invention and classified by a logistic regression classifier, a naive Bayes classifier, a support vector machine classifier, a decision tree classifier, a KNN classifier and the present model;

[0032] Figure 4 This is a ROC comparison result diagram obtained when the data set is classified using the model of the present invention after 7-3 division;

[0033] Figure 5 This is the ROC comparison result and average ROC comparison result diagram obtained after ten-fold cross validation using this model. DETAILED DESCRIPTION

[0034] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] The present invention provides a diabetes prediction algorithm based on the KS model, which is implemented in the following steps:

[0036] Step 1: Obtain the Pima Indian Diabetes Dataset from the UCI Machine Learning Repository.

[0037] The Pima India Diabetes Dataset includes nine attribute values ​​for 768 patients, eight of which are related to diabetes diagnosis and one label attribute. The label attribute uses '0' to represent healthy people and '1' to represent patients.

[0038] Step 2: For the obtained Pima India diabetes dataset, delete all content except each attribute data. There are missing values ​​in blood sugar, blood pressure, skin thickness, insulin level and BMI attributes. Replace the missing data with the average value of each attribute data.

[0039] Step 3: Use the quartile analysis method to detect and delete extreme values ​​and outliers in the normalized data set, streamline the data set, and improve the data quality.

[0040] Step 4: For the processed data, since there are eight attributes related to diabetes diagnosis in this data set, in order to ensure that irrelevant attributes will not have a negative impact on the model, the data set is reduced in dimension, and the Person correlation coefficient analysis method is used to select four attributes with high correlation with diabetes prediction; the specific process is as follows:

[0041] Take the data X in each attribute i ,

[0042] Calculate the Person correlation coefficient, the formula is:

[0043]

[0044] in, represents the standard score of the data in the attribute, represents the average value of the data in the attribute, σ X Represents the standard deviation of the data in the attribute;

[0045] The Person correlation coefficients obtained for each attribute are selected and sorted by size, four Person correlation coefficients with larger values ​​are selected, and four attributes are obtained according to the Person correlation coefficients.

[0046] Step 5: Standardize the corresponding data of the four attributes to improve the processing speed of the model; the specific process is as follows:

[0047] The data in the four attributes are standardized respectively, and the formula is:

[0048]

[0049] Among them, X represents the data in the attribute, represents the mean value of the data in the attribute, σ represents the standard deviation of the data in the attribute, and X * Represents the data obtained after standardization.

[0050] Step 6. Use the KMEANS++ algorithm to cluster the data set. The KMEANS++ algorithm is an improved version of the KMEANS algorithm. The main difference between it and the KMEANS algorithm is the selection of the initial cluster center point, which will affect the results of the subsequent clustering and the number of iterations. While increasing the number of correct clustered data, the time complexity is also reduced. The KMEANS++ algorithm uses the Euclidean distance to calculate the distance between the data point and the cluster center, so as to divide the data point into the correct cluster. The standardized data set is input into the KMEANS++ algorithm for unsupervised learning to obtain the correct clustered data;

[0051] Step 7: Input the data into SVM for prediction and evaluate its performance.

[0052] The KMEANS++ algorithm is used to cluster the correct data as the input of the SVM (Support Vector Machine). As a supervised learning algorithm, SVM (Support Vector Machine) is data-driven, model-free, and has strong classification and discrimination capabilities. SVM (Support Vector Machine) is increasingly used for disease detection and performs well in solving classification problems in bioinformatics. SVM (Support Vector Machine) predicts the correct data clustered by KMEANS++, thereby improving the accuracy of the entire diabetes prediction. Through actual training and testing, the accuracy of the prediction has been significantly improved.

[0053] Example

[0054] The Pima Indian Diabetes Dataset obtained from the UCI Machine Learning Repository consists of 768 female patient samples from Arizona, USA, all aged over 21 years old. The dataset includes nine attribute values, eight of which are related to diabetes diagnosis and one label attribute, '0' represents healthy people and '1' represents patient population. The dataset includes 268 test positive examples and 500 test negative examples.

[0055] Use steps 2, 3, 4, and 5 to process the above data, input the data set into the KMEANS++ algorithm for unsupervised learning (the KMEANS++ algorithm is referenced in 'k-means++: The Advantages of Careful Seeding'), and use the correct clustering results as the input of the next model to enhance the accuracy of model recognition, strengthen the reliability of the model, and make the final classification results more accurate. Then input the data into the SVM (support vector machine) for prediction and evaluate its performance. In order to prevent a certain data from generating prediction errors, the data set is divided into 7-3 and 10-fold cross validation is used, and the average value is taken to evaluate the model. Subsequently, the new diabetes data set is input into the KS model to continuously learn new features and improve the accuracy of the model.

[0056] The pseudo code of a diabetes prediction algorithm based on the KS model proposed in the present invention is shown in Table 1:

[0057] Table 1

[0058]

[0059] In order to verify the performance of the diabetes prediction algorithm based on the KS model, the model proposed in the present invention is compared with five popular machine learning algorithms, namely support vector machine (SVM), logistic regression, naive Bayes, decision tree and K nearest neighbor. The comparison results are shown in Figure 2. Figure 3 In order to prevent the influence of a certain data on the model, the model was cross-validated 10 times. The average area under the ROC curve of different models is shown in Table 2. Figure 3 It can be seen from Table 2 that the performance of the model proposed in the present invention is better than that of other comparison algorithms.

[0060] Table 2

[0061]

[0062] In order to further evaluate the performance of the model proposed in the present invention and prevent the influence of a certain specific data on the model, 7-3 data set partitioning and 10-fold cross validation were adopted to verify the performance of the model on the data set. Figure 4 and Figure 5 They are model 7-3 data set partitioning and 10-fold cross validation ROC curve. Figure 5 In the present invention, the average area under the ROC curve is calculated, and the performance is better than other comparative algorithms. The model proposed in the present invention is a reliable diabetes prediction algorithm.

[0063] Through the above method, the present invention provides a diabetes prediction algorithm based on the KS model, and the data set uses the Pima India diabetes data set; the data set is cleaned and standardized, etc., so that the data quality is higher. The Person correlation coefficient is used to select the data set attributes, remove unimportant attributes, speed up the processing speed of the later model, and improve the performance of the model. The KS model is trained using the data set, and the data is classified and predicted. The model is evaluated using the correlation index, and when verifying its accuracy, it is determined by label control. Experiments have shown that the model proposed in the present invention has better performance than others, indicating that the results of this study are very promising. Its research results are mainly used in the clinical screening and early warning of early diabetes. In the future, it may involve the early diagnosis of other diseases and is widely used in the field of bioinformatics.

Claims

1. A diabetes prediction algorithm based on the KS model, characterized in that: Follow the steps below to implement it: Step 1: Obtain the Pima Indian Diabetes Dataset from the UCI Machine Learning Repository. Step 2: Clean the acquired Pima India diabetes dataset to obtain a normalized dataset; the specific process is: delete the content other than each attribute data, fill the missing values ​​in the dataset, and the specific process of filling the missing values ​​in the dataset is: replace the missing data with the average value of each attribute data; Step 3: Use the quartile analysis method to detect and delete extreme values ​​and outliers in the normalized data set; Step 4: Use the Person correlation coefficient analysis method to select four attributes that are highly correlated with diabetes prediction; the specific process is as follows: Get the data in each attribute X i , Calculate the Person correlation coefficient, the formula is: in, represents the standard score of the data in the attribute, Represents the average value of the data in the attribute, Represents the standard deviation of the data in the attribute; Sort the Person correlation coefficients obtained for each attribute by size, select the four Person correlation coefficients with larger values, and obtain four attributes based on the Person correlation coefficients; Step 5: Standardize the corresponding data of the four attributes respectively; Step 6: Input the standardized data set into the KMEANS++ algorithm for unsupervised learning to obtain the correct clustered data; Step 7: Input the data into SVM for prediction and evaluate its performance.

2. A diabetes prediction algorithm based on the KS model according to claim 1, characterized in that: The Pima India diabetes dataset has nine attribute values, eight of which are related to diabetes diagnosis and one label attribute. In the label attribute, '0' represents a healthy population and '1' represents a patient population.

3. The diabetes prediction algorithm based on the KS model according to claim 1, characterized in that: The specific process of step 5 is as follows: The data in the four attributes are standardized respectively, and the formula is: Among them, X represents the data in the attribute, Represents the average value of the data in the attribute, The standard deviation of the data in the attribute, Represents the data obtained after standardization.

Citation Information

Patent Citations

  • Type diabetes predicting and early warning method based on machine learning

    CN107403072A

  • Diabetes risk factor cause and effect discovery method based on improved function cause and effect likelihood

    CN112233802A