Data grading and classifying method based on artificial intelligence
By introducing concepts such as interpolation step size, direction, deviation value and credibility value in medical data classification, combined with K-nearest neighbor interpolation and convolutional neural network, the problem of inaccurate data classification in the existing technology is solved, and the accuracy and reliability of data classification grading is improved.
Patent Information
- Application Number
- CN202510504357.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The prior art fails to fully consider the dynamic changing characteristics and multi-dimensional information of the data in the classification of medical data, resulting in inaccurate classification results and ineffective support of medical decision-making.
By introducing concepts such as interpolation step size, pointing, deviation value and trustworthiness numerical values, combined with K-nearest neighbor interpolation and convolutional neural network, a data prediction model is built to carry out real-time correction and classification of data.
The prediction performance of the data classification grading model is improved, the deviation caused by data loss is reduced, and the interpolated data is closer to the true value, thereby improving the accuracy and reliability of data classification grading.
Smart Images

Figure CN120030392A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data classification, and in particular to a data grading and classification method based on artificial intelligence. Background Art
[0002] In today's digital age, data is growing explosively, with a wide range of sources, diverse types and large scale. How to efficiently and accurately process and analyze these massive data to mine the valuable information contained in them has become a key issue faced by many fields. Data classification, as an important part of data processing and analysis, aims to divide data into different levels and categories based on factors such as data characteristics, value and risk, so as to achieve refined management, efficient use and effective protection of data. In the medical field, the importance of data is particularly prominent. Medical data covers many aspects such as patients' basic information, symptoms, diagnosis results, treatment plans, drug use, allergy history, etc. These data are not only crucial for patients' disease diagnosis, treatment decisions and prognosis evaluation, but also provide strong support for medical research, public health management and medical policy formulation. However, medical data often has problems such as incompleteness and inaccuracy, among which data missing is more common. For example, due to various reasons, some physiological indicators of patients may not be recorded, or the recorded data may have errors. These missing or inaccurate data will seriously affect the accuracy and reliability of medical data analysis, and thus have an adverse impact on the scientific nature of medical decision-making.
[0003] For example, Chinese patent application No. 202411110376.8 discloses a data classification and grading method based on artificial intelligence, including: collecting original medical data, preprocessing the original medical data with a k-means clustering algorithm to obtain a clustered medical data set; completing the clustered medical data set with a K-nearest neighbor interpolation method to obtain a complete clustered medical data set; using a convolutional neural network and a regularization technology based on a reward function to construct a medical data prediction model; using a complete clustered medical data set to train the medical data prediction model; combining the Bayesian optimization algorithm to perform hyperparameter tuning on the medical data prediction model; inputting the medical data to be classified and graded into the medical data prediction model to obtain a classification and grading result. The present invention adopts global data automatic completion technology, model regularization technology, and hyperparameter space effective retrieval technology, and can perform model training according to preset data classification and grading standards, thereby improving the accuracy of the data classification and grading model.
[0004] In existing patent documents, most existing methods are based on simple rules or statistical features, and fail to fully consider the dynamic characteristics of data and multi-dimensional information. For example, in the classification of medical data, classification based only on a single indicator or a simple threshold cannot fully reflect the complexity of the patient's condition and individual differences. Moreover, with the continuous development of medical technology and the sharp increase in data volume, traditional classification methods are difficult to adapt to new data types and classification requirements, resulting in inaccurate and unreasonable classification results, and unable to provide effective support for medical decision-making. Summary of the invention
[0005] The present application provides a data grading and classification method based on artificial intelligence. By introducing concepts such as interpolation step, direction, deviation value and credibility value, as well as corresponding quantitative relationships, it can more accurately evaluate the reliability of interpolation and make reasonable corrections to real-time data, reduce the deviation caused by missing data, and improve the predictive performance of the data classification and grading model.
[0006] The present application provides a data grading and classification method based on artificial intelligence, including: S101, collect real-time data, use an algorithm to calculate and interpolate missing data in the real-time data, record the interpolation step length, use the algorithm to determine the associated perspective, and set the secondary weight according to the obtained associated perspective; S102, calculating the credibility value according to the interpolation step length, and calculating the deviation value between the real-time data and the interpolation; Collect initial data, form an initial data cluster and a current data cluster according to the initial data and the real-time data respectively, calculate the numerical deviation according to the initial data cluster and the current data cluster, use the numerical deviation to match the transition cluster between the initial data cluster and the current data cluster, and construct the initial data cluster, the transition cluster and the current data cluster into a development chain; Obtain the target area according to the transition cluster, and the target area guides the missing data; S103, calculating a correction coefficient according to the credibility value and the deviation value, and using the correction coefficient to correct the real-time data and the credibility value; S104, establishing a data prediction model based on the corrected real-time data, and using the data prediction model to classify and grade the data.
[0007] Preferably, for each missing data in the data, the distance between each missing data and the sample is calculated, the calculated distances are sorted, and the K samples closest to the missing value are selected as nearest neighbor samples. Based on the obtained K nearest neighbor samples, a first-level weight is assigned to each neighbor point, and the interpolation is calculated using the values of the selected K nearest neighbor samples and the corresponding weights.
[0008] Preferably, the method for calculating the interpolation is: calculating the interpolation by weighted average method, the formula is: ,in, represents interpolation, K represents the value of K nearest neighbor points, represents the value of the i-th neighbor point, represents the first-level weight of the i-th neighbor point. The interpolation step size refers to the absolute difference between the interpolation value and the reference value, that is, s=| -R|, where is interpolation, R is the reference value, the interpolation step reflects the closeness between the interpolation and the reference value, the smallest interpolation step indicates the highest credibility, and the credibility value is calculated. The interpolation direction refers to recording the number of times each reference value is used during the interpolation process.
[0009] Preferably, the calculation formula of the credibility value is: , where C is the credibility value and T is the threshold. When s≤T, the credibility value C increases linearly with the decrease of step size s, and the maximum value is 1; when s>T, C is 0, and the threshold T is dynamically adjusted according to the distribution characteristics of the data set. The deviation value is calculated by the formula ,in, is the deviation value, is interpolation, For real-time data, > When , d is a positive number, indicating that the interpolation value is greater than the true value; when < When , d is negative, indicating that the interpolated value is less than the true value; when = When d=0, it indicates that the interpolation is completely accurate, and the positive or negative value of the deviation directly reflects the offset direction of the interpolation.
[0010] Preferably, the formula for calculating the correction coefficient is: , where k is the correction coefficient, ranging from [0, 1], C is the credibility value, d is the deviation value, indicating the difference between the interpolated value and the true value, and D is the normalization constant used to control the sensitivity of the correction coefficient.
[0011] Preferably, the formula for calculating the secondary weight is: ,in, represents the weight of the i-th associated perspective, n represents the number of selected perspectives, and wk represents the weight corresponding to the k-th factor when it is not normalized.
[0012] Preferably, the data space between the initial data cluster and the current data cluster is divided into multiple intervals, and a deviation threshold and a direction threshold are set. The multiple intervals are set according to the deviation threshold and the direction threshold. According to the deviation and direction of each indicator, the intervals are matched to corresponding transition clusters. The initial data cluster is located at the starting point of the development chain, the current data cluster is located at the end point of the development chain, and the transition clusters are arranged in sequence between the initial data cluster and the current data cluster in chronological order of disease development. According to the determined connection order, the initial data cluster, the matched transition cluster and the current data cluster are connected in sequence to form a development chain.
[0013] Preferably, the target area is obtained according to the transition cluster, and the method for the target area to guide the missing data is: a similarity threshold is set according to the characteristics of the data in the development chain, the Euclidean distance is used to calculate the similarity between the transition cluster and the missing data, and the calculated similarity is compared with the similarity threshold. When the similarity is greater than the similarity threshold, the interval in the transition cluster is the target area for obtaining the missing data, and features are extracted from the target area to identify the development trend of the features, and the development trend is input into the set prediction model. The prediction result output by the prediction model is used to guide the missing data.
[0014] Preferably, the development chain is identified: the first development chain and the second development chain are identified using software, and for each development chain, its first data point is identified as the starting point and the last data point as the end point, and the starting point value of the first development chain is compared with the end point value. If the end point value is greater than the starting point value, it is judged to be in an ascending direction; if the end point value is less than the starting point value, it is judged to be in a descending direction, and it is identified whether the directions of other data points in the first development chain are consistent with the directions of the starting point and the end point. If there are data points with inconsistent directions, these data points that are inconsistent with the directions of the starting point and the end point are removed from the first development chain, and the remaining values in the first development chain are calculated using the credibility value formula to screen out the highest credibility value in the first development chain, and similarly calculate the highest credibility value on the second development chain, and calculate the deviation between the highest credibility value in the first development chain and the highest credibility value in the second development chain.
[0015] Preferably, the weights are allocated according to the calculated deviation and the credibility values of the first development chain and the second development chain. The specific formula is: , ,in, represents the weight of the first development chain, is a constant used to adjust the sensitivity of weight distribution to ensure that the sum of the weights is 1. The set weights are substituted into step S101.
[0016] One or more technical solutions provided in the present application have at least the following technical effects or advantages: by introducing concepts such as interpolation step, direction, deviation value and credibility value, as well as corresponding quantitative relationships, the reliability of interpolation can be evaluated more accurately, and reasonable corrections can be made to real-time data, thereby reducing deviations caused by missing data, improving the prediction performance of data classification and grading models, improving the accuracy of medical data interpolation, and ensuring that the interpolated data is closer to the true value, thereby improving the accuracy and reliability of data classification and grading. By recording directions and limiting the number of references, deviations caused by over-reliance on a single reference value can be prevented. Using more accurate interpolation data to train classification and grading models can improve the prediction performance and generalization ability of the model. By introducing secondary weights and selecting values from different perspectives, when the credibility of the first embodiment is low, it is possible to comprehensively consider data information from more different sources and perspectives, reducing errors caused by individual abnormal data or limitations of a single perspective. For example, when processing medical data, the judgment of a cluster of disease data may be biased due to the abnormality of certain adjacent values. After the introduction of secondary weights, the data obtained from different feature spaces and time series relationships and related perspectives can balance the impact of such abnormalities, obtain more accurate disease assessment values, and avoid a cluster of values being too dependent on a few reference values, which in turn leads to the cluster values dominating the values of a certain type of data during training, thereby improving the accuracy and robustness of data processing and making the trained model more generalizable. By constructing a development chain between the patient's initial data cluster, transition cluster and the cluster where the current diagnosed disease is located, and guided by the development chain, the direction of obtaining missing data can be determined more accurately according to the actual dynamics of the patient's disease development, so that the supplementary data is more in line with the individual situation of the patient, and the direction of obtaining missing values (missing data) can be guided, thereby improving the patient's data set, improving the accuracy of disease analysis and the rationality of treatment plan formulation, and obtaining a complete patient data set. It can more comprehensively reflect the development process and characteristics of the patient's disease, provide more reliable data support for subsequent disease diagnosis, treatment and prognosis evaluation, and the supplementary data is more in line with the individual situation of the patient and the actual needs of disease development, making the grading and classification of data more accurate; Accurately identify the impact of interpolation calculations of development chains in different directions, make interpolation calculations more accurate, improve the efficiency of data classification, and obtain more accurate and reliable analysis results by reasonably allocating weights and comprehensively considering the information of development chains in different directions. Obtain a final result that comprehensively considers the information of data clusters in different directions, which more accurately reflects the actual development trend of data or events. Weights are allocated according to the deviation and credibility of the values, making the impact of data clusters in different directions on the final result more reasonable and accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1A schematic diagram of a process of a data grading and classification method based on artificial intelligence of the present invention; Figure 2 A schematic diagram of a process for setting secondary weights according to an embodiment of the present invention; Figure 3 A schematic diagram of a process for constructing a development chain according to an embodiment of the present invention. DETAILED DESCRIPTION
[0018] To facilitate the understanding of the present invention, the present application will be described more comprehensively below with reference to the relevant drawings; the drawings show preferred embodiments of the present invention, but the present invention can be implemented in many different forms and is not limited to the embodiments described herein; on the contrary, the purpose of providing these embodiments is to enable a more thorough and comprehensive understanding of the disclosed content of the present invention.
[0019] It should be noted that the terms “vertical”, “horizontal”, “up”, “down”, “left”, “right” and similar expressions used in this document are only for illustrative purposes and do not represent the only implementation method.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by technicians in the technical field to which the present invention belongs; the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention; the term "and / or" used herein includes any and all combinations of one or more related listed items.
[0021] Embodiment 1: Figure 1 The following is a flow chart of a data classification method based on artificial intelligence according to an embodiment of the present invention, including: S101, collecting data, interpolating missing data in the collected data using K-nearest neighbor interpolation, and recording the interpolation step length during the interpolation calculation process; Specifically, medical data is collected from the hospital's database, electronic health record (EHR) system, laboratory test results, and imaging diagnostic reports. The hospital's database usually includes the patient's basic information, such as name, age, gender, contact information, etc., and the patient's medical records, including visit time, department, doctor, diagnosis results, etc.; the electronic health record (EHR) system includes symptoms, signs, diagnosis, treatment plan, medication use, allergy history, etc. The EHR system is usually connected to various departments and divisions of the hospital and can update the patient's medical information in real time; laboratory test results include laboratory test results (such as blood routine, biochemical indicators, etc.) and imaging test results (such as X-ray, CT, MRI, etc.).
[0022] In the K-nearest neighbor imputation method, the value of K represents the number of nearest neighbor samples used for imputing missing values. Cross-validation is used to determine the value of K. The dataset is randomly divided into a training set and a validation set. For different values of K (e.g., from 1 to a certain maximum value), K-nearest neighbor imputation is performed on the training set, and the accuracy of the imputation results is evaluated on the validation set (e.g., using metrics such as mean squared error, mean absolute error, etc.). The value of K that performs best on the validation set is selected as the optimal K value. For each missing value in the data, the distance between each missing value and other samples is calculated. The Euclidean distance is used to calculate the distance between the missing value and the samples. The calculated distances are sorted, and the K samples closest to the missing value are selected as the nearest neighbor samples. According to the obtained K nearest neighbor samples, a first-level weight is assigned to each neighbor point. The magnitude of the weight reflects the influence degree of the neighbor point on the interpolation. Using the values of the selected K nearest neighbor samples and the corresponding weights, interpolation is calculated by the weighted average method. The formula is: , where, represents the interpolation, K represents the values of K neighbor points, represents the value of the i-th neighbor point, represents the first-level weight of the i-th neighbor point. The interpolation step size refers to the magnitude of the difference between the interpolation and a certain reference value during the calculation of interpolation. The interpolation direction refers to recording the number of times each reference value (i.e., the value of the nearest neighbor sample) is used during the imputation process. If a certain reference value is used too many times, its weight is reduced to improve the accuracy and stability of the imputation results.
[0023] S102. Calculate the credibility value according to the recorded interpolation step size, obtain real-time data, and calculate the deviation value between the real-time data and the interpolation; Furthermore, the interpolation step size (s) is the absolute difference between the interpolation and the reference value, i.e., s = ∣ - R∣, where is the interpolation and R is the reference value. The interpolation step size reflects the proximity between the interpolation and the reference value. A smaller step size means that the interpolation is more likely to be close to the true value and thus has a higher credibility. The formula for the credibility value is:
[0024] where C is the credibility value, T is the threshold. When s ≤ T, the credibility value C increases linearly with the decrease of the step size s, and the maximum value is 1 (when s = 0); when s > T, C is forced to be set to 0, indicating that the interpolation is not credible because it deviates too much from the reference value. The threshold T is dynamically adjusted according to the distribution characteristics of the dataset. The deviation value is calculated by the formula , where, is the deviation value, is the interpolation, is the real-time data. When > When , d is a positive number, indicating that the interpolation overestimates the true value; when < When , d is negative, indicating that the interpolation underestimates the true value; when = When d=0, it indicates that the interpolation is completely accurate. The positive or negative sign of the deviation value directly reflects the offset direction of the interpolation. A positive deviation indicates an overestimation, and a negative deviation indicates an underestimation. The absolute value of the deviation value indicates the degree of difference between the interpolation and the real-time data. The larger the absolute value, the more significant the difference.
[0025] Specific implementation example: There is a medical data set that contains missing values for a physiological indicator of a patient (such as blood sugar level). In order to fill these missing values, we use the K-nearest neighbor interpolation method (K=3), that is, select the values of the 3 known samples closest to each missing value for interpolation. The following is the specific calculation process, including the calculation of the credibility value and the deviation value. The blood sugar level data of patient ID 101 is missing, and the nearest neighbor samples are: sample A blood sugar value 5.8 mmol / L, distance 0.2, sample B blood sugar value 6.2 mmol / L, distance 0.3, sample C blood sugar value 6.0 mmol / L, distance 0.5, , total weight W=25+11.11+4=40.11, , using the average of the nearest neighbor samples as the reference value: , calculate the interpolation step: s = | 5.93-6.0 | = 0.07mmol / L, calculate the credibility value: set the threshold T = 0.2mmol / L, Through actual measurement, the blood glucose value of patient 101 was 6.1 mmol / L, and the calculated deviation value was: d=5.93-6.1=-0.17mmol / L.
[0026] S103, calculating a correction coefficient according to the calculated credibility value and the deviation value, and using the calculated correction coefficient to correct the real-time data and the credibility value; Specifically, the correction coefficient is calculated according to the reliability value and the deviation value. The formula for calculating the correction coefficient is:
[0027] Where k is the correction coefficient, which ranges from [0, 1], but can be extended to a reasonable range by adjusting D. C is the credibility value, which reflects the initial reliability of the interpolation. d is the deviation value, which indicates the difference between the interpolation and the true value. D is the normalization constant, which is used to control the sensitivity of the correction coefficient. , so that the maximum value of k does not exceed 1. When |d| is large (significant deviation), the denominator increases and k decreases, but the weight effect of C makes k still inversely proportional to the deviation; when C is small (low credibility), even if |d| is small, k will increase due to the amplification effect of C, strengthening the correction of low credibility interpolation. The introduction of D avoids k due to |d| being too small or too large, ensuring that the correction amplitude is controllable.
[0028] Use the calculated correction coefficient to correct the real-time data. The formula is: ,in, is the corrected real-time data. When k is large, Closer to interpolation, reducing the impact of deviation, when k is small, retaining more original values of real-time data; using the correction coefficient to correct the credibility value, the formula is: ,in, is the corrected credibility value. If k is large (such as k=0.5), the corrected =0.5C, reduce the weight of low-confidence interpolation, if k is small (such as k=0.1), after correction ≈0.9C, preserving the stability of high-confidence interpolation.
[0029] S104, establishing a data prediction model based on the corrected real-time data, and using the data prediction model to classify and grade the data; Specifically, a convolutional neural network (CNN) is used as the basic architecture of the model, and a regularization technology based on a reward function is adopted to prevent the model from overfitting. The reward function gives a higher reward when the difference between the model prediction result and the true value is small, and at the same time imposes a certain penalty on the model complexity, so that the model will not be too complex and lose its generalization ability while pursuing accurate prediction. The interpolated and corrected data is used as the input of the data prediction model, and the data prediction model outputs the prediction result; the corrected data set is divided into a training set and a validation set in a ratio of (7:3). The training set is used for model parameter learning, and the validation set is used in the training process. The performance of the model is evaluated in order to adjust the training strategy in time. The model is iteratively trained using the training set. In each iteration, the training data is input into the model, and the predicted output of the model is calculated through forward propagation. Then, the loss function value between the predicted output and the true label is calculated. Next, the backpropagation algorithm is used to calculate the gradient of the loss function with respect to the model parameters, and the parameters of the model are updated according to the optimization algorithm (stochastic gradient descent) to minimize the loss function. During the training process, the model is regularly evaluated using the validation set, and the validation set data is input into the trained model to calculate the performance indicators of the model on the validation set, such as accuracy, loss value, etc. By observing the performance changes on the validation set, we can determine whether the model is overfitting or underfitting. If overfitting occurs, we can consider adjusting the complexity of the model, increasing the regularization strength, or reducing the number of training rounds. If underfitting occurs, we can increase the complexity of the model, increase the training data, or adjust the learning rate. The actual medical data is input into the model, and the model classifies and grades the data according to the learned features and patterns. For example, the severity of the patient's condition is graded into different levels such as mild, moderate, and severe. Diseases are classified to determine which disease the patient has.
[0030] The technical solutions in the above-mentioned embodiments of the present application have at least the following technical effects or advantages: by introducing concepts such as interpolation step, direction, deviation value and credibility value, as well as corresponding quantitative relationships, it is possible to more accurately evaluate the reliability of interpolation and make reasonable corrections to real-time data, thereby reducing deviations caused by missing data, improving the predictive performance of data classification and grading models, improving the accuracy of medical data interpolation, ensuring that the interpolated data is closer to the true value, thereby improving the accuracy and reliability of data classification and grading, and preventing deviations caused by over-reliance on a single reference value by recording directions and limiting the number of references. Using more accurate interpolated data to train classification and grading models can improve the predictive performance and generalization ability of the model.
[0031] Embodiment 2: Based on the method of taking values of adjacent points and setting different weights in Embodiment 1, i.e., the first weight, this embodiment finds values from different perspectives, and sets weights for the values found from different perspectives, i.e., the second-level weight. Not only the adjacent values are considered, when the credibility calculated in Embodiment 1 is low, the second-level weight is entered, such as Figure 2 shown.
[0032] S201, discretize the collected data and use an algorithm to determine the associated perspective; Furthermore, for the collected continuous data, the intervals are divided according to data distribution and business logic, and the patient age is divided into 0-18 years old (children), 19-40 years old (youth), 41-60 years old (middle-aged), and 61 years old and above (elderly); the Apriori algorithm is used to determine the association perspective. First, the minimum support and minimum confidence are set. The minimum support refers to the lowest frequency of an item set appearing in all transactions, which reflects the prevalence of the item set. In medical data analysis, if there are a total of 1,000 patient records, then an item set must appear in at least 100 records to meet the requirements. The minimum confidence refers to the credibility of an association rule, which means the probability of the result item set appearing when the premise item set appears. The minimum confidence is set to 0.6. Only when the confidence of an association rule reaches 60%, it is considered valid. Scan the entire data set, count the frequency of each single item, find the single item that meets the minimum support, and form a frequent item set. In medical data, statistics show that "smoking" appears in 300 records. records, "drinking" appears in 200 records, and "lung cancer" appears in 150 records. If the minimum support is 0.1, then "smoking" and "drinking" will enter the frequent item set, while "lung cancer" may be eliminated because it does not meet the support requirement. Use the frequent item set to generate the candidate item set, scan the data set again, count the occurrence frequency of the candidate item set, find the frequent item set that meets the minimum support, and so on, and continuously iterate to generate higher-order frequent item sets until no new frequent item sets can be generated. For example, by combining the items in the frequent item set, generate the "smoking and drinking" waiting option set, and then count its occurrence frequency to determine whether it becomes a frequent item set, and generate association rules from the frequent item set. For each frequent item set, its non-empty true subset is used as the premise item set, and the remaining part is used as the result item set to form an association rule. For example, for the frequent item set "smoking and drinking", the association rules "smoking→drinking", "drinking→smoking", and "smoking and drinking→lung cancer" can be generated. For each generated association rule, its confidence is calculated. The confidence calculation formula is: , for the association rule "smoking and drinking → lung cancer", if the support of "smoking and drinking and lung cancer" is 0.1 and the support of "smoking and drinking" is 0.15, then the confidence of the rule is ≈0.67, according to the set minimum confidence, filter out the valid association rules, and only retain the rules whose confidence is greater than or equal to the minimum confidence.
[0033] Analyze the selected association rules to find out the rules that are strongly associated with the target data. Strong associations are usually manifested as rules with high support and confidence. When analyzing the factors related to lung cancer, it is found that the support of the association rule "smoking and drinking → lung cancer" is 0.1 and the confidence is 0.6, indicating that there is a strong association between smoking and drinking and lung cancer. According to the strong association rules, extract the factors that are strongly associated with the target data. For example, extract the factor "smoking and drinking" from the rule "smoking and drinking → lung cancer", and use the factors that are strongly associated with the target data as the association perspective. "Smoking and drinking" is the association perspective related to "lung cancer". Based on the determined association perspective, select the relevant indicators as the new value selection perspective. For example, if the association perspective is other diseases related to a certain disease, then select the indicators of these related diseases (such as incidence, cure rate, etc.) as the new value selection perspective.
[0034] S202, setting a secondary weight for the obtained new selected value perspective; Specifically, we select the analytic hierarchy process to determine the weights according to the characteristics of the data, and construct a hierarchical model. The target layer is the accuracy of disease prediction, which is the core purpose of the entire analysis. The middle layer is a new value selection perspective, such as the "smoking and drinking" perspective and the gene module perspective. In medical research, smoking and drinking are known to be important factors that may affect the occurrence of diseases, and the gene module can provide information from a genetic level. The bottom layer is the indicators under each perspective, such as the amount of smoking and drinking frequency under the "smoking and drinking" perspective, and the gene expression under the gene module perspective. These indicators are specific and quantifiable data, which are used to support the value selection perspective of the middle layer and construct a judgment matrix. Medical experts are invited to compare the value selection perspectives of the middle layer in pairs. Assume that there are n value selection perspectives in the middle layer, and the constructed judgment matrix A= ,in It indicates the importance of the i-th selected value perspective relative to the j-th selected value perspective. Its value selection rule is as follows: when i=j, =1, indicating that the same selected perspective is equally important compared to itself; when i When j is selected, the value is assigned according to the expert's judgment. For example, if the expert thinks that the i-th selected perspective is slightly more important, obviously more important, or strongly more important than the j-th selected perspective, the values are assigned as 3, 5, 7, etc., respectively. Otherwise, the values are assigned as 1 / 3, 1 / 5, 1 / 7, etc. Experts make judgments based on their rich clinical experience and professional knowledge and construct a judgment matrix. The elements in the matrix reflect the relative importance of each selected perspective. The judgment matrix is normalized, and then the eigenvector and eigenvalue are calculated. The elements in the eigenvector correspond to the weights of each selected perspective. The weight formula for calculating the selected perspective is: ,in, represents the weight of the i-th value selection perspective, n represents the number of value selection perspectives, and wk represents the weight corresponding to the k-th factor when it is not normalized.
[0035] The technical solutions in the above-mentioned embodiments of the present application have at least the following technical effects or advantages: by introducing secondary weights and selecting values from different perspectives, when the credibility of the first embodiment is low, it is possible to comprehensively consider data information from more different sources and perspectives, and reduce errors caused by individual abnormal data or limitations of a single perspective. For example, when processing medical data, the judgment of a cluster of disease data may be biased due to the abnormality of certain adjacent values. After the introduction of secondary weights, the data obtained from different feature spaces and time series relationships and related perspectives can balance the impact of such abnormalities, obtain more accurate disease assessment values, avoid a cluster of values being too dependent on a few reference values, and then cause the cluster values to dominate the values of a certain type of data during training, improve the accuracy and robustness of data processing, and make the trained model more generalizable.
[0036] Embodiment 3: Based on the first and second embodiments, obtaining values for adjacent points and obtaining values from multiple perspectives only obtains values from point pairs, which cannot accurately reflect the process of disease development of patients. This embodiment constructs a development chain through initial data clusters, transition clusters and current data clusters for data of patients in different time periods, and can more accurately determine the direction of obtaining missing data, so that the supplementary data is more in line with the individual situation of the patient, such as Figure 3 shown.
[0037] S301, collecting initial data, forming an initial data cluster and a current data cluster according to the initial data and the real-time data, respectively, calculating a numerical deviation amount according to the initial data cluster and the current data cluster, and determining a numerical deviation direction; Specifically, initial data is collected, and the initial data is the basic physiological indicators, symptom manifestations, past medical history and genetic test results of the patient when he is just diagnosed or first admitted to the hospital. An initial data cluster is formed based on the above initial data, and a current data cluster is formed based on the collected real-time data. Key indicators are selected from the initial data cluster and the current data cluster for comparison. The selection of key indicators is based on the characteristics of the disease, treatment needs and the purpose of data analysis. The key indicators include physiological indicators and symptom levels. The change of each indicator from the initial state to the current state is calculated based on the physiological indicators and the symptom level. The calculation formula of the numerical deviation is: ,in, is the corresponding indicator value in the current data cluster, is the corresponding index value in the initial data cluster, is the data deviation. When Δx > 0, it is a positive deviation, indicating that the index value increases, that is, the patient's physiological state or symptom severity worsens during the development of the disease; when Δx < 0, it is a negative deviation, indicating that the index value decreases, that is, the patient's physiological state or symptom severity improves during the development of the disease; when Δx = 0, it is no deviation, indicating that the index value does not change, that is, the patient's physiological state or symptom severity remains stable during the development of the disease; Determine the direction of the deviation: when the deviations of multiple key indicators are all positive, it means that the patient's disease state has an overall tendency to worsen, and more active treatment measures need to be taken; when the deviations of multiple key indicators are all negative, it means that the patient's disease state has an overall tendency to improve, and the current treatment plan can be continued or appropriately adjusted; when the deviations of multiple key indicators are close to zero or the positive and negative deviations offset each other, it means that the patient's disease state remains stable overall, and the current treatment plan can be maintained and regularly monitored.
[0038] S302, matching the initial data cluster with the current data cluster according to the numerical deviation amount and the deviation direction, and constructing a development chain based on the initial data cluster, the transition cluster and the current data cluster; Furthermore, a deviation threshold and a direction threshold are set according to the importance and variation range of the indicator. The deviation threshold is used to divide the variation degree of the indicator value, and the direction threshold is used to judge the variation direction of the indicator value, i.e., increase, decrease or no change. The data space between the initial data cluster and the current data cluster is divided into a plurality of intervals. The plurality of intervals can be set according to the deviation threshold and the direction threshold. The data space is divided into seven intervals, i.e., slight increase, moderate increase, significant increase, slight decrease, moderate decrease, significant decrease and no change. According to the deviation and direction of each indicator, it is matched to the corresponding transition cluster, and the relationship between the deviation and the threshold of the indicator value is compared to achieve If the deviation of an indicator is less than the deviation threshold and is in the positive direction, it will be matched to the slightly increased transition cluster; if the deviation of an indicator is greater than the deviation threshold and is in the negative direction, it will be matched to the significantly reduced transition cluster; based on the development time and logical order of the patient's disease, the connection order between each data cluster and the transition cluster is determined. The initial data cluster should be located at the starting point of the development chain, the current data cluster should be located at the end point of the development chain, and the transition cluster is arranged in sequence between the two according to the chronological order of the disease development. According to the determined connection order, the initial data cluster, the matched transition cluster and the current data cluster are connected in sequence to form a development chain that reflects the development process of the patient's disease.
[0039] S303, based on the obtained transition cluster, calculate the similarity between the transition cluster and the missing data, when the similarity is greater than a set threshold, the area near the transition cluster is the target area for obtaining the missing data, and the missing data is guided according to the obtained target area; Specifically, a similarity threshold is set according to the characteristics of the data in the development chain. The similarity threshold is used to determine the degree of similarity between the transition cluster and the missing features of the missing data. The Euclidean distance is used to calculate the similarity between the transition cluster and the missing data, and the calculated similarity is compared with the similarity threshold. When the similarity is greater than the similarity threshold, the vicinity of the transition cluster is the target area for obtaining the missing data. The target area is marked in the development chain, and the missing data is guided according to the obtained target area. Features related to the missing data, such as numerical trends, change patterns, etc., are extracted from the target area. The extracted features are matched with known features of the missing data (such as features of adjacent data points) to find the most similar feature combination. According to the matching results, the values or patterns corresponding to similar features in the target area are mapped to the missing data as a reference for interpolation. The development trend of the data in the target area is analyzed, such as rising, falling or stable, and the possible values or change patterns of the missing data are predicted according to the trend analysis results. The prediction results are used as a guide for the missing data to help it interpolate or correct it.
[0040] The technical solutions in the above-mentioned embodiments of the present application have at least the following technical effects or advantages: by constructing a development chain between the patient's initial data cluster, the transition cluster and the cluster where the currently diagnosed disease is located, and through the guidance of the development chain, it is possible to more accurately determine the direction of obtaining missing data according to the actual dynamics of the patient's disease development, so that the supplemented data is more in line with the individual situation of the patient, and guide the direction of obtaining missing values (missing data), thereby improving the patient's data set, improving the accuracy of disease analysis and the rationality of treatment plan formulation, and obtaining a complete patient data set, which can more comprehensively reflect the development process and characteristics of the patient's disease, and provide more reliable data support for subsequent disease diagnosis, treatment and prognosis evaluation. The supplemented data is more in line with the individual situation of the patient and the actual needs of disease development, making the grading and classification of data more accurate.
[0041] Embodiment 4: Based on the development chain in the above embodiment 3, if there are several development chains in different directions, the data clusters in each development chain are adjacent, resulting in inaccurate interpolation calculation. This embodiment solves the data processing problem generated when the data clusters in the development chains in different directions are adjacent, so that the interpolation calculation is more accurate.
[0042] The development chain constructed in step S302 is identified, and Python is used to identify the first development chain and the second development chain. For each development chain, its first data point is identified as the starting point and the last data point as the end point. The starting value of the first development chain is compared with the end point value. If the end point value is greater than the starting value, it is judged to be in an ascending direction; if the end point value is less than the starting value, it is judged to be in a descending direction. The judged direction is recorded, and it is identified whether the directions of other data points in the first development chain are consistent with the directions of the starting point and the end point. If data points with inconsistent directions are found, these data points inconsistent with the directions of the starting point and the end point are removed from the first development chain. The remaining values in the first development chain are calculated using the credibility formula in step S102, and the highest credibility value in the first development chain is screened out. Similarly, the highest credibility value on the second development chain is calculated, and the deviation between the highest credibility value in the first development chain and the highest credibility value on the second development chain is calculated. The deviation is calculated using the deviation calculation formula in step S301. The weight is allocated according to the calculated deviation and the credibility of the first development chain and the second development chain. The specific formula is: , ,in, represents the weight of the first development chain, is a constant used to adjust the sensitivity of weight distribution, ensure that the sum of the weights is 1, and avoid calculation errors when the denominator is zero. The set weights are substituted into step S101.
[0043] The technical solution in the above-mentioned embodiment of the present application has at least the following technical effects or advantages: accurately identifying the influence of interpolation calculation of development chains in different directions, making the calculation of interpolation more accurate, improving the efficiency of data classification, and obtaining more accurate and reliable analysis results by reasonably allocating weights and comprehensively considering the information of development chains in different directions. A final result that comprehensively considers the information of data clusters in different directions is obtained, and the result more accurately reflects the actual development trend of data or events. Weights are allocated according to the deviation and credibility of the values, so that the influence of data clusters in different directions on the final result is more reasonable and accurate.
[0044] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A data classification method based on artificial intelligence, characterized in that: include: S101, collect real-time data, use an algorithm to calculate and interpolate missing data in the real-time data, record the interpolation step length, use the algorithm to determine the associated perspective, and set the secondary weight according to the obtained associated perspective; S102, calculating the credibility value according to the interpolation step length, and calculating the deviation value between the real-time data and the interpolation; Collect initial data, form an initial data cluster and a current data cluster according to the initial data and the real-time data respectively, calculate the numerical deviation according to the initial data cluster and the current data cluster, use the numerical deviation to match the transition cluster between the initial data cluster and the current data cluster, and construct the initial data cluster, the transition cluster and the current data cluster into a development chain; Obtain the target area according to the transition cluster, and the target area guides the missing data; S103, calculating a correction coefficient according to the credibility value and the deviation value, and using the correction coefficient to correct the real-time data and the credibility value; S104, establishing a data prediction model based on the corrected real-time data, and using the data prediction model to classify and grade the data.
2. The data classification method based on artificial intelligence as claimed in claim 1, characterized in that: For each missing data in the data, calculate the distance between each missing data and the sample, sort the calculated distances, select the K samples closest to the missing value as the nearest neighbor samples, assign a first-level weight to each neighbor point based on the obtained K nearest neighbor samples, and use the values of the selected K nearest neighbor samples and the corresponding weights to calculate the interpolation.
3. The data classification method based on artificial intelligence as claimed in claim 2, characterized in that: The method for calculating interpolation is: to calculate interpolation by weighted average method, the formula is: ,in, represents interpolation, K represents the value of K nearest neighbor points, represents the value of the i-th neighbor point, represents the first-level weight of the i-th neighbor point. The interpolation step size refers to the absolute difference between the interpolation value and the reference value, that is, s=| -R|, where is interpolation, R is the reference value, the interpolation step reflects the closeness between the interpolation and the reference value, the smallest interpolation step indicates the highest credibility, and the credibility value is calculated. The interpolation direction refers to recording the number of times each reference value is used during the interpolation process.
4. The data classification method based on artificial intelligence as claimed in claim 3, characterized in that: The calculation formula of the credibility value is: , where C is the credibility value and T is the threshold. When s≤T, the credibility value C increases linearly with the decrease of step size s, and the maximum value is 1; when s>T, C is 0, and the threshold T is dynamically adjusted according to the distribution characteristics of the data set. The deviation value is calculated by the formula ,in, is the deviation value, is interpolation, For real-time data, > When , d is a positive number, indicating that the interpolation value is greater than the true value; when < When , d is negative, indicating that the interpolated value is less than the true value; when = When d=0, it indicates that the interpolation is completely accurate, and the positive or negative value of the deviation directly reflects the offset direction of the interpolation.
5. The data classification method based on artificial intelligence as claimed in claim 1, characterized in that: The formula for calculating the correction coefficient is: , where k is the correction coefficient, ranging from [0, 1], C is the credibility value, d is the deviation value, indicating the difference between the interpolated value and the true value, and D is the normalization constant used to control the sensitivity of the correction coefficient.
6. The data classification method based on artificial intelligence as claimed in claim 1, characterized in that: The formula for calculating the secondary weight is: ,in, represents the weight of the i-th associated perspective, n represents the number of selected perspectives, and wk represents the weight corresponding to the k-th factor when it is not normalized.
7. The data classification method based on artificial intelligence as claimed in claim 1, characterized in that: The data space between the initial data cluster and the current data cluster is divided into multiple intervals, and a deviation threshold and a direction threshold are set. The multiple intervals are set according to the deviation threshold and the direction threshold. According to the deviation and direction of each indicator, the intervals are matched to corresponding transition clusters. The initial data cluster is located at the starting point of the development chain, the current data cluster is located at the end point of the development chain, and the transition clusters are arranged in sequence between the initial data cluster and the current data cluster according to the chronological order of disease development. According to the determined connection order, the initial data cluster, the matched transition cluster and the current data cluster are connected in sequence to form a development chain.
8. The data classification method based on artificial intelligence as claimed in claim 1, characterized in that: The target area is obtained according to the transition cluster. The method for the target area to guide the missing data is as follows: a similarity threshold is set according to the characteristics of the data in the development chain, the similarity between the transition cluster and the missing data is calculated using the Euclidean distance, and the calculated similarity is compared with the similarity threshold. When the similarity is greater than the similarity threshold, the interval in the transition cluster is the target area for obtaining the missing data, and features are extracted from the target area to identify the development trend of the features. The development trend is input into the set prediction model, and the prediction result output by the prediction model is used to guide the missing data.
9. The data classification method based on artificial intelligence as claimed in claim 7, characterized in that: Identify the development chain: use the software to identify the first development chain and the second development chain. For each development chain, identify its first data point as the starting point and the last data point as the end point. Compare the starting value of the first development chain with the end point value. If the end point value is greater than the starting value, it is judged to be in an ascending direction; if the end point value is less than the starting value, it is judged to be in a descending direction. Identify whether the directions of other data points in the first development chain are consistent with the directions of the starting point and the end point. If there are data points with inconsistent directions, remove these data points that are inconsistent with the starting point and the end point from the first development chain. Use the credibility value formula to calculate the remaining values in the first development chain, screen out the most credible value in the first development chain, and similarly calculate the most credible value on the second development chain. Calculate the deviation between the most credible value in the first development chain and the most credible value on the second development chain.
10. The data classification method based on artificial intelligence as claimed in claim 9, characterized in that: The weights are allocated according to the calculated deviation and the credibility values of the first development chain and the second development chain. The specific formula is: , ,in, represents the weight of the first development chain, is a constant used to adjust the sensitivity of weight distribution to ensure that the sum of the weights is 1. The set weights are substituted into step S101.
Citation Information
Patent Citations
Improved algorithm for missing value interpolation
CN108197079A
Multi-view data missing completion method for multi-manifold regularization non-negative matrix factorization
CN111368254A
Metabonomics data processing method, device and equipment and readable storage medium
CN118212994A
Data classification and grading method based on artificial intelligence
CN118626918A
Method of handling a missing value in a data mining system
IN201941023805A
Cited By
Photovoltaic module energy efficiency data online processing method based on cloud computing
CN120470246A