Partially ordered data classification method and system based on order relevancy
By calculating the order correlation degree to identify ordered features and using monotonic neural network optimization classification method, the classification problem of some ordered data sets is solved, and the classification accuracy and robustness of the model are improved.
Patent Information
- Application Number
- CN202510342693.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-01
AI Technical Summary
When processing part of the ordered data set, the existing data classification method fails to fully consider the order relationship of ordered features, resulting in a degradation of model performance and it is difficult to accurately identify and utilize ordered features to improve classification accuracy and robustness.
By calculating the order correlation degree (OC), and using monotonic neural network (MNN) combined with disordered feature weighting, partial orderly neural network (PONN) is designed, and the punishment terms are added during the training process to maintain the monotonic constraints of ordered features to optimize classification performance.
High-precision classification on some ordered data sets is realized, which improves the classification accuracy and robustness of the model, and is better than traditional methods.
Smart Images

Figure CN120234671A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of research on the classification of partially ordered data, and specifically relates to a method and system for classifying partially ordered data based on order correlation degree. Background Art
[0002] The complexity of a data set is usually reflected in the diversity of the data. The data not only has differences in format, but also may have significant differences in characteristics under different application scenarios. There are both continuous variables and discrete variables, and even in some cases, features with specific sequence relationships are included. These complex feature relationships often make data processing more difficult, especially when performing tasks such as classification, clustering, and regression. In the classification task of machine learning, it is usually necessary to assign the samples in the data set to predefined categories according to different features.
[0003] General classification methods usually adopt a unified way to process all features in the data set, ignoring the significant differences that may exist between different features. When dealing with the features in a partially ordered data set, such methods do not consider the order relationship between features, which may lead to a decline in model performance. The order relationship refers to a potential sequential association between features and categories, that is, there is a clear order or hierarchical relationship between the values of certain features and the categories. Such features are usually called ordered features. On the contrary, features for which no such order relationship is presented between the category and the feature are called unordered features. In real-world applications, many data sets are often partially ordered data sets, that is, they contain both ordered features and unordered features. For example, in medical diagnosis, biomarker such as blood glucose level and blood pressure usually show an obvious order relationship with the severity of the disease; in educational data analysis, there is often a certain sequence between students' test scores and academic performance levels; in financial risk management, the relationship between credit scores and default risks often has an obvious order. The identification and reasonable utilization of these features are of great significance for improving the performance and interpretability of classification models.
[0004] However, in traditional data classification methods, the particularity of ordered features is usually not fully considered. Many classic classification algorithms, such as support vector machines (SVM), decision trees, and K-nearest neighbors (KNN), often assume that all features are independent and equal, and take the same approach to different types of features. This "one-size-fits-all" approach ignores the potential information contained in ordered features, resulting in the model not being able to achieve the best effect in classification tasks. Ordered features, as an important bridge between data and categories, can provide classification models with more discriminative and explanatory feature information. How to effectively identify these ordered features in partially ordered data sets and make full use of their characteristics for optimization is an important problem that needs to be solved in the current field of data mining and machine learning. In order to effectively solve this problem, researchers have proposed different solutions. For example, some studies try to identify ordered features by improving feature selection methods, while others design new learning algorithms, especially deep learning algorithms, to automatically discover the order relationship between features. However, there are still relatively few specialized studies on ordered features, and most studies focus on data preprocessing, feature selection, or the design of deep learning models, lacking research on in-depth mining and application of ordered features themselves.
[0005] The identification and utilization of ordered features have broad significance in practical applications. In the medical field, physiological characteristics of patients such as blood sugar level, body temperature, and blood pressure can often reflect the severity and development trend of the disease. If these ordered features can be accurately identified and utilized, it can help doctors make more accurate diagnoses and provide patients with personalized treatment plans. Therefore, how to effectively identify and utilize ordered features from partially ordered data sets is an important direction for improving the performance and interpretability of classification models. At present, although some studies have explored the processing methods of ordered features, there are still certain challenges. First, how to accurately identify ordered features from complex data sets and effectively distinguish the relationship between them and disordered features is a key issue. Secondly, how to make full use of these ordered features in the classification model to improve the classification accuracy and ensure the robustness of the model is also a difficult problem faced by current technology. Summary of the invention
[0006] In order to accurately identify ordered features from complex data sets and solve the problem of classifying ordered data, the present invention provides a method and system for classifying partially ordered data based on order correlation. The method identifies ordered features and their order directions, and while maintaining the intrinsic order relationship of ordered features, considers the influence of disordered features. Finally, a monotonic neural network is used to combine ordered features and disordered features to achieve classification of partially ordered data.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] A method for classifying partially ordered data based on order correlation includes the following steps:
[0009] S1. Select a data set. For the features in the data set, calculate the order correlation (OC) to distinguish ordered and unordered features and determine the order direction;
[0010] Further, the specific operation of step S1 is as follows:
[0011] S1.1. Determine the data set U = {x i , y i}, and set the threshold θ for judging ordered features to 0.2;
[0012] S1.2. Read the data set {x i , y i};
[0013] S1.3. Respectively set 3 empty lists A m+ , A m- , A n to store positive-order features, reverse-order features, and unordered features respectively;
[0014] S1.4. Use a i to represent the feature to be recognized, represents the remaining features except a i ;
[0015] S1.5. Judge whether there is collinearity between the feature a i and the features in . When there is collinearity, remove the collinearity, and then calculate the residual set of the feature a i and the variable Y, and calculate the order correlation OC value of the feature;
[0016] S1.5.1. Use the Pearson correlation coefficient to judge whether there is collinearity between the feature a i and the features in . The calculation formula of the Pearson correlation coefficient is as follows:
[0017]
[0018] In the formula, X i and Y i respectively represent the i-th observation values of two variables; respectively represent the means of variables X and Y; n is the sample size;
[0019] S1.5.2. Use the following method to remove collinearity:
[0020] Taking the feature ai Taking a as the independent variable and A as the dependent variable, a linear regression model is constructed:
[0021] A = β1·a i + e1 # (2)
[0022] In the formula, β1 is the regression coefficient, which is used to describe the linear contribution of feature a i to the dependent variable A; e1 is the residual term, indicating the part in the dependent variable A that cannot be explained by feature a i ;
[0023] The residual e1 is calculated through the regression equation:
[0024] e1 = A - β1·a i # (3)
[0025] The residual e1 represents the independent part remaining in the dependent variable A after removing the linear influence of feature a i ;
[0026] Replacing the dependent variable A with the residual e1 eliminates the collinearity between feature a i and the dependent variable A:
[0027] A = e1 # (4)
[0028] S1.5.3. Calculate the residual set of feature a i and variable Y; Using the feature respectively perform regression on feature a i and variable Y:
[0029]
[0030] In the formula, β i represents the weight coefficient of the feature i when performing regression on feature a ; β j represents the weight coefficient of the feature when performing regression on variable Y; represents the predicted value of feature a i ; represents the predicted value of variable Y; α i represents the bias term when performing regression on feature a i ; α j represents the bias term when performing regression on variable Y;
[0031]
[0032]
[0033] In the formula, e i and e jThey are the variable Y and the feature a respectively i not covered in the unexplained residuals;
[0034] S1.5.4. Calculate the order correlation degree OC value of the feature, and the formula is as follows:
[0035]
[0036] In the formula, are respectively the ranks of e i and e j ;
[0037] S1.6. When then:
[0038] A n ← a i # (10)
[0039] When then:
[0040] A m+ ← a i # (11)
[0041] When then:
[0042] A m- ← a i # (12)
[0043] S1.7. Return the set A m+ , A m- , A n , are the positive-order feature, the reverse-order feature, and the disordered feature respectively;
[0044] S1.8. Divide the data set into an unordered data set and an ordered data set Normalize and standardize the data set. For each feature x i , standardize it through the following formula:
[0045]
[0046] In the formula, μ is the mean of the feature, and σ is the standard deviation;
[0047] Use the min-max scaling method for normalization, and the formula is as follows:
[0048]
[0049] S2. Obtain the unordered features of the data set from step S1, and for the unordered features, obtain the weight coefficients of the unordered features by calculating their weight functions;
[0050] Furthermore, the specific operation of step S2 is as follows:
[0051] S2.1. Obtain the data set U containing unordered features through the order correlation degree OC value of the features n ;
[0052] S2.2. Use hierarchical clustering to divide the data set to obtain the clustering result;
[0053] S2.3. Initialize an empty list s to store the silhouette coefficient of each clustering result;
[0054] S2.4. For each clustering result, calculate its silhouette coefficient through the following formula and store the silhouette coefficient in the list s:
[0055]
[0056] In the formula, a represents the average distance between the target sample and other samples in the same cluster, and b represents the average distance between the target sample and all points in the nearest different cluster; when there is only one sample in a certain cluster, that is, an outlier appears, s = 0;
[0057] S2.5. Select the division with a high silhouette coefficient in the empty list s as the clustering result;
[0058] S2.6. For each clustering result, perform the following steps:
[0059] S2.6.1. Calculate the clustering mean;
[0060] S2.6.2. For each sample in this clustering result, calculate the Euclidean distance using the following formula:
[0061]
[0062] In the formula, represents the Euclidean distance of the sample , represents the i-th sample of the unordered feature, is the cluster mean containing ;
[0063] S2.6.3. Calculate the weight of using the following formula:
[0064]
[0065] S2.7. For the unordered data set The finally obtained set of weight coefficients is S3. Obtain the ordered features of the dataset from step S1, and use a Monotonic Neural Network (MNN) to model the ordered features;
[0066] Further, the specific operation of step S3 is as follows:
[0067] S3.1. Construct a monotonic neural network, which is a fully connected four-layer neural network with I inputs, the first hidden layer has H nodes, the second hidden layer has L nodes, and the output layer is a single output. The output layer is defined as shown in Equation (18):
[0068]
[0069] In the formula, represents the final output value of the network; ω b , ω l represent the weights and bias terms of the second hidden layer respectively; ω b,l represents the bias term of the first hidden layer; ω lh represents the weights of the first hidden layer; ω b,h represents the bias term of the input layer; ω hi represents the weights of the input layer; θ1 and θ2 represent the activation functions of the first and second hidden layers respectively;
[0070] S3.2. Construct the input layer; the input layer has i nodes, and each node represents an input feature; the input feature is the ordered feature in U m , and is passed to the hidden layer through weighting;
[0071] S3.3. Construct the first hidden layer; this layer includes multiple nodes for processing the input data. The output value of the node is calculated through the activation function tanh, and the expression of the tanh function is as shown in Equation (19):
[0072]
[0073] The output expression of the hidden layer is as shown in (20):
[0074]
[0075] In the formula, W1 is the weight and b1 is the bias term, is the i-th input of the input layer;
[0076] S3.4. Construct the second hidden layer, which includes multiple nodes The output value of the node is calculated using the tanh activation function for output, and the expression is as shown in (21):
[0077]
[0078] Among them, W2 is the weight of the second layer, and b2 is the bias term of the second layer. is the result output from the first hidden layer;
[0079] S3.5. Construct the output layer; the output layer has one node is the predicted output of the entire network; Softmax is used to convert the output into probability values, and the expression is as shown in (22):
[0080]
[0081] In the formula, is the output of the second hidden layer;
[0082] S3.6. Perform forward propagation. Forward propagation is the process of calculating the output of the neural network; for each layer, calculate the weighted sum of the input and the weights, and process it through the activation function; the specific steps are as follows:
[0083] S3.6.1. From the input layer to the first hidden layer; for the input through the weights W1 and the bias b1, calculate the output of each hidden node
[0084] S3.6.2. From the first hidden layer to the second hidden layer; use the output of the previous layer as the input, and through the weights W2 and the bias b2, calculate the output of the second hidden layer nodes
[0085] S3.6.3. From the second hidden layer to the output layer; pass the output of the second hidden layer to the output layer, and obtain the final output through the Softmax activation function;
[0086] S3.7. Perform error backpropagation; evaluate the difference between the predicted output and the actual label by calculating the loss of the network. The backpropagation algorithm is used to adjust the weights in the network to minimize the error. The steps are as follows:
[0087] S3.7.1. Calculate the output error; the error formula for calculating the output layer is formula (23):
[0088]
[0089] In the formula, is the network predicted output, and y is the true label;
[0090] S3.7.2. Calculate the hidden layer error; use the chain rule to calculate the hidden layer error and pass the error to the previous layer until the input layer;
[0091] S3.7.3. Update the weights and biases; use gradient descent to update the weights W and biases, as shown in formulas (24) and (25):
[0092]
[0093] In the formula, η is the learning rate, is the gradient of the weight, is the bias gradient.
[0094] S4. On the basis of the monotonic neural network obtained in S3, add a custom loss function to implement monotonicity constraints to optimize the classification performance;
[0095] Furthermore, the specific operation of step S4 is as follows:
[0096] S4.1. Obtain the gradient of the model output with respect to each feature;
[0097] S4.2. The gradients of the monotonically increasing features and monotonically decreasing features are represented by vectors g i and g d respectively;
[0098] S4.3. For g i , the formula of the penalty function p i is as shown in formula (26):
[0099]
[0100] In the formula, the sigmoid function expression is as shown in formula (27):
[0101]
[0102] S4.4. For g d , the formula of the penalty function p d is as shown in formula (28):
[0103]
[0104] S4.5. Combine the original network loss function, and the finally obtained loss function is as shown in formula (29):
[0105] loss = l o + λ(p i + p d ) # (29)
[0106] In the formula, l o represents the original network loss function, and λ represents the penalty term weight.
[0107] S5.1. Obtain the output of the monotonic neural network model, and the result is as shown in formula (30):
[0108]
[0109] S5.2. Combine the model output results with the unordered feature weights to calculate the probability vector, representing the probabilities corresponding to each category, as shown in Equation (31):
[0110]
[0111] In the formula, represents the weight of the unordered features corresponding to each sample.
[0112] S5.3. The category corresponding to the maximum probability value in
[0113] A partially ordered data classification system based on order correlation degree, the system includes a computer processor and memory, a partially ordered data preprocessing unit, a selection model unit, a partially ordered data model training unit, and a partially ordered data model testing unit; the partially ordered data preprocessing unit preprocesses the data set input to the network and loads it into the computer memory; the selection model unit selects an existing classification model as the classifier for the data set, sets the corresponding parameters and loads them into the computer memory; the partially ordered data model training and prediction unit trains the selected model on the imported data set and displays the loss curve graph, and finally saves the trained model to the computer memory; the partially ordered data model testing unit uses the trained classification model to test the test data set, and displays the test results as well as the classification accuracy and mean square error; the specific data processing and calculation work in all units is completed by the computer processor, and all units interact with the data in the computer memory;
[0114] The partially ordered data preprocessing unit executes steps S1 to S2;
[0115] The selection model unit executes steps S3.1 to S3.5;
[0116] The partially ordered data model training unit executes steps S3.6 to step S5.
[0117] Compared with the prior art, the present invention has the following advantages:
[0118] (1) An order correlation degree (OC) method based on mining the order relationship between features and classes is invented to accurately identify ordered features in a dataset. By dividing the feature set into target features and the remaining features, linear regression is used to remove the feature collinearity that affects the partial correlation calculation. Then, the remaining features are used to perform regression on the target features and labels respectively. After eliminating the mutual influence between features, two sets of residuals are obtained. The Spearman correlation coefficient of the two sets of residuals is calculated, and the final result represents the true order relationship and order direction between the target features and labels. Finally, a threshold is set to distinguish ordered features from disordered features. Compared with other identification methods, this method can accurately distinguish the true ordered and disordered features in a partially ordered dataset.
[0119] (2) Based on the traditional monotonic neural network, a partially ordered neural network (PONN) is designed using the scheme of weighting disordered features. First, OC is used to distinguish ordered and disordered features. For disordered features, after optimal clustering division, the distance between samples and the cluster center is calculated and converted into a corresponding set of weight coefficients. For ordered features, a monotonic constraint is imposed by initializing the weight signs of the input network. At the same time, a penalty term is added to the original loss function to correct the features that violate the monotonic constraint during the training process, ensuring that the ordered features obey the monotonic constraint during the training process. Finally, the network output is convolved with the disordered feature weight coefficients to obtain the final output result. The present invention can solve the classification problem of partially ordered data, overcome the problem of fully considering the influence of disordered features while imposing a monotonic constraint on ordered features in a partially ordered dataset, and this method is superior to other existing algorithms in terms of performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0120] Figure 1 It is an overall framework diagram of a partially ordered data classification method based on order correlation degree;
[0121] Figure 2 It is a flowchart of the order correlation degree (OC) of the ordered feature recognition method;
[0122] Figure 3 It is a structure diagram of the PONN network;
[0123] Figure 4 It is an operation flowchart of a partially ordered data classification system based on order correlation degree. DETAILED DESCRIPTION OF THE INVENTION
[0124] To gain a deep understanding of the present invention, we will describe it comprehensively and meticulously. However, the present invention has multiple implementation manners and is not limited to the specific examples listed herein. The presentation of these examples aims to deepen the comprehensive understanding of the disclosed content of the present invention.
[0125] Example 1
[0126] The method for classifying partially ordered data based on order correlation degree according to the present invention is implemented through a computer program. The specific implementation manners of the technical solution proposed by the present invention will be described in detail according to the Figure 1 flow shown below. Through the technical solution of the present invention, classification modeling is performed on ten real datasets, and these datasets are from UCI (University of California, Irvine datasets) and Weka (a machine learning software). The detailed information is shown in Table 1.
[0127] Table 1 Standard datasets
[0128]
[0129]
[0130] S1. Select a dataset. For the features in the dataset, calculate the order correlation degree (OC) to distinguish ordered and unordered features and determine the order direction;
[0131] Further, the specific operation of step S1 is as follows:
[0132] S1.1. Determine the dataset U = {x i , y i}, and set the threshold θ for judging ordered features to 0.2;
[0133] S1.2. Read the dataset {x i , y i};
[0134] S1.3. Respectively set 3 empty lists A m+ , A m- , A n , which store positive-order features, reverse-order features, and unordered features respectively;
[0135] S1.4. Use a i to represent the feature to be recognized, represents the remaining features except a i ;
[0136] S1.5. Judge whether there is collinearity between the feature a i and the features in . When there is collinearity, remove the collinearity. Then calculate the residual set of the feature a i and the variable U, and calculate the order correlation degree OC value of the feature;
[0137] S1.5.1. Use the Pearson correlation coefficient to judge whether there is collinearity between the feature a i and the features in . The calculation formula of the Pearson correlation coefficient is as follows:
[0138]
[0139] Wherein, X i and Y i respectively represent the i-th observation values of two variables; respectively represent the means of variables X and Y; n is the sample size;
[0140] S1.5.2. Remove collinearity by the following method:
[0141] Taking feature a i as the independent variable and A as the dependent variable, construct a linear regression model:
[0142] A = β1·a i + e1# (2)
[0143] Wherein, β1 is the regression coefficient, used to describe the linear contribution of feature a i to the dependent variable A; e1 is the residual term, indicating the part in the dependent variable A that cannot be explained by feature a i ;
[0144] Calculate the residual e1 through the regression equation:
[0145] e1 = A - β1·a i # (3)
[0146] The residual e1 represents the independent part remaining in the dependent variable A after removing the linear influence of feature a i ;
[0147] Replace the dependent variable A with the residual e1 to eliminate the collinearity between feature a i and the dependent variable A:
[0148] A = e1#(4)
[0149] S1.5.3. Calculate the residual set of feature a i and variable Y; Use the feature to perform regression on feature a i and variable Y respectively:
[0150]
[0151]
[0152] Wherein, β i represents the weight coefficient of the feature i when performing regression on feature a ; β j represents the weight coefficient of the feature when performing regression on variable Y; represents the predicted value of feature a i ; Represents the predicted value of variable Y; α i Represents the bias term when performing regression on feature a i α j Represents the bias term when performing regression on variable Y;
[0153]
[0154]
[0155] In the formula, e i and e j are the residuals of variable Y and feature a i that are not explained respectively;
[0156] S1.5.4. Calculate the order correlation degree OC value of the feature. The formula is as follows:
[0157]
[0158] In the formula, are the ranks of e i and e j respectively;
[0159] S1.6. When then:
[0160] A n ← a i # (10)
[0161] When then:
[0162] A m+ ← a i # (11)
[0163] When then:
[0164] A m- ← a i # (12)
[0165] S1.7. Return the set A m+ A m- A n are the positive-order feature, reverse-order feature, and disordered feature respectively;
[0166] S1.8. Divide the dataset into an unordered dataset and an ordered dataset Normalize and standardize the dataset. For each feature x i , standardize it through the following formula:
[0167]
[0168] In the formula, μ is the mean of this feature, and σ is the standard deviation;
[0169] Normalization uses the min-max scaling method, and the formula is as follows:
[0170]
[0171] S2. Obtain the unordered features of the data set from step S1. For the unordered features, calculate their weight coefficients by calculating their weight functions;
[0172] Furthermore, the specific operation of step S2 is as follows:
[0173] S2.1. Obtain the data set U containing unordered features through the order correlation degree OC value of the features n ;
[0174] S2.2. Use hierarchical clustering to divide the data set to obtain the clustering result;
[0175] S2.3. Initialize an empty list s to store the silhouette coefficient of each clustering result;
[0176] S2.4. For each clustering result, calculate its silhouette coefficient through the following formula and store the silhouette coefficient in the list s:
[0177]
[0178] In the formula, a represents the average distance between the target sample and other samples in the same cluster, and b represents the average distance between the target sample and all points in the nearest different cluster; when there is only one sample in a certain cluster, that is, an outlier appears, s = 0;
[0179] S2.5. Select the division with a high silhouette coefficient in the empty list s as the clustering result;
[0180] S2.6. For each clustering result, perform the following steps:
[0181] S2.6.1. Calculate the cluster mean;
[0182] S2.6.2. For each sample in this clustering result, calculate the Euclidean distance using the following formula:
[0183]
[0184] In the formula, represents the sample 's Euclidean distance, represents the i-th sample of the unordered feature, is included cluster mean;
[0185] S2.6.3. Calculate using the following formula weight of:
[0186]
[0187] S2.7. For the unordered data set The finally obtained set of weight coefficients is S3. Obtain the ordered features of the data set from step S1. For the ordered features, use a monotonic neural network (MNN) for modeling;
[0188] Furthermore, the specific operation of step S3 is as follows:
[0189] S3.1. Construct a monotonic neural network, which is a fully connected four-layer neural network with I inputs, H nodes in the first hidden layer, L nodes in the second hidden layer, and a single output in the output layer. The output layer is defined as shown in Equation (18):
[0190]
[0191] In the formula, represents the final output value of the network; ω b , ω l represent the weights and bias terms of the second hidden layer respectively; ω b,l represents the bias term of the first hidden layer; ω lh represents the weights of the first hidden layer; ω b,h represents the bias term of the input layer; ω hi represents the weights of the input layer; θ1 and θ2 represent the activation functions of the first and second hidden layers respectively;
[0192] S3.2. Construct the input layer; the input layer has i nodes, and each node represents an input feature; the input features are the ordered features in U m and are passed to the hidden layer through weighting;
[0193] S3.3. Construct the first hidden layer; this layer includes multiple nodes for processing the input data. The output value of the node is calculated through the activation function tanh, and the expression of the tanh function is as shown in Equation (19):
[0194]
[0195] The output expression of the hidden layer is as shown in (20):
[0196]
[0197] Wherein, W1 is the weight, b1 is the bias term, is the i-th input of the input layer;
[0198] S3.4. Construct a second hidden layer, which includes multiple nodes The output value of the node is calculated using the tanh activation function, and the expression is as shown in (21):
[0199]
[0200] Wherein, W2 is the weight of the second layer, b2 is the bias term of the second layer, is the result output from the first hidden layer;
[0201] S3.5. Construct an output layer; the output layer has one node is the predicted output of the entire network; Softmax is used to convert the output into probability values, and the expression is as shown in (22):
[0202]
[0203] Wherein, is the output of the second hidden layer;
[0204] S3.6. Perform forward propagation. Forward propagation is the process of calculating the output of the neural network; for each layer, calculate the weighted sum of the input and the weight, and process it through the activation function; the specific steps are as follows:
[0205] S3.6.1. From the input layer to the first hidden layer; for the input calculate the output of each hidden node through the weight W1 and the bias n1
[0206] S3.6.2. From the first hidden layer to the second hidden layer; use the output of the previous layer as the input, and calculate the output of the second hidden layer nodes through the weight W2 and the bias b2
[0207] S3.6.3. From the second hidden layer to the output layer; pass the output of the second hidden layer to the output layer, and obtain the final output through the Softmax activation function;
[0208] S3.7. Perform error backpropagation; evaluate the difference between the predicted output and the actual label by calculating the loss of the network, and the backpropagation algorithm is used to adjust the weights in the network to minimize the error, and the steps are as follows:
[0209] S3.7.1. Calculate the output error; the error formula for calculating the output layer is formula (23):
[0210]
[0211] In the formula, is the network prediction output, and y is the true label;
[0212] S3.7.2. Calculate the hidden layer error; use the chain rule to calculate the error of the hidden layer and transfer the error to the previous layer until the input layer;
[0213] S3.7.3. Update the weights and biases; use the gradient descent method to update the weights W and biases, and the formulas are shown in (24) and (25):
[0214]
[0215]
[0216] In the formula, η is the learning rate, is the gradient of the weight, is the bias gradient.
[0217] S4. On the basis of the monotonic neural network obtained in S3, add a custom loss function to implement monotonicity constraints to optimize the classification performance;
[0218] Furthermore, the specific operation of step S4 is as follows:
[0219] S4.1. Obtain the gradient of the model output with respect to each feature;
[0220] S4.2. The gradients of the monotonically increasing feature and the monotonically decreasing feature are represented by vectors g i and g d respectively;
[0221] S4.3. For g i , the formula of the penalty function p i is shown in formula (26):
[0222]
[0223] In the formula, the expression of the sigmoid function is shown in formula (27):
[0224]
[0225] S4.4. For g d , the formula of the penalty function p d is shown in formula (28):
[0226]
[0227] S4.5. Combine the original network loss function, and the finally obtained loss function is shown in formula (29):
[0228] loss = l o + λ(p i + p d ) # (29)
[0229] In the formula, l o represents the original network loss function, and λ represents the penalty term weight.
[0230] S5.1. Obtain the output of the monotonic neural network model, and the result is shown in formula (30):
[0231]
[0232] S5.2. Combine the model output result with the unordered feature weights to calculate the probability vector, representing the probabilities corresponding to each category, as shown in formula (31):
[0233]
[0234] In the formula, represents the weight of the unordered features corresponding to each sample.
[0235] S5.3 The category corresponding to the maximum probability value in
[0236] is used as the final classification result.
[0237] Technical effect evaluation:
[0238]
[0239]
[0240] In the formula, y(x i ) is the true label of sample x i , is the predicted label. Tables 2 and 3 list the ACC and MAE of each classifier on different datasets respectively. For the highest ACC and the lowest MAE of each dataset, they are represented in bold.
[0241] Table 2 Comparative analysis of Acc
[0242]
[0243] Table 3 Comparative Analysis of MAE
[0244]
[0245] As can be seen from the table, the present invention achieves the highest classification accuracy in 9 tasks, while other algorithms only achieve better accuracy in 1 task. At the same time, the present invention also obtains the lowest MAE in 9 tasks. The results show that for most tasks, the present invention is superior to other classifiers in terms of ACC and MAE.
[0246] To verify its average performance, we use statistical tests to determine whether the present invention significantly improves the classification effect. We use the t-test to pairwise compare the average performance of all algorithms, as shown in Table 4 and Table 5.
[0247] Table 4 Significance Levels of Average Accuracy of Different Classifiers
[0248]
[0249]
[0250] Table 5 Significance Levels of Mean Absolute Error of Different Classifiers
[0251]
[0252] As shown in Table 4 and Table 5, the significant differences between the two algorithms are indicated in bold, and the p-value is less than 0.05. The results show that the present invention is significantly different from other algorithms and is superior to other algorithms with a 95% confidence level.
[0253] In summary, the excellent performance of the present invention in various evaluation indicators is mainly due to the proposed order correlation metric that can accurately distinguish the ordered and disordered features in the dataset; the weighting of disordered features can fully incorporate the influence of disordered features into the monotonic neural network; the improvement of the penalty function in the network model of the present invention can effectively maintain the monotonic constraint of the network for ordered features. Therefore, in the classification problem of partially ordered datasets, the present invention can achieve better results.
[0254] Example 2
[0255] As Figure 4 shown, a partially ordered data classification system based on order correlation degree includes a computer processor and memory, a partially ordered data preprocessing unit, a selection model unit, a partially ordered data model training unit, and a partially ordered data model testing unit.
[0256] The partial ordered data preprocessing unit executes steps S1 to S2. For the features in the dataset, it first distinguishes ordered and unordered features; then calculates the weight coefficients of the unordered features through a weight function.
[0257] The selected model unit executes steps S3.1 to S3.5, loading the selected model into the computer memory.
[0258] The partial ordered data model training unit executes steps S3.6 to step S5, training the selected model on the dataset and saving it to the computer memory.
[0259] The partial ordered data model testing unit uses the trained classification model to test the test dataset, and displays the test results, as well as the classification accuracy and mean square error.
[0260] The specific data processing and calculation work in all units is completed by the computer processor, and all units interact with the data in the computer memory.
[0261] The content not described in detail in the specification of the present invention belongs to the prior art well-known to those skilled in the art. Although the illustrative specific embodiments of the present invention have been described above for the understanding of those skilled in the art of the present technology, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
Claims
1. A method for classifying partially ordered data based on order correlation, characterized in that: The following steps are involved: S1. Select a data set, calculate the order correlation of the features in the data set to distinguish ordered and disordered features, and determine the order direction; S2, obtaining the disordered features of the data set from step S1, and obtaining the weight coefficient of the disordered features by calculating the weight function of the disordered features; S3, obtaining the ordered features of the data set from step S1, and modeling the ordered features using a monotone neural network; S4. Based on the monotonic neural network obtained in step S3, a custom loss function is added to implement the monotonicity constraint to optimize the classification performance; S5. Combine the disordered feature weight coefficient obtained in step S2 with the monotone neural network obtained in step S4 to construct a final classification model and classify the data set.
2. A method for partially ordered data classification based on order correlation according to claim 1, characterized in that: The step S1 comprises the following specific steps: S1.
1. Determine the data set U = {x i ,y i }, set the threshold θ for judging ordered features = 0.2; S1.2, read the data set {x i ,y i }; S1.3, set up 3 empty lists respectively A n , respectively store the positive sequence features, reverse sequence features and disordered features; S1.4, use a i represents the features to be identified, Indicates that except a i Other features besides S1.
5. Determine feature a i With features Is there collinearity among the features in the equation? If so, remove the collinearity and then calculate feature a. i The residual set with variable Y, and calculate the ordinal correlation OC value of the feature; S1.5.
1. Using Pearson correlation coefficient to determine feature a i With features Whether there is collinearity among the features in the Pearson correlation coefficient is calculated as follows: Where, X i and Y i Respectively represent the i-th observation value of the two variables; and Represent the means of variables X and Y respectively; n is the sample size; S1.5.
2. Use the following method to remove collinearity: With feature a i As the independent variable and A as the dependent variable, a linear regression model is constructed: A=β1·a i +e1#(2) In the formula, β1 is the regression coefficient, which is used to describe feature a i The linear contribution to the dependent variable A; e1 is the residual term, indicating that the dependent variable A cannot be obtained by feature a i The explanation part; Calculate the residual e1 through the regression equation: e1=A-β1·a i #(3) The residual e1 indicates that feature a is removed from the dependent variable A. i The independent part remaining after the linear influence of ; Replace the dependent variable A with the residual e1 and eliminate feature a i Collinearity with dependent variable A: A=e1#(4) S1.5.
3. Calculate feature a i The residual set with the variable Y; using the feature For feature a i Regression with variable Y: In the formula, β i Represents feature a i When doing regression features The weight coefficient of j Indicates the characteristics when regressing variable Y The weight coefficient of Indicates feature a i The predicted value of Represents the predicted value of variable Y; α i Represents feature a i Bias term for regression, α j Represents the bias term when regressing variable Y; In the formula, e i and e j They are variable Y and feature a respectively i Not Explained residuals; S1.5.
4. Calculate the OC value of the feature order correlation. The formula is as follows: In the formula, and e i and e j rank; S1.6, when or but: A n ←a i #(10) when but: when but: S1.
7. Return collection A n , respectively, positive sequence feature, reverse sequence feature and disordered feature; S1.
8. Divide the data set into an unordered data set based on the recognition results and ordered datasets Standardize and normalize the data set. For each feature x i , normalized by the following formula: In the formula, μ is the mean of the feature and σ is the standard deviation; Normalization uses the minimum and maximum scaling method, the formula is as follows:
3. A method for classifying partially ordered data based on order correlation according to claim 2, characterized in that: The step S2 comprises the following specific steps: S2.
1. Obtaining a dataset U containing unordered features through the order correlation OC value of the features n ; S2.2, use hierarchical clustering to divide the data set and obtain clustering results; S2.3, initialize an empty list s to store the silhouette coefficient of each clustering result; S2.
4. For each clustering result, calculate its silhouette coefficient using the following formula and store the silhouette coefficient in list s: Where a represents the average distance between the target sample and other samples in the same cluster, and b represents the average distance between the target sample and all points in the nearest different clusters. When there is only one sample in a cluster, that is, when an outlier occurs, s = 0. S2.5, select the partition with high silhouette coefficient in the empty list s as the clustering result; S2.
6. For each clustering result, perform the following steps: S2.6.
1. Calculate the cluster mean; S2.6.
2. For each sample in the clustering result, use the following formula to calculate the Euclidean distance: In the formula, Representation sample The Euclidean distance of represents the i-th sample of the disordered feature, is included The cluster mean of S2.6.
3. Use the following formula to calculate Weight: S2.
7. For unordered datasets The final set of weight coefficients is 4. A method for classifying partially ordered data based on order correlation according to claim 3, characterized in that: The step S3 comprises the following specific steps: S3.
1. Construct a monotone neural network. The network is a fully connected four-layer neural network with I inputs, H nodes in the first hidden layer, L nodes in the second hidden layer, and a single output in the output layer. The output layer is defined as shown in formula (18): In the formula, Represents the final output value of the network; ω b ,ω l Respectively represent the weight and bias term of the second hidden layer; ω b,l represents the bias term of the first hidden layer; ω lh represents the weight of the first hidden layer; ω b,h Represents the bias term of the input layer; ω hi represents the weight of the input layer; θ1, θ2 represent the activation functions of the first and second hidden layers respectively; S3.2, construct the input layer; the input layer has i nodes, each node represents an input feature; The input feature is U m The ordered features in are passed to the hidden layer through weighting; S3.
3. Construct the first hidden layer; this layer includes multiple nodes It is used to process input data. The output value of the node is calculated by the activation function tanh. The expression of the tanh function is shown in formula (19): The output expression of the hidden layer is shown in (20): Where W1 is the weight, b1 is the bias term, is the i-th input of the input layer; S3.
4. Construct a second hidden layer, which includes multiple nodes The output value of the node is calculated using the tanh activation function, as shown in (21): Among them, W2 is the weight of the second layer, b2 is the bias term of the second layer, is the output from the first hidden layer; S3.
5. Construct the output layer; the output layer has a node is the predicted output of the entire network; Softmax is used to convert the output into a probability value, as shown in (22): In the formula, is the output of the second hidden layer; S3.
6. Perform forward propagation. Forward propagation is the process of calculating the output of the neural network. For each layer, the weighted sum of the input and the weight is calculated and processed through the activation function. The specific steps are as follows: S3.6.1, input layer to the first hidden layer; for input By using weight W1 and bias b1, the output of each hidden node is calculated S3.6.2, from the first hidden layer to the second hidden layer; use the output of the previous layer as input, and calculate the output of the second hidden layer node through weight W2 and bias b2 S3.6.3, from the second hidden layer to the output layer; the output of the second hidden layer is passed to the output layer, and the final output is obtained through the Softmax activation function; S3.7, perform error back propagation; evaluate the difference between the predicted output and the actual label by calculating the loss of the network. The back propagation algorithm is used to adjust the weights in the network to minimize the error. The steps are as follows: S3.7.
1. Calculate the output error. The error formula for calculating the output layer is formula (23): In the formula, is the network prediction output, and y is the true label; S3.7.
2. Calculate the hidden layer error. Use the chain rule to calculate the hidden layer error and pass the error to the previous layer until the input layer. S3.7.
3. Update weights and biases. Use the gradient descent method to update weights W and biases. The formulas are shown in (24) and (25): Where η is the learning rate, is the gradient of the weight, is the bias gradient.
5. A method for classifying partially ordered data based on order correlation according to claim 4, characterized in that: The step S4 comprises the following specific steps: S4.
1. Obtain the gradient of the model output relative to each feature; S4.2, the gradients of monotonically increasing features and monotonically decreasing features are represented by vector g i and g d express; S4.3, for g i , penalty function p i The formula is shown in formula (26): In the formula, the sigmoid function expression is shown in formula (27): S4.4, for g d , penalty function p d The formula is shown in formula (28): S4.
5. Combined with the original network loss function, the final loss function is shown in formula (29): loss=l o +λ(p i +p d )#(29) In the formula, l o Represents the original network loss function, and λ represents the penalty item weight.
6. A method for classifying partially ordered data based on order correlation according to claim 5, characterized in that: The step S5 comprises the following specific steps: S5.
1. Obtain the output of the monotone neural network model. The result is shown in formula (30): S5.
2. The model output is combined with the unordered feature weights to calculate the probability vector, which represents the probability corresponding to each category, as shown in formula (31): In the formula, Indicates the weight of each sample corresponding to the disordered feature; S5.3, The category corresponding to the maximum probability value is taken as the final classification result.
7. A partially ordered data classification system based on order correlation, characterized in that: The system includes a computer processor and memory, a partially ordered data preprocessing unit, a model selection unit, a partially ordered data model training unit, and a partially ordered data model testing unit; the partially ordered data preprocessing unit preprocesses the data set input to the network and loads it into the computer memory; the model selection unit selects an existing classification model as a classifier of the data set, sets corresponding parameters and loads it into the computer memory; The partially ordered data model training prediction unit trains the selected model on the imported data set and displays the loss curve graph, and finally saves the trained model to the computer memory; the partially ordered data model testing unit uses the trained classification model to test the test data set, and displays the test results as well as the classification accuracy and mean square error; the specific data processing and calculation work in all units is completed by the computer processor, and all units interact with the data in the computer memory; The partially ordered data preprocessing unit executes steps S1 to S2; The model selection unit executes steps S3.1 to S3.5; The partially ordered data model training unit executes steps S3.6 to S5.
Citation Information
Cited By
Order rule-based earth surface natural electric field risk feature decoupling method and system
CN122288378A