Multi-label text classification method based on positive and negative label learning and label correlation
By constructing a dual-hidden layer feedforward neural network model and a composite error function, the problems of insufficient label correlation modeling and category imbalance in multi-label text classification are solved, and high-precision and efficient multi-label text classification is achieved.
Patent Information
- Application Number
- CN202510740485.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-19
AI Technical Summary
Existing multi-label text classification methods suffer from poor classification performance due to insufficient label correlation modeling, imbalanced multi-label data categories, and high model complexity.
A dual-hidden-layer feedforward neural network model is constructed, combined with a composite error function and adaptive optimization strategy. Through positive and negative label learning and label correlation constraints, the correlation between text features and labels is improved to achieve high-precision classification.
It significantly improves the accuracy and efficiency of multi-label text classification, solves the problems of weak nonlinear modeling capabilities, sensitivity to category imbalance, and insufficient utilization of label correlation, and demonstrates better stability and accuracy.
Smart Images

Figure CN120670595A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-label text classification method based on positive and negative label learning and label correlation, and belongs to the technical field of multi-label classification. Background Art
[0002] With the rapid development of information technology, massive amounts of textual information are growing exponentially, encompassing diverse domains such as news, social media, medical records, legal documents, and scientific research. As an important extension of text classification, multi-label text classification involves mapping text instances to multiple labels. It has been widely applied to learning from semantically rich real-world text objects. Compared to the "either-or" paradigm of traditional single-label classification, multi-label classification better reflects the intertwined nature of real-world text semantics. The current evolution of multi-label classification technology primarily follows two main paths: traditional strategies based on problem transformation and algorithm adaptation, and modern methodological frameworks based on deep learning. Among these traditional approaches, binary association methods achieve label decoupling by constructing independent binary classifiers, while label power set methods attempt to transform the multi-label problem into a multi-class classification task. While these methods are feasible in basic scenarios, they suffer from two inherent drawbacks: first, the linear modeling paradigm struggles to capture the complex nonlinear relationships between the text feature space and the label space; and second, potential correlations between labels are systematically ignored.
[0003] It's worth noting that in real-world scenarios, the relationship between text features and labels is often extremely complex and nonlinear, which significantly limits traditional methods in terms of semantic capture accuracy and model generalization capabilities. Advanced methods, such as deep neural networks, leverage their modeling advantages of multi-layer nonlinear transformations to provide new insights into multi-label text classification. However, existing deep learning solutions still face the following challenges: high model complexity leads to poor adaptability to scenarios with small samples or unstructured text data; a lack of explicit modeling of label correlations makes it difficult to capture label co-occurrence dependencies; and classification thresholds rely on manual experience, making them incapable of addressing class imbalance.
[0004] Therefore, the present invention provides a multi-label text classification method based on positive and negative label learning and label correlation. The technology of the present invention is funded by the Yunnan Basic Research Program General Project (Project Approval Number: Grant No.202501AT070299). Summary of the Invention
[0005] The technical problem solved by the present invention is: the present invention provides a multi-label text classification method based on positive and negative label learning and label correlation, so as to solve the problems in the field of text classification, namely, insufficient label correlation modeling, imbalance of multi-label data categories and poor classification performance caused by high model complexity; the present invention constructs a dual-hidden layer feedforward neural network model and implements a novel composite error function to fully mine high-order semantic features in text data and establish correlations between text labels, thereby achieving high-precision and high-efficiency classification of complex text data.
[0006] The technical solution of the present invention is: a multi-label text classification method based on positive and negative label learning and label correlation, the method comprising:
[0007] Step 1: Data preprocessing and dividing the dataset into training set and test set;
[0008] Step 2: Construct a feedforward neural network model with two hidden layers, including an input layer, a first hidden layer, a second hidden layer, and an output layer;
[0009] Step 3: Input the training set and initialize the double hidden layer feedforward neural network model component; read the features and label information of the samples in the training set to generate the feature matrix and label matrix; randomly initialize the weight matrix and bias parameters of the double hidden layer feedforward neural network model;
[0010] Step 4, forward propagation: input the feature matrix into the input layer of the double hidden layer feedforward neural network model, and calculate the neuron output layer by layer;
[0011] Step 5, back propagation: Use the composite error function to calculate the gradient of the weight matrix and bias parameters between each layer in the double hidden layer feedforward neural network model, and use the Adam optimization algorithm to dynamically adjust the gradient, with the goal of minimizing the composite error function, and update the weight matrix and bias parameters;
[0012] Step 6: Set the error threshold and the maximum number of iterations as the training termination conditions. When the error change amplitude is lower than the threshold or reaches the preset number of iterations, the training is terminated. Otherwise, repeat steps 4 to 5 until the convergence judgment criteria are met.
[0013] Step 7: After the model converges, the test set is predicted and the prediction output is binarized using a fixed threshold to form the final multi-label classification result.
[0014] Furthermore, the step 1 includes:
[0015] Step 1.1. Data preprocessing: perform word segmentation, stop word removal, and word stemming on the text data; generate a TF-IDF matrix and a Word2Vec embedding matrix, and concatenate them to form a text feature vector;
[0016] Step 1.2: Obtain a text dataset. The text dataset is divided into a training set G and a test set T using a five-fold cross-validation method.
[0017] Furthermore, the step 2 includes:
[0018] Step 2.1: Build a double hidden layer feedforward neural network model:
[0019] Input layer: Configure neurons that match the dimension d of the text feature vector;
[0020] The first hidden layer: sets N hidden neurons for preliminary feature extraction and transmission;
[0021] The second hidden layer: configures M hidden neurons for deep feature extraction and abstract representation;
[0022] Output layer: Set Q neurons, each neuron corresponds to a candidate category label, and outputs the predicted score;
[0023] Step 2.2, the weight matrix and bias parameters are defined as follows:
[0024] The input layer and the first hidden layer are fully connected, and the weight matrix from the input layer to the first hidden layer is: Z = [z qp ](1≤q≤d,1≤p≤N), The bias parameter is δ p (1≤p≤N),
[0025] The first hidden layer and the second hidden layer are fully connected, and the weight matrix from the first hidden layer to the second hidden layer is: V = [v ps ](1≤p≤N,1≤s≤M), The bias parameter is γ s (1≤s≤M),
[0026]
[0027] The second hidden layer and the output layer are fully connected, and the weight matrix from the second hidden layer to the output layer is: W = [w sj ](1≤s≤M,1≤j≤Q), The bias parameter is θ j (1≤j≤Q),
[0028] Furthermore, in step 3, the initializing the dual hidden layer feedforward neural network model component includes initializing the following components: a classifier f(X) and a nonlinear mapper g(·);
[0029] The classifier f(X) is used to extract sample features and output the classification prediction value C k (x), C k (x) represents the predicted output of the k-th class label of sample x in the dataset;
[0030] The network architecture of the classifier f(X) is: input layer, two hidden layers, and output layer; each layer is fully connected, and the weight matrix and bias parameters between each layer are (Z, V, W) and (δ p ,γ s ,θ j );
[0031] The nonlinear mapper g(·) is located in the hidden layer and the output layer. The activation function of the first hidden layer adopts a linear activation function, and the activation functions of the second hidden layer and the output layer adopt a "Sigmoid" function.
[0032] Furthermore, the step 4 includes:
[0033] The feature matrix Input the double hidden layer feedforward neural network model, calculate the neuron output layer by layer, and obtain the neuron output through weighted calculation and activation function processing, where d is the dimension of the text feature vector, n is the number of samples, and x is the number of samples. n is the nth sample; repeat the above steps until the output layer. The specific steps are:
[0034] Step 4.1: The predicted output of the pth neuron in the first hidden layer is: b p =f(netb p +δ p ), where δ p is the bias parameter of the pth hidden neuron, netb p is the input signal of the pth hidden neuron, which is: a q is the input feature x i The qth component of Z qp is the weight matrix connecting the qth input neuron to the pth hidden neuron, and the activation function is set to a linear function;
[0035] Step 4.2: Output the prediction of the sth neuron in the second hidden layer as: h s =f(neth s +γ s ), where γ sis the bias parameter of the sth hidden neuron, neth s is the input signal of the pth hidden neuron, which is: v ps is the weight matrix connecting the pth hidden neuron and the sth hidden neuron, and the activation function is set to the "sigmoid" function;
[0036] Step 4.3: The predicted output of the jth neuron in the output layer is: C j =f(netC j +θ j ), where θ j is the bias parameter of the j-th output neuron, netC j is the input signal of the j-th output neuron, which is: w sj is the weight matrix connecting the sth hidden neuron and the jth hidden neuron, and the activation function is set to the "sigmoid" function.
[0037] Furthermore, the specific steps of step 5 include:
[0038] Step 5.1. Design a composite error function E i , introduce dynamic weight coefficient to adjust the error term adaptively, in the composite error function E i Embed label correlation constraints in
[0039] Step 5.2: Based on the composite error function E i Calculation error;
[0040] Step 5.3, calculate the weight matrix (Z, V, W) and bias parameters (δ p ,γ s ,θ j )’s gradient;
[0041] Step 5.4: Update the weight matrix and bias parameters: Use the Adam optimizer to update the weight matrix and bias parameters.
[0042] Furthermore, in step 5.1, the composite error function is defined as a weighted sum of the following three parts, specifically including:
[0043] (1) The composite error function integrates the positive and negative label learning items, and introduces two dynamic weight coefficients λ1 and λ2 into the composite error function; the specific form of the positive and negative label learning items is:
[0044]
[0045] in, Represents sample x iThe predicted output on the kth label class, Y i is a sequence that contains only sample x i The set of positive labels, and is the sample x i The set of negative labels, exp is an exponential function used to amplify the penalty for extreme errors, Q is the number of sample labels, and is also the number of neurons in the output layer of the double hidden layer feedforward neural network model;
[0046] (2) Embed a label relevance constraint term in the composite error function. The label relevance constraint term is constructed based on the label co-occurrence matrix. The specific form of the label relevance constraint term is:
[0047]
[0048] in, It measures the similarity of different category labels, β is the weight coefficient of the label correlation term, R kl is the label co-occurrence matrix, label co-occurrence matrix R kl The calculation formula is:
[0049]
[0050] in, is the label matrix, Q is also the number of sample labels, n is the number of samples, Y .,k represents the k-th column label vector;
[0051] (3) The final composite error function is:
[0052]
[0053] Among them, E i is the i-th error term, indicating the network's relative error to sample x i error.
[0054] Furthermore, the step 5.2 includes:
[0055] (1) The error of the jth neuron in the output layer is defined as:
[0056] (2) Considering that the activation function of the output layer is "sigmoid", the sigmoid activation function f'(netC j +θ j )=(1-C j )(C j ),get:
[0057]
[0058] (3) Similarly, the error of the sth neuron in the second hidden layer is defined as:
[0059] (4) Since the activation function of the second hidden layer is also "sigmoid", the sigmoid activation function f'(neth s +γ s )=(1-h s )(h s ), and we get
[0060] (5) The error of the pth hidden neuron in the first hidden layer is similarly defined as:
[0061] (6) Since the activation function of the first hidden layer is "linear", the activation function f'(netb p +δ p )=1, we get:
[0062] Furthermore, in step 5.3, the gradient of the weight matrix (Z, V, W) is:
[0063]
[0064] Bias parameter (δ p ,γ s ,θ j ) is:
[0065] in, is the gradient of the weight matrices W, V, and Z, is the bias parameter θ j , γ s , δ p gradient.
[0066] Furthermore, the step 5.4 includes:
[0067] Step 5.4.1. Update the moment estimates of the weight matrix and bias parameters:
[0068] First, introduce the biased first-order moment estimate to perform exponential weighted averaging on the weight matrix gradient, denoted as m t , the calculation formula is:
[0069]
[0070] in, is the biased first-order moment estimate of the weight matrices W, V, and Z;
[0071] Update the biased first-order moment estimate of the bias parameter gradient of each layer:
[0072]
[0073] Where t is the current iteration number, and β1 is a hyperparameter used to control the influence of the gradient; is the bias parameter θ j , γ s , δ p The biased first moment estimate of ;
[0074] The learning rate of the weight matrix and bias parameters is dynamically adjusted based on the biased second-order moment estimation to avoid large parameter update steps or oscillations in the early stages of training. The biased second-order moment estimation is an exponentially weighted average of the square of the gradient, denoted as v t , the calculation formula is:
[0075]
[0076] Among them, β2 is a hyperparameter used to control the influence of the square of the gradient; is the biased second moment estimate of the weight matrices W, V, and Z; is the bias parameter θ j , γ s , δ p The biased second moment estimate of ;
[0077] Since the biased first moment estimate m t and the biased second moment estimate v t The estimates of all start from zero, which may lead to underestimates in the first few time steps, which will be biased towards 0, resulting in biased estimates; therefore, bias correction is required, which involves scaling the estimates of the moving average to compensate for the bias introduced by zero initialization;
[0078] To m t and v t Perform bias correction, that is, calculate the first-order moment estimate of the weight matrix and bias parameter bias correction and the second moment estimate
[0079] Calculate the bias-corrected first-order moment estimates of the weight matrix and bias parameters for each layer:
[0080]
[0081] in, is the bias-corrected first-order moment estimate of the weight matrices W, V, and Z, is the bias parameter θ j, γ s , δ p The bias-corrected first-order moment estimate of ; is the hyperparameter after the tth iteration;
[0082] Calculate the bias-corrected second-order raw moment estimates of the weight matrix and bias parameters for each layer:
[0083]
[0084] in, is the bias-corrected second-order moment estimate of the weight matrices W, V, and Z, is the bias parameter θ j , γ s , δ p The bias-corrected second moment estimates of ; is the hyperparameter after the tth iteration;
[0085] Step 5.4.2, update the weight matrix and bias parameters:
[0086]
[0087] Where α is the learning rate, ∈ is set to 10 -8 , which ensures that divide-by-zero errors are not encountered.
[0088] Furthermore, in step 6, the training termination condition is:
[0089] The preset range of iterations is: 50≤epochs≤100;
[0090] Based on the composite error function E i Calculate the error and the error change using the mean square error form:
[0091]
[0092] When the error change is less than Or terminate the training when the preset number of iterations is reached;
[0093] If the above conditions are not met, repeat steps 4 and 5 until convergence.
[0094] Furthermore, the step 7 includes:
[0095] Make predictions on the test set T, input the feature matrix X of the test set T into the classifier f(X), and obtain the classification prediction value C k (x), according to the threshold value, the threshold selector is defined as:
[0096] h(x)={k∈Y:C k(x)>t(x)}
[0097] Set the fixed threshold t(x) to 0.5, if C k (x)>0.5, then the sample x is judged to belong to the kth category label.
[0098] The beneficial effects of the present invention are:
[0099] 1. The multi-label text classification method proposed in this paper based on positive and negative label learning and label correlation significantly improves the accuracy and efficiency of multi-label text classification;
[0100] 2. The method of the present invention effectively solves the problems of weak nonlinear modeling ability, sensitivity to class imbalance, and insufficient utilization of label correlation in multi-label classification methods by constructing a double-hidden layer feedforward neural network model, combining a composite error function and an adaptive optimization strategy;
[0101] 3. Compared with existing advanced multi-label text classification methods, the present invention demonstrates better stability and accuracy on different existing text datasets. Experimental results show that key indicators such as average precision, 1-error rate and ranking loss of the present invention are significantly improved compared with traditional methods, fully verifying the superior performance of the model in complex application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] Figure 1 This is a training flow chart of the present invention;
[0103] Figure 2 This is a diagram of the architecture of the dual-hidden-layer feedforward neural network model in the present invention. DETAILED DESCRIPTION
[0104] Example 1: Figure 1-Figure 2 As shown, a multi-label text classification method based on positive and negative label learning and label correlation includes:
[0105] Step 1: Data preprocessing and dividing the dataset into training set and test set;
[0106] Furthermore, the step 1 includes:
[0107] Step 1.1. Data preprocessing: perform word segmentation, stop word removal, and word stemming on the text data; generate a TF-IDF matrix and a Word2Vec embedding matrix, and concatenate them to form a text feature vector;
[0108] Step 1.2: Obtain a text dataset. The text dataset is divided into a training set G and a test set T using a five-fold cross-validation method.
[0109] Step 2: Construct a feedforward neural network model with two hidden layers, including an input layer, a first hidden layer, a second hidden layer, and an output layer;
[0110] Furthermore, the step 2 includes:
[0111] Step 2.1: Build a double hidden layer feedforward neural network model:
[0112] Input layer: Configure neurons that match the dimension d of the text feature vector;
[0113] The first hidden layer: sets N hidden neurons for preliminary feature extraction and transmission;
[0114] The second hidden layer: configures M hidden neurons for deep feature extraction and abstract representation;
[0115] Output layer: Set Q neurons, each neuron corresponds to a candidate category label, and outputs the predicted score;
[0116] Step 2.2, the weight matrix and bias parameters are defined as follows:
[0117] The input layer and the first hidden layer are fully connected, and the weight matrix from the input layer to the first hidden layer is: Z = [z qp ](1≤q≤d,1≤p≤N), The bias parameter is δ p (1≤p≤N),
[0118] The first hidden layer and the second hidden layer are fully connected, and the weight matrix from the first hidden layer to the second hidden layer is: V = [v ps ](1≤p≤N,1≤s≤M), The bias parameter is γ s (1≤s≤M),
[0119] The second hidden layer and the output layer are fully connected, and the weight matrix from the second hidden layer to the output layer is: W = [w sj ](1≤s≤M,1≤j≤Q), The bias parameter is θ j (1≤j≤Q),
[0120] Step 3: Input the training set and initialize the double hidden layer feedforward neural network model component; read the features and label information of the samples in the training set to generate the feature matrix and label matrix; randomly initialize the weight matrix and bias parameters of the double hidden layer feedforward neural network model;
[0121] Feature Matrix Label Matrix Q is also the number of sample labels, d is the dimension of the text feature vector, n is the number of samples, x n is the nth sample;
[0122] Furthermore, in step 3, the initializing the dual hidden layer feedforward neural network model component includes initializing the following components: a classifier f(X) and a nonlinear mapper g(·);
[0123] The classifier f(X) is used to extract sample features and output the classification prediction value C k (x), C k (x) represents the predicted output of the k-th class label of sample x in the dataset;
[0124] The network architecture of the classifier f(X) (such as Figure 2 As shown in Figure 2, the network consists of an input layer (the dimension matches the feature vector), two hidden layers (containing N and M neurons respectively), and an output layer (containing Q label neurons). The layers are fully connected, and the weight matrices and bias parameters between the layers are (Z, V, W) and (δ p ,γ s ,θ j );
[0125] The nonlinear mapper g(·) is located in the hidden layer and the output layer. The activation function of the first hidden layer adopts a linear activation function, and the activation functions of the second hidden layer and the output layer adopt a "Sigmoid" function.
[0126] Step 4, forward propagation: input the feature matrix into the input layer of the double hidden layer feedforward neural network model, and calculate the neuron output layer by layer;
[0127] Furthermore, the step 4 includes:
[0128] The feature matrix Input the double hidden layer feedforward neural network model, calculate the neuron output layer by layer, and obtain the neuron output through weighted calculation and activation function processing, where d is the dimension of the text feature vector, n is the number of samples, and x is the number of samples. n is the nth sample; repeat the above steps until the output layer. The specific steps are:
[0129] Step 4.1: The predicted output of the pth neuron in the first hidden layer is: b p =f(netb p +δ p ), where δ p is the bias parameter of the pth hidden neuron, netb p is the input signal of the pth hidden neuron, which is: a qis the input feature x i The qth component of Z qp is the weight matrix connecting the qth input neuron to the pth hidden neuron, and the activation function is set to a linear function;
[0130] Step 4.2: Output the prediction of the sth neuron in the second hidden layer as: h s =f(neth s +γ s ), where γ s is the bias parameter of the sth hidden neuron, neth s is the input signal of the pth hidden neuron, which is: v ps is the weight matrix connecting the pth hidden neuron and the sth hidden neuron, and the activation function is set to the "sigmoid" function;
[0131] Step 4.3: The predicted output of the jth neuron in the output layer is: C j =f(netC j +θ j ), where θ j is the bias parameter of the j-th output neuron, netC j is the input signal of the j-th output neuron, which is: w sj is the weight matrix connecting the sth hidden neuron and the jth hidden neuron, and the activation function is set to the "sigmoid" function.
[0132] Step 5, back propagation: A new composite error function is used to calculate the gradient of the weight matrix and bias parameters between each layer in the double hidden layer feedforward neural network model, and the Adam optimization algorithm is used to dynamically adjust the gradient to minimize the composite error function and update the weight matrix and bias parameters;
[0133] Furthermore, the specific steps of step 5 include:
[0134] Step 5.1: To measure the difference between the predicted label and the true label, design a composite error function E i At the same time, in order to solve the problem of class imbalance, a dynamic weight coefficient is introduced to adaptively adjust the error term. In addition, in order to capture the high-order correlation characteristics between labels, the composite error function E i The label correlation constraint term is embedded in the , which is constructed based on the label co-occurrence matrix and is used to force the network to learn the dependency relationship between labels during training;
[0135] Step 5.2: Based on the composite error function E i Calculation error;
[0136] Step 5.3, calculate the weight matrix (Z, V, W) and bias parameters (δ p ,γ s ,θ j )’s gradient;
[0137] Step 5.4: Update the weight matrix and bias parameters: Use the Adam optimizer to update the weight matrix and bias parameters.
[0138] Furthermore, in step 5.1, the composite error function is defined as a weighted sum of the following three parts, specifically including:
[0139] (1) The composite error function integrates the positive and negative label learning items. At the same time, in order to alleviate the imbalance problem of positive and negative labels in multi-label datasets, two dynamic weight coefficients λ1 and λ2 are introduced into the composite error function to balance the outputs of the positive and negative classes in the two error terms and maintain the consistency of their error outputs. The specific form of the positive and negative label learning items is:
[0140]
[0141] in, Represents sample x i The predicted output on the kth label class, Y i is a sequence that contains only sample x i The set of positive labels, and is the sample x i The set of negative labels, exp is an exponential function used to amplify the penalty for extreme errors, Q is the number of sample labels, and is also the number of neurons in the output layer of the double hidden layer feedforward neural network model;
[0142] (2) In order to capture the high-order correlation characteristics between labels, a label correlation constraint term is embedded in the composite error function. The label correlation constraint term is constructed based on the label co-occurrence matrix and is used to force the network to learn the dependency relationship between labels during training. The specific form of the label correlation constraint term is:
[0143]
[0144] in, It measures the similarity of different category labels, β is the weight coefficient of the label correlation term, R kl is the label co-occurrence matrix, label co-occurrence matrix R kl The calculation formula is:
[0145]
[0146] in, is the label matrix, Q is also the number of sample labels, n is the number of samples, Y.,k represents the k-th column label vector;
[0147] (3) The final composite error function is:
[0148]
[0149] Among them, E i is the i-th error term, indicating the network's relative error to sample x i error.
[0150] This composite loss function can quantify the difference between the predicted label matrix and the initial label matrix. That is, for a positive label with a label of 1, when the model's predicted output is less than 0.5, it means that the model mistakenly classifies it as a negative label. In this case, the exponential function in the positive label error term will significantly amplify the error value. This error value is then passed back to each layer of the model through the backpropagation algorithm to guide the update of the model parameters to minimize the composite error function to ensure the consistency of the predicted label matrix and the initial label matrix.
[0151] Furthermore, the step 5.2 includes:
[0152] (1) The error of the jth neuron in the output layer is defined as:
[0153] (2) Considering that the activation function of the output layer is "sigmoid", the sigmoid activation function f'(netC j +θ j )=(1-C j )(C j ),get:
[0154]
[0155] (3) Similarly, the error of the sth neuron in the second hidden layer is defined as:
[0156] (4) Since the activation function of the second hidden layer is also "sigmoid", the sigmoid activation function f'(neth s +γ s )=(1-h s )(h s ), substitute this equation into equation (3), and we get
[0157] (5) The error of the pth hidden neuron in the first hidden layer is similarly defined as:
[0158] (6) Since the activation function of the first hidden layer is "linear", the activation function f'(netb p +δ p )=1, we get:
[0159] Furthermore, in step 5.3, the gradient of the weight matrix (Z, V, W) is:
[0160]
[0161] Bias parameter (δ p ,γ s ,θ j ) is:
[0162] in, is the gradient of the weight matrix W, V, Z, is the bias parameter θ j , γ s , δ p gradient.
[0163] When optimizing the objective function, the model is prone to falling into local minima and cannot guarantee convergence to the global optimum. In addition, the gradient descent method is extremely sensitive to the configuration of the learning rate. If the learning rate is set too high, the algorithm may hover near the optimal solution and it will be difficult to achieve stable convergence. Therefore, in this invention, the Adam optimizer is used to update the parameters.
[0164] Furthermore, the step 5.4 includes:
[0165] Step 5.4.1. Update the moment estimates of the weight matrix and bias parameters:
[0166] First, introduce the biased first-order moment estimate to perform exponential weighted averaging on the weight matrix gradient, denoted as m t , the calculation formula is:
[0167]
[0168] in, is the biased first-order moment estimate of the weight matrices W, V, and Z;
[0169] Update the biased first-order moment estimate of the bias parameter gradient of each layer:
[0170]
[0171] Where t is the current iteration number, β1 is a hyperparameter (set to 0.9) used to control the influence of the gradient; is the bias parameter θ j , γs , δ p The biased first moment estimate of ;
[0172] The learning rate of the weight matrix and bias parameters is dynamically adjusted based on the biased second-order moment estimation to avoid large parameter update steps or oscillations in the early stages of training. The biased second-order moment estimation is an exponentially weighted average of the square of the gradient, denoted as v t , the calculation formula is:
[0173]
[0174] Among them, β2 is a hyperparameter (set to 0.999) used to control the influence of the square of the gradient; is the biased second moment estimate of the weight matrices W, V, and Z; is the bias parameter θ j , γ s , δ p The biased second moment estimate of ;
[0175] Since there is a partial first-order moment m t and the partial second moment v t The estimates all start from zero, and β1 and β2 are close to 1, m t and v t The values are biased towards 0 in the first few time steps, which may lead to biased first-order moment estimates m in the first few time steps. t and the biased second moment estimate v t There is a bias; therefore, bias correction is required, which involves scaling the estimate of the moving average to compensate for the bias introduced by zero initialization;
[0176] To m t and v t Perform bias correction, that is, calculate the first-order moment estimate of the weight matrix and bias parameter bias correction and the second moment estimate
[0177] Calculate the bias-corrected first-order moment estimates of the weight matrix and bias parameters for each layer:
[0178]
[0179] in, is the bias-corrected first-order moment estimate of the weight matrices W, V, and Z, is the bias parameter θ j , γ s , δ p The bias-corrected first-order moment estimate of ; is the hyperparameter after the tth iteration;
[0180] Calculate the bias-corrected second-order raw moment estimates of the weight matrix and bias parameters for each layer:
[0181]
[0182] in, is the bias-corrected second-order moment estimate of the weight matrices W, V, and Z, is the bias parameter θ j , γ s , δ p The bias-corrected second moment estimates of ; is the hyperparameter after the tth iteration;
[0183] Step 5.4.2, update the weight matrix and bias parameters:
[0184]
[0185] Where α is the learning rate, which is set to 0.01, and ∈ is set to 10 -8 , which ensures that divide-by-zero errors are not encountered.
[0186] Step 6: Set the error threshold and the maximum number of iterations as the training termination conditions. When the error change amplitude is lower than the threshold or reaches the preset number of iterations, the training is terminated. Otherwise, repeat steps 4 to 5 until the convergence judgment criteria are met.
[0187] Furthermore, in step 6, the training termination condition is:
[0188] The preset range of iterations is: 50≤epochs≤100;
[0189] Based on the composite error function E i Calculate the error and the error change using the mean square error form:
[0190]
[0191] When the error change is less than Or terminate the training when the preset number of iterations is reached;
[0192] If the above conditions are not met, repeat steps 4 and 5 until convergence.
[0193] Step 7: After the model converges, the test set is predicted and the prediction output is binarized using a fixed threshold to form the final multi-label classification result.
[0194] Furthermore, the step 7 includes:
[0195] Make predictions on the test set T, input the feature matrix X of the test set T into the classifier f(X), and obtain the classification prediction value C k (x), according to the threshold value, the threshold selector is defined as:
[0196] h(x)={k∈Y:C k (x)>t(x)}
[0197] Set the fixed threshold t(x) to 0.5, if C k (x)>0.5, then the sample x is judged to belong to the kth category label.
[0198] The present invention proposes a multi-label text classification method based on positive and negative label learning and label correlation, aiming to break through the limitations of traditional text classification through deep learning technology and achieve high-precision and high-efficiency classification of complex text data. Its innovation lies in the construction of a composite error function. The composite error function has unique properties: first, it uses a dual constraint strategy to strictly divide the positive and negative label output thresholds, abandons empirical or inapplicable threshold setting methods, and effectively improves the accuracy of single label prediction; second, it introduces two weight factors to give different importance to the positive and negative label error terms, achieve error term balance, and enhance the prediction performance on multi-label data sets with unbalanced label categories; third, it incorporates label correlation constraints to help the network mine label category co-occurrence relationships. In terms of network training, the Adam algorithm is used to dynamically adjust the learning rate according to the historical gradient value of the parameters, and the backpropagation algorithm is combined to update the parameters, effectively improving the multi-label classification performance.
[0199] To verify the effectiveness of our method, we conducted the following experiments: We used 1-error rate, ranking loss, and average precision to measure the performance of multi-label classification methods. A higher average precision indicates better model performance. 1-error rate, in contrast to ranking loss, indicates better model performance as it decreases. The methods used in this comparative experiment included lazy learning-based methods (ML-KNN), neural network-based methods (LSIC-PS, BP-MLL, CLIF), and a method based on higher-order label correlation (HOMI).
[0200] MLNN-SOT (implemented method) was compared with five benchmark methods using five-fold cross-validation on five multi-label text datasets: Computers, Yelp, Education, Reference, and Recreation. Average precision, 1-error rate, and ranking loss were used as evaluation metrics. The experimental results, shown in Tables 1-3, show that the MLNN-SOT method, based on positive and negative label learning and label correlation, demonstrates significant advantages across all five datasets, achieving the highest average precision (e.g., 0.895 on Yelp), the lowest 1-error rate (0.145 on Yelp), and the best ranking loss (only 0.061 on Reference), comprehensively outperforming benchmark methods such as ML-KNN and BP-MLL. MLNN-SOT mitigates class imbalance by adjusting the positive and negative label error weights (λ1, λ2), enhances co-occurrence dependency modeling by combining label cosine similarity constraints, and utilizes a four-layer deep network to extract nonlinear features from high-dimensional data, enabling accurate classification in complex scenarios. In contrast, traditional methods (such as ML-KNN) are limited by linear assumptions, high-order labeling methods (such as HOMI) lack deep feature learning, and existing deep models (such as CLIF, BP-MLL, and LSIC-PS) ignore error adjustment for different categories. The stability of this method across multiple indicators and datasets demonstrates the effectiveness and generalization capabilities of its innovative design.
[0201] Table 1 shows the experimental results of the six models in average precision
[0202] ML-KNN BP-MLL LSIC CLIF HOMI Example Methods Computers 0.660 0.684 0.679 0.659 0.608 0.701 Education 0.524 0.579 0.595 0.560 0.529 0.626 Yelp 0.852 0.853 0.823 0.888 0.743 0.895 Reference 0.626 0.685 0.642 0.678 0.569 0.719 Recreation 0.499 0.588 0.592 0.607 0.593 0.654
[0203] Table 2 shows the experimental results of the six models in 1-error rate
[0204] ML-KNN BP-MLL LSIC CLIF HOMI Example Methods Computers 0.413 0.398 0.382 0.391 0.471 0.374 Education 0.629 0.549 0.519 0.513 0.648 0.484 Yelp 0.210 0.215 0.208 0.155 0.341 0.145 Reference 0.473 0.396 0.440 0.380 0.526 0.366 Recreation 0.645 0.532 0.513 0.472 0.523 0.446
[0205] Table 3 shows the experimental results of the six models in ranking loss
[0206] ML-KNN BP-MLL LSIC CLIF HOMI Example Methods Computers 0.081 0.091 0.096 0.143 0.098 0.068 Education 0.099 0.109 0.093 0.196 0.097 0.083 Yelp 0.145 0.139 0.107 0.111 0.173 0.103 Reference 0.086 0.092 0.095 0.072 0.105 0.061 Recreation 0.175 0.158 0.149 0.187 0.141 0.116
[0207] The specific embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in this field without departing from the scope of the present invention.
Claims
1. A multi-label text classification method based on positive and negative label learning and label correlation, characterized by: The method comprises: Step 1: Data preprocessing and dividing the dataset into training set and test set; Step 2: Construct a feedforward neural network model with two hidden layers, including an input layer, a first hidden layer, a second hidden layer, and an output layer; Step 3: Input the training set and initialize the double hidden layer feedforward neural network model component; read the features and label information of the samples in the training set to generate the feature matrix and label matrix; randomly initialize the weight matrix and bias parameters of the double hidden layer feedforward neural network model; Step 4, forward propagation: input the feature matrix into the input layer of the double hidden layer feedforward neural network model, and calculate the neuron output layer by layer; Step 5, back propagation: Use the composite error function to calculate the gradient of the weight matrix and bias parameters between each layer in the double hidden layer feedforward neural network model, and use the Adam optimization algorithm to dynamically adjust the gradient, with the goal of minimizing the composite error function, and update the weight matrix and bias parameters; Step 6: Set the error threshold and the maximum number of iterations as the training termination conditions. When the error change amplitude is lower than the threshold or reaches the preset number of iterations, the training is terminated. Otherwise, repeat steps 4 to 5 until the convergence judgment criteria are met. Step 7: After the model converges, the test set is predicted and the prediction output is binarized using a fixed threshold to form the final multi-label classification result.
2. The multi-label text classification method based on positive and negative label learning and label correlation according to claim 1, characterized in that: The step 1 comprises: Step 1.
1. Data preprocessing: perform word segmentation, stop word removal, and word stemming on the text data; generate a TF-IDF matrix and a Word2Vec embedding matrix, and concatenate them to form a text feature vector; Step 1.2: Obtain a text dataset. The text dataset is divided into a training set G and a test set T using a five-fold cross-validation method.
3. The multi-label text classification method based on positive and negative label learning and label correlation according to claim 1, characterized in that: The step 2 includes: Step 2.1: Build a double hidden layer feedforward neural network model: Input layer: configure neurons that match the dimension d of the text feature vector; The first hidden layer: N hidden neurons are set for preliminary feature extraction and transmission; The second hidden layer: configures M hidden neurons for deep feature extraction and abstract representation; Output layer: Set Q neurons, each neuron corresponds to a candidate category label, and outputs the predicted score; Step 2.2, the weight matrix and bias parameters are defined as follows: The input layer and the first hidden layer are fully connected, and the weight matrix from the input layer to the first hidden layer is: Z = [z qp ](1≤q≤d,1≤p≤N), The bias parameter is δ p (1≤p≤N), The first hidden layer and the second hidden layer are fully connected, and the weight matrix from the first hidden layer to the second hidden layer is: V = [v ps ](1≤p≤N,1≤s≤M), The bias parameter is γ s (1≤s≤M), The second hidden layer and the output layer are fully connected, and the weight matrix from the second hidden layer to the output layer is: W = [w sj ](1≤s≤M,1≤j≤Q), The bias parameter is θ j (1≤j≤Q), 4. The multi-label text classification method based on positive and negative label learning and label correlation according to claim 1, characterized in that: In step 3, the initialization of the dual hidden layer feedforward neural network model component includes initializing the following components: a classifier f(X) and a nonlinear mapper g(·); The classifier f(X) is used to extract sample features and output the classification prediction value C k (x), C k (x) represents the predicted output of the k-th class label of sample x in the dataset; The network architecture of the classifier f(X) is: input layer, two hidden layers, and output layer; each layer is fully connected, and the weight matrix and bias parameters between each layer are (Z, V, W) and (δ p ,γ s ,θ j ); The nonlinear mapper g(·) is located in the hidden layer and the output layer. The activation function of the first hidden layer adopts a linear activation function, and the activation functions of the second hidden layer and the output layer adopt a "Sigmoid" function.
5. The multi-label text classification method based on positive and negative label learning and label correlation according to claim 1, characterized in that: The step 4 comprises: The feature matrix Input the double hidden layer feedforward neural network model, calculate the neuron output layer by layer, and obtain the neuron output through weighted calculation and activation function processing, where d is the dimension of the text feature vector, n is the number of samples, and x is the number of samples. n is the nth sample; repeat the above steps until the output layer. The specific steps are: Step 4.1: The predicted output of the pth neuron in the first hidden layer is: b p =f(netb p +δ p ), where δ p is the bias parameter of the pth hidden neuron, netb p is the input signal of the pth hidden neuron, which is: a q is the input feature x i The qth component of Z qp is the weight matrix connecting the qth input neuron to the pth hidden neuron, and the activation function is set to a linear function; Step 4.2: Output the prediction of the sth neuron in the second hidden layer as: h s =f(neth s +γ s ), where γ s is the bias parameter of the sth hidden neuron, neth s is the input signal of the pth hidden neuron, which is: v ps is the weight matrix connecting the pth hidden neuron and the sth hidden neuron, and the activation function is set to the "sigmoid" function; Step 4.3: The predicted output of the jth neuron in the output layer is: C j =f(netC j +θ j ), where θ j is the bias parameter of the j-th output neuron, netC j is the input signal of the j-th output neuron, which is: w sj is the weight matrix connecting the sth hidden neuron and the jth hidden neuron, and the activation function is set to the "sigmoid" function.
6. The multi-label text classification method based on positive and negative label learning and label correlation according to claim 1, characterized in that: The specific steps of step 5 include: Step 5.
1. Design a composite error function E i , introduce dynamic weight coefficient to adjust the error term adaptively, in the composite error function E i Embed label correlation constraints in Step 5.2: Based on the composite error function E i Calculation error; Step 5.3, calculate the weight matrix (Z, V, W) and bias parameters (δ p ,γ s ,θ j )’s gradient; Step 5.4: Update the weight matrix and bias parameters: Use the Adam optimizer to update the weight matrix and bias parameters.
7. The multi-label text classification method based on positive and negative label learning and label correlation according to claim 6, characterized in that: In step 5.1, the composite error function is defined as the weighted sum of the following three parts, specifically including: (1) The composite error function integrates the positive and negative label learning items, and introduces two dynamic weight coefficients λ1 and λ2 into the composite error function; the specific form of the positive and negative label learning items is: in, Represents sample x i The predicted output on the kth label class, Y i is a sequence that contains only sample x i The set of positive labels, and is the sample x i The set of negative labels, exp is an exponential function used to amplify the penalty for extreme errors, Q is the number of sample labels, and is also the number of neurons in the output layer of the double hidden layer feedforward neural network model; (2) Embed a label relevance constraint term in the composite error function. The label relevance constraint term is constructed based on the label co-occurrence matrix. The specific form of the label relevance constraint term is: in, It measures the similarity of different category labels, β is the weight coefficient of the label correlation term, R kl is the label co-occurrence matrix, label co-occurrence matrix R kl The calculation formula is: in, is the label matrix, Q is also the number of sample labels, n is the number of samples, Y .,k represents the k-th column label vector; (3) The final composite error function is: Among them, E i is the i-th error term, indicating the network's relative error to sample x i error.
8. The multi-label text classification method based on positive and negative label learning and label correlation according to claim 6, characterized in that: The step 5.2 includes: (1) The error of the jth neuron in the output layer is defined as: (2) Considering that the activation function of the output layer is "sigmoid", the sigmoid activation function f'(netC j +θ j )=(1-C j )(C j ),get: (3) Similarly, the error of the sth neuron in the second hidden layer is defined as: (4) Since the activation function of the second hidden layer is also "sigmoid", the sigmoid activation function f'(neth s +γ s )=(1-h s )(h s ), and we get (5) The error of the pth hidden neuron in the first hidden layer is similarly defined as: (6) Since the activation function of the first hidden layer is "linear", the activation function f'(netb p +δ p )=1, we get:
9. The multi-label text classification method based on positive and negative label learning and label correlation according to claim 6, characterized in that: In step 5.3, the gradient of the weight matrix (Z, V, W) is: Bias parameter (δ p ,γ s ,θ j ) is: in, is the gradient of the weight matrix W, V, Z, is the bias parameter θ j , γ s , δ p gradient.
10. The multi-label text classification method based on positive and negative label learning and label correlation according to claim 6, characterized in that: The step 5.4 includes: Step 5.4.
1. Update the moment estimates of the weight matrix and bias parameters: First, introduce the biased first-order moment estimate to perform exponential weighted averaging on the weight matrix gradient, denoted as m t , the calculation formula is: in, is the biased first-order moment estimate of the weight matrices W, V, and Z; Update the biased first-order moment estimate of the bias parameter gradient of each layer: Where t is the current iteration number, and β1 is a hyperparameter used to control the influence of the gradient; is the bias parameter θ j , γ s , δ p The biased first moment estimate of ; The learning rate of the weight matrix and bias parameters is dynamically adjusted based on the biased second-order moment estimation to avoid large parameter update steps or oscillations in the early stages of training. The biased second-order moment estimation is an exponentially weighted average of the square of the gradient, denoted as v t , the calculation formula is: Among them, β2 is a hyperparameter used to control the influence of the square of the gradient; is the biased second moment estimate of the weight matrices W, V, and Z; is the bias parameter θ j , γ s , δ p The biased second moment estimate of ; Since the biased first moment estimate m t and the biased second moment estimate v t The estimates of all start from zero, which may lead to underestimates in the first few time steps, which will be biased towards 0, resulting in biased estimates; therefore, bias correction is required, which involves scaling the estimates of the moving average to compensate for the bias introduced by zero initialization; To m t and v t Perform bias correction, that is, calculate the first-order moment estimate of the weight matrix and bias parameter bias correction and the second moment estimate Calculate the bias-corrected first-order moment estimates of the weight matrix and bias parameters for each layer: in, is the bias-corrected first-order moment estimate of the weight matrices W, V, and Z, is the bias parameter θ j , γ s , δ p The bias-corrected first-order moment estimate of ; is the hyperparameter after the tth iteration; Calculate the bias-corrected second-order raw moment estimates of the weight matrix and bias parameters for each layer: in, is the bias-corrected second-order moment estimate of the weight matrices W, V, and Z, is the bias parameter θ j , γ s , δ p The bias-corrected second moment estimates of ; is the hyperparameter after the tth iteration; Step 5.4.2, update the weight matrix and bias parameters: Where α is the learning rate, ∈ is set to 10 -8 , which ensures that divide-by-zero errors are not encountered.
Citation Information
Patent Citations
Method for carrying out paper multi-label classification by utilizing deep neural network
CN112434159A
Multi-label text classification method, device and equipment
CN119829769A
Method and device for training multi-label classification model
WO2019100723A1
Cited By
Prediction method for separating and extracting lignin based on deep neural network model
CN122346667A