Online Machine Learning Method Based on Interval Distribution for Text Multi-Label Classification Task
The method optimizes gap distribution in online text multi-label classification by simultaneously learning score and threshold models, enhancing performance and reducing computational costs.
Patent Information
- Application Number
- CN202210877759.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-07-25
AI Technical Summary
The existing online text multi-label classification method ignores the association between the two when learning threshold model and score model, resulting in poor classification performance and inability to effectively handle text multi-label classification tasks in large-scale or streaming modes.
By optimizing interval distribution, the score model and threshold model are learned simultaneously by using the online gradient descent method, combined with Mercer's core technique to achieve nonlinear multi-label prediction, and the online gradient descent method is used to solve the first-order approximation problem, and the efficient update of the multi-label classifier is achieved.
Improves the generalization performance of multi-label classifiers, reduces time and space costs, and is suitable for text multi-label classification tasks in large-scale or streaming modes.
Smart Images

Figure CN115081448B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of machine learning and text multi-label classification, and particularly relates to an online machine learning method. Background Art
[0002] With the development of the information age, the text multi-label classification problem has become a very important classical problem, which is widely applied in fields such as news recommendation, public opinion monitoring, and online review sentiment analysis. For the text multi-label classification problem, each text instance can be associated with multiple labels, namely the so-called "relevant labels". The task of text multi-label classification is to identify all relevant labels of the input text instance from a predefined label set.
[0003] In most offline methods, to obtain text multi-label prediction, first a scoring model needs to be trained to output a real-valued score for each label of the text instance, and then an additional threshold model needs to be trained or simply a constant is used as the threshold to convert the multi-label score into the final multi-label classification result. Once a new batch of training data arrives, the offline method incurs an expensive cost due to the need to retrain the model. Therefore, the offline method is not suitable for large-scale or potentially streaming-mode text multi-label classification applications. Different from the offline method, the online method processes data one by one, updates the multi-label classifier incrementally, and makes online multi-label predictions, which can provide an effective solution for such tasks. To achieve online multi-label prediction, the online method should learn the scoring model and the threshold model simultaneously in an online manner.
[0004] Existing online text multi-label classification methods such as OSML-ELM (Online Sequential Multi-label Extreme Learning Machines) use an offline post-processing program to determine its global label threshold. The ELM-OMLL (ELM based Online Multi-Label Learning) algorithm directly fixes its global label threshold to 0. There are also some methods that separately adjust tree-based, clustering-based, and Bayesian-based techniques to make them suitable for the online text multi-label classification setting, and then calculate the label cardinality on a fixed or variable-length sample window as the number of relevant labels to be predicted. The above methods generally have the problem that the learning of the threshold model is independent of the learning of the scoring model, thus ignoring the association between the two, which misses the opportunity to use the scoring model to learn an improved threshold model. To overcome this defect, the FALT (First-order Adaptive label thresholding) method is based on the ALT (Adaptive label thresholding) framework, extends the idea of maximizing the minimum margin of SVM (Support Vector Machine) to multi-label applications, and takes both the scoring model and the threshold model as important components of the online multi-label classifier, and learns them simultaneously in an online manner. The FALT method can achieve relatively excellent classification performance, but this method only considers the margin at a single point and ignores the margin distribution.
[0005] Recent research has shown that optimizing the margin distribution is more crucial for generalization performance. mlODM (multi-label Optimal Margin Distribution Machine) addresses this discovery by using the definition of the ranking margin in Rank-SVM to optimize the mean and variance of the margins for all label pairs to improve the prediction performance of the multi-label classifier. However, this method has many limitations. First, since mlODM uses the ranking margin, it needs to optimize a large number of variables, thus having higher requirements for time and space costs. Second, to convert the label ranking into a classification result, mlODM requires a post-processing program to learn the threshold function. Finally, mlODM is an offline method and cannot incrementally update its ranking model and make online predictions. Therefore, this method cannot handle large-scale or potentially streaming-mode text multi-label classification tasks. Summary of the Invention
[0006] In view of the deficiencies in the prior art, the present invention provides an online machine learning method based on interval distribution for text multi-label classification tasks. By utilizing the concept of optimizing interval distribution, the score model and the threshold model are learned simultaneously in an online manner. When updating the multi-label classifier, the online gradient descent method is used to solve the first-order approximation problem of the original problem, so as to efficiently update the score model and the threshold model incrementally. Finally, the kernel trick is used to further improve and generalize the present invention.
[0007] The object of the present invention is achieved as follows: An online machine learning method based on interval distribution for text multi-label classification tasks, comprising the following steps:
[0008] 1) In the th round, receive the feature vector of the document in this round ;
[0009] 2) Enter the multi-label prediction program, and use the text multi-label classifier to predict the relevant label set; the classifier contains score predictors for calculating label scores and an additional predictor for determining label thresholds , each score predictor assigns a real-valued score to each label in the label set , and at the same time calculates the threshold of the instance through the threshold predictor . Predict the labels with scores higher than the threshold as 's relevant labels, and predict the labels with label scores lower than the threshold as non-relevant labels, that is, the predicted relevant label set is: ; ; ;
[0010] 3) After the prediction, receive 's true relevant label set , enter the online update program of the multi-label classifier, update the current multi-label classifier , and obtain the multi-label classifier for the next round;
[0011] 4) Return to step 1), wait to receive the text feature vector of the th round; if no new text feature vectors are input, exit the program.
[0012] As a further limitation of the present invention, the specific update steps of the multi-label prediction program in step 2) are as follows:
[0013] a) In the Round, the program receives complete text data information ;
[0014] b) For any calculate respectively 、 、 、 ,
[0015] wherein, , , , if is true, then value is 1, otherwise 0;
[0016] c) For any ,calculate respectively ,
[0017] where is about at partial derivative, that is ;
[0018] d) According to the following formula, improve the current multi-label classifier to a new multi-label classifier ,the formula is:
[0019]
[0020] where , is used to control the trade-off between the first term and the last two terms, is a predefined parameter, used to trade off different deviations, , and respectively represent the number of relevant labels and non-relevant labels of the instance .
[0021] As a further limitation of the present invention, the online machine learning method based on interval distribution is promoted by using Mercer kernels:
[0022] For any ,from accumulation can obtain ,therefore each classifier in can be expressed as ,where
[0023] Introduce a Mercer kernel , through the non-linear feature mapping corresponding to this Mercer kernel , a non-linear classifier can be obtained , where is obtained by replacing the appearing in with . Using the kernel trick, all label scores and thresholds can be efficiently calculated by the following formula:
[0024] ;
[0025] where is the round of text instance after feature vector mapping, used to calculate the and inner product. This process does not require explicit calculation of , but instead realizes non-linear multi-label prediction by replacing the inner product in the feature space with a simple kernel function operation.
[0026] As a further limitation of the present invention, before running the method of the present invention, initialization is required. The initialization method is: for any , let , and at the same time, hyperparameters , and need to be determined in advance. In addition, for the kernelization method, in the case of using the RBF kernel, the hyperparameter also needs to be determined in advance.
[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0028] The present invention considers the interval distribution and has better generalization performance. Therefore, the results on multiple multi-label performance evaluation indicators are better. At the same time, due to the characteristics of the online method, the method of the present invention has lower time and space costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0030] Figure 1 is the flowchart of the present invention.
[0031] Figure 2 It is the convex surrogate function of the 0-1 loss function in the present invention. Detailed implementation manners
[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0033] As Figure 1 shown, an online machine learning method based on interval distribution for text multi-label classification tasks is as follows:
[0034] (1) In the th round, the program of the present invention receives the feature vector of the document in this round ;
[0035] (2) Enter the multi-label prediction program, and use the multi-label classifier to predict the relevant label set. The classifier contains score predictors for calculating label scores and an additional predictor for determining the label threshold. Each score predictor assigns a real-valued score to each label in the label set of size , and at the same time calculates the threshold of the instance through the threshold predictor . Labels with scores higher than the threshold are predicted as relevant labels, while those labels with scores lower than the threshold are predicted as non-relevant labels, that is, the predicted relevant label set is: . .
[0036] (3) After the prediction is completed, receive the true relevant label set , enter the online update program of the multi-label classifier, and update the current multi-label classifier to obtain the multi-label classifier for the next round;
[0037] (4) Return to step 1), and wait to receive the text feature vector of the th round; if no new text feature vector is input, exit the program.
[0038] According to the above process, the core program of the method of the present invention is the multi-label classifier online update program. In the multi-label prediction program, according to its prediction rules, the following definition of the margin will naturally be obtained:
[0039]
[0040] Among them , if a relevant label is correctly predicted, then ; if an irrelevant label is correctly predicted, then ; the method of the present invention improves the current multi-label classifier into a new multi-label classifier ; specifically, is obtained by solving the following problem:
[0041]
[0042] Among them, the first term of the objective function is to ensure that is as close as possible to so as to retain the information learned from the previous, represents the Frobenius norm of the matrix. The second and third terms are to minimize the standard deviation of the margin. is used to control the trade-off between the first term and the latter two terms. and respectively represent the number of relevant labels and non-relevant labels of the instance . By assuming that the mean of the margin is 1, the deviation between each margin and the mean of the margin is either or , and for any , and at most only one of them is positive. At the same time, and values can be regarded as the penalty for deviating from the mean of the margin. If , the deviation at this time will result in a smaller margin, thus more likely to produce incorrect predictions, so this deviation should be more punished. On the contrary, if , the deviation at this time will bring a larger margin and should be given a smaller penalty. Therefore, the latter two terms of the objective function are converted into
[0043]
[0044] Among them, is a predefined parameter used to weigh different deviations. To avoid overly strict penalties that would turn all samples into support vectors, we draw on the idea of the insensitive loss function and tolerate deviations less than . Only deviations greater than are penalized, where . At the same time, to ensure that is a proper convex surrogate for the 0-1 loss, it is further transformed into , as shown in Figure 2 .
[0045] Therefore, the classifier is finally updated by solving the following problem:
[0046]
[0047] Next, by slightly transforming and simplifying Problem (2), an efficient closed-form solution can be derived. Specifically, all the constraints in Problem (2) can be eliminated by setting the following:
[0048]
[0049] In this way, Problem (2) can be transformed into an unconstrained problem:
[0050]
[0051] where is defined as
[0052]
[0053] Since directly solving Problem (3) is relatively complex, the method of the present invention uses the first-order approximation of to replace , where is the partial derivative of with respect to at , that is, . After eliminating the constant terms unrelated to , Problem (3) is transformed into:
[0054]
[0055] The advantage of this simplification is that an efficient closed-form solution can be derived. By solving Problem (5), the following update can be obtained:
[0056]
[0057] where ,
[0058] and , , , , where, if is true, then has a value of 1, otherwise 0.
[0059] Therefore, the steps of the multi-label classifier online update program of the present invention are as follows:
[0060] (1) In the round, the program receives complete text data information ;
[0061] (2) For any , calculate respectively , , , ;
[0062] (3) For any , calculate respectively ;
[0063] (4) According to formula (6), improve the current multi-label classifier into a new multi-label classifier .
[0064] Thus, the margin distribution based adaptive label thresholding online learning method MDALT (Margin Distribution based Adaptive Label Thresholding) for text multi-label classification proposed by the method of the present invention is obtained.
[0065] To further improve the method of the present invention, MDALT can be generalized using Mercer kernels. For any , from accumulation, we can get , so each classifier in can be expressed as , where
[0066] Introduce a Mercer kernel , through the non-linear feature mapping corresponding to this Mercer kernel, we can obtain the non-linear classifier , where It is obtained by replacing appearing in with . Then, using the kernel trick, all the label scores and thresholds can be efficiently calculated by the following formula:
[0067]
[0068] where is the feature vector after mapping the text instance in the -th round, which is used to calculate the inner product of and . Therefore, this process does not require explicit calculation of , but instead realizes non-linear multi-label prediction by replacing the inner product in the feature space with a simple kernel function operation.
[0069] Before running the method of the present invention, initialization is required. A common initialization method is: for any , let , and at the same time, the hyperparameters , and need to be determined in advance. In addition, for the kernelized method, in the case of using the RBF kernel, the hyperparameter also needs to be determined in advance.
[0070] Embodiment
[0071] To make the technical solutions and advantages of the present invention clearer, a preferred embodiment of the MDALT algorithm proposed by the present invention is given below. The specific implementation steps are as follows:
[0072] 1. Set the hyperparameters required for running this method: , , . The parameters corresponding to different text datasets are also different. Changing the hyperparameters will result in different classification effects.
[0073] 2. Initialization operation: for any , let , and obtain the initial classifier .
[0074] 3. For the -th text entering the program, where , perform the following steps:
[0075] 3.1. Obtain the feature vector corresponding to the current text data ;
[0076] 3.2. Online multi-label classification prediction: Let the predicted relevant label set be: ;
[0077] 3.3. Query and receive the true relevant label set , and then perform online update on the multi-label classifier:
[0078] 3.3.1. For any , respectively let , , , .
[0079] 3.3.2. For any , calculate the partial derivative of the first-order approximation of the loss function with respect to at :
[0080]
[0081] 3.3.3. According to the online gradient descent method, improve the current multi-label classifier to a new multi-label classifier :
[0082]
[0083] 3.4. Obtain the next-round data, and go back to step 3.1. If no new data is input, exit the loop.
[0084] Two benchmark datasets Tmc2007 and medical for text multi-label classification are selected in the experiment to verify the effectiveness and superiority of the method proposed in the present invention. Each row vector of the dataset is a feature vector obtained after feature extraction of a document. The comparison methods selected in the experiment are as follows:
[0085] OSML-ELM: Online Sequential Multi-Label Extreme Learning Machine.
[0086] ELM-OMLL: Extreme Learning Machine based Online Multi-Label Learning
[0087] mlODM: multi-label Optimal margin Distribution Machine。
[0088] FALT(l) and FALT(k): linear ALT algorithm and its kernel extension.
[0089] MDALT (l) and MDALT (k): The algorithm proposed in the present invention for updating the multi-label classifier using the OGD method and its kernel extension.
[0090] For the regularization parameter of the mlODM algorithm, and FALT(l) and MDALT (l) algorithm's step size parameter , from search among them. For the mlODM algorithm and the MDALT algorithm proposed in the present invention, the hyperparameters and are both searched within . For all kernelization methods, use the RBF kernel with hyperparameter , that is , and search for the kernel hyperparameter from . All hyperparameter optimizations are obtained through ten-fold cross-validation on each training set.
[0091] To simulate a longer data stream, the evaluation model in the online method is obtained after multiple passes through the training data. For the Medical dataset with sparse features, the number of passes through the training data is set to 30, while for the Tmc2007 dataset, the number of passes through the training data is set to 10. For the two ELM-based algorithms, the result of one pass is sometimes better than that of multiple passes, so we always report the better result. Each method is executed 20 times on 2 datasets, and the order of the training set data is randomly shuffled each time. Finally, the average of the performance metric values obtained by running each method 20 times on the test set data, as well as the average training and test times, are recorded.
[0092] According to the paired t-test at the 95% confidence level, the best result and the comparable results for each metric are shown in bold:
[0093] Table 1: Performance evaluation results of each algorithm on the Tmc2007 dataset ("-" indicates that the result could not be obtained within 48 hours)
[0094]
[0095] Table 2: Performance evaluation results of each algorithm on the medical dataset
[0096]
[0097] Tables 1 and 2 show the performance results of MDALT proposed in the present invention and the algorithms OSML-ELM, ELM-OMLL, mlODM, and FALT on the benchmark datasets Tmc2007 and medical for text multi-label classification respectively. Among them, the multi-label classification performance evaluation metrics include:
[0098] Sample-based evaluation metrics: Precision (Psn), Recall (Rcal), F1-measure (F1), Hamming loss (Hl), and Ranking loss (Rl). Denote as the total number of samples for evaluation, and the operator is used to measure the "symmetric difference" between two sets, represents the real-valued score assigned to the label corresponding to the instance , then the calculation methods of the above-mentioned various measurement metrics are as follows: , , , , .
[0099] Label-based evaluation metrics: MacroF1 and MicroF1. Denote , , as the number of true positives, false positives, and false negatives of the test sample with respect to the label respectively, then , .
[0100] It can be found from Tables 1 and 2 that for most performance metrics, the performance of OSML-ELM, ELM-OMLL, and mlODM is significantly lower than that of MDALT. At the same time, MDALT proposed in the present invention is superior to or at least comparable to FALT in most metrics.
[0101] Table 3: Average time taken for algorithm training and testing
[0102]
[0103] Table 3 summarizes the time spent on training and testing by five algorithms, namely MDALT, OSML-ELM, ELM-OMLL, mlODM, and FALT, on the benchmark datasets Tmc2007 and medical.
[0104] As can be seen from Table 3, compared with other online algorithms, the offline mlODM is very inefficient in terms of both training and testing time, indicating that mlODM can only be applied to very small-scale datasets. In terms of training time, the MDALT and FALT proposed in the present invention are more effective than the two ELM-based methods, and the time spent by FALT and MDALT is comparable. In terms of testing time, MDALT and FALT require similar time. In the linear version, they predict faster than the two ELM-based methods, but slower in the kernelized version because the number of support vectors they retain is larger.
[0105] The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. An online machine learning method based on interval distribution for text multi-label classification tasks, characterized in that, including the following steps: 1) In the round, the feature vector of the document for this round is received ; 2) Enter the multi-label prediction program and use the text multi-label classifier right The relevant label set is predicted; classifier Included A score predictor for computing label scores and an additional predictor for determining label thresholds , each score predictor is of size Label set Each tag in Assign a real-valued score , while passing the threshold predictor Calculate the instance threshold , predict the labels with scores above the threshold as The relevant tags are predicted as follows: , the specific steps of updating the multi-label prediction program are as follows: a) In the round, this program received the complete text data information ; b) For any calculate respectively , , , , Among them, , , , , if is true, then has a value of 1, otherwise 0; c) For any , calculate respectively , wherein is with respect to at i.e., the partial derivative ; d) Improve the current multi-label classifier into a new multi-label classifier according to the following formula , and the formula is: wherein , is used to control the trade-off between the first item and the last two items, is a predefined parameter for weighing different deviations, , and respectively represent the numbers of relevant tags and irrelevant tags of the instance ; 3) After the prediction ends, receive the true relevant tag set , enter the online update program of the multi-label classifier to update the current multi-label classifier , and obtain the multi-label classifier for the next round ; 4) Return to step 1) and wait to receive the text feature vectors of the n-th round; if no new text feature vectors are input, then exit the program.
2. The online machine learning method based on interval distribution for text multi-label classification tasks according to claim 1, characterized in that, generalizing the online machine learning method based on the interval distribution by using Mercer kernels: For any , by accumulative addition, we can get . Therefore, each classifier in can be expressed as , where Introduce a Mercer kernel , through the non-linear feature mapping corresponding to this Mercer kernel , a non-linear classifier can be obtained , where is obtained by replacing appearing in with . Using the kernel trick, all label scores and thresholds can be efficiently computed by the following formula: ; Among them is the round of text instances feature vectors after mapping, used to calculate the inner product with This process does not require explicit calculation of , but instead realizes non-linear multi-label prediction by replacing the inner product in the feature space with a simple kernel function operation.
3. The online machine learning method based on interval distribution for text multi-label classification tasks according to claim 1 or 2, characterized in that Before running the method of the present invention, initialization is required. The initialization method is as follows: For any , let . At the same time, hyperparameters , and need to be determined in advance. In addition, for the kernelization method, in the case of using the RBF kernel, the hyperparameter also needs to be determined in advance.