Text classification method based on adversarial maximization metric divergence

By constructing an adversarial divergence maximization model, combining the cross entropy loss and the confidence of the outlier detector, synergistically maximizes the measurement divergence and calibrates the abnormal data detection, the problem of the model's error classification out of text distribution is solved, and the accuracy and generalization ability of text classification are improved.

CN120336532APending Publication Date: 2025-07-18BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510403443.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing models may incorrectly classify texts from outside the distribution into a known category, resulting in a large amount of mislabeled data used during model training, affecting the classification accuracy.

Method used

A adversarial divergence maximization model is constructed, combining the cross entropy loss on the classifier in the distribution and the confidence on the outlier detector, and synergistically maximizes the divergence between the two measurements through adversarial learning methods, and abnormal data detection method is used to calibrate the divergence maximization process.

Benefits of technology

Effectively identify and exclude data from distribution, ensure that the model is trained on pure data sets, and improve the accuracy and generalization ability of text classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336532A_ABST
    Figure CN120336532A_ABST
Patent Text Reader

Abstract

The invention provides a text classification method based on adversarial maximization metric divergence. The method comprises the following steps: constructing an adversarial maximization model for an input text; the method comprises the following steps: firstly, carrying out standard measurement, and appointing two measurements by using cross entropy loss on a classifier in distribution and confidence on an outlier detector; then, the bifurcation of the two measurements is maximized in a collaborative manner by adopting an adversarial learning method; and finally, an abnormal data detection method is adopted to calibrate the divergence maximization process. The model first specifies two measurements using cross entropy loss on an intra-distribution classifier and confidence on an outlier detector. And then providing an adversarial learning method to collaboratively maximize the divergence of the two measurements. In order to ensure the optimization quality, an abnormal data detection method is provided to calibrate the divergence maximization process, so that a new normal form of text classification is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computers, and more particularly, to a text classification method based on adversarial maximization of metric disagreement. Background Art

[0002] Text classification is a fundamental task in natural language processing, which involves assigning text data to predefined categories. With the development of modern deep learning techniques, text classification has made significant progress, especially in areas such as sentiment analysis, topic classification, and information retrieval. However, deep learning models typically require a large amount of labeled data for training, which is an expensive and time-consuming task in many real-world applications. To alleviate this problem, semi-supervised text classification methods have been proposed to reduce the dependence on a large amount of labeled data.

[0003] Semi-supervised text classification combines the use of limited labeled text and a large amount of unlabeled text to train the model. This method improves the generalization ability of the model by developing various regularization techniques and consistency training methods. For example, through entropy regularization, consistency constraints, or pseudo-labeling techniques, semi-supervised learning can utilize the unlabeled text information during the training process, thus reducing the expensive need for data annotation. Nevertheless, traditional semi-supervised text classification methods usually assume that all unlabeled text comes from the same distribution as the labeled data, which is often unrealistic in real-world applications because the unlabeled dataset may contain samples that do not belong to any known category.

[0004] Therefore, researchers have recently started to explore open-set semi-supervised text classification, which is a more practical but not fully studied task. Open-set semi-supervised learning assumes that there is out-of-distribution data in the unlabeled text set, and these data do not belong to any predefined category. The main challenge of this task is the false positive inference problem, that is, the model may misclassify out-of-distribution text as a certain known category, resulting in the use of a large amount of mislabeled data during the model training process. This mislabeled data will interfere with the training process of the model and ultimately affect the classification accuracy of in-distribution text.

[0005] To address this issue, recent research work has proposed a series of methods. These methods first utilize different outlier detection techniques, such as MSP, DOC, LMCL, and LSoftmax, to identify and filter out the out-of-distribution samples in the unlabeled dataset. In this way, only those texts that are confirmed to be in-distribution will be used for the subsequent training process. The key to this method lies in the effective identification and exclusion of out-of-distribution data, thereby ensuring that the model can be trained on a more pure dataset and improving the performance of the final classification task. The successful application of these methods not only brings a new perspective to semi-supervised text classification but also provides a powerful tool for dealing with complex data distributions in the real world. Summary of the Invention

[0006] An object of an embodiment of the present disclosure is to provide a text classification method based on adversarial maximization of metric divergence, which solves the problem that the existing model may misclassify out-of-distribution text as a certain known category, resulting in the use of a large amount of mislabeled data in the model training process.

[0007] In a general aspect, a text classification method based on adversarial maximization of metric divergence is provided, including: constructing an adversarial divergence maximization model for the input text. During inference, the model comprehensively judges the text category according to different metric criteria, uses the output 0 or 1 of the sigmoid outlier detector to determine whether the sample is out-of-distribution data, and uses the output probability of the softmax classifier to assign a specific category to the sample identified as in-distribution data, while completing outlier detection and text classification.

[0008] First, perform canonical measurement, and use the cross-entropy loss on the in-distribution classifier and the confidence on the outlier detector to specify two measurements;

[0009] Then, adopt an adversarial learning method to collaboratively maximize the divergence of the two measurements;

[0010] Finally, adopt an outlier data detection method to calibrate the divergence maximization process.

[0011] The specific method of the canonical measurement is: define the measurement as M(f; x, Y), where f is the internal function, which is the process of obtaining the text representation using a pre-trained language model. The input text data x ∈ U and the in-distribution class set Y are input parameters. A softmax classifier is constructed on the in-distribution class set Y to classify the in-distribution data, and a sigmoid outlier detector is constructed to identify the out-of-distribution data. The cross-entropy loss and the confidence are specified as the measurements for the text input;

[0012] Then the measurement of the cross-entropy loss is:

[0013]

[0014] where θ is a parameter in the softmax classifier;

[0015] The measurement of the outlier detection confidence is:

[0016] M2(f, Φ; x) = σ(Φ T f(x))

[0017] where Φ is the outlier detector parameter and σ is the logistic function. The implementation of the adversarial learning method is as follows: Define the adversarial training process as:

[0018]

[0019] where λ i is a binary indicator that satisfies:

[0020]

[0021] In the minimization step, the outlier detection confidence M2 serves as the weight for the cross-entropy loss M1, and λ i serves as a switch to determine whether to maximize or minimize M2; when λ i = 1, x i is regarded as in-distribution data, and the model minimizes its corresponding loss Otherwise, when λ i = 0, x i is regarded as out-of-distribution data, and the model minimizes the negative loss i.e., the model maximizes M1;

[0022] In the maximization step, M1 becomes the weight, and the outlier detection confidence M2 is maximized or minimized accordingly according to λ i When λ i = 1, x i is regarded as in-distribution data, and the model maximizes the outlier detection confidence M2; when λ i = 0, x i is regarded as out-of-distribution data, and the model minimizes the outlier detection confidence M2. The above process keeps measuring M1 and M2 updated in opposite maximize-minimize directions, maximizing the divergence in the measurement and The specific outlier data detection method is: Specify the cross-measurement divergence as:

[0023]

[0024] Measurement M1 is normalized by max normalization to be consistent with the range [0, 1] of measurement M2; Modify the adversarial learning objective to:

[0025]

[0026] where α i is a binary indicator indicating whether x i is abnormal data, and its value is determined by whether x i satisfies the in-data consistency; since M1 and M2 are opposite measurements, that is, larger M1 and smaller M2 indicate out-of-distribution data, when the data x i is detected as abnormal data, that is, α i = 0, then reverse the optimization direction in the learning objective.

[0027] The technical effects to be achieved by the embodiments of the present invention are as follows:

[0028] First, by reexamining the outlier detection, a formal definition and assumption of the measurement disagreement are made. On this basis, an adversarial disagreement maximization model is proposed, which directly maximizes the measurement disagreement by combining measurement calibration. The model first uses the cross-entropy loss on the in-distribution classifier and the confidence on the outlier detector to specify the two measurements. Then an adversarial learning method is proposed to collaboratively maximize the disagreement between the two measurements. To ensure the optimization quality, an outlier data detection method is proposed to calibrate the disagreement maximization process. The model proposed by the present invention realizes out-of-distribution data detection for the first time by directly maximizing the metric disagreement, forming a new technical paradigm in this direction. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The above and other objects and features of the present disclosure will become more apparent from the following description with reference to the drawings.

[0030] Figure 1 is a schematic diagram showing the architecture of a text classification method based on adversarial maximization of metric disagreement according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] The following detailed description is provided to assist the reader in obtaining a comprehensive understanding of the methods, devices, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will be apparent after understanding the disclosure of the present application. For example, the order of operations described herein is merely exemplary and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of the present application, except for operations that must occur in a specific order. In addition, descriptions of features known in the art may be omitted for greater clarity and conciseness.

[0032] The features described herein can be implemented in various forms and should not be construed as limited to the examples described herein. Instead, the examples described herein are provided only to illustrate some of the many viable ways of implementing the methods, devices, and / or systems described herein, which will be apparent after understanding the disclosure of the present application.

[0033] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more thereof.

[0034] Although terms such as "first", "second", and "third" may be used herein to describe various components, components, regions, layers, or parts, these components, components, regions, layers, or parts should not be limited by these terms. Instead, these terms are only used to distinguish one component, component, region, layer, or part from another component, component, region, layer, or part. Thus, the first component, first component, first region, first layer, or first part referred to in the examples described herein may also be referred to as the second component, second component, second region, second layer, or second part without departing from the teachings of the examples.

[0035] In the specification, when an element (such as a layer, region, or substrate) is described as "on", "connected to", or "coupled to" another element, the element may be directly "on", directly "connected to", or "coupled to" the other element, or there may be one or more other elements in between. Conversely, when an element is described as "directly on", "directly connected to", or "directly coupled to" another element, there may be no other elements in between.

[0036] The terms used herein are only for describing various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. The terms "comprising", "including", and "having" specify the presence of the recited features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.

[0037] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs after understanding the present disclosure. Unless explicitly defined as such herein, terms (such as those defined in a general dictionary) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and should not be interpreted in an idealized or overly formal manner.

[0038] In addition, in the description of the examples, when a detailed description of a related structure or function that is considered to be well-known will cause an ambiguous interpretation of the present disclosure, such a detailed description will be omitted.

[0039] Figure 1 is a schematic diagram showing the architecture of a text classification method based on adversarial maximum measure divergence according to an embodiment of the present disclosure.

[0040] Optimization framework based on adversarial divergence maximization:

[0041] To solve the above problems, the present invention first makes a formal definition and assumption of the measurement divergence by reexamining the outlier detection. And on this basis, an adversarial divergence maximization (ADM) model is proposed, which directly maximizes the measurement divergence by combining measurement calibration. Its overall framework is as Figure 1 shown. In ADM, first, the cross-entropy loss on the in-distribution classifier and the confidence on the outlier detector are used to specify two measurements. Then an adversarial learning method is proposed to synergistically maximize the divergence of the two measurements. To ensure the optimization quality, an outlier data detection method is proposed to calibrate the divergence maximization process. During inference, the text category is comprehensively judged according to different metrics. The sigmoid outlier detector is used to judge whether it is out-of-distribution data, and the softmax classifier is used to assign a specific category to the data identified as in-distribution data, while completing outlier detection and text classification.

[0042] Method definition:

[0043] Re-examine outlier detection from the perspective of measurement divergence:

[0044] Formally, a measurement M(f; x, Y) is defined, which involves an internal function f and two inputs: text data x ∈ U and the in-distribution class set Y. Based on the specified measurement M, existing outlier detection methods introduce a threshold for distinguishing in-distribution and out-of-distribution data. Specifically, the measurement formula and the out-of-distribution data recognition conditions of each outlier detection method are given in Table 1. The above analysis shows that out-of-distribution data can be detected under the following assumptions: for the in-distribution data in x + ∈ U + and the out-of-distribution data in x - ∈ U - if the measurements satisfy the following inconsistency conditions and reach a specified threshold η, then it is considered that the index can detect out-of-distribution data:

[0045]

[0046] Table 1 Comparison of classification effects of ADM and baseline models on the dataset

[0047]

[0048] Formal definition of measurement divergence:

[0049] For the convenience of formal analysis, the following definitions are made:

[0050] Measurement divergence: For any in-distribution data x + ∈U + and out-of-distribution data x - ∈U - , given a specified measurement M(f; x, Y), the measurement divergence between the two data x + and x - is defined as follows:

[0051] d M (x + , x - ) = |M(f; x + , Y) - M(f; x - , Y)|

[0052] Cross-measurement divergence: For any data x ∈ U and two different specified measurements M1(f; x, Y) and M2(f; x, Y), the cross-measurement divergence of the data under these two measurements is defined as follows:

[0053] d(x, M1, M2) = |M1(f; x, Y) - M2(f; x, Y)|

[0054] ∈-bounded divergence: For a real number ∈ ≥ 0 and a given measurement M(f; x, Y), in-distribution data x + ∈U + and out-of-distribution data x - ∈U - are said to have ∈-bounded divergence if every pair (x + , x_) satisfies:

[0055] d M (x + , x - ) - ∈ ≥ 0

[0056] According to the above definition, the larger d M (x + , x - ), the easier it is for the model to distinguish x + and x - . However, since the training objectives of existing outlier detection methods are not formulated to directly increase measurement divergence, the potential advantage of expanding measurement divergence is ignored.

[0057] In addition, if d M (x + , x- ) If it satisfies ∈-bounded divergence, the worst-case divergence it can achieve is ∈. Therefore, if one wants to increase the measurement divergence of the pair (x + , x_), the divergence bound ∈ can be maximized accordingly. Existing outlier detection methods can only satisfy 0-bound divergence, which greatly limits the model's ability to distinguish in-distribution and out-of-distribution data.

[0058] Additionally, it can leverage the advantage of cross-measurement divergence in outlier detection for open-set semi-supervised text classification. To understand the possibility of cross-measurement divergence, the present solution makes the following assumptions:

[0059] Comparison consistency: For two data x, x' ∈ U + and given two different measurements M1(f; x, Y) and M2(f; x, Y), assume that these two measurements have comparison consistency. That is, if M1(f; x, Y) > M1(f; x', Y), then M2(f; x, Y) > M2(f; x', Y).

[0060] Under the above assumption, when comparing two different data, different measurement results exhibit similar behaviors. This property indicates that when optimizing one measurement on a set of data, the other measurement can be optimized accordingly. Therefore, two different measurements can be jointly optimized in a collaborative manner to maximize the measurement divergence mutually. However, existing methods only consider a single measurement and ignore the mutual enhancement between different measurements.

[0061] Intra-data consistency: For any data x ∈ U + and given two different measurements M1(f; x, Y) and M2(f; x, Y), assume that these two measurements satisfy intra-data consistency. That is, there exists a small real value δ such that for all x, d(x, M1, M2) ≤ δ.

[0062] Under the intra-data consistency assumption, it is expected that the different measurement results for each data are consistent. However, when optimizing with misidentified in-distribution or out-of-distribution data, this consistency may not be guaranteed, and such data is called abnormal data. This motivates the present invention to utilize cross-measurement divergence to detect abnormal data and calibrate the measurements during training. However, existing methods ignore this and lack a reliable mechanism to correct misidentified out-of-distribution data during training.

[0063] Measurement specification:

[0064] To perform open-set semi-supervised classification, a softmax classifier is constructed on Y to classify in-distribution data, and a sigmoid outlier detector is constructed to identify out-of-distribution data. Since out-of-distribution data may cause a large cross-entropy loss and a low outlier detection confidence for the classifier, the cross-entropy loss and the confidence are naturally regarded as measurements. First, a pre-trained language model is used to obtain text representations. This text representation process is expressed as a function f. The measurement of the cross-entropy loss can be defined as

[0065]

[0066] where θ are the parameters in the softmax classifier. The measurement of the outlier detection confidence is defined as

[0067] M2(f, Φ; x) = σ(Φ T f(x))

[0068] where Φ are the outlier detector parameters and σ is the logistic function. Now two measurements M1 and M2 are specified.

[0069] Adversarial learning maximizes the disagreement:

[0070] To synergistically maximize the disagreement between the two measurements, the present invention proposes an adversarial learning method to iteratively amplify the two disagreements. The following adversarial training process is defined:

[0071]

[0072] where λ i is a binary indicator that satisfies:

[0073]

[0074] The above objective adopts a min-max optimization process. In the minimization step, the outlier detection confidence M2 serves as the weight for the cross-entropy loss M1, and λ i serves as a switch to determine whether to maximize or minimize M2. When λ i = 1, x i is regarded as in-distribution data, and the model minimizes its corresponding loss Otherwise, when λ i = 0, x i is regarded as out-of-distribution data, and the model minimizes the negative loss i.e., the model maximizes M1. Similarly, in the maximization step, M1 becomes the weight, and the outlier detection confidence M2 is maximized or minimized accordingly according to λ i . In this way, the measurements M1 and M2 are updated in opposite directions, maximizing the disagreement in the measurements and Thus increasing the divergence bound.

[0075] Anomaly detection and measurement calibration:

[0076] Binary indicator λ i depends on the value of the outlier detector M2(f, Φ; x i ). During adversarial training, since the outlier detector has no supervision signal, incorrect indicators are inevitably generated. The incorrect indicator λ i will reverse the optimization direction and lead to performance degradation. However, there is no information to guide the discrimination of data that may lead to incorrect indicators.

[0077] Fortunately, the consistency assumption in the data provides an opportunity to alleviate this problem. When the outlier detector misclassifies data, it is very likely that it fails to satisfy the within-data consistency. This property is used to detect outlier examples and reverse the corresponding indicator λ i to maintain the correct optimization direction. To achieve outlier data detection, the cross-measurement divergence is specified as:

[0078]

[0079] The measurement M1 is normalized by maximum normalization to be consistent with the range [0,1] of the measurement M2. Under this specification, the adversarial learning objective is modified to:

[0080]

[0081] where α i is a binary indicator indicating whether x i is outlier data, and its value is determined by whether x i satisfies the within-data consistency. Since M1 and M2 are opposite measurements, i.e., larger M1 and smaller M2 indicate out-of-distribution data. When there is a large margin between M1 and M2, they are more consistent. Under this design, when the data x i is detected as outlier data, i.e., α i = 0, the model will reverse the optimization direction in the learning objective. For outlier data, at this time λ i = 0, and it is necessary to reduce M1 through the minimization process and increase M2 through the maximization process to narrow the margin between the two, which makes the consistency between measurements match the outlier data and calibrates the measurement to the correct optimization direction, ensuring the effectiveness of maximizing the adversarial divergence.

[0082] Optimization process:

[0083] Pre-stage1: To provide an initial model, a classifier is pre-trained using labeled text and cross-entropy loss, considering the labeled text as in-distribution data and the unlabeled text as out-of-distribution data, and pre-training the data of the outlier detector using BCE loss.

[0084] Pre-stage2: Use the model trained in Pre-stage1 to assign pseudo-labels to the unlabeled text, and further refine the classifier and outlier detector using the pseudo-labeled text.

[0085] ADM: Perform adversarial divergence maximization to iteratively optimize the two metrics, and further update the classifier and outlier detector by combining measurement calibration.

[0086] Experimental study:

[0087] Three benchmark datasets created using existing text classification datasets in previous work are used to evaluate the open-set semi-supervised classification task, including AGNews, DBPedia, and Yahoo. These datasets are widely used for text classification and contain a sufficient number of common classes to evaluate the classification model. The results are shown in Tables 2 and 3. ADM outperforms the pipeline model and the EM algorithm model comprehensively on all nine sub-datasets. The results show that the adversarial divergence maximization method is effective and more effective in challenging environments.

[0088] Another observation is that in most settings, the outlier detection (Out) accuracy results have a significant advantage. These results indicate that the superiority of ADM lies mainly in its good performance in detecting out-of-distribution data, although it sacrifices the classification results of in-distribution data in some settings.

[0089] To analyze the contributions of each component and each training stage in ADM, the present invention conducts an ablation study. The ablation results are shown in Table 4. The ablation results show that Pre-stage1 provides an initial model for ADM, and Pre-stage2 further improves on the basis of Pre-stage1. However, when ADM is trained without outlier data detection, the performance of the ADM 0 threshold is even worse than that of Pre-stage2. When an appropriate outlier data detection threshold is adopted, ADM significantly improves the classification results. This result shows that outlier detection and measurement calibration ensure the effectiveness of ADM.

[0090] Table 2 Comparison of the classification effects of ADM and baseline models on the dataset

[0091]

[0092] Table 3 In-distribution data (In) and out-of-distribution data monitoring accuracy (Out) of ADM

[0093]

[0094] Table 4 Ablation study of ADM on the AGNews dataset

[0095]

[0096] Although some embodiments of the technical solution of the present invention have been shown and described, those skilled in the art should understand that these embodiments can be modified without departing from the principles and spirit of the present disclosure as defined by the claims and their equivalents.

Claims

1. A text classification method based on adversarial maximization of metric divergence, characterized in that Including: Construct an adversarial divergence maximization model for the input text; First, perform canonical measurement, using the cross-entropy loss on the in-distribution classifier and the confidence on the outlier detector to specify two measurements; Then, adopt an adversarial learning method to jointly maximize the divergence of the two measurements; Finally, adopt an outlier data detection method to calibrate the divergence maximization process. During inference, comprehensively judge the text category according to different metrics, use the sigmoid outlier detector to judge whether it is out-of-distribution data, use the softmax classifier to assign a specific category to the data identified as in-distribution data, and simultaneously complete outlier detection and text classification.

2. The text classification method based on adversarial maximum metric disagreement according to claim 1, wherein The specific method of the canonical measurement is: Define the measurement as M(f; x, Y), where f is the internal function, which is the process of obtaining the text representation using the pre-trained language model. The input text data x ∈ U and the in-distribution class set Y are input parameters. Construct a softmax classifier on the in-distribution class set Y to classify the in-distribution data, and construct a sigmoid outlier detector to identify the out-of-distribution data. Specify the cross-entropy loss and the confidence as the measurements for the text input; Then the measurement of the cross-entropy loss is: where θ is the parameter in the softmax classifier; The measurement of the outlier detection confidence is: M2(f, Φ; x) = σ(Φ T f(x)) where Φ is the outlier detector parameter and σ is the logistic function.

3. The text classification method based on adversarial maximum measure divergence according to claim 2, characterized in that, The implementation method of the adversarial learning method is: Define the adversarial training process as: where λ i is a binary index satisfying: In the minimization step, the outlier detection confidence M2 serves as the weight of the cross-entropy loss M1, λ i serves as a switch to determine whether to maximize or minimize M2; when λ i = 1, x i is regarded as in-distribution data, and the model minimizes its corresponding loss Otherwise, when λ i = 0, x i is regarded as out-of-distribution data, and the model minimizes the negative loss i.e., the model maximizes M1; In the maximization step, M1 becomes the weight, and the outlier detection confidence M2 is maximized or minimized according to λ i Accordingly, when λ i = 1, x i is regarded as in-distribution data, and the model maximizes the outlier detection confidence M2; when λ i = 0, x i is regarded as out-of-distribution data, and the model minimizes the outlier detection confidence M2. The above process keeps measuring M1 and M2 updated in opposite directions to maximize the disagreement in the measurement and 4. The text classification method based on adversarial maximum measure divergence according to claim 3, characterized in that The specific outlier data detection method is: Specify the cross-measurement divergence as: The measurement M1 is normalized by maximum normalization to keep it consistent with the range [0,1] of the measurement M2; Modify the adversarial learning objective to: where α i is a binary indicator indicating whether x i is abnormal data, and its value is determined by whether x i satisfies the in-data consistency; since M1 and M2 are opposite measurements, i.e., larger M1 and smaller M2 indicate out-of-distribution data, when the data x i is detected as abnormal data, i.e., α i = 0, then reverse the optimization direction in the learning objective.