Confidence-Guided Text Classification Method, Device, and Computer Equipment
By utilizing the confidence-guided method in text classification, the confidence of the target text and constructing a loss function optimization model, the adaptability problem of the deep learning model under distribution offset is solved, and the robustness and speed improvement of text classification is achieved.
Patent Information
- Application Number
- CN202210992878.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-08-18
AI Technical Summary
In the text processing field, deep learning models are difficult to effectively adapt to actual text classification applications due to distribution offsets. Existing methods such as fine-tuning, domain adaptation and test time adaptation have limitations, especially in passive domain adaptation, it is difficult to effectively improve the robustness of text classification.
By inputting the target text to be classified into the pre-trained model, calculate the confidence of it being divided into each text category, and construct a loss function based on the confidence, the number of text categories and batch size, optimize the text classification model, and use the updated model for classification.
Without relying on additional information, the robustness and speed of text classification are significantly improved, ensuring the accuracy and efficiency of classification.
Smart Images

Figure CN115292496B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, and in particular, to a text classification method, device, and computer device based on confidence guidance. Background Art
[0002] Deep learning models have shown good results in the field of text processing. However, due to distribution shift, that is, the training text distribution is different from the test text distribution, it is still difficult to deploy deep learning models to actual text classification applications, and this problem is a fundamental task in the field of text processing. To solve this problem, many sub - fields have been proposed under different settings, such as fine - tuning, domain adaptation, and test - time adaptation.
[0003] Recently, researchers have proposed full test - time adaptation, that is, adapting the source pre - trained model by learning from unlabeled test texts at test time. Test - time adaptation is also known as source - free domain adaptation. Different from domain adaptation which requires access to the source domain and the target domain, source - free domain adaptation does not need to obtain any text data from the source domain for adaptation. Some existing works use generative models to support feature alignment without source texts. Another popular direction is to fine - tune the source pre - trained model without explicitly performing domain alignment. For example, Test Entropy Minimization (TENT) adopts a pre - trained model and adapts to test data by updating the trainable parameters of the Batchnorm layer using entropy minimization; Source Hypothesis Transfer (SHOT) adapts using both entropy minimization and a diversity regularizer. SHOT needs to use source texts to train a dedicated source model, uses label smoothing techniques with weight normalization layers; TTT needs to retrain the source text to promote supervised adaptation of the target text and has an additional auxiliary rotation prediction branch, making it impossible to reuse existing pre - trained models. Summary of the Invention
[0004] Based on this, in view of the above - mentioned technical problems, it is necessary to provide a text classification method, device, and computer device based on confidence guidance to improve the robustness of text classification.
[0005] A text classification method based on confidence guidance, the method includes:
[0006] Input the target text to be classified into a pre - trained text classification model, and obtain the confidence that the target text is assigned to each text category respectively; the confidence is obtained by taking the logarithm of the softmax function value corresponding to the target text;
[0007] Construct a loss function according to the confidence, the number of text categories, and the product of the batch size of the input target text, and update the text classification model by optimizing the loss function;
[0008] Classify the target text by using the updated text classification model.
[0009] Preferably, construct a loss function according to the product of the confidence level, the number of text categories, and the batch size of the input target text, including:
[0010] Construct the first loss function according to the product of the confidence level, the number of text categories, and the batch size of the input target text as:
[0011]
[0012] where, L conf (f θ (x t ),y t ) is the first loss function, θ, t are text classification model parameters, N is the batch size of the target text, C is the number of text categories for classification, z i is the output value corresponding to the i-th text category.
[0013] Preferably, before constructing the loss function according to the product of the confidence level, the number of text categories, and the batch size of the input target text, further include:
[0014] Calculate the difference between the first largest confidence level and the second largest confidence level corresponding to each of the target texts;
[0015] When the difference is less than a preset threshold, assign a first attention coefficient to the corresponding target text;
[0016] When the difference is not less than the preset threshold, assign a second attention coefficient to the corresponding target text.
[0017] Preferably, construct a loss function according to the product of the confidence level, the number of text categories, and the batch size of the input target text, including:
[0018] Construct a second loss function according to the first attention coefficient, the second attention coefficient, and the first loss function:
[0019]
[0020] where, L cgtta (f θ (x t ),y t ) is the second loss function, α is the second attention coefficient, and β is the first attention coefficient.
[0021] Preferably, calculating the difference between the first largest confidence level and the second largest confidence level corresponding to each of the target texts includes:
[0022] diff(f θ (x t )) = conf(cos θ (x t )) 1st -conf(f θ (x t )) 2nd
[0023] Among them, diff(f θ (x t ) is the confidence difference, conf(f θ (x t ) is the confidence, conf(f θ (x t ) 1st is the first largest confidence, conf(f θ (x t ) 2nd is the second largest confidence.
[0024] Preferably, when the difference is less than a preset threshold, a first attention coefficient is assigned to the corresponding target text, and when the difference is not less than the preset threshold, a second attention coefficient assigned to the corresponding target text is:
[0025]
[0026] Among them, y′ t is the output of the text classification model after the attention coefficient is assigned, is the output of the text classification model before the attention coefficient is assigned when the difference is not less than the preset threshold, is the output of the text classification model before the attention coefficient is assigned when the difference is less than the preset threshold, g th is the preset threshold.
[0027] A text classification device based on confidence guidance, the device includes:
[0028] A confidence calculation module, configured to input a target text to be classified into a pre-trained text classification model, and respectively obtain the confidence of the target text being classified into each text category; the confidence is obtained by taking the logarithm of the softmax value corresponding to the target text;
[0029] A model update module, configured to construct a loss function according to the product of the confidence, the number of text categories, and the batch size of the input target text, and update the text classification model by optimizing the loss function;
[0030] A text classification module, configured to classify the target text by using the updated text classification model.
[0031] A computer device includes a memory and a processor. The memory stores a computer program. It is characterized in that when the processor executes the computer program, the following steps are implemented:
[0032] Input the target text to be classified into a pre-trained text classification model to obtain the confidence levels of the target text being classified into each text category respectively; the confidence level is obtained by taking the logarithm of the softmax function value corresponding to the target text;
[0033] Construct a loss function according to the confidence level and the product of the number of text categories and the batch size of the input target text, and update the text classification model by optimizing the loss function;
[0034] Use the updated text classification model to classify the target text.
[0035] The above text classification method, device and computer device based on confidence-guided first input the target text to be classified into a pre-trained text classification model to obtain the confidence levels of the target text being classified into each text category respectively, where the confidence level is obtained by taking the logarithm of the softmax function value corresponding to the target text, then construct a loss function according to the confidence level and the product of the number of text categories and the batch size of the input target text, update the text classification model by optimizing the loss function, and finally use the updated text classification model to classify the target text. The present invention enables an off-the-shelf source pre-trained text classification model to adapt to the target text in an online manner, combines the batch size of the input text and the number of text categories, and uses the confidence information of the target text to guide the gradient optimization of the loss function, greatly improving the text classification speed on the premise of ensuring text classification accuracy. Using the present invention can greatly improve the robustness of text classification. Description of the Drawings
[0036] Figure 1 It is a schematic flowchart of a text classification method based on confidence-guided in an embodiment;
[0037] Figure 2 It is a structural block diagram of a text classification device based on confidence-guided in an embodiment;
[0038] Figure 3 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments
[0039] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0040] First, review the scenario of test-time adaptation: Given a model f pre-trained on source text data (x, y), with parameters θ, the goal of test-time adaptation is to reduce the gap in test text data by optimizing the model parameters during the test phase without touching the source data. Generally, test-time adaptation also cannot access the test-phase text data labels, but only the model parameters and the unlabeled target text data. θ (x), with parameters θ, the goal of test-time adaptation is to reduce the gap in test text data by optimizing the model parameters during the test phase without touching the source data. Generally, test-time adaptation also cannot access the test-phase text data labels, but only the model parameters and the unlabeled target text data.
[0041] In one embodiment, as Figure 1 shown, a confidence-guided text classification method is provided, including the following steps:
[0042] Step 102, input the target text to be classified into the pre-trained text classification model, and obtain the confidence levels of the target text being assigned to each text category respectively.
[0043] Among them, the confidence level is obtained by taking the logarithm of the softmax function value corresponding to the target text.
[0044] Consider a C-classification problem, where the target data and label can be represented as x t , y t , and represent the logit output by the model as The softmax equation can be expressed as:
[0045]
[0046] The confidence level calculation formula is specifically as follows:
[0047]
[0048] Among them, conf(f θ (x t ) is the confidence level, θ, t are the text classification model parameters, C is the number of text categories for classification, and z i is the output value corresponding to the i-th text category.
[0049] The target text can be input into the pre-trained text classification model in batches.
[0050] Step 104, construct a loss function based on the confidence level, the product of the number of text categories, and the batch size of the input target text, and update the text classification model by optimizing the loss function.
[0051] The most commonly used loss function in test-time domain adaptation is entropy minimization (tent), expressed as
[0052]
[0053] Although entropy can represent the uncertainty of prediction and entropy minimization can increase the model confidence to a certain extent, compared with the confidence-guided text classification method, entropy minimization is a sub-optimal choice. This method adapts the source pre-trained text classification model to the target text in an online manner, directly uses the confidence information to guide the gradient optimization of the loss function, making the text classification model converge faster and perform better, and thus greatly improving the text classification speed on the premise of ensuring the text classification accuracy.
[0054] Step 106: Classify the target text using the updated text classification model.
[0055] For the above-mentioned confidence-guided text classification method, first, the target text to be classified is input into the pre-trained text classification model to obtain the confidence levels of the target text being assigned to each text category respectively, where the confidence level is obtained by taking the logarithm of the softmax function value corresponding to the target text. Then, a loss function is constructed based on the confidence level, the number of text categories, and the product of the batch size of the input target text. The text classification model is updated by optimizing the loss function. Finally, the updated text classification model is used to classify the target text. The present invention adapts the off-the-shelf source pre-trained text classification model to the target text in an online manner, combines the batch size of the input text and the number of text categories, and uses the confidence information of the target text to guide the gradient optimization of the loss function, greatly improving the text classification speed on the premise of ensuring the text classification accuracy. Using this method can greatly improve the robustness of text classification.
[0056] In one embodiment, the first loss function is constructed based on the product of the confidence level, the number of text categories, and the batch size of the input target text as follows:
[0057]
[0058] where, L conf (f θ (x t ),y t ) is the first loss function, θ, t are the text classification model parameters, N is the batch size of the target text, C is the number of text categories for classification, and z i is the output value corresponding to the i-th text category.
[0059] It can be seen that the confidence loss function provides a direct connection between model confidence and model optimization. Noting that the gradient of log(·) is its reciprocal, which is much larger than the gradient of entropy-based methods. Therefore, when the model starts training, the confidence of the C-1 classes will quickly drop to 0, and the maximum confidence will quickly dominate the model optimization. This means that confidence-related optimization strategies can achieve faster convergence. It will encourage those confused samples to follow the better-classified samples. Since the model prediction is directly associated with confidence, entropy can only be regarded as a sub-optimal choice for model updates.
[0060] In one embodiment, before constructing the loss function according to the product of the confidence, the number of text categories, and the batch size of the input target text, it further includes:
[0061] Calculating the difference between the first-largest confidence and the second-largest confidence corresponding to each target text respectively:
[0062] diff(f θ (x t ))=conf(f θ (x t )) 1st -conf(f θ (x t )) 2nd
[0063] where diff(f θ (x t ) is the confidence difference, conf(f θ (x t ) is the confidence, conf(f θ (x t ) 1st is the first-largest confidence, and conf(f θ (x t ) 2nd is the second-largest confidence.
[0064] Sort the confidences of each target text obtained by calculation for each category it is assigned to, calculate the difference between the first-largest confidence and the second-largest confidence corresponding to each target text, and further define a difference threshold to distinguish easily classified samples and difficult-to-classify samples. Note that the test-time domain adaptation is carried out in an unsupervised setting. When a target text passes through the text classification model, if the difference is greater than the predefined difference threshold, the model will regard it as an easily classified sample. When the difference is less than the predefined difference threshold, it means that the model is confused about this sample, and additional attention needs to be given to the confused text, that is, dynamically weight the unlabeled test target text. It can be understood that the first attention coefficient should be greater than the second attention coefficient.
[0065] When the difference is less than the preset threshold, assign the first attention coefficient to the corresponding target text. When the difference is not less than the preset threshold, assign the second attention coefficient to the corresponding target text.
[0066] Calculate the difference between the first highest confidence and the second highest confidence corresponding to each target text. When the difference is less than the preset threshold, assign the first attention coefficient to the corresponding target text. When the difference is not less than the preset threshold, assign the second attention coefficient to the corresponding target text:
[0067]
[0068] where y′ t is the output of the text classification model after assigning the attention coefficient, is the output of the text classification model when the difference is not less than the preset threshold before assigning the attention coefficient, is the output of the text classification model when the difference is less than the preset threshold before assigning the attention coefficient. α is the second attention coefficient, β is the first attention coefficient, and g th is the preset threshold.
[0069] In one embodiment, construct a loss function according to the product of the confidence, the number of text categories, and the batch size of the input target text, including:
[0070] Construct a second loss function according to the first attention coefficient, the second attention coefficient, and the first loss function:
[0071]
[0072] where L cgtta (f θ (x t ),y t ) is the second loss function.
[0073] It can be understood that this embodiment is equivalent to defining the concepts of good classification texts and confused classification texts for unlabeled test target texts, and proposing a dynamic weighting mechanism, which further improves the classification performance of the model, thereby greatly improving the classification speed on the premise of ensuring the accuracy of text classification.
[0074] This method has no additional requirements for the source pre-trained model. It uses an off-the-shelf source pre-trained model and adapts it to the target text in an online manner. No additional auxiliary information is introduced here because this auxiliary information needs to be adjusted according to the text dataset, which may be difficult to achieve in real scenarios. This method is a general method that can adapt to any source pre-trained text classification model. Introducing effective additional information will correspondingly improve the performance of the text classification model and further improve the accuracy of text classification.
[0075] It should be understood that although Figure 1 the steps in the flowchart are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in
[0076] In one embodiment, as Figure 2 shown, a text classification device based on confidence guidance is provided, including: a confidence calculation module, an attention assignment module, a model update module, and a text classification module, where:
[0077] The confidence calculation module is configured to input the target text to be classified into a pre-trained text classification model, and respectively obtain the confidence levels of the target text being classified into each text category; the confidence level is obtained by taking the logarithm of the softmax value corresponding to the target text;
[0078] The model update module is configured to construct a loss function according to the product of the confidence level, the number of text categories, and the batch size of the input target text, and update the text classification model by optimizing the loss function;
[0079] The text classification module is configured to classify the target text by using the updated text classification model.
[0080] In one embodiment, the model update module is further configured to construct a first loss function according to the product of the confidence level, the number of text categories, and the batch size of the input target text:
[0081]
[0082] where, L conf (f θ (x t ), y t ) is the first loss function, θ, t are text classification model parameters, N is the batch size of the target text, C is the number of text categories for classification, and z i is the output value corresponding to the i-th text category.
[0083] In one embodiment, the model update module is further configured to construct a second loss function according to the first attention coefficient, the second attention coefficient, and the first loss function:
[0084]
[0085] Among them, L cgtta (f θ (x t ), y t ) is the second loss function, α is the second attention coefficient, and β is the first attention coefficient.
[0086] For the specific definition of the confidence-guided text classification device, please refer to the definition of the confidence-guided text classification method above, which will not be repeated here. Each module in the above-mentioned confidence-guided text classification device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0087] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store text data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a text classification method based on confidence guidance is implemented.
[0088] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0089] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method in the above embodiment when executing the computer program.
[0090] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method in the above embodiment are implemented.
[0091] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0092] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0093] The above-described embodiments merely represent several implementation manners of this application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.
Claims
1. A text classification method based on confidence guidance, characterized in that The method includes: Inputting the target text to be classified into a pre-trained text classification model to obtain the confidence levels of the target text being classified into each text category respectively; the confidence level is obtained by taking the logarithm of the softmax function value corresponding to the target text; Constructing a loss function based on the confidence level, the product of the number of text categories, and the batch size of the input target text, and updating the text classification model by optimizing the loss function; Classifying the target text by using the updated text classification model; Constructing a loss function based on the confidence level, the product of the number of text categories, and the batch size of the input target text, including: Constructing a first loss function based on the confidence level, the product of the number of text categories, and the batch size of the input target text; Among them, ) is the first loss function, is the parameter of the text classification model, is the batch size of the target text, is the number of text categories for classification, is the output value corresponding to the text category; Before constructing a loss function based on the confidence level, the product of the number of text categories, and the batch size of the input target text, it further includes: Calculating the difference between the first largest confidence level and the second largest confidence level corresponding to each target text respectively; When the difference is less than a preset threshold, assigning a first attention coefficient to the corresponding target text; When the difference is not less than the preset threshold, assigning a second attention coefficient to the corresponding target text; Constructing a loss function based on the confidence level, the product of the number of text categories, and the batch size of the input target text, including: Constructing a second loss function based on the first attention coefficient, the second attention coefficient, and the first loss function; Among them, ) is the second loss function, is the second attention coefficient, is the first attention coefficient.
2. The method according to claim 1, wherein Calculating the difference between the first largest confidence level and the second largest confidence level corresponding to each target text, including: Calculating the difference between the first largest confidence level and the second largest confidence level corresponding to each target text as: Among them, is the confidence difference, is the confidence, is the largest confidence, is the second largest confidence.
3. The method according to claim 1, characterized in that, When the difference is less than a preset threshold, assigning a first attention coefficient to the corresponding target text, and when the difference is not less than the preset threshold, assigning a second attention coefficient to the corresponding target text, including: When the difference is less than a preset threshold, assigning a first attention coefficient to the corresponding target text, and when the difference is not less than the preset threshold, assigning a second attention coefficient to the corresponding target text as: Among them, is the output of the text classification model after the attention coefficient is assigned, is the output of the text classification model before the attention coefficient is assigned when the difference is not less than the preset threshold, is the output of the text classification model before the attention coefficient is assigned when the difference is less than the preset threshold, is the preset threshold.
4. A confidence-guided text classification device applied to the method according to any one of claims 1-3, characterized in that, The device includes: A confidence level calculation module, configured to input the target text to be classified into a pre-trained text classification model to obtain the confidence levels of the target text being classified into each text category respectively; the confidence level is obtained by taking the logarithm of the softmax value corresponding to the target text; A model update module, configured to construct a loss function based on the confidence level, the product of the number of text categories, and the batch size of the input target text, and update the text classification model by optimizing the loss function; A text classification module, configured to classify the target text by using the updated text classification model.
5. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Classification model training method and device, computer equipment and storage medium
CN113849648A
Text classification method and apparatus, electronic device, and storage medium
WO2022160449A1