Text statement classification method, classification device, electronic device and storage medium

By introducing a combination of feature constraints and multiple models into the text classification model, the problems of small sample size and poor relevance of text labels in industrial scenarios are solved, and high-precision text statement classification is achieved.

CN115510232BActive Publication Date: 2025-06-20PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211201108.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-28
Publication Date
2025-06-20
Estimated Expiration
2042-09-28

AI Technical Summary

Technical Problem

In industrial scenarios, the size of the annotated sample is small, the sample distribution is uneven, and the correlation between the content of the text and the label is not obvious, resulting in the problem of poor classification results of the existing text classification model.

Method used

A text statement classification method is proposed. By obtaining the training sample set and the initial text statement classification model, combining the pre-trained sub-model, feature constraint sub-model and text classification sub-model, feature extraction, feature constraints and text classification processing is performed, and model parameters are adjusted to improve classification accuracy.

Benefits of technology

In the case of small sample size and poor pre-label correlation of text statements, high-precision text statement classification can be fully classified for recognized statements to improve the classification accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115510232B_ABST
    Figure CN115510232B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method and apparatus for classifying text statements, an electronic device, and a storage medium, belonging to the field of artificial intelligence technology. The method includes: inputting each sample pair data in a training sample set into an initial text statement classification model, respectively extracting features of the sample statements to obtain first text features and second text features; performing feature constraint on the first text features to update the first text features; performing text classification processing on the first text features to obtain multiple sample prediction probability values; performing numerical comparison on the multiple sample prediction probability values to determine a target sample label; adjusting the model parameters of the initial text statement classification model according to the positive sample label and the target sample label to obtain a target text statement classification model; and classifying the obtained initial text statement through the target text statement classification model to obtain a target category. The embodiment of the present application can improve the accuracy of classifying text statements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and particularly to a method and device for classifying text sentences, an electronic device, and a storage medium. Background Art

[0002] Currently, the text classification task is a basic and important task in natural language processing. This task has extremely many applications in industrial scenarios, such as negative sentiment recognition, intent recognition, etc. However, in actual industrial scenarios, problems such as a small number of labeled samples, uneven sample distribution, and unclear correlation between the text content and the label make the classification effect of existing text classification models not good. Therefore, how to improve the accuracy of text sentence classification in a few-shot scenario has become a technical problem to be solved urgently. Summary of the Invention

[0003] The main purpose of the embodiments of the present application is to propose a method and device for classifying text sentences, an electronic device, and a storage medium, aiming to improve the accuracy of the model in classifying text sentences.

[0004] To achieve the above object, a first aspect of the embodiments of the present application proposes a method for classifying text sentences, and the method includes:

[0005] Obtain a training sample set, where the training sample set includes a plurality of sample pair data, and each sample pair data includes a positive sample sentence, a positive sample label corresponding to the positive sample sentence, a negative sample sentence, and a negative sample label corresponding to the negative sample sentence;

[0006] Obtain an initial text sentence classification model, where the initial text sentence classification model includes a pre-trained sub-model, a feature constraint sub-model, and a text classification sub-model;

[0007] Input the positive sample sentence and the negative sample sentence of each sample pair data into the initial text sentence classification model, and respectively extract features of the positive sample sentence and the negative sample sentence through the pre-trained sub-model to obtain a first text feature of the positive sample sentence and a second text feature of the negative sample sentence;

[0008] Perform feature constraint on the first text feature through the feature constraint sub-model and the second text feature to update the first text feature;

[0009] Perform text classification processing on the first text feature through the text classification sub-model to obtain a sample prediction probability value of the positive sample sentence belonging to each category label;

[0010] Perform numerical comparison on the multiple sample prediction probability values of the positive sample sentence to determine the target sample label of the positive sample sentence;

[0011] Adjust the model parameters of the initial text statement classification model according to the positive sample labels and the target sample labels of the positive sample statements, and continue to train the adjusted initial text statement classification model based on the training sample set until the model loss value of the initial text statement classification model meets the preset training end condition, so as to obtain a target text statement classification model;

[0012] Obtain an initial text statement to be classified, and classify the initial text statement through the target text statement classification model to obtain a target category.

[0013] In some embodiments, after the text classification sub-model performs text classification processing on the first text feature to obtain the sample prediction probability values of the positive sample statement belonging to each category label, the method further includes:

[0014] Perform feature constraint calculation on the first text feature, the second text feature, the positive sample label, and the negative sample label through the feature constraint sub-model to obtain a contrast loss value;

[0015] Obtain a cross-entropy loss value according to the positive sample label and the sample prediction probability value;

[0016] Obtain a model loss value according to the contrast loss value and the cross-entropy loss value.

[0017] In some embodiments, the pre-training sub-model includes feature encoding processing and self-attention processing. Inputting the positive sample statement and the negative sample statement of each sample pair data into the initial text statement classification model, and respectively performing feature extraction on the positive sample statement and the negative sample statement through the pre-training sub-model to obtain the first text feature of the positive sample statement and the second text feature of the negative sample statement includes:

[0018] Input the positive sample statement and the negative sample statement of each sample pair data into the initial text statement classification model;

[0019] Perform the feature encoding processing on each text word in the positive sample statement and each text word in the negative sample statement respectively to obtain the positive sample word features corresponding to the positive sample statement and the negative sample word features corresponding to the negative sample statement;

[0020] Perform the self-attention processing on all the positive sample word features and all the negative sample word features respectively to obtain the first text feature of the positive sample statement and the second text feature of the negative sample statement.

[0021] In some embodiments, before inputting the positive sample statement and the negative sample statement of each sample pair data into the initial text statement classification model, the method further includes:

[0022] Compare the lengths of the positive sample statement and the negative sample statement respectively according to a preset text length threshold. When the text length of the positive sample statement is less than the text length threshold, perform a zero-padding operation on the positive sample statement according to the text length threshold, and update the positive sample statement until the text length of the positive sample statement is equal to the text length threshold;

[0023] When the text length of the negative sample statement is less than the text length threshold, perform a zero-padding operation on the negative sample statement according to the text length threshold, and update the negative sample statement until the text length of the negative sample statement is equal to the text length threshold.

[0024] In some embodiments, the calculating the contrast loss value by performing feature constraint calculation on the first text feature, the second text feature, the positive sample label, and the negative sample label through the feature constraint sub-model includes:

[0025] Compare the positive sample label and the negative sample label according to a preset indicator function to determine a contrast coefficient;

[0026] Perform feature constraint calculation according to the first text feature, the second text feature, and the contrast coefficient to obtain a contrast loss value.

[0027] In some embodiments, the obtaining the cross-entropy loss value according to the positive sample label and the sample prediction probability value includes:

[0028] Obtain the number of characters in the positive sample statement;

[0029] Calculate according to a preset cross-entropy loss function using the number of characters, the positive sample label, and the sample prediction probability value to obtain a cross-entropy loss value.

[0030] To achieve the above object, a second aspect of the embodiments of the present application proposes a text statement classification device, and the device includes:

[0031] A sample set acquisition module, configured to acquire a training sample set, where the training sample set includes a plurality of sample pair data, and each sample pair data includes a positive sample statement, a positive sample label corresponding to the positive sample statement, a negative sample statement, and a negative sample label corresponding to the negative sample statement;

[0032] An initial model acquisition module, configured to acquire an initial text statement classification model, where the initial text statement classification model includes a pre-trained sub-model, a feature constraint sub-model, and a text classification sub-model;

[0033] A feature extraction module, configured to input the positive sample statement and the negative sample statement of each sample pair data into the initial text statement classification model, and respectively perform feature extraction on the positive sample statement and the negative sample statement through the pre-trained sub-model to obtain a first text feature of the positive sample statement and a second text feature of the negative sample statement;

[0034] A feature constraint module, configured to perform feature constraint on the first text feature through the feature constraint sub-model and the second text feature to update the first text feature;

[0035] A sample text classification module, configured to perform text classification processing on the first text feature through the text classification sub-model to obtain a sample prediction probability value of the positive sample statement belonging to each category label;

[0036] A numerical comparison module, configured to perform numerical comparison on multiple sample prediction probability values of the positive sample statement to determine a target sample label of the positive sample statement;

[0037] A target model construction module, configured to adjust model parameters of the initial text statement classification model according to the positive sample label and the target sample label of the positive sample statement, and continue to train the adjusted initial text statement classification model based on the training sample set until a model loss value of the initial text statement classification model meets a preset training end condition, so as to obtain a target text statement classification model;

[0038] A target text classification module, configured to acquire an initial text statement to be classified, and classify the initial text statement through the target text statement classification model to obtain a target category.

[0039] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, where the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0040] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, where the storage medium is a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.

[0041] A text statement classification method, classification device, electronic device, and storage medium proposed in an embodiment of the present application can achieve high-precision text statement classification for statements to be recognized even when the sample size is small and the relevance of pre-labels of text statements is not strong by combining positive and negative samples and feature constraints. First, a target text statement classification model is constructed. Specifically, a training sample set is obtained. The training sample set includes multiple sample pair data, and each sample pair data includes a positive sample statement, a positive sample label corresponding to the positive sample statement, a negative sample statement, and a negative sample label corresponding to the negative sample statement. An initial text statement classification model is obtained. The initial text statement classification model includes a pre-training sub-model, a feature constraint sub-model, and a text classification sub-model. Then, the positive sample statement and the negative sample statement of each sample pair data are input into the initial text statement classification model. The pre-training sub-model extracts features from the positive sample statement and the negative sample statement respectively to obtain the first text feature of the positive sample statement and the second text feature of the negative sample statement. After that, the first text feature is feature-constrained by the feature constraint sub-model and the second text feature to update the first text feature. The text classification sub-model performs text classification processing on the first text feature to obtain the sample prediction probability values of the positive sample statement belonging to each category label. The multiple sample prediction probability values obtained for the positive sample statement are numerically compared to determine the target sample label of the positive sample statement. Finally, the model parameters of the initial text statement classification model are adjusted according to the positive sample label and the target sample label of the positive sample statement, and the adjusted initial text statement classification model is continuously trained based on the training sample set until the model loss value of the initial text statement classification model meets the preset training end condition to obtain the target text statement classification model. An initial text statement to be classified is obtained, and the initial text statement is classified by the target text statement classification model to obtain the target category. The embodiment of the present application can improve the accuracy of the model in classifying text statements. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is the first flowchart of the text statement classification method provided by an embodiment of the present application;

[0043] Figure 2 is Figure 1 the flowchart of step S103 in

[0044] Figure 3 is the second flowchart of the text statement classification method provided by an embodiment of the present application;

[0045] Figure 4 is the third flowchart of the text statement classification method provided by an embodiment of the present application;

[0046] Figure 5 is Figure 4 the flowchart of step S401 in

[0047] Figure 6 is Figure 4 a flowchart of step S402 in

[0048] Figure 7 a schematic structural diagram of a text statement classification device provided by an embodiment of the present application;

[0049] Figure 8 a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0050] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0051] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application, and are not intended to limit this application.

[0053] First, several nouns involved in the present application are analyzed:

[0054] Artificial Intelligence (AI): It is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theories, methods, technologies and application systems.

[0055] Natural Language Processing (NLP): NLP uses computers to process, understand, and apply human languages (such as Chinese, English, etc.). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. Natural language processing includes syntactic analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intent recognition, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and linguistic research related to language computing, etc.

[0056] Unified Language Model (UniLM model): It is a pre-trained language model that can be fine-tuned for natural language understanding and generation tasks. The model is pre-trained using three types of language modeling tasks: unidirectional, bidirectional, and sequence-to-sequence prediction. It can use a shared Transformer network and utilize specific attention masks to control the context of the prediction conditions.

[0057] Metric Learning: A metric (or distance function) is a function that defines the distance between elements in a set. A set with a metric is called a metric space. Metric learning is also commonly known as similarity learning. The purpose of distance measure learning is to measure the degree of similarity between samples.

[0058] BERT (Bidirectional Encoder Representation from Transformers) model: used to further increase the generalization ability of the word vector model, fully describe the characteristics of character level, word level, sentence level and even inter-sentence relationship, and is built based on Transformer. There are three types of Embedding in BERT, namely Token Embedding, Segment Embedding, and Position Embedding; Token Embeddings is a word vector, and the first word is the CLS mark, which can be used for subsequent classification tasks; Segment Embeddings is used to distinguish between two sentences, because pre-training not only does LM but also does classification tasks with two sentences as input; Position Embeddings, the position word vector here is not the trigonometric function in Transfor, but is learned by BERT through training. However, BERT directly trains a Position Embedding to retain the position information, randomly initializes a vector for each position, adds model training, and finally obtains an Embedding containing position information. Finally, in the combination of this Position Embedding and word Embedding, BERT chooses direct splicing.

[0059] Contrastive Loss: It is mainly used for dimensionality reduction, that is, similar samples, after dimensionality reduction (feature extraction), make the two samples still similar in the feature space, while the originally dissimilar samples are still dissimilar in the feature space. Similarly, this loss function can also well express the matching degree of paired samples.

[0060] Cross Entropy Loss: It is the most commonly used loss function in classification. Cross entropy is used to measure the difference between two probability distributions to measure the difference between the distribution learned by the model and the true distribution.

[0061] At present, text classification is a basic and important task in natural language processing. This task has many applications in industrial scenarios, such as negative emotion recognition and intent recognition. With the emergence of pre-trained models such as BERT, existing technical solutions often only complete classification directly through pre-trained models. However, in actual industrial scenarios, the existing text classification models have poor classification effects due to the small number of labeled samples, uneven sample distribution, and unclear correlation between text content and labels. Therefore, how to improve the accuracy of text sentence classification in a few-sample scenario has become a technical problem that needs to be solved urgently.

[0062] Based on this, the embodiments of the present application provide a text statement classification method, a classification device, an electronic device, and a storage medium, aiming to improve the accuracy of the model in classifying text statements.

[0063] The text statement classification method, classification device, electronic device, and storage medium provided by the embodiments of the present application will be specifically described through the following embodiments. First, the text statement classification method in the embodiments of the present application will be described.

[0064] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0065] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0066] The text statement classification method provided by the embodiments of the present application relates to the field of artificial intelligence technology. The text statement classification method provided by the embodiments of the present application can be applied to a terminal, can also be applied to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms; the software can be an application that implements the text statement classification method, etc., but is not limited to the above forms.

[0067] This application can be used in numerous general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0068] Figure 1 is an optional flowchart of the text statement classification method provided by an embodiment of this application, Figure 1 The method in may include but is not limited to steps S101 to S108.

[0069] Step S101, obtain a training sample set, where the training sample set includes multiple sample pair data, and each sample pair data includes a positive sample statement, a positive sample label corresponding to the positive sample statement, a negative sample statement, and a negative sample label corresponding to the negative sample statement;

[0070] Step S102, obtain an initial text statement classification model, where the initial text statement classification model includes a pre-trained sub-model, a feature constraint sub-model, and a text classification sub-model;

[0071] Step S103, input the positive sample statement and the negative sample statement of each sample pair data into the initial text statement classification model, and respectively extract features of the positive sample statement and the negative sample statement through the pre-trained sub-model to obtain a first text feature of the positive sample statement and a second text feature of the negative sample statement;

[0072] Step S104, perform feature constraint on the first text feature through the feature constraint sub-model and the second text feature to update the first text feature;

[0073] Step S105, perform text classification processing on the first text feature through the text classification sub-model to obtain sample prediction probability values of the positive sample statement belonging to each category label;

[0074] Step S106, perform numerical comparison on multiple sample prediction probability values of the positive sample statement to determine the target sample label of the positive sample statement;

[0075] Step S107: Adjust the model parameters of the initial text statement classification model according to the positive sample labels and target sample labels of the positive sample statements, and continue to train the adjusted initial text statement classification model based on the training sample set until the model loss value of the initial text statement classification model meets the preset training end condition, so as to obtain the target text statement classification model;

[0076] Step S108: Obtain the initial text statement to be classified, and classify the initial text statement through the target text statement classification model to obtain the target category.

[0077] Steps S101 to S108 shown in the embodiments of the present application, by combining positive and negative samples and feature constraints, can achieve high-precision text statement classification of the statements to be recognized even when the sample size is small and the relevance of the pre-labels of the text statements is not strong. First, construct the target text statement classification model. Specifically, obtain the training sample set, which includes multiple sample pair data. Each sample pair data includes a positive sample statement, the positive sample label corresponding to the positive sample statement, a negative sample statement, and the negative sample label corresponding to the negative sample statement. Obtain the initial text statement classification model, which includes a pre-trained sub-model, a feature constraint sub-model, and a text classification sub-model. Then, input the positive sample statement and negative sample statement of each sample pair data into the initial text statement classification model, and respectively extract the features of the positive sample statement and negative sample statement through the pre-trained sub-model to obtain the first text feature of the positive sample statement and the second text feature of the negative sample statement. Then, perform feature constraint on the first text feature through the feature constraint sub-model and the second text feature to update the first text feature; perform text classification processing on the first text feature through the text classification sub-model to obtain the sample prediction probability values of the positive sample statement belonging to each category label. Compare the multiple sample prediction probability values obtained for the positive sample statement numerically to determine the target sample label of the positive sample statement. Finally, adjust the model parameters of the initial text statement classification model according to the positive sample labels and target sample labels of the positive sample statements, and continue to train the adjusted initial text statement classification model based on the training sample set until the model loss value of the initial text statement classification model meets the preset training end condition, so as to obtain the target text statement classification model. Obtain the initial text statement to be classified, and classify the initial text statement through the target text statement classification model to obtain the target category. The embodiments of the present application can improve the accuracy of the model in classifying text statements.

[0078] In step S101 of some embodiments, the positive sample statement and the negative sample statement in the sample pair data may be sample statements obtained from the same document or from different documents, that is, the embodiments of the present application do not limit the relationship between the sample statements in the sample pair data. The positive sample label and the negative sample label may be specified by a technician or extracted from the corresponding sample statement, and are not limited thereto. The positive sample label is used to represent the classification of the positive sample statement, and the negative sample label is used to represent the classification of the negative sample statement.

[0079] Exemplarily, if x represents the positive sample statement and y represents the negative sample statement, then a sample pair data can be expressed as (x, y), and the training sample set can be expressed as ((x1, y1), (x2, y2),..., (xm, ym)), where m represents the number of sample pair data. For example, when extracting positive sample statements from multiple statements in the movie review text, the positive sample statement can be "This movie is very good-looking", and the positive sample label of this positive sample statement is "good-looking". When extracting negative sample statements from multiple statements in the movie review text, the negative sample statement can be "This movie is really ugly", and the negative sample label of this negative sample statement is "ugly". Therefore, the positive sample statement and the negative sample statement extracted above and their corresponding labels can form a sample pair data.

[0080] It should be noted that in order to solve problems such as a small number of labeled samples, unbalanced sample distribution, and unclear correlation between the content of the text and the label in the existing application scenarios, when the label category data distribution in the training sample set is relatively balanced, the obtained positive sample statements and negative sample statements can be cross-matched to expand the training sample pair data.

[0081] It should be noted that when the label category data distribution in the training sample set is unbalanced, in the security monitoring problem, most of the positive samples are normal people, and the available negative samples are quite scarce. If the full amount of samples is used to train a simple and highly accurate binary classification model, the result will be severely biased towards normal people, resulting in the failure of the model. In order to balance the sample pair data of different category labels in the training sample set, the full amount of sample pairs can be used to match the training sample set with a small number of labeled samples. For example, in the case of binary classification, assuming the number of positive sample statements is n1 and the number of negative sample statements is n2, then the total number of samples in the training sample set is n1 + n2. The first sample statement can be matched with all other sample statements except itself to generate (n1 + n2) * (n1 + n2 - 1) / 2 sample pair data, so as to make the labels more diverse, enrich the training sample set at the same time, and improve the accuracy of model training.

[0082] In step S102 of some embodiments, in order to improve the accuracy of text statement classification, an initial text statement classification model is obtained. The initial text statement classification model includes a pre-training sub-model, a feature constraint sub-model, and a text classification sub-model.

[0083] It should be noted that the pre-training sub-model is used to extract text information from the input text statement. The feature constraint sub-model is used to constrain and adjust the representation of the text statement features based on the metric learning of the extracted text statement features. The text classification sub-model is used to form the final classification result according to the obtained features.

[0084] It should be noted that the initial text statement classification model can be a Bert model, a UniLM model, an ELECTRA model, etc.

[0085] In step S103 of some embodiments, the pre-training sub-model is used to extract text information from the input text statement. The positive sample statement and the negative sample statement of each sample pair data are input into the initial text statement classification model. The pre-training sub-model respectively extracts features from the positive sample statement and the negative sample statement to obtain the first text feature of the positive sample statement and the second text feature of the negative sample statement. The first text feature and the second text feature are used to represent the deep semantic information features of the corresponding sample statements.

[0086] Please refer to Figure 2 , in some embodiments, the pre-training sub-model includes feature encoding processing and self-attention processing. Step S103 may include but is not limited to steps S201 to S203:

[0087] Step S201, input the positive sample statement and the negative sample statement of each sample pair data into the initial text statement classification model;

[0088] Step S202, respectively perform feature encoding processing on each text word in the positive sample statement and each text word in the negative sample statement to obtain the positive sample word features corresponding to the positive sample statement and the negative sample word features corresponding to the negative sample statement;

[0089] Step S203, respectively perform self-attention processing on all the positive sample word features and all the negative sample word features to obtain the first text feature of the positive sample statement and the second text feature of the negative sample statement.

[0090] In steps S201 to S202 of some embodiments, after inputting the positive sample statement and the negative sample statement of each sample pair of data into the initial text statement classification model, each text word in the positive sample statement is subjected to feature encoding processing through a pre-trained sub-model to obtain positive sample word features corresponding to the positive sample statement; each text word in the negative sample statement is subjected to feature encoding processing through a pre-trained sub-model to obtain negative sample word features corresponding to the negative sample statement.

[0091] It should be noted that when using the Bert model as the initial text statement classification model, each text word in the positive sample statement is numbered through the Bert dictionary, and through feature encoding processing of each number, positive sample word features corresponding to each statement word in the positive sample statement are obtained, that is, token feature vectors corresponding to each statement word are obtained. Similarly, each text word in the negative sample statement is numbered through the Bert dictionary, and through feature encoding processing of each number, negative sample word features corresponding to each statement word in the negative sample statement are obtained, that is, token feature vectors corresponding to each statement word are obtained, and the token feature vectors are used to represent the feature meanings of the corresponding statement words.

[0092] In step S203 of some embodiments, in order to extract more deeply informative feature vectors while reducing the data volume, the self-attention mechanism of the pre-trained sub-model is used to perform self-attention processing on all positive sample word features and all negative sample word features respectively to obtain a first text feature of the positive sample statement and a second text feature of the negative sample statement. The first text feature and the second text feature are used to represent the text features after deeply fusing the feature information between text words through the self-attention mechanism.

[0093] It should be noted that the first text feature is obtained by summing the token feature vectors of all positive sample sub-features of the corresponding positive sample statement, and the second text feature is obtained by summing the token feature vectors of all positive sample sub-features of the corresponding negative sample statement.

[0094] Please refer to Figure 3 , in some embodiments, before step S103, the text statement classification method provided by the embodiments of the present application may further include but is not limited to steps S301 to S302:

[0095] Step S301, respectively compare the lengths of the positive sample statement and the negative sample statement according to a preset text length threshold. When the text length of the positive sample statement is less than the text length threshold, perform a zero-padding operation on the positive sample statement according to the text length threshold, and update the positive sample statement until the text length of the positive sample statement is equal to the text length threshold;

[0096] Step S302: When the text length of the negative sample statement is less than the text length threshold, perform zero-padding on the negative sample statement according to the text length threshold, and update the negative sample statement until the text length of the negative sample statement is equal to the text length threshold.

[0097] In step S301 of some embodiments, to improve the training efficiency of the model, the text lengths of the positive sample statements and negative sample statements in a batch input into the initial text statement classification model are adjusted to ensure that the text lengths are consistent. Specifically, the text lengths of the positive sample statements and negative sample statements are respectively compared with the preset text length threshold. When the text length of the positive sample statement is less than the text length threshold, perform zero-padding on the positive sample statement according to the text length threshold, that is, append zeros at the end of the original sample statement until the text length of the positive sample statement is equal to the text length threshold, and use the concatenated sample statement as the new positive sample statement.

[0098] In step S302 of some embodiments, similarly, when the text length of the negative sample statement is less than the text length threshold, perform zero-padding on the negative sample statement according to the text length threshold, that is, append zeros at the end of the original sample statement until the text length of the negative sample statement is equal to the text length threshold, and use the concatenated sample statement as the new negative sample statement.

[0099] It should be noted that the preset text length threshold can be set according to the maximum lengths of the positive sample statements and negative sample statements in the training sample set, or can be set manually, and no specific limitation is made here.

[0100] In step S104 of some embodiments, since it is expected that the text feature vectors between sample statements of the same category are more similar, that is, the feature distances are closer, while the text feature vectors between different categories are more distant, that is, the feature distances are larger, so as to continuously adjust the distribution of text features during the training process. Specifically, the feature constraint sub-model is used to constrain and adjust the representation of the text statement features based on the metric learning of the extracted text statement features, and the first text feature is feature-constrained through the feature constraint sub-model and the second text feature to update the first text feature.

[0101] In step S105 of some embodiments, after updating the first text feature of the positive sample statement, perform text classification processing on the first text feature through the text classification sub-model to obtain the sample prediction probability values of the positive sample statement belonging to each category label.

[0102] It should be noted that the text classification sub-model can be a model constructed by using a fully connected neural network, and this model is combined with the softmax activation function to obtain the sample prediction probability values of the positive sample statement belonging to each category label.

[0103] Please refer to Figure 4 , in some embodiments, after step S105, the text statement classification method provided by the embodiments of the present application may further include but is not limited to steps S401 to S403:

[0104] Step S401, performing feature constraint calculation on the first text feature, the second text feature, the positive sample label, and the negative sample label through a feature constraint sub-model to obtain a contrast loss value;

[0105] Step S402, obtaining a cross-entropy loss value according to the positive sample label and the sample prediction probability value;

[0106] Step S403, obtaining a model loss value according to the contrast loss value and the cross-entropy loss value.

[0107] In steps S401 to S403 of some embodiments, in the feature constraint sub-model of the embodiments of the present application, feature constraint calculation is performed on the first text feature, the second text feature, the positive sample label, and the negative sample label through a contrast loss function to obtain a contrast loss value, so as to guide text features of the same category to be closer and text features of different categories to be farther away, so that the model can effectively learn the corresponding distribution even when the relevance between the content of the text and the label is not obvious. And a cross-entropy loss value is obtained according to the positive sample label and the sample prediction probability value, so as to obtain a model loss value according to the contrast loss value and the cross-entropy loss value, realizing the combination of the method of deep metric learning and the method of cross-entropy loss, continuously adjusting the distribution of text statements and token feature vectors, and improving the training ability of the model.

[0108] Please refer to Figure 5 , in some embodiments, step S401 may include but is not limited to steps S501 to S502:

[0109] Step S501, comparing the positive sample label and the negative sample label according to a preset indication function to determine a contrast coefficient;

[0110] Step S502, performing feature constraint calculation according to the first text feature, the second text feature, and the contrast coefficient to obtain a contrast loss value.

[0111] In steps S501 and S502 of some embodiments, in order to perform feature constraint calculations on the first text feature, the second text feature, the positive sample label, and the negative sample label through a contrast loss function. Specifically, the positive sample label and the negative sample label are compared according to a preset indicator function 1{·} to determine the contrast coefficient. That is, the contrast coefficient includes 0 and 1. When the positive sample label and the negative sample label are the same, the contrast coefficient is 1; when the positive sample label and the negative sample label are different, the contrast coefficient is 0. As shown in formula (1), feature constraint calculations are performed based on the first text feature, the second text feature, and the contrast coefficient to obtain the contrast loss value L CL .

[0112]

[0113] where f x represents the first text feature, f y represents the second text feature, L x represents the positive sample label, L y represents the negative sample label, and a represents a preset constant.

[0114] Please refer to Figure 6 , in some embodiments, step S402 may include but is not limited to steps S601 to S602:

[0115] Step S601, obtaining the text word count of the positive sample statement;

[0116] Step S602, calculating the cross-entropy loss value according to a preset cross-entropy loss function for the text word count, the positive sample label, and the sample prediction probability value.

[0117] In steps S601 to S602 of some embodiments, for the solution of the cross-entropy loss value, specifically, first obtain the text word count of the positive sample statement, and then as shown in formula (2), calculate the cross-entropy loss value L CEL .

[0118] L CEL = ∑ m y · log(p) (2)

[0119] where m represents the text word count of the sample statement, represents the positive sample label number of each text statement, and p represents the corresponding sample prediction probability value.

[0120] It should be noted that the contrast loss value L CL and the cross-entropy loss value L CEL are summed to obtain the model loss value Loss sum .

[0121] In step S106 of some embodiments, after obtaining the sample prediction probability values of the positive sample statement belonging to each category label, the sample prediction probability values of multiple category labels are compared to determine the target sample label to which the positive sample statement belongs.

[0122] In step S107 of some embodiments, the model parameters of the initial text statement classification model are adjusted according to the positive sample label and the target sample label of the positive sample statement, that is, the model parameters are adjusted according to the similarity accuracy rate of the positive sample label and the target sample label. And the adjusted initial text statement classification model is continuously trained based on the training sample set until the model loss value of the initial text statement classification model meets the preset training end condition, that is, it can be considered that the performance of the initial text statement classification model at this time can meet the requirements, and then the target text statement classification model can be determined according to the model parameters and network structure of the initial text statement classification model.

[0123] It should be noted that the preset training end condition can be that when the total model loss value is less than the preset loss value threshold, or when the similarity accuracy rate of the obtained target sample label and the positive sample label is greater than or equal to the preset accuracy rate threshold.

[0124] In step S108 of some embodiments, in a specific application scenario, the application scenario includes a client device and a server device. The client device is used to send the initial text statement to be classified obtained by itself to the server device, and the server device is used to execute the text statement classification method provided in the embodiments of the present application. After identifying the initial text statement to be classified sent by the client device, the initial text statement is classified by the target text statement classification model to obtain the target category.

[0125] It should be noted that when the user needs to obtain information related to the initial text statement to be recognized by determining the named entities included in the initial text statement to be classified, the user can input the initial text statement to be classified in the text input field for classification provided on the client device. Furthermore, after the client device obtains the initial text statement to be classified input by the user, the initial text statement is sent to the server device.

[0126] It should be noted that the text statement classification method in the embodiments of the present application can be applied to, for example, the sentiment classification of movie reviews, or the positive and negative direction judgment of APP usage evaluations, etc. That is, after obtaining a certain movie review, the text statement classification method in the embodiments of the present application is used to determine the sentiment classification category of the user towards the movie, or after obtaining a certain usage evaluation of an APP, the text statement classification method in the embodiments of the present application is used to determine the recognition degree classification category of the user towards the APP.

[0127] A text statement classification method provided by an embodiment of the present application first constructs a target text statement classification model. To solve the problem of high-precision text statement classification for the statement to be recognized even when the sample size is small and the relevance of the pre-labels of the text statements is not strong, the embodiment of the present application obtains a training sample set by collecting sample pairs of data. Each sample pair of data includes a positive sample statement, a positive sample label corresponding to the positive sample statement, a negative sample statement, and a negative sample label corresponding to the negative sample statement. Then, during the training process of the model, first obtain an initial text statement classification model, which includes a pre-training sub-model, a feature constraint sub-model, and a text classification sub-model. The pre-training sub-model performs feature encoding processing on each text word in the positive sample statement and each text word in the negative sample statement respectively to obtain a positive sample word feature corresponding to the positive sample statement and a negative sample word feature corresponding to the negative sample statement. And perform self-attention processing on all the positive sample word features and all the negative sample word features respectively to obtain a first text feature of the positive sample statement and a second text feature of the negative sample statement. After that, perform feature constraint on the first text feature through the feature constraint sub-model and the second text feature to update the first text feature. Perform text classification processing on the first text feature through the text classification sub-model to obtain the sample prediction probability value of the positive sample statement belonging to each category label. Compare the multiple sample prediction probability values obtained for the positive sample statement numerically to determine the target sample label of the positive sample statement. Perform feature constraint calculation on the first text feature, the second text feature, the positive sample label, and the negative sample label through the feature constraint sub-model to obtain a contrast loss value, obtain a cross-entropy loss value according to the positive sample label and the sample prediction probability value, and obtain a model loss value according to the contrast loss value and the cross-entropy loss value. Finally, adjust the model parameters of the initial text statement classification model according to the positive sample label and the target sample label of the positive sample statement, and continue to train the adjusted initial text statement classification model based on the training sample set until the model loss value of the initial text statement classification model meets the preset training end condition to obtain the target text statement classification model. Obtain the initial text statement to be classified, and classify the initial text statement through the target text statement classification model to obtain the target category. The embodiment of the present application can improve the accuracy of the model in classifying text statements. The embodiment of the present application obtains a training sample set by collecting sample pairs of data, effectively realizing the expansion of the training samples and being able to flexibly control the number of training samples in each category to achieve sample balance. The embodiment of the present application introduces deep metric learning to add constraints to the features of the text, so as to be able to guide the sample sentence features of the same category to be closer and the sentence features of different categories to be farther away, realizing that even when the relevance between the content of the text and the label is not obvious, the model can effectively learn the corresponding distribution.Meanwhile, in the embodiments of the present application, by combining the contrastive loss function of deep metric learning with the cross-entropy loss function, the distribution of the sentence and token feature vectors is continuously adjusted, so as to guide the acquisition of a high-precision target text statement classification model, thereby improving the accuracy of the model in classifying text statements.

[0128] Please refer to Figure 7 , the embodiments of the present application also provide a text statement classification device, which can implement the above text statement classification method. The device includes:

[0129] A sample set acquisition module 701, configured to acquire a training sample set, where the training sample set includes a plurality of sample pair data, and each sample pair data includes a positive sample statement, a positive sample label corresponding to the positive sample statement, a negative sample statement, and a negative sample label corresponding to the negative sample statement;

[0130] An initial model acquisition module 702, configured to acquire an initial text statement classification model, where the initial text statement classification model includes a pre-trained sub-model, a feature constraint sub-model, and a text classification sub-model;

[0131] A feature extraction module 703, configured to input the positive sample statement and the negative sample statement of each sample pair data into the initial text statement classification model, and respectively extract features of the positive sample statement and the negative sample statement through the pre-trained sub-model to obtain a first text feature of the positive sample statement and a second text feature of the negative sample statement;

[0132] A feature constraint module 704, configured to perform feature constraint on the first text feature through the feature constraint sub-model and the second text feature to update the first text feature;

[0133] A sample text classification module 705, configured to perform text classification processing on the first text feature through the text classification sub-model to obtain sample prediction probability values of the positive sample statement belonging to each category label;

[0134] A numerical comparison module 706, configured to perform numerical comparison on the multiple sample prediction probability values of the positive sample statement to determine the target sample label of the positive sample statement;

[0135] A target model construction module 707, configured to adjust the model parameters of the initial text statement classification model according to the positive sample label and the target sample label of the positive sample statement, and continue to train the adjusted initial text statement classification model based on the training sample set until the model loss value of the initial text statement classification model meets a preset training end condition to obtain a target text statement classification model;

[0136] A target text classification module 708, configured to acquire an initial text statement to be classified, and classify the initial text statement through the target text statement classification model to obtain a target category.

[0137] The specific implementation manner of the text statement classification device is basically the same as the specific embodiment of the above text statement classification method, and will not be elaborated here.

[0138] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above text statement classification method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0139] Please refer to Figure 8 , Figure 8 , which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0140] A processor 801, which can be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0141] A memory 802, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 802 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 802, and the processor 801 is called to execute the text statement classification method of the embodiments of the present application;

[0142] An input / output interface 803, which is used to implement information input and output;

[0143] A communication interface 804, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);

[0144] A bus 805, which transmits information between various components of the device (such as the processor 801, the memory 802, the input / output interface 803, and the communication interface 804);

[0145] Among them, the processor 801, the memory 802, the input / output interface 803, and the communication interface 804 are communicatively connected to each other inside the device through the bus 805.

[0146] The embodiment of the present application further provides a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned text statement classification method is implemented.

[0147] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0148] A text statement classification method, classification device, electronic device, and storage medium provided by an embodiment of the present application obtain a training sample set by collecting sample pairs of data. Each sample pair of data includes a positive sample statement, a positive sample label corresponding to the positive sample statement, a negative sample statement, and a negative sample label corresponding to the negative sample statement. Then, during the training process of the model, first obtain an initial text statement classification model, which includes a pre-training sub-model, a feature constraint sub-model, and a text classification sub-model. The pre-training sub-model performs feature encoding processing on each text character in the positive sample statement and each text character in the negative sample statement respectively to obtain a positive sample character feature corresponding to the positive sample statement and a negative sample character feature corresponding to the negative sample statement. Self-attention processing is performed on all positive sample character features and all negative sample character features respectively to obtain a first text feature of the positive sample statement and a second text feature of the negative sample statement. Then, the first text feature is feature-constrained by the feature constraint sub-model and the second text feature to update the first text feature. The text classification sub-model performs text classification processing on the first text feature to obtain a sample prediction probability value of the positive sample statement belonging to each category label. Numerical comparison is performed on the multiple sample prediction probability values obtained for the positive sample statement to determine the target sample label of the positive sample statement. The feature constraint sub-model performs feature constraint calculation on the first text feature, the second text feature, the positive sample label, and the negative sample label to obtain a contrast loss value, obtains a cross-entropy loss value according to the positive sample label and the sample prediction probability value, and obtains a model loss value according to the contrast loss value and the cross-entropy loss value. Finally, the model parameters of the initial text statement classification model are adjusted according to the positive sample label and the target sample label of the positive sample statement, and the adjusted initial text statement classification model is continuously trained based on the training sample set until the model loss value of the initial text statement classification model meets a preset training end condition to obtain a target text statement classification model. An initial text statement to be classified is obtained, and the initial text statement is classified by the target text statement classification model to obtain a target category. The embodiment of the present application can improve the accuracy of the model in classifying text statements. The embodiment of the present application obtains a training sample set by collecting sample pairs of data, effectively realizing the expansion of training samples and being able to flexibly control the number of training samples in each category to achieve sample balance. The embodiment of the present application introduces deep metric learning to impose constraints on the features of the text, so as to be able to guide the sample sentence features of the same category to be closer and the sentence features of different categories to be farther away, realizing that even when the correlation between the content of the text and the label is not obvious, the model can effectively learn the corresponding distribution. At the same time, the embodiment of the present application combines the contrast loss function of deep metric learning with the cross-entropy loss function to continuously adjust the distribution of the sentence and token feature vectors, thereby guiding the acquisition of a high-precision target text statement classification model, and further improving the accuracy of the model in classifying text statements.

[0149] The embodiments described in the embodiments of the present application are to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. As those skilled in the art know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0150] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0151] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0152] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0153] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0154] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0155] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0156] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0157] In addition, the functional units in each embodiment of this application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0158] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0159] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, which does not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of the rights of the embodiments of this application.

Claims

1. A method for classifying text statements, characterized in that, The method includes: Obtaining a training sample set, where the training sample set includes multiple sample pair data, and each sample pair data includes a positive sample statement, a positive sample label corresponding to the positive sample statement, a negative sample statement, and a negative sample label corresponding to the negative sample statement; Obtaining an initial text statement classification model, where the initial text statement classification model includes a pre-trained sub-model, a feature constraint sub-model, and a text classification sub-model; Inputting the positive sample statement and the negative sample statement of each sample pair data into the initial text statement classification model, and respectively performing feature extraction on the positive sample statement and the negative sample statement through the pre-trained sub-model to obtain a first text feature of the positive sample statement and a second text feature of the negative sample statement; Performing feature constraint on the first text feature through the feature constraint sub-model and the second text feature to update the first text feature; Performing text classification processing on the first text feature through the text classification sub-model to obtain sample prediction probability values of the positive sample statement belonging to each category label; Performing numerical comparison on the multiple sample prediction probability values of the positive sample statement to determine the target sample label of the positive sample statement; Adjusting the model parameters of the initial text statement classification model according to the positive sample label and the target sample label of the positive sample statement, and continuously training the adjusted initial text statement classification model based on the training sample set until the model loss value of the initial text statement classification model meets a preset training end condition to obtain a target text statement classification model; Obtaining an initial text statement to be classified, and classifying the initial text statement through the target text statement classification model to obtain a target category.

2. The method according to claim 1, characterized in that, After performing text classification processing on the first text feature through the text classification sub-model to obtain sample prediction probability values of the positive sample statement belonging to each category label, the method further includes: Performing feature constraint calculation on the first text feature, the second text feature, the positive sample label, and the negative sample label through the feature constraint sub-model to obtain a contrast loss value; Obtaining a cross-entropy loss value according to the positive sample label and the sample prediction probability value; Obtaining a model loss value according to the contrast loss value and the cross-entropy loss value.

3. The method according to claim 2, characterized in that, The pre-trained sub-model includes feature encoding processing and self-attention processing. Inputting the positive sample statement and the negative sample statement of each sample pair data into the initial text statement classification model, and respectively performing feature extraction on the positive sample statement and the negative sample statement through the pre-trained sub-model to obtain a first text feature of the positive sample statement and a second text feature of the negative sample statement includes: Inputting the positive sample statement and the negative sample statement of each sample pair data into the initial text statement classification model; Perform the feature encoding process on each text character in the positive sample statement and each text character in the negative sample statement respectively, to obtain the positive sample character features corresponding to the positive sample statement and the negative sample character features corresponding to the negative sample statement; Perform the self-attention process on all the positive sample character features and all the negative sample character features respectively, to obtain the first text feature of the positive sample statement and the second text feature of the negative sample statement.

4. The method according to claim 3, characterized in that, Before inputting the positive sample statement and the negative sample statement of each sample pair data into the initial text statement classification model, the method further includes: Compare the lengths of the positive sample statement and the negative sample statement respectively according to a preset text length threshold. When the text length of the positive sample statement is less than the text length threshold, perform zero-padding operation on the positive sample statement according to the text length threshold, and update the positive sample statement until the text length of the positive sample statement is equal to the text length threshold; When the text length of the negative sample statement is less than the text length threshold, perform zero-padding operation on the negative sample statement according to the text length threshold, and update the negative sample statement until the text length of the negative sample statement is equal to the text length threshold.

5. The method according to claim 2, characterized in that, The feature constraint calculation of the first text feature, the second text feature, the positive sample label and the negative sample label by the feature constraint sub-model to obtain the contrast loss value includes: Compare the positive sample label and the negative sample label according to a preset indicator function to determine the contrast coefficient; Perform feature constraint calculation according to the first text feature, the second text feature and the contrast coefficient to obtain the contrast loss value.

6. The method according to any one of claims 2 to 5, characterized in that, The obtaining of the cross-entropy loss value according to the positive sample label and the sample prediction probability value includes: Obtain the number of text characters of the positive sample statement; Calculate the text character number, the positive sample label and the sample prediction probability value according to a preset cross-entropy loss function to obtain the cross-entropy loss value.

7. A text statement classification device, characterized in that, The device includes: A sample set acquisition module, configured to acquire a training sample set, where the training sample set includes multiple sample pair data, and each sample pair data includes a positive sample statement, a positive sample label corresponding to the positive sample statement, a negative sample statement, and a negative sample label corresponding to the negative sample statement; An initial model acquisition module, configured to acquire an initial text statement classification model, where the initial text statement classification model includes a pre-training sub-model, a feature constraint sub-model, and a text classification sub-model; A feature extraction module, configured to input the positive sample statement and the negative sample statement of each sample pair data into the initial text statement classification model, and respectively extract features of the positive sample statement and the negative sample statement through the pre-training sub-model to obtain the first text feature of the positive sample statement and the second text feature of the negative sample statement; A feature constraint module, configured to perform feature constraint on the first text feature through the feature constraint sub-model and the second text feature to update the first text feature; A sample text classification module for performing text classification processing on the first text feature through the text classification sub-model to obtain the sample prediction probability values of the positive sample statement belonging to each category label; A numerical comparison module for numerically comparing the multiple sample prediction probability values of the positive sample statement to determine the target sample label of the positive sample statement; A target model construction module for adjusting the model parameters of the initial text statement classification model according to the positive sample label and the target sample label of the positive sample statement, and continuing to train the adjusted initial text statement classification model based on the training sample set until the model loss value of the initial text statement classification model meets the preset training end condition to obtain a target text statement classification model; A target text classification module for obtaining an initial text statement to be classified and classifying the initial text statement through the target text statement classification model to obtain a target category.

8. The device according to claim 7, wherein, Before the sample text classification module for performing text classification processing on the first text feature through the text classification sub-model to obtain the sample prediction probability values of the positive sample statement belonging to each category label, the apparatus further includes: A contrast loss value acquisition module for performing feature constraint calculation on the first text feature, the second text feature, the positive sample label, and the negative sample label through the feature constraint sub-model to obtain a contrast loss value; A cross-entropy loss value acquisition module for obtaining a cross-entropy loss value according to the positive sample label and the sample prediction probability value; A model loss value acquisition module for obtaining a model loss value according to the contrast loss value and the cross-entropy loss value.

9. An electronic device, wherein, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements a text statement classification method according to any one of claims 1 to 6.

10. A storage medium, the storage medium being a computer-readable storage medium, the computer-readable storage medium storing a computer program, wherein, When the computer program is executed by the processor, it implements a text statement classification method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Classification model construction method and device, text statement classification method and device and storage medium

    CN112966102A

  • Efficient updating of a model used for data learning

    US20180075351A1