Method and device for obtaining text encoder and text classification method and device

By obtaining and training multi-level tag data, text encoders suitable for specific fields are obtained, which solves the accuracy of multi-level tag classification in fields such as government work orders, and realizes the ability to automatically identify and adapt to changes in the tag system.

CN120216690APending Publication Date: 2025-06-27CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510138227.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When dealing with multi-level label classification, especially in the field of government work tickets, the existing technology faces the problems of excessive number of labels and irregular changes in the label system, resulting in a decrease in classification accuracy.

Method used

A method of obtaining text encoder is proposed. By obtaining multiple training data, including historical text and its corresponding classification tags, the encoder is trained until the parameters converge, and a text encoder suitable for a specific field is obtained.

Benefits of technology

In the multi-level label classification scenario, the multi-level label category of text is automatically identified through the calculation of feature similarity between text and labels, which improves the accuracy of classification and reduces the dependence on changes in the label system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216690A_ABST
    Figure CN120216690A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for obtaining a text encoder and a text classification method and device. The invention aims to solve the problem of text classification, for example, the practical problem of multi-level label classification in a government affair work order. The problem of text classification is abstracted into the problem of calculation of the similarity between the text and the tag. Therefore, under the condition of classifying the texts, the labels of the texts are determined by calculating the similarity between the texts and the labels. For example, in the field of government affair work orders, multi-level label categories to which government affair work order contents belong are automatically identified. And secondly, the problems that the number of labels is too large and a label system is irregularly changed in the label classification field of the text can be effectively solved, for example, the problems that the number of multi-level labels is too large and the label system is irregularly changed in the government affair work order field are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method and apparatus for obtaining a text encoder, and a method and apparatus for text classification. Background Art

[0002] The technology of natural language processing has been widely applied in various fields, and an important application direction is text understanding and classification.

[0003] Currently, the commonly adopted method is: for a specific task, collect the sample data in the field where the specific task is located and the annotation data of the sample data, use the sample data in the field where the specific task is located and the annotation data of the sample data to fine-tune the model, obtain the fine-tuned model, and then the fine-tuned model can be used to process the specific task, such as text understanding and classification, etc. Summary of the Invention

[0004] This application discloses a method and apparatus for obtaining a text encoder, and a method and apparatus for text classification.

[0005] In a first aspect, this application discloses a method for obtaining a text encoder, the method comprising:

[0006] Obtain a to-be-trained encoder;

[0007] Obtain a plurality of first training data, where a first training data includes a first training text and a first annotation text. The first training text includes historical texts that have appeared in the historical process and classification labels obtained by classifying the historical texts in the historical process. The first annotation text includes a first positive annotation sub-text and a first negative annotation sub-text. The first positive annotation sub-text includes the feature similarity between the label feature of the classification label and the text feature of the historical text. The first negative annotation sub-text includes: the feature similarity between the label feature of the classification label and the text feature of the historical text in at least part of the first training texts in the plurality of first training data other than the first training data where the historical text is located;

[0008] Use the plurality of first training data to train the to-be-trained encoder until the parameters in the to-be-trained encoder converge, obtaining a text encoder.

[0009] In an optional implementation manner, the historical texts in the first training text include historical texts in a specific field.

[0010] In an optional implementation manner, the method further comprises:

[0011] Store the corresponding relationship between the domain identifier of the specific field and the text encoder.

[0012] In an alternative implementation, the obtaining of the encoder to be trained includes:

[0013] Obtain an initial encoder;

[0014] Obtain a plurality of second training data. A second training data includes a second training text and a second annotation text. The second training text includes historical text pairs in the general domain. The historical text pairs include a first historical text and a second historical text with an associated relationship therebetween. The second annotation text includes at least the feature similarity between the text features of the first historical text and the text features of the second historical text in the historical text pair;

[0015] Use the plurality of second training data to train the initial encoder until the parameters in the initial encoder converge, and obtain the encoder to be trained applicable to the general domain.

[0016] In an alternative implementation, the second annotation text includes a second positive annotation sub - text and a second negative annotation sub - text;

[0017] The second positive annotation sub - text includes the feature similarity between the text features of the first historical text and the text features of the second historical text in the historical text pair in the one second training text;

[0018] The second negative annotation sub - text includes: the feature similarity between the text features of the first historical text in the historical text pair in the one second training text and the text features of the first historical text or the text features of the second historical text in the historical text pair in the second training text of the additional second training data. The additional second training data includes at least part of the second training data other than the second training data where the one second training text is located among the plurality of second training data.

[0019] In an alternative implementation, the method further includes:

[0020] Store the encoder to be trained applicable to the general domain.

[0021] In an alternative implementation, the obtaining of the initial encoder includes:

[0022] Obtain a plurality of third training data. A third training data includes a third training text and a third annotation text. The third training text includes: a masked text obtained by masking at least one word in the sample text. The third annotation text includes: the masked word in the masked text;

[0023] Construct the network structure of the initial model. The initial model includes an encoder and a decoder. The output end of the encoder is connected to the input end of the decoder. The encoder is used to obtain the features of the text, and the decoder is used to predict the masked words in the text at least based on the features of the text;

[0024] Use multiple third training data to train the initial model until the parameters in the initial model converge to obtain a text processing model;

[0025] Obtain an initial encoder according to the encoder in the text processing model.

[0026] In an optional implementation, the masked text in the third training text includes: a first masked text and a second masked text. The first masked text is obtained by masking at least one word in the sample text, and the second masked text is obtained by masking at least one word in the sample text. The number of masked words in the first masked text is less than the number of masked words in the second masked text;

[0027] The second annotation text is obtained according to the masked words in the second masked text;

[0028] The first masked text is used to input the encoder, and the second masked text is used to input the decoder.

[0029] In a second aspect, the present application discloses a text classification method, and the method includes:

[0030] Obtain the text to be classified, and obtain the text features of the text to be classified based on the text encoder;

[0031] Obtain multiple preset classification labels, and obtain the label features of each preset classification label based on the text encoder respectively;

[0032] Obtain the feature similarities between the text features of the text to be classified and the label features of each preset classification label respectively;

[0033] Among the label features of each preset classification label, select the label feature with the largest feature similarity with the text features of the text to be classified;

[0034] Determine the preset classification label corresponding to the selected label feature as the classification label of the text to be classified.

[0035] In an optional implementation, the text encoder is obtained based on the method shown in the first aspect.

[0036] In a third aspect, the present application discloses a device for obtaining a text encoder, and the device includes:

[0037] A first obtaining module, configured to obtain an encoder to be trained;

[0038] A second acquisition module, configured to acquire a plurality of first training data. A first training data includes a first training text and a first annotation text. The first training text includes a historical text that has appeared in a historical process and a classification label obtained by classifying the historical text in the historical process. The first annotation text includes a first positive annotation sub-text and a first negative annotation sub-text. The first positive annotation sub-text includes a feature similarity between the label feature of the classification label and the text feature of the historical text. The first negative annotation sub-text includes: a feature similarity between the label feature of the classification label and the text feature of the historical text in at least some of the first training data other than the first training data where the historical text is located among the plurality of first training data;

[0039] A training module, configured to train a to-be-trained encoder using the plurality of first training data until the parameters in the to-be-trained encoder converge, so as to obtain a text encoder.

[0040] In an optional implementation manner, the historical text in the first training text includes historical texts in a specific domain.

[0041] In an optional implementation manner, the apparatus further includes:

[0042] A storage module, configured to store the correspondence between the domain identifier of the specific domain and the text encoder.

[0043] In an optional implementation manner, the first acquisition module includes:

[0044] A first acquisition unit, configured to acquire an initial encoder;

[0045] A second acquisition unit, configured to acquire a plurality of second training data. A second training data includes a second training text and a second annotation text. The second training text includes historical text pairs in a general domain. The historical text pairs include a first historical text and a second historical text that have an association relationship therebetween. The second annotation text at least includes a feature similarity between the text feature of the first historical text and the text feature of the second historical text in the historical text pair;

[0046] A training unit, configured to train the initial encoder using the plurality of second training data until the parameters in the initial encoder converge, so as to obtain a to-be-trained encoder applicable to the general domain.

[0047] In an optional implementation manner, the second annotation text includes a second positive annotation sub-text and a second negative annotation sub-text;

[0048] The second positive annotation sub - text includes the feature similarity between the text features of the first historical text and the text features of the second historical text in the historical text pair in the one second training text;

[0049] The second negative annotation sub - text includes: the feature similarity between the text features of the first historical text in the historical text pair in the one second training text and the text features of the first historical text or the text features of the second historical text in the historical text pair in the second training text in the additional second training data, where the additional second training data includes at least part of the second training data other than the second training data where the one second training text is located among the multiple second training data.

[0050] In an optional implementation manner, the first acquisition module further includes:

[0051] A storage unit for storing the encoder to be trained applicable to the general domain.

[0052] In an optional implementation manner, the first acquisition unit includes:

[0053] A first acquisition sub - unit for acquiring multiple third training data, where one third training data includes a third training text and a third annotation text. The third training text includes: a masked text obtained by masking at least one vocabulary in the sample text, and the third annotation text includes: the masked vocabulary in the masked text;

[0054] A construction sub - unit for constructing the network structure of the initial model. The initial model includes an encoder and a decoder. The output end of the encoder is connected to the input end of the decoder. The encoder is used to acquire the features of the text, and the decoder is used to predict at least the masked vocabulary in the text according to the features of the text;

[0055] A training sub - unit for training the initial model using multiple third training data until the parameters in the initial model converge to obtain a text processing model;

[0056] A second acquisition sub - unit for acquiring the initial encoder according to the encoder in the text processing model.

[0057] In an optional implementation manner, the masked text in the third training text includes: a first masked text and a second masked text. The first masked text is obtained by masking at least one vocabulary in the sample text, and the second masked text is obtained by masking at least one vocabulary in the sample text. The number of masked vocabularies in the first masked text is less than the number of masked vocabularies in the second masked text;

[0058] The second annotation text is obtained according to the masked vocabulary in the second masked text;

[0059] The first masked text is for the input encoder, and the second masked text is for the input decoder.

[0060] In a fourth aspect, the present application discloses a text classification device, which includes:

[0061] A third acquisition module, configured to acquire the text to be classified and obtain the text features of the text to be classified based on a text encoder;

[0062] A fourth acquisition module, configured to acquire a plurality of preset classification labels and obtain the label features of each preset classification label based on the text encoder respectively;

[0063] A fifth acquisition module, configured to obtain the feature similarity between the text features of the text to be classified and the label features of each preset classification label respectively;

[0064] A selection module, configured to select, from the label features of each preset classification label, the label feature with the maximum feature similarity to the text features of the text to be classified;

[0065] A determination module, configured to determine the preset classification label corresponding to the selected label feature as the classification label of the text to be classified.

[0066] In an optional implementation, the text encoder is obtained based on the method shown in the first aspect; or, the text encoder is obtained based on the device shown in the third aspect.

[0067] In a fifth aspect, the present application discloses an electronic device, which includes: a processor; and a memory for storing processor-executable instructions; wherein, the processor is configured to execute the method described in any of the above aspects.

[0068] In a sixth aspect, the present application discloses a non-transitory computer-readable storage medium, which, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method described in any of the above aspects.

[0069] In a seventh aspect, the present application discloses a computer program product, which, when the instructions in the computer program product are executed by a processor of an electronic device, enables the electronic device to execute the method described in any of the above aspects.

[0070] The technical solution provided by the present application may include the following beneficial effects:

[0071] This application aims to solve the problem of text classification. For example, it solves the practical problem of multi-level label classification in government work orders. The problem of text classification is abstracted into the problem of calculating the similarity between text and labels (such as calculating the feature similarity between the text features of the text and the label features of the labels). Thus, when classifying text, the label of the text is determined by calculating the similarity between the text and the label (such as calculating the feature similarity between the text features of the text and the label features of the labels). For example, in the field of government work orders, the multi-level label category to which the content of the government work order belongs is automatically identified.

[0072] Secondly, this application can effectively overcome the problems of excessive number of labels and irregular changes in the label system in the field of text label classification. For example, it effectively overcomes the problems of excessive number of multi-level labels and irregular changes in the label system in the field of government work orders.

[0073] Subsequently, even if labels are added, deleted, or modified, it does not affect obtaining the text features of the text and the label features of the labels based on the text encoder, nor does it affect calculating the feature similarity between the text features of the text and the label features of the labels, nor does it affect the accuracy of text classification. There is no need to retrain the encoder, and the training cost will not increase. Description of the Drawings

[0074] Figure 1 It is a flowchart of the steps of a method for obtaining a text encoder in this application.

[0075] Figure 2 It is a flowchart of the steps of a method for obtaining an encoder to be trained in this application.

[0076] Figure 3 It is a flowchart of the steps of a method for obtaining an initial encoder in this application.

[0077] Figure 4 It is a flowchart of the steps of a text classification method in this application.

[0078] Figure 5 It is a block diagram of the structure of a device for obtaining a text encoder in this application.

[0079] Figure 6 It is a block diagram of the structure of a text classification device in this application.

[0080] Figure 7 It is a block diagram of an electronic device in this application.

[0081] Figure 8 It is a block diagram of an electronic device in this application. Detailed Embodiments

[0082] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0083] In the scenario of natural language processing, the problem of text classification can be abstracted to determine which label among multiple labels (such as multi-level labels) a text belongs to.

[0084] However, when there are numerous label categories, the label system changes, or the label system is not enumerable, the accuracy of determining which label among multiple labels a text belongs to drops sharply.

[0085] With the advent of generative language models, large language models provide more options for problems such as language understanding and generation. In particular, they can handle the problem of non-enumerable category labels in the above-mentioned text classification problem. However, limited by the context length that large language models can handle and the inference time under long context conditions, large language models currently cannot handle some real-time classification problems of long text content well.

[0086] The present application aims to solve the actual problems of label classification in government work orders (such as the actual problems of multi-level label classification in government work orders).

[0087] For this reason, the solution of the present application is proposed.

[0088] Before introducing the solution of the present application, the technical terms that the present application may involve will be explained first.

[0089] Natural language understanding: refers to the process of using computer technology to analyze and understand natural language. It involves multiple aspects such as semantic analysis, syntactic analysis, and context understanding, aiming to convert human language into a form that can be understood and processed by a computer.

[0090] Large language model: refers to a large-scale neural network language model trained on a large amount of text data, usually including hundreds of millions to billions of parameters. Large models generally require a large amount of computing resources and data sets for training, but have high accuracy and generalization ability in complex tasks.

[0091] Autoencoder model: refers to an unsupervised learning model used to learn data representations. It is mainly used for feature extraction, dimensionality reduction, denoising, and generation tasks.

[0092] Encoder: It refers to a component or model structure that converts input data into a specific representation form. In the field of natural language processing, the main role of the encoder is to convert the original text data (such as sentences, paragraphs, etc.) into low-dimensional, dense vector representations. This vector representation not only contains the key information of the original data but also can, to a certain extent, solve the ambiguity problem of natural language.

[0093] Decoder: The reverse process of the encoder. It gradually generates the final target data based on the output of the encoder (i.e., the encoded representation) and the partially generated target sequence. In NLP tasks, the decoder is usually used to generate text, speech, or other forms of output data. The working principle of the decoder is based on the encoding information provided by the encoder and possibly other inputs (such as the partially generated sequence).

[0094] Token: A basic concept in NLP (Natural Language Processing). It refers to the basic unit in text. This unit can be a word, punctuation mark, sub-word, or character, depending on the requirements and methods of text processing. In natural language processing, dividing text into several tokens is the first step in text processing, and this process is called tokenization.

[0095] Prompt: It refers to a mechanism used to help the model understand the text context. Usually, it uses context information to determine the meaning of a certain word. The guiding word / prompt can be a single word, phrase, or the whole sentence, and its role is to provide additional information for the model to better understand and process the input text.

[0096] Government Affairs Work Order: It refers to various forms of documents or electronic forms received by government agencies or relevant departments, such as work requests, complaints, consultations, etc., which usually need to be processed and replied to.

[0097] Multi-level Labels: It refers to a set of predefined, hierarchical categories or label collections used in text classification tasks. This system provides a structured framework for text classification and helps us classify text content into specific categories.

[0098] Refer to Figure 1 , which shows the step flowchart of a method for obtaining a text encoder in the present application. Among them, the method includes:

[0099] In step S101, obtain the encoder to be trained.

[0100] The encoder to be trained may include the encoder in an auto-encoding model, and the auto-encoding model includes a Transformer model, etc. For example, it includes a bert model (Bidirectional Encoder Representations from Transformers, a bidirectional semantic encoding representation model built based on the transformer), etc.

[0101] The encoder to be trained may include an encoder that has not been trained (for example, not pre-trained), or may also include an encoder obtained after training (for example, pre-training) the encoder that has not been trained. The specific acquisition method of the encoder to be trained can be seen in the embodiments shown later and will not be elaborated here.

[0102] In step S102, a plurality of first training data are obtained. A first training data includes a first training text and a first annotation text. The first training text includes historical texts that have appeared in the historical process and classification labels obtained by classifying the historical texts in the historical process. The first annotation text includes a first positive annotation sub-text and a first negative annotation sub-text. The first positive annotation sub-text includes the feature similarity between the label feature of the classification label and the text feature of the historical text. The first negative annotation sub-text includes: the feature similarity between the label feature of the classification label and the text features of the historical texts in at least some of the first training texts in the plurality of first training data other than the first training data where the historical text is located.

[0103] Thus, for any one of the first training data, the following operations can be performed:

[0104] Input the classification label obtained by classifying the historical text in the first training text of the arbitrary first training data into the initial encoder, so that the initial encoder encodes the classification label obtained by classifying the historical text in the first training text of the arbitrary first training data, and obtains the label feature of the classification label.

[0105] And input the historical text in the first training text of the arbitrary first training data into the initial encoder, so that the initial encoder encodes the historical text in the first training text of the arbitrary first training data, and obtains the text feature of the historical text.

[0106] Then calculate the feature similarity between the label feature of the classification label and the text feature of the historical text.

[0107] For example, the inner product between the feature vector of the classification label and the feature vector of the historical text can be calculated, the norm value of the feature vector of the classification label can be calculated, the norm value of the feature vector of the historical text can be calculated, the product between the norm value of the feature vector of the classification label and the norm value of the feature vector of the historical text can be calculated, and the ratio between the inner product and the product can be calculated and used as the feature similarity between the label feature of the classification label and the text feature of the historical text.

[0108] Moreover, the historical texts in the first training texts of the multiple first training data except for the first training data where the historical text is located are input into the initial encoder, so that the initial encoder encodes the historical texts in the first training texts of the multiple first training data except for the first training data where the historical text is located, and the text features of the historical texts in the first training texts of the multiple first training data except for the first training data where the historical text is located are obtained. Then, the feature similarity between the label feature of the classification label and the text features of the historical texts in the first training texts of the multiple first training data except for the first training data where the historical text is located is calculated. The specific calculation method can refer to the method of "calculating the feature similarity between the label feature of the classification label and the text feature of the historical text", which will not be elaborated here.

[0109] For each of the other first training data, the above operations are similarly performed.

[0110] The historical texts in the first training texts of different first training data are different.

[0111] The historical texts in the first training texts of each first training data are all texts that have appeared in the historical process.

[0112] The historical text in the first training text is a text that has been classified in the historical process. For example, it is a text that needs to be classified and submitted by a large number of users in the historical process.

[0113] The classification label in the first training text is obtained after classifying the historical text in the historical process. For example, it is obtained after classifying the text that needs to be classified and submitted by a large number of users received in the historical process, and the classification label of the text and the corresponding text are stored locally in the electronic device.

[0114] Therefore, the classification label in the first training text can be directly obtained from the electronic device and does not require separate manual collection.

[0115] The first negative annotation sub - text includes: the feature similarity between the label feature of the classification label and the text features of the historical text in the first training texts of all the first training data except the first training data where the historical text is located among the multiple first training data.

[0116] Alternatively, the first negative annotation sub - text includes: the feature similarity between the label feature of the classification label and the text features of the historical text in the first training texts of some of the first training data except the first training data where the historical text is located among the multiple first training data.

[0117] In an embodiment of the present application, the feature similarity between the label feature of the classification label and the text features of the historical text in the first training texts of some of the first training data except the first training data where the historical text is located among the multiple first training data is obtained in the following manner.

[0118] Obtain the text features of the historical text in the first training texts of each of the first training data except the first training data where the historical text is located among the multiple first training data.

[0119] Calculate the feature similarity between the label feature of the classification label and the text features of the historical text in the first training texts of each of the first training data except the first training data where the historical text is located among the multiple first training data, and then select the top N feature similarities in the order from largest to smallest feature similarity.

[0120] N can include 1, 2, 3, 4, 5, 6, 10, or 11, etc., and can be determined according to the actual situation specifically. The present application does not limit this.

[0121] In step S103, use multiple first training data to train the encoder to be trained until the parameters in the encoder to be trained converge, and obtain the text encoder.

[0122] Bring the first training data into the contrast loss function loss_contrast, and optimize the parameters in the encoder to be trained by minimizing the contrast loss function loss_contrast.

[0123] For example, for any one of the multiple first training data, calculate the ratio of the first positive annotation sub-text "the feature similarity between the label feature of the classification label and the text feature of the historical text" in the first annotation text in the first training text of the any one to the temperature coefficient, calculate the power of e to this ratio, and obtain the positive value corresponding to the any one of the first training data. The temperature coefficient is usually between 0 and 1 and is adjusted as the encoder is optimized, or is a preset fixed value. Specifically, which value between 0 and 1 is selected can be determined according to the actual situation, and the present application does not limit this.

[0124] Calculate the ratios of the first negative annotation sub-text "the feature similarity between the label feature of the classification label and the text feature of the historical text in the first training text in at least part of the first training data other than the first training data where the historical text is located among the multiple first training data" in the first annotation text in the first training text of the any one to the temperature coefficient respectively, calculate the power of e to each ratio respectively, and sum up the powers of e to each ratio to obtain the negative value corresponding to the any one of the first training data.

[0125] Calculate the sum value between the positive value corresponding to the any one of the first training data and the negative value corresponding to the any one of the first training data, and calculate the ratio between the positive value corresponding to the any one of the first training data and this sum value to obtain the ratio corresponding to the any one of the first training data.

[0126] Optimize the parameters in the initial encoder by minimizing the value of the contrast loss function loss_contrast (the ratio corresponding to the any one of the first training data).

[0127] The same applies to other batches of first training data.

[0128] Until the parameters in the encoder to be trained converge to obtain a text encoder.

[0129] In an embodiment of the present application, the historical text in the first training text includes historical texts in a specific field, so that the obtained text encoder is a text encoder applicable to the specific field.

[0130] So that the text encoder has higher performance for texts in the specific field, and the accuracy and comprehensiveness of the text features of the texts in the specific field extracted are higher.

[0131] The specific field includes fields such as government affairs work order field, search term classification field, question and answer matching field, and product and evaluation matching field, etc., and the present application does not limit this.

[0132] In another embodiment of the present application, the corresponding relationship between the domain identifier of a specific domain and the text encoder obtained in step S103 can be stored. For example, the corresponding relationship between the domain identifier of a specific domain and the text encoder is stored in an electronic device. This facilitates subsequent calling of the text encoder applicable to a specific domain according to the domain identifier of the specific domain in this corresponding relationship when the text encoder applicable to a specific domain is needed. The domain identifiers of different domains are different.

[0133] Among them, the inventors found that in order for the text encoder to have high performance for texts in a specific domain, the quantity of the first training data in the specific domain required for training the text encoder often needs to be large.

[0134] However, a specific domain is not a general domain. Sometimes a specific domain is a very small domain, and sometimes the texts that have appeared in a very small domain during the historical process are often few. Therefore, it is difficult and costly to collect a large amount of the first training data in the specific domain.

[0135] For this reason, in order to reduce the difficulty and cost of collecting the first training data in the specific domain, the quantity of the first training data in the specific domain that needs to be collected can be reduced.

[0136] In another embodiment of the present application, in order to reduce the quantity of the first training data in the specific domain that needs to be collected, the encoder to be trained can include an encoder obtained after training (such as pre-training) an initial encoder. That is, the initial encoder can be trained first to optimize the initial encoder and obtain an optimized encoder, and the optimized encoder is used as the encoder to be trained.

[0137] For example, in another embodiment of the present application, refer to Figure 2 , step S101 includes:

[0138] In step S201, an initial encoder is obtained.

[0139] The initial encoder can include the encoder in an autoencoder model. The autoencoder model includes a Transformer model, etc. For example, it includes the encoder in a bert model, etc.

[0140] The initial encoder can include an encoder that has not been trained (such as not pre-trained), or can also include an encoder obtained after training (such as pre-training) an encoder that has not been trained. The specific obtaining method of the initial encoder can be referred to the embodiments shown later and will not be elaborated here.

[0141] In step S202, a plurality of second training data are obtained. A second training data includes a second training text and a second annotation text. The second training text includes historical text pairs in the general domain, and the historical text pairs include a first historical text and a second historical text that are related to each other. The second annotation text includes at least the feature similarity between the text features of the first historical text and the text features of the second historical text in the historical text pair.

[0142] In this embodiment, the historical text pairs in the general domain included in the collected second training text can be unlabeled data. Subsequently, the electronic device automatically calculates the feature similarity between the text features of the first historical text and the text features of the second historical text as the labeled data, and the process of obtaining the labeled data can be without human participation. Therefore, this embodiment does not involve labor costs for the labeled data and does not form a dependence on the labeled data.

[0143] The first historical text and the second historical text in a historical text pair are different texts.

[0144] The first historical texts in different historical text pairs are different and / or the second historical texts are different.

[0145] The general domain is not for a specific domain, but involves multiple different domains, even a large number of different domains.

[0146] For example, the historical texts in the second training text of each of the plurality of second training data are respectively related to their own domains, and there are many domains in the union of the domains of the historical texts in the second training text of each of the plurality of second training data.

[0147] The historical text pairs in different second training data are different.

[0148] The historical text pairs in the second training text of each second training data are all text pairs that have appeared in the historical process. In this way, text pairs that have appeared in the historical process can be directly and automatically collected without arranging manual collection, and then different second training texts are generated according to each historical text pair.

[0149] In an example, the relationship between the first historical text and the second historical text can be understood as: the first historical text and the second historical text are respectively the main body and the title of the same article. In this way, the relationship is the relationship between the main body and the title of the same article.

[0150] In another example, the association relationship between the first historical text and the second historical text can be understood as follows: the first historical text is a question, and the second historical text is the answer corresponding to the question of the first historical text. Thus, the association relationship is the relationship between a question and its answer.

[0151] In yet another example, the association relationship between the first historical text and the second historical text can be understood as follows: the first historical text is a commodity, and the second historical text is the evaluation corresponding to the commodity of the first historical text. Thus, the association relationship is the relationship between a commodity and its evaluation.

[0152] Thus, multiple second training data can be multiple batches of second training data. Each batch of second training data has at least two second training data. The association relationship between the first historical text and the second historical text in the historical text pairs in the same batch of second training data is of the same type, and the association relationship between the first historical text and the second historical text in the historical text pairs in different batches of second training data is of different types.

[0153] In another embodiment of the present application, the second labeled text includes a second positive labeled sub - text and a second negative labeled sub - text.

[0154] The second positive labeled sub - text includes the feature similarity between the text features of the first historical text in the historical text pair in this one second training text and the text features of the second historical text in the historical text pair in this one second training text.

[0155] The second negative labeled sub - text includes: the feature similarity between the text features of the first historical text in the historical text pair in this one second training text and the text features of the first historical text or the text features of the second historical text in the historical text pairs in the second training texts in the additional second training data. The additional second training data includes at least part of the second training data other than the second training data where this one second training text is located among the multiple second training data.

[0156] Thus, for any batch of second training data, the following operations can be performed:

[0157] For any one second training data in this batch of second training data, input the first historical text and the second historical text in the historical text pair in this one second training data into the initial encoder respectively, so that the initial encoder encodes the first historical text to obtain the text features of the first historical text, and, enables the initial encoder to encode the second historical text to obtain the text features of the second historical text.

[0158] Then calculate the feature similarity between the text features of the first historical text and the text features of the second historical text.

[0159] For example, the inner product between the feature vector of the first historical text and the feature vector of the second historical text can be calculated, the norm value of the feature vector of the first historical text can be calculated, the norm value of the feature vector of the second historical text can be calculated, the product between the norm value of the feature vector of the first historical text and the norm value of the feature vector of the second historical text can be calculated, and the ratio between the inner product and the product can be calculated and used as the feature similarity between the text features of the first historical text and the text features of the second historical text.

[0160] In addition, the first historical text or the second historical text in the historical text pair in the second training text in the additional second training data is input into the initial encoder, so that the initial encoder encodes the first historical text or the second historical text in the historical text pair in the second training text in the additional second training data, and the text features of the first historical text or the text features of the second historical text in the historical text pair in the second training text in the additional second training data are obtained. Then, the feature similarity between the text features of the first historical text in the historical text pair in one second training text and the text features of the first historical text or the text features of the second historical text in the historical text pair in the second training text in the additional second training data is calculated. The specific calculation method can refer to the method of "the feature similarity between the text features of the first historical text and the text features of the second historical text", which will not be elaborated here.

[0161] For each of the other second training data in this batch of second training data, the above operations are similarly performed.

[0162] The same applies to other batches of second training data.

[0163] The text features of the first historical text include the feature vector of the first historical text.

[0164] The text features of the second historical text include the feature vector of the second historical text.

[0165] The dimensions of the text features of the first historical text and the dimensions of the text features of the second historical text can be the same. For example, the dimensions of the feature vector of the first historical text and the dimensions of the feature vector of the second historical text can be the same.

[0166] In step S203, multiple second training data are used to train the initial encoder until the parameters in the initial encoder converge, and a to-be-trained encoder applicable to the general domain is obtained.

[0167] For example, the second training data is brought into the contrast loss function loss_contrast, and the parameters in the initial encoder are optimized by minimizing the contrast loss function loss_contrast.

[0168] For example, for any one of the second training data in any batch of second training data, calculate the ratio of the second positive annotation sub-text "the feature similarity between the text features of the first historical text and the text features of the second historical text in the historical text pair in this second training text" in the second positive annotation text in this second training text to the temperature coefficient, and calculate the power of e to this ratio to obtain the positive value corresponding to this second training data. The temperature coefficient often ranges from 0 to 1 and is adjusted as the encoder is optimized, or is a preset fixed value. Specifically, which value between 0 and 1 is selected can be determined according to the actual situation, and this application does not limit this.

[0169] For the second negative annotation sub-text "the feature similarity between the text features of the first historical text in the historical text pair in this second training text and the text features of the first historical text or the text features of the second historical text in the historical text pairs in each additional second training data" in the second annotation text in this second training data, calculate the ratio to the temperature coefficient respectively, calculate the power of e to each ratio respectively, and sum the powers of e to each ratio to obtain the negative value corresponding to this second training data.

[0170] Calculate the sum value between the positive value corresponding to this second training data and the negative value corresponding to this second training data, and calculate the ratio between the positive value corresponding to this second training data and this sum value to obtain the ratio corresponding to this second training data.

[0171] Do the same for every other second training data in this batch of second training data.

[0172] Sum the ratios corresponding to each second training data in each batch of second training data to obtain the total sum value, and optimize the parameters in the initial encoder by minimizing the value of the contrast loss function loss_contrast (the total sum value).

[0173] Do the same for other batches of second training data.

[0174] In this way, the encoder to be trained obtained through this embodiment can be an encoder to be trained applicable to the general field.

[0175] Subsequently, the encoder to be trained applicable to the general field can be stored for subsequent invocation.

[0176] For example, if it is necessary to obtain a text encoder applicable to a specific field later, the encoder to be trained applicable to the general field can be directly called, and then the encoder to be trained applicable to the general field can be optimized using the training data applicable to the specific field to obtain a text encoder applicable to the specific field.

[0177] In another embodiment of the present application, referring to Figure 3 , step S201 includes:

[0178] In step S301, a plurality of third training data are obtained. A third training data includes a third training text and a third annotation text. The third training text includes: a masked text obtained by masking at least one vocabulary in the sample text. The third annotation text includes: the masked vocabulary in the masked text.

[0179] In this embodiment, the collected is the sample text. The process of masking the sample text can be automatically executed by the electronic device without human participation. The sample text can be unannotated data, and the third annotation text can be automatically obtained by the electronic device after masking the sample text, which is not manually annotated. Therefore, this embodiment does not involve human costs for the annotated data and does not rely on the annotated data.

[0180] The unmasked sample texts corresponding to the masked texts in the third training texts in different third training data are different.

[0181] There are multiple vocabularies in the sample text.

[0182] Among them, after masking at least one vocabulary in the sample text, a masked text can be obtained, and the masked text can be used as the third training text.

[0183] Regarding "which vocabularies in the sample text to mask", the present application does not limit this. For example, the number of masked vocabularies is not limited, and which vocabularies to mask is not limited.

[0184] In step S302, the network structure of the initial model is constructed. The initial model includes an encoder and a decoder. The output end of the encoder is connected to the input end of the decoder. The encoder is used to obtain the features of the text. The decoder is used to predict at least the masked vocabulary in the text according to the features of the text.

[0185] The initial model can include a Transformer model. For example, it includes a bert model, etc., or other models with an encoder and an encoder, etc.

[0186] In step S303, the initial model is trained using a plurality of third training data until the parameters in the initial model converge to obtain a text processing model.

[0187] For example, for any third training data, the masked text in the third training text of the third training data is input into the encoder in the initial model, so that the encoder encodes the masked text to obtain the text features (feature vectors) of the masked text. Then, the text features of the masked text are input into the decoder, so that the decoder predicts the masked words in the masked text according to the text features of the masked text. Then, the predicted masked words in the masked text and the masked words in the masked text of the third annotation text in the third training data are input into the cross-entropy loss function, and the parameters of the encoder and the parameters of the decoder in the initial model are optimized by using the gradient descent method.

[0188] The same applies to each of the other third training data.

[0189] In step S304, an initial encoder is obtained according to the encoder in the text processing model.

[0190] In an embodiment of the present application, the encoder in the text processing model can be determined as the initial encoder.

[0191] Alternatively, in another embodiment of the present application, the encoder in the text processing model can be further optimized, and the further optimized encoder is determined as the initial encoder.

[0192] In another embodiment of the present application, the masked text in the third training text includes: a first masked text and a second masked text. The first masked text is obtained by masking at least one word in the sample text, and the second masked text is obtained by masking at least one word in the sample text. The number of masked words in the first masked text is less than the number of masked words in the second masked text.

[0193] The second annotation text is obtained according to the masked words in the second masked text.

[0194] The first masked text is used to be input into the encoder, and the second masked text is used to be input into the decoder.

[0195] For example, the ratio between the number of masked words in the first masked text and the total number of words in the sample text can be 10% - 20%.

[0196] For example, the ratio between the number of masked words in the second masked text and the total number of words in the sample text can be 50% - 70%.

[0197] Thus, for any third training data, the first masked text in the third training text of the third training data is input into the encoder in the initial model, so that the encoder encodes the first masked text to obtain the text features (feature vectors) of the first masked text, and then the text features of the first masked text are input into the decoder. The second masked text in the third training text of the third training data is input into the decoder. So that the decoder predicts the masked vocabulary in the masked text according to the text features of the first masked text and the second masked text. Then, the predicted masked vocabulary in the masked text and the masked vocabulary in the second masked text of the third labeled text in the third training data are input into the cross-entropy loss function, and the parameters of the encoder and the parameters of the decoder in the initial model are optimized by using the gradient descent method.

[0198] The same applies to each of the other third training data.

[0199] Referring to Figure 4 , a step flowchart of a text classification method of the present application is shown. This method is applied to an electronic device. Wherein, the method includes:

[0200] In step S401, the text to be classified is obtained, and the text features of the text to be classified are obtained based on the text encoder.

[0201] The text to be classified can be text that needs to be classified submitted by a large number of users to the electronic device, etc.

[0202] In this step, the text to be classified can be input into the text encoder, so that the text encoder obtains the text features of the text to be classified and outputs the text features of the text to be classified, so that the electronic device obtains the text features of the text to be classified output by the text encoder.

[0203] The text features of the text to be classified can be feature vectors of the text to be classified, etc.

[0204] In step S402, multiple preset classification labels are obtained, and the label features of each preset classification label are obtained based on the text encoder.

[0205] The multiple preset classification labels are a classification label pool set in advance according to business requirements or other requirements. When classifying any text, the classification label of the text can be selected from the multiple preset classification labels.

[0206] The multiple preset classification labels can be multi-level classification labels.

[0207] The multiple preset classification labels can be increased, decreased, changed, etc. according to the actual situation later.

[0208] In this step, for any one of the multiple preset classification labels, the preset classification label can be input into the text encoder so that the text encoder obtains the label feature of the preset classification label and outputs the label feature of the preset classification label, enabling the electronic device to obtain the label feature of the preset classification label output by the text encoder.

[0209] The label feature of the preset classification label can be the feature vector of the preset classification label, etc.

[0210] The label feature of the preset classification label can also be stored locally in the electronic device for subsequent direct acquisition.

[0211] The same applies to each of the other preset classification labels among the multiple preset classification labels.

[0212] In step S403, the feature similarities between the text feature of the text to be classified and the label features of each preset classification label are obtained.

[0213] The dimension of the text feature of the text to be classified is the same as the dimension of the label feature of each preset classification label.

[0214] For example, the dimension of the feature vector of the text to be classified is the same as the dimension of the feature vector of each preset classification label.

[0215] For any one of the multiple preset classification labels, when the text feature of the text to be classified is the feature vector of the text to be classified and the label feature of the preset classification label is the feature vector of the preset classification label, the inner product between the feature vector of the text to be classified and the feature vector of the preset classification label can be calculated, the modulus value of the feature vector of the text to be classified can be calculated, the modulus value of the feature vector of the preset classification label can be calculated, the product between the modulus value of the feature vector of the text to be classified and the modulus value of the feature vector of the preset classification label can be calculated, and the ratio between the inner product and the product can be calculated and used as the feature similarity between the text feature of the text to be classified and the label feature of the preset classification label.

[0216] The same applies to each of the other preset classification labels among the multiple preset classification labels.

[0217] In step S404, among the label features of each preset classification label, the label feature with the largest feature similarity to the text feature of the text to be classified is selected.

[0218] In step S405, the preset classification label corresponding to the selected label feature is determined as the classification label of the text to be classified.

[0219] In an embodiment of the present application, the text encoder used in step S401 and step S402 may be obtained based on any one of the foregoing embodiments.

[0220] The encoder is pre-trained using a large number of unlabeled training texts (masked texts) (corresponding to the Figure 3 illustrated embodiment), to obtain Encoder1.

[0221] Among them, the process of pre-training the encoder may not use labeled training texts, so that it is not necessary to rely on manual annotation of the training texts, reducing the dependence on labeled data and saving the resources consumed by data annotation (such as time resources and human resources, etc.).

[0222] A large number of contrastive learning text pairs are constructed using the structural information of the text. The contrastive learning text pairs may be text pairs in the general domain, or may be text pairs that are not limited to a specific domain but cover a large number of domains. The two texts in the text pair are related to each other. The encoder Encoder1 is fine-tuned using a large number of contrastive learning text pairs to obtain a general encoder Encoder2 (corresponding to the Figure 2 illustrated embodiment).

[0223] Among them, the process of fine-tuning the encoder Encoder1 may use unlabeled contrastive learning texts, so that it is not necessary to rely on manual annotation of the contrastive learning texts, reducing the dependence on labeled data and saving the resources consumed by data annotation (such as time resources and human resources, etc.).

[0224] Secondly, the encoder Encoder2 fine-tuned by general data can assist the cold start of the encoder in a specific domain. For example, the obtained encoder Encoder2 may be a to-be-trained encoder applicable to the general domain. Subsequently, the encoder Encoder2 applicable to the general domain may be stored for subsequent invocation. For example, if an encoder applicable to a specific domain needs to be obtained later, the encoder Encoder2 applicable to the general domain can be directly invoked, and then the encoder Encoder2 applicable to the general domain is optimized using the training data applicable to the specific domain to obtain the text encoder for the specific domain.

[0225] For example, by combining the encoder Encoder2 in the general domain with business data for negative example mining, an encoder Encoder3 for dynamic label classification in a specific domain (such as the government work order domain) is obtained (corresponding to the Figure 1The embodiments shown). Alternatively, after fine-tuning a relatively small amount of labeled service data, an encoder Encoder3 for dynamic label classification in a specific field (such as the field of government work orders) is obtained. When there is no specific business data related to Greene, it provides preliminary assistance for classification in a specific field.

[0226] The application of contrastive learning technology to unlabeled data with a specific structure in this application does not rely on a pre-defined label system, and can more flexibly adapt to the actual situation where the label systems in specific fields (such as the field of government work orders) are diverse and not enumerable. It can dynamically adapt to the adjustment of the label system and has strong generalization performance.

[0227] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily essential to this application.

[0228] Referring to Figure 5 , a device for obtaining a text encoder according to this application is shown. The device includes:

[0229] A first acquisition module 11, configured to acquire an encoder to be trained;

[0230] A second acquisition module 12, configured to acquire a plurality of first training data. One first training data includes a first training text and a first annotation text. The first training text includes historical texts that have appeared in the historical process and classification labels obtained by classifying the historical texts in the historical process. The first annotation text includes a first positive annotation sub-text and a first negative annotation sub-text. The first positive annotation sub-text includes the feature similarity between the label feature of the classification label and the text feature of the historical text. The first negative annotation sub-text includes: the feature similarity between the label feature of the classification label and the text feature of the historical text in at least part of the first training texts in the plurality of first training data other than the first training data where the historical text is located;

[0231] A training module 13, configured to train the encoder to be trained using the plurality of first training data until the parameters in the encoder to be trained converge, and obtain a text encoder.

[0232] In an optional implementation manner, the historical texts in the first training text include historical texts in a specific field.

[0233] In an optional implementation manner, the device further includes:

[0234] A storage module for storing the correspondence between the domain identifier of the specific domain and the text encoder.

[0235] In an optional implementation manner, the first acquisition module includes:

[0236] A first acquisition unit for acquiring an initial encoder;

[0237] A second acquisition unit for acquiring a plurality of second training data. One second training data includes a second training text and a second annotation text. The second training text includes historical text pairs in the general domain. The historical text pairs include a first historical text and a second historical text with an associated relationship therebetween. The second annotation text at least includes the feature similarity between the text features of the first historical text and the text features of the second historical text in the historical text pair.

[0238] A training unit for training the initial encoder using the plurality of second training data until the parameters in the initial encoder converge to obtain a to-be-trained encoder applicable to the general domain.

[0239] In an optional implementation manner, the second annotation text includes a second positive annotation sub-text and a second negative annotation sub-text;

[0240] The second positive annotation sub-text includes the feature similarity between the text features of the first historical text and the text features of the second historical text in the historical text pair in the one second training text.

[0241] The second negative annotation sub-text includes: the feature similarity between the text features of the first historical text in the historical text pair in the one second training text and the text features of the first historical text or the text features of the second historical text in the historical text pair in the second training text in the additional second training data. The additional second training data includes at least part of the second training data other than the second training data where the one second training text is located among the plurality of second training data.

[0242] In an optional implementation manner, the first acquisition module further includes:

[0243] A storage unit for storing the to-be-trained encoder applicable to the general domain.

[0244] In an optional implementation manner, the first acquisition unit includes:

[0245] A first acquisition subunit, configured to acquire a plurality of third training data, where a third training data includes a third training text and a third annotation text, the third training text includes: a masked text obtained by masking at least one word in a sample text, and the third annotation text includes: the masked words in the masked text;

[0246] A construction subunit, configured to construct a network structure of an initial model, where the initial model includes an encoder and a decoder, an output end of the encoder is connected to an input end of the decoder, the encoder is configured to acquire features of a text, and the decoder is configured to predict at least the masked words in the text according to the features of the text;

[0247] A training subunit, configured to train the initial model using the plurality of third training data until the parameters in the initial model converge, to obtain a text processing model;

[0248] A second acquisition subunit, configured to acquire an initial encoder according to the encoder in the text processing model.

[0249] In an optional implementation manner, the masked text in the third training text includes: a first masked text and a second masked text, the first masked text is obtained by masking at least one word in the sample text, the second masked text is obtained by masking at least one word in the sample text, and the number of masked words in the first masked text is less than the number of masked words in the second masked text;

[0250] The second annotation text is obtained according to the masked words in the second masked text;

[0251] The first masked text is used to be input into the encoder, and the second masked text is used to be input into the decoder.

[0252] Refer to Figure 6 , which shows a text classification device of the present application, the device includes:

[0253] A third acquisition module 21, configured to acquire a text to be classified, and acquire text features of the text to be classified based on a text encoder;

[0254] A fourth acquisition module 22, configured to acquire a plurality of preset classification labels, and acquire label features of each preset classification label based on the text encoder respectively;

[0255] A fifth acquisition module 23, configured to acquire feature similarities between the text features of the text to be classified and the label features of each preset classification label respectively;

[0256] A selection module 24, configured to select, from the label features of each preset classification label, the label feature with the maximum feature similarity to the text features of the text to be classified;

[0257] A determination module 25, configured to determine a preset classification label corresponding to the selected label feature as the classification label of the text to be classified.

[0258] In an optional implementation manner, the text encoder is based on Figures 1 - 3 any of the illustrated embodiments; or, the text encoder is based on Figure 5 the one shown and obtained.

[0259] Optionally, an embodiment of the present application further provides an electronic device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0260] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (Read-Only Memory, abbreviated as ROM), a random access memory (Random Access Memory, abbreviated as RAM), a magnetic disk, or an optical disc, etc.

[0261] Figure 7 is a block diagram of an electronic device 800 shown in the present application. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0262] Referring to Figure 7 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0263] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0264] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, and the like. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0265] The power supply component 806 provides power to various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.

[0266] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions but also monitor the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0267] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.

[0268] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.

[0269] The sensor assembly 814 includes one or more sensors for providing an assessment of the status of various aspects of the electronic device 800. For example, the sensor assembly 814 can monitor the on / off state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also monitor a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0270] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0271] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.

[0272] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions that can be executed by a processor 820 of the electronic device 800 to complete the above-described method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0273] Figure 8is a block diagram of an electronic device 1900 shown in the present application. For example, the electronic device 1900 can be provided as a server.

[0274] Referring to Figure 8 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0275] The electronic device 1900 may further include a power component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.

[0276] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0277] Through the description of the above embodiments, those skilled in the art can clearly understand that the above method of the embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0278] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

[0279] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0280] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0281] In the embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical or other form.

[0282] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0283] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0284] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0285] As described above, the foregoing are only specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

1. A method for obtaining a text encoder, characterized in that: The method comprises: Get the encoder to be trained; Acquire multiple first training data, one of the first training data includes a first training text and a first annotated text, the first training text includes a historical text that has appeared in a historical process and a classification label obtained after classifying the historical text in the historical process, the first annotated text includes a first positive annotated subtext and a first negative annotated subtext, the first positive annotated subtext includes a feature similarity between a label feature of the classification label and a text feature of the historical text, and the first negative annotated subtext includes: a feature similarity between a label feature of the classification label and a text feature of a historical text in the first training text in at least part of the first training data other than the first training data where the historical text is located among the multiple first training data; The encoder to be trained is trained using the plurality of first training data until the parameters of the encoder to be trained converge, thereby obtaining a text encoder.

2. The method according to claim 1, characterized in that The historical texts in the first training text include historical texts in a specific field.

3. The method according to claim 2, characterized in that The method further comprises: The correspondence between the domain identifier of the specific domain and the text encoder is stored.

4. The method according to claim 1 or 2, characterized in that: The step of obtaining the encoder to be trained comprises: Get the initial encoder; Acquire multiple second training data, one of the second training data includes a second training text and a second annotated text, the second training text includes a historical text pair in a general field, the historical text pair includes a first historical text and a second historical text that have an association relationship, and the second annotated text includes at least a feature similarity between a text feature of the first historical text and a text feature of the second historical text in the historical text pair; The initial encoder is trained using a plurality of second training data until the parameters in the initial encoder converge, thereby obtaining an encoder to be trained that is applicable to a general field.

5. The method according to claim 4, characterized in that The second annotated text includes a second positive annotated subtext and a second negative annotated subtext; The second positively annotated subtext includes feature similarity between text features of a first historical text and text features of a second historical text in a pair of historical texts in the second training text; The second negatively annotated sub-text includes: feature similarity between text features of the first historical text in the historical text pair in the one second training text and text features of the first historical text or the second historical text in the historical text pair in the second training text in additional second training data, the additional second training data including at least part of the second training data among multiple second training data except the second training data where the one second training text is located.

6. The method according to claim 4, characterized in that The method further comprises: Stores encoders to be trained for general domains.

7. The method according to claim 4, characterized in that The obtaining of the initial encoder comprises: Acquire multiple third training data, one of the third training data includes a third training text and a third annotated text, the third training text includes: a masked text obtained by masking at least one word in the sample text, and the third annotated text includes: the masked words in the masked text; Constructing a network structure of an initial model, wherein the initial model includes an encoder and a decoder, wherein an output end of the encoder is connected to an input end of the decoder, the encoder is used to obtain features of the text, and the decoder is used to predict masked words in the text based on at least the features of the text; Using a plurality of third training data to train the initial model until the parameters in the initial model converge, thereby obtaining a text processing model; Get the initial encoder from the encoder in the text processing model.

8. The method according to claim 7, characterized in that The masked text in the third training text includes: a first masked text and a second masked text, the first masked text is obtained by masking at least one word in the sample text, the second masked text is obtained by masking at least one word in the sample text, and the number of masked words in the first masked text is less than the number of masked words in the second masked text; The second annotated text is obtained based on the masked words in the second masked text; The first mask text is used as input to the encoder, and the second mask text is used as input to the decoder.

9. A text classification method, characterized in that: The method comprises: Obtaining the text to be classified, and obtaining the text features of the text to be classified based on the text encoder; Obtain multiple preset classification labels, and obtain label features of each preset classification label based on a text encoder; Obtaining feature similarities between the text features of the text to be classified and the label features of each preset classification label; Among the label features of each preset classification label, select the label feature with the greatest feature similarity with the text feature of the text to be classified; The preset classification label corresponding to the selected label feature is determined as the classification label of the text to be classified.

10. The method according to claim 9, characterized in that The text encoder is obtained based on any one of claims 1-8.

11. A device for obtaining a text encoder, characterized in that: The device comprises: A first acquisition module is used to acquire an encoder to be trained; a second acquisition module, for acquiring a plurality of first training data, wherein one first training data includes a first training text and a first annotated text, the first training text includes a historical text that has appeared in a historical process and a classification label obtained after classifying the historical text in the historical process, the first annotated text includes a first positive annotated subtext and a first negative annotated subtext, the first positive annotated subtext includes a feature similarity between a label feature of the classification label and a text feature of the historical text, and the first negative annotated subtext includes: a feature similarity between a label feature of the classification label and a text feature of a historical text in the first training text in at least part of the first training data other than the first training data where the historical text is located among the plurality of first training data; The training module is used to train the encoder to be trained using a plurality of first training data until the parameters in the encoder to be trained converge to obtain a text encoder.

12. A text classification device, characterized in that: The device comprises: A third acquisition module is used to acquire the text to be classified, and acquire the text features of the text to be classified based on the text encoder; A fourth acquisition module is used to acquire multiple preset classification labels, and acquire label features of each preset classification label based on a text encoder; A fifth acquisition module is used to obtain feature similarities between text features of the text to be classified and label features of each preset classification label; A selection module is used to select, from the label features of each preset classification label, the label feature with the greatest feature similarity to the text feature of the text to be classified; The determination module is used to determine the preset classification label corresponding to the selected label feature as the classification label of the text to be classified.

13. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 10 when executed by the processor.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.