Text classification method and device
By segmenting and semantic extraction of classified texts, and classifying them with fragments and text semantic vectors, the problem of distinguishing long and short texts in the existing technology is solved, and the efficiency and accuracy of text classification are improved.
Patent Information
- Application Number
- CN202110955119.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-08-19
AI Technical Summary
The prior art usually distinguishes long and short texts in text classification, which affects the efficiency and accuracy of text classification, and is difficult to apply to text sets that mix long and short texts.
By segmenting the classified text, several text fragments are obtained, and fragment semantics are extracted and classified respectively. If the classification result of a text fragment meets the preset confidence requirements, the text classification of the text to be classified is determined; if the classification result of all text fragments does not meet the confidence requirements, text semantic extraction and classification are combined with all fragment semantic vectors.
This method takes into account text classification of long and short texts, improves the efficiency and accuracy of text classification, and is suitable for text sets that mix long and short texts.
Smart Images

Figure CN113626602B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and device for text classification. Background Art
[0002] Internet technology has penetrated into all aspects of social life. In order to better understand users' evaluation of various commodities and services and inappropriate comments on the Internet, it has become a general trend to implement sentiment analysis, public opinion analysis and risk monitoring of online texts based on natural language processing technology. When classifying text, current related technologies usually distinguish between long and short texts, which greatly affects the efficiency and accuracy of text classification. Summary of the invention
[0003] In view of this, one or more embodiments of the present specification provide a method and device for text classification.
[0004] To achieve the above objectives, the technical solutions provided by one or more embodiments of this specification are as follows:
[0005] According to a first aspect of one or more embodiments of this specification, a method for text classification is proposed, the method comprising:
[0006] Segment the text to be classified to obtain several text fragments;
[0007] For each text segment, the text segment is input as an input parameter into a trained segment semantic extraction model to perform semantic extraction on the text segment, and a segment semantic vector corresponding to the text segment is obtained;
[0008] Inputting the segment semantic vector as an input parameter into a trained first classification model to classify the text segment, thereby obtaining a classification result of the text segment;
[0009] If the classification result of any text segment meets the preset confidence requirement, then determine the text category to which the to-be-classified text belongs according to the classification result that meets the preset confidence requirement;
[0010] If the classification results of all text segments do not meet the confidence requirement, the segment semantic vectors corresponding to the segment semantic vectors are input as input parameters into the trained text semantic extraction model to perform semantic extraction on the text to be classified, so as to obtain the text semantic vectors corresponding to the text to be classified;
[0011] The text semantic vector is input as an input parameter into a trained second classification model to classify the text to be classified, and the text category to which the text to be classified belongs is determined.
[0012] According to a second aspect of one or more embodiments of this specification, a device for text classification is provided, the device comprising:
[0013] A text segmentation unit segments the text to be classified into several text segments;
[0014] A segment semantic extraction unit, for each text segment, inputs the text segment as an input parameter into a trained segment semantic extraction model to perform semantic extraction on the text segment, and obtains a segment semantic vector corresponding to the text segment;
[0015] A segment classification unit, inputting the segment semantic vector as an input parameter into a trained first classification model to classify the text segment, and obtaining a classification result of the text segment;
[0016] If the classification result of any text segment meets the preset confidence requirement, then determine the text category to which the to-be-classified text belongs according to the classification result that meets the preset confidence requirement;
[0017] A text extraction unit, if the classification results of all text segments do not meet the confidence requirement, inputs the multiple segment semantic vectors corresponding to the multiple text segments as input parameters into a trained text semantic extraction model to perform semantic extraction on the text to be classified, and obtains the text semantic vector corresponding to the text to be classified;
[0018] The text classification unit inputs the text semantic vector as an input parameter into a trained second classification model to classify the text to be classified and determine the text category to which the text to be classified belongs.
[0019] According to a third aspect of one or more embodiments of this specification, an electronic device is provided, comprising a processor and a memory for storing machine-executable instructions;
[0020] Among them, by reading and executing the machine executable instructions corresponding to the logic of text classification stored in the memory, the processor implements the steps of the method described in the first aspect above.
[0021] According to a fourth aspect of one or more embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0022] It can be seen from the above description that in this specification, first, the fragment semantics of several text fragments obtained by segmenting the text to be classified are extracted respectively, and each text fragment is classified based on the fragment semantic vector of each extracted text fragment. If the classification result of a certain text fragment can meet the preset confidence requirement, it means that the fragment semantics of the text fragment can represent the semantics of the text to be classified. According to the classification result of the text fragment, the text category to which the text to be classified belongs can be determined. Since the short text has the above-mentioned characteristic that the semantics of the text can be represented by the fragment semantics of the text fragment, the text category to which the short text belongs can be determined in the above-mentioned way.
[0023] If the classification results of all text segments obtained by segmenting the text to be classified do not meet the preset confidence requirements, it means that the segment semantics of only one text segment is not sufficient to characterize the semantics of the text to be classified, and the text category to which the text to be classified belongs cannot be determined by the classification result of a certain text segment. Therefore, the text semantics are extracted by combining the segment semantic vectors of all text segments, and the text to be classified is classified based on the extracted text semantic vectors of the text to be classified, thereby determining the text category to which the long text with semantics distributed in multiple text segments belongs. The text classification scheme provided in this specification takes into account the text classification of long and short texts at the same time, thereby improving the efficiency and accuracy of text classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a flowchart of a text classification method shown in an exemplary embodiment of this specification.
[0025] Figure 2 It is a schematic diagram of segmenting text to be classified, shown in an exemplary embodiment of this specification.
[0026] Figure 3 It is a schematic diagram of text preprocessing for text to be classified, shown in an exemplary embodiment of this specification.
[0027] Figure 4 It is a schematic diagram of a fragment semantic extraction model shown in an exemplary embodiment of this specification.
[0028] Figure 5 It is a schematic diagram of a text semantic extraction model shown in an exemplary embodiment of this specification.
[0029] Figure 6 It is a schematic diagram of the overall structure of a text classification model shown in an exemplary embodiment of this specification.
[0030] Figure 7 It is a flowchart of a method for training a text classification model shown in an exemplary embodiment of this specification.
[0031] Figure 8It is a structural diagram of an electronic device where a text classification device is located, shown as an exemplary embodiment of this specification.
[0032] Fig. 9 It is a block diagram of a text classification device shown in an exemplary embodiment of this specification. DETAILED DESCRIPTION
[0033] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with this specification. Instead, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0034] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. The singular forms "a", "the" and "the" used in this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0035] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0036] The Internet is filled with a large amount of text containing rich information. Classifying texts on the Internet based on natural language processing (NLP) technology can be effectively applied to various scenarios such as text sentiment analysis, public opinion analysis, and risk monitoring.
[0037] For example, by classifying texts such as product reviews and service reviews, we can determine users' emotional tendencies towards the products and services for market analysis; and by classifying texts such as news, current affairs, and personal social updates posted on major platforms, we can determine the public's opinion on a certain thing or event for public opinion analysis; in addition, text classification can also be used to distinguish the field to which the text to be published belongs and to detect whether the text to be published is compliant.
[0038] In the field of computer technology, text can generally be divided into long text and short text. The main difference between the two is the length of the text. There is currently no unified definition for long text and short text. In Microsoft's database system ACCESS, short text is defined as text with a length of less than 256 characters, and long text is defined as text with a length of more than 256 characters. Common short texts include text messages, emails, and document abstracts, while common long texts include news, current affairs, and long reviews of various products and services.
[0039] In the related technology, the text classification of long and short texts is generally distinguished from each other and processed independently. That is, different text classification models are used during training to train long text sets and short text sets respectively. In actual application, it will first be determined whether the text to be classified is a long text or a short text, and then the corresponding text classification model is selected to output the corresponding classification result.
[0040] The text classification method used in the related art cannot learn comprehensive text semantics during training, which affects the efficiency and accuracy of text classification and is not suitable for text collections of mixed long and short texts. The text classification method shown in this embodiment is applicable to long text collections, short text collections, and text collections of mixed long and short texts, and has higher text classification efficiency and accuracy.
[0041] Figure 1 FIG. 1 is a flowchart of a text classification method shown in an exemplary embodiment of this specification. The text classification method may include the following specific steps:
[0042] Step 102, segment the text to be classified to obtain a number of text segments.
[0043] Considering that the texts to be classified of different lengths are not easy to be processed uniformly by a computer, in this embodiment, the texts to be classified may be segmented to obtain several text segments.
[0044] There are multiple optional implementation methods for segmenting the text to be classified to obtain a number of text segments.
[0045] In an example, Figure 2 As shown, the text to be classified can be segmented by a sliding window according to a preset window length and a preset number of text segments.
[0046] The preset window length, i.e., the number of characters of the text to be classified included in the new window after each sliding window, can be a Chinese character, an English word, or a vocabulary unit in other forms;
[0047] The preset number of text segments refers to the number of text segments expected to be obtained after the text is segmented.
[0048] For example, assuming that the preset window length is 200 and the preset number of text segments is 5, the specific process of segmenting the text to be classified using the sliding window method is as follows:
[0049] The first window includes the 1st to 200th characters of the text to be classified. The preset segmentation marks such as punctuation marks, spaces, and line breaks are searched in sequence from the last character of the current window, that is, from the 200th character of the text to be classified, to segment the first text segment. Assuming that the preset segmentation mark is found at the 110th character, the first text segment obtained by segmentation is composed of the 1st to 110th characters of the text to be classified.
[0050] Slide the window backward from the 111th character of the text to be classified to obtain a second window including the 111th to 310th characters of the text to be classified, and continue segmentation as described above to obtain the second text segment of the text to be classified, and so on. No further details will be given.
[0051] When the sliding window method is used to segment the text to be classified, due to the different lengths of the text to be classified, the following two segmentation results may be produced:
[0052] (1) The number of text segments obtained by segmenting the text to be classified cannot reach the preset number of text segments. In this case, the preset text segments can be used to supplement the segmentation results;
[0053] The number of text segments obtained by segmenting short texts generally cannot reach the preset number of text segments, and the segmentation results need to be supplemented with preset text segments.
[0054] (2) The number of text segments obtained by segmenting the text to be classified exceeds the preset number of text segments. In this case, the text segments exceeding the preset number of text segments may be discarded according to the semantic order;
[0055] The number of text segments obtained by segmenting a long text may exceed the preset number of text segments, and the text segments exceeding the preset number of text segments need to be discarded.
[0056] In actual implementation, after each sliding window performs text segmentation, it will determine whether the text to be classified has been segmented and the number of text segments currently obtained by segmentation. The specific process is as follows:
[0057] After any sliding window is used to segment the text, it is determined whether the segmentation of the text to be classified is complete.
[0058] If the text to be classified has been segmented, it can be determined whether the number of text segments currently segmented has reached the preset number of text segments:
[0059] If yes, it means that the number of text segments obtained by segmenting the text to be classified is exactly the preset number of text segments, and no additional or discarding is required;
[0060] If not, it means that the number of text segments obtained by segmenting the text to be classified cannot reach the preset number of text segments. The preset text segments can be used to supplement the segmentation results to obtain a preset number of text segments. The preset text segments can be composed of preset meaningless characters according to a preset window length.
[0061] After any sliding window is used to segment the text, if the text to be classified has not been segmented, it can be determined whether the number of text segments obtained by the current segmentation reaches the preset number of text segments:
[0062] If yes, it means that the number of text segments obtained by segmenting the text to be classified exceeds the preset number of text segments, and the text segmentation is no longer performed, and the text to be classified after the text segment obtained by the current segmentation is discarded in the semantic order;
[0063] If not, continue with window sliding and text segmentation.
[0064] For example, for a short text to be classified, assuming that the first window after sliding has segmented it, the preset text segments can be used to supplement 4 text segments for the short text to be classified;
[0065] For a long text to be classified, assuming that the fifth window after sliding has not completed the segmentation, and the fifth text segment obtained by segmenting the fifth window consists of the 601st to 780th characters of the long text to be classified, then the text after the 781th character of the long text to be classified is discarded.
[0066] In this embodiment, before segmenting the text to be classified, the text to be classified may be subjected to a first text preprocessing to improve the efficiency of subsequent text processing, wherein the first text preprocessing includes text cleaning and text segmentation.
[0067] The text cleaning includes deleting invalid characters such as emoji, URL, etc. in the text to be classified, as well as performing spelling correction and grammar checking on the text to be classified; the text segmentation includes the segmentation of Chinese words and the segmentation of English affixes. The specific methods of the text cleaning and text segmentation refer to the relevant technology and will not be repeated here.
[0068] Step 104: for each text segment, input the text segment as an input parameter into a trained segment semantic extraction model to perform semantic extraction on the text segment, and obtain a segment semantic vector corresponding to the text segment.
[0069] In this embodiment, semantic extraction is performed twice. The first semantic extraction uses a segment semantic extraction model to extract segment semantics from each text segment.
[0070] Each text segment obtained by segmentation in step 102 is input as an input parameter into the trained segment semantic extraction model to obtain a segment semantic vector corresponding to each text segment. That is, multiple different text segments are input as input parameters into the same trained segment semantic extraction model to obtain corresponding multiple different segment semantic vectors.
[0071] For example, assuming that five text segments are obtained by segmenting the text to be classified, namely text segments 1 to 5, then text segments 1 to 5 are respectively input as input parameters into the trained segment semantic extraction model to obtain segment semantic vectors 1 to 5 corresponding to text segments 1 to 5 respectively.
[0072] It should be noted that if Figure 3 As shown, in order to facilitate unified processing by computers, each text segment is usually subjected to a second text preprocessing before being input into the segment semantic extraction model.
[0073] The second text preprocessing includes adding tags and padding the length, etc., so that each text segment can be input into the segment semantic extraction model in the same format and length.
[0074] The adding of the mark includes adding a classification mark CLS before the first character of the text segment to indicate the start of the original text segment, and adding a separation mark SEP after the last character of the text segment to indicate the end of the original text segment.
[0075] After completing the marker addition, it can be determined whether the current length of the text segment reaches the preset text segment length, and if not, the length of the text segment to which the marker addition is completed can be padded; the separator SEP, after the length is padded, is used to separate the original text segment and the filled characters.
[0076] The length padding, i.e. padding the text segment according to a preset text segment length, includes padding the text segment to the preset text segment length using preset meaningless characters.
[0077] For example, the segmented text fragment 1 includes the 1st to 110th characters of the text to be classified. The label CLS can be added before the 1st character, and the label SEP can be added after the 110th character. According to the preset text fragment length 512, the 113th to 512th characters of the text fragment 1 to which the label is added are padded with preset meaningless characters, and then the text fragment 1 after the label is added and the length is padded is input into the fragment semantic extraction model.
[0078] The following describes the structure of the segment semantic extraction model and the specific process of extracting a segment semantic vector of a text segment by the segment semantic extraction model.
[0079] like Figure 4 As shown, the fragment semantic extraction model includes an embedding layer and several serially connected fragment semantic extraction layers.
[0080] The embedding layer is used to convert the text fragment into a corresponding number of embedding vectors.
[0081] Specifically, the embedding layer will perform conversion on each character contained in the input text fragment, and convert each character contained in the text fragment into a corresponding embedding vector; based on the previous example, the embedding layer converts the input preprocessed text fragment 1 into an embedding vector corresponding to each of the 512 characters.
[0082] The plurality of serially connected segment semantic extraction layers, i.e., the output of the previous segment semantic extraction layer is the input of the next segment semantic extraction layer. Each segment semantic extraction layer is used to perform semantic extraction based on the vector output by the previous layer, and output the intermediate segment semantic vector extracted by this layer. Based on the intermediate segment semantic vector output by the last segment semantic extraction layer, the segment semantic extraction model will determine the segment semantic vector corresponding to the text segment.
[0083] Specifically, refer to Figure 4 The segment semantic extraction layer 1 takes the embedding vectors corresponding to the characters output by the embedding layer as input, performs the first segment semantic extraction through this layer, and outputs the intermediate segment semantic vectors corresponding to the characters extracted for the first time;
[0084] The segment semantic extraction layer 2 takes the intermediate segment semantic vector extracted for the first time output by the segment semantic extraction layer 1 as input, performs a second segment semantic extraction through this layer, and outputs the intermediate segment semantic vector corresponding to the plurality of characters extracted for the second time;
[0085] This process is repeated in this way until the last segment semantic extraction layer outputs the intermediate segment semantic vectors corresponding to the plurality of characters extracted for the last time.
[0086] The last segment semantic extraction layer outputs the intermediate segment semantic vectors corresponding to the several characters extracted last time, and the intermediate segment semantic vector corresponding to the identifier CLS extracted last time can be used as the segment semantic vector corresponding to the text segment.
[0087] Based on the previous example, the several serially connected fragment semantic extraction layers can output the last extracted intermediate fragment semantic vector corresponding to the 512 embedding vectors of the input text fragment 1 in the last fragment semantic extraction layer, wherein the last extracted intermediate fragment semantic vector corresponding to the identifier CLS can be used as the final fragment semantic vector corresponding to the text fragment 1.
[0088] There are multiple selectable implementation models for the fragment semantic extraction model described in this embodiment.
[0089] In one example, the fragment semantic extraction model can be constructed based on a BERT model (Bidirectional Encoder Representations from Transformers) or an ALBERT model (A Lite Bidirectional Encoder Representations from Transformers).
[0090] Specifically, the embedding layer in the fragment semantic extraction model can be implemented using the embedding layer in the BERT model or the ALBERT model, and several series-connected fragment semantic extraction layers in the fragment semantic extraction model can be implemented using several series-connected encoder layers in the BERT model or the ALBERT model.
[0091] Step 106: input the segment semantic vector as an input parameter into the trained first classification model to classify the text segment, and obtain a classification result of the text segment.
[0092] In this embodiment, text classification can be performed twice. The first text classification uses the first classification model to classify each text segment based on the segment semantic vector obtained by the first semantic extraction. Its purpose is to determine the text category to which the short text belongs based on the classification result of the text segment.
[0093] The semantic vectors of each fragment obtained by the first semantic extraction are respectively input as input parameters into the trained first classification model to obtain the classification results corresponding to each text fragment. That is, multiple different fragment semantic vectors are respectively input as input parameters into the same trained first classification model to obtain corresponding multiple different classification results.
[0094] For example, assuming that after the first semantic extraction, the fragment semantic vectors 1 to 5 corresponding to the text fragments 1 to 5 are obtained, then the fragment semantic vectors 1 to 5 are respectively input as input parameters into the trained first classification model to obtain the classification results 1 to 5 corresponding to the text fragments 1 to 5.
[0095] In this embodiment, there is no restriction on the specific types and quantities of text categories. Text categories can be simply divided into two categories: positive text and negative text, or divided into more different types according to actual application scenarios.
[0096] There are multiple selectable implementation models for the first classification model. For example, the first classification model can be a classification model implemented based on a neural network model such as LSTM (Long Short-Term Memory) and CNN (Convolutional Neural Networks).
[0097] The classification results of each text segment output by the first classification model for each segment semantic vector are determined by the specific first classification model adopted. For example, if the first classification model based on CNN is adopted, the confidence that each text segment belongs to each text classification can be obtained.
[0098] Step 108: If the classification result of any text segment meets the preset confidence requirement, the text category to which the to-be-classified text belongs is determined according to the classification result that meets the preset confidence requirement.
[0099] In this embodiment, for the classification results of each text fragment obtained after the first text classification in step 106, it is judged whether the classification results meet the preset confidence requirements. If there is a classification result that can meet the confidence requirements among the classification results of the several text fragments, the text category to which the text to be classified belongs can be determined based on the classification result of this text fragment.
[0100] For example, assuming that text is classified into two categories, positive text and negative text, and the preset confidence requirement is that the confidence that the text to be classified belongs to any text category reaches a preset confidence threshold of 0.9, the classification result output by the first classification model for text fragment 1 is: the confidence level for positive text is 0.91, and the confidence level for negative text is 0.09. Then, we can directly determine that the text to be classified belongs to positive text based on the classification result output by the first classification model for text fragment 1.
[0101] Since there is a text segment that includes all text semantics or highlights text semantics among the several text segments obtained by segmenting short texts and texts with prominent local semantics, the short texts and texts with prominent local semantics can use the classification results of the text segment that includes all text semantics or highlights text semantics to directly determine the text to which the text to be classified belongs. The text classification of the short text can be achieved by the method described in step 108 without executing subsequent steps.
[0102] Step 110, if the classification results of all text fragments do not meet the confidence requirement, the several fragment semantic vectors corresponding to the several text fragments are input as input parameters into the trained text semantic extraction model to perform semantic extraction on the text to be classified, and obtain the text semantic vector corresponding to the text to be classified.
[0103] If the classification results of all text fragments obtained by segmenting the text to be classified cannot meet the preset confidence requirements, it means that the fragment semantics of only one text fragment cannot represent the overall semantics of the text to be classified, and the text category to which the text to be classified belongs cannot be determined based on the classification result of one text fragment. It is necessary to combine the fragment semantics of all text fragments to represent the text to be classified and extract the text semantics of the text to be classified for a second text classification.
[0104] According to the semantic order, the text to be classified is represented in combination with the several segment semantic vectors obtained in step 104 and a second semantic extraction is performed, that is, the text semantics is extracted from the several segment semantic vectors using a text semantic extraction model.
[0105] The fragment semantic vectors corresponding to all text fragments extracted by the fragment semantic extraction model are input into the trained text semantic extraction model in semantic order as input parameters to obtain the text semantic vector corresponding to the text to be classified. That is, multiple different fragment semantic vectors are input into a trained text semantic extraction model in semantic order as input parameters to obtain a text semantic vector corresponding to the text to be classified.
[0106] For example, assuming that the segment semantic vectors 1 to 5 corresponding to the text segments 1 to 5 are extracted in step 104, the segment semantic vectors 1 to 5 are input into the trained text semantic extraction model in semantic order as input parameters to obtain the text semantic vector corresponding to the text to be classified.
[0107] As described above, before the plurality of segment semantic vectors are input into the text semantic extraction model, a third text preprocessing including adding identifiers and length padding may be performed.
[0108] Based on the previous example, the identifier CLS can be added before the fragment semantic vector 1, the identifier SEP can be added after the fragment semantic vector 5, and according to the preset text length 8, a preset meaningless fragment semantic vector can be added to the fragment semantic vectors 1 to 5 that have completed the identifier addition. Then, the several fragment semantic vectors after the identifier addition and the length padding are input into the text semantic extraction model in semantic order.
[0109] The following describes the structure of the text semantic extraction model and the specific process of extracting the text semantic vector of the text to be classified by the text semantic extraction model.
[0110] like Figure 5 As shown, the text semantic extraction model includes several serially connected text semantic extraction layers.
[0111] Among them, the output of the several serially connected text semantic extraction layers, that is, the output of the previous text semantic extraction layer is the input of the next text semantic extraction layer. Each of the text semantic extraction layers is used to perform semantic extraction based on the vector input to this layer, and output the intermediate text semantic vector obtained by the extraction of this layer. Based on the intermediate text semantic vector output by the last text semantic extraction layer, the text semantic extraction model will determine the text semantic vector corresponding to the text to be classified.
[0112] Specifically, refer to Figure 5 The text semantic extraction layer 1 takes the semantic vectors of the segments after adding identifiers and padding the lengths in the semantic order as input, performs the first text semantic extraction through this layer, and outputs the intermediate text semantic vectors corresponding to the semantic vectors of the segments extracted for the first time;
[0113] The text semantic extraction layer 2 takes the intermediate text semantic vector extracted for the first time output by the text semantic extraction layer 1 as input, performs a second text semantic extraction through this layer, and outputs the intermediate text semantic vector corresponding to the plurality of segment semantic vectors extracted for the second time;
[0114] This process is repeated in this way until the last text semantic extraction layer outputs the intermediate text semantic vectors corresponding to the plurality of segment semantic vectors extracted last time.
[0115] The last text semantic extraction layer outputs the intermediate text semantic vector corresponding to the several fragment semantic vectors extracted for the last time, and the intermediate text semantic vector corresponding to the identifier CLS extracted for the last time can be used as the text semantic vector corresponding to the text to be classified.
[0116] Based on the previous example, the several serially connected text semantic extraction layers can output the last extracted intermediate text semantic vector corresponding to the 8 input segment semantic vectors in the last text semantic extraction layer, wherein the last extracted intermediate text semantic vector corresponding to the identifier CLS can be used as the text semantic vector finally corresponding to the text to be classified.
[0117] As mentioned above, there are also multiple selectable implementation models for the text semantic extraction model described in this embodiment.
[0118] In one example, the fragment semantic extraction model can be constructed according to the BERT model or the ALBERT model, and several serially connected encoder layers in the BERT model or the ALBERT model can be used to implement several serially connected text semantic extraction layers in the text semantic extraction model.
[0119] Step 112: input the text semantic vector as an input parameter into a trained second classification model to classify the text to be classified, and determine the text category to which the text to be classified belongs.
[0120] After the text semantics are extracted, a second text classification will be performed, that is, the second classification model is used to classify the text to be classified based on the text semantic vector obtained by the second semantic extraction, thereby realizing text classification of long texts.
[0121] The text semantic vector obtained by the second semantic extraction in step 110 is input as an input parameter into the trained second classification model to obtain the classification result of the text to be classified, and the text category to which the text to be classified belongs is determined according to the classification result of the text to be classified.
[0122] The specific types and quantities of the text categories are consistent with the text categories in the first classification model.
[0123] There are also multiple optional implementation models for the second classification model. For example, the second classification model can be a classification model implemented based on a neural network model such as LSTM, CNN, etc.
[0124] According to the classification result output by the second classification model for the text semantic vector, the text category to which the text to be classified belongs is determined, which is determined by the specific second classification model adopted. For example, if a CNN-based classification model is adopted, the text category to which the text to be classified belongs can be determined based on the confidence level output by the second classification model that the text to be classified belongs to each text category.
[0125] It can be seen from the above description that in this specification, first, the fragment semantics of several text fragments obtained by segmenting the text to be classified are extracted respectively, and each text fragment is classified based on the fragment semantic vector of each extracted text fragment. If the classification result of a certain text fragment can meet the preset confidence requirement, it means that the fragment semantics of the text fragment can represent the semantics of the text to be classified. According to the classification result of the text fragment, the text category to which the text to be classified belongs can be determined. Since the short text has the above-mentioned characteristic that the semantics of the text can be represented by the fragment semantics of the text fragment, the text category to which the short text belongs can be determined in the above-mentioned way.
[0126] If the classification results of all text segments obtained by segmenting the text to be classified do not meet the preset confidence requirements, it means that the segment semantics of only one text segment is not sufficient to characterize the semantics of the text to be classified, and the text category to which the text to be classified belongs cannot be determined by the classification result of a certain text segment. Therefore, the text semantics are extracted by combining the segment semantic vectors of all text segments, and the text to be classified is classified based on the extracted text semantic vectors of the text to be classified, thereby determining the text category to which the long text with semantics distributed in multiple text segments belongs. The text classification scheme provided in this specification takes into account the text classification of long and short texts at the same time, thereby improving the efficiency and accuracy of text classification.
[0127] In this embodiment, the segment semantic extraction model, the text semantic extraction model, the first classification model and the second classification model are trained together as a whole in an end-to-end manner. Figure 5 , which is a schematic diagram of the overall structure of the text classification model shown in this embodiment.
[0128] In one example, a BERT model or an ALBERT model can be selected to construct an original segment semantic extraction model and a text semantic extraction model, and a classification model based on a CNN neural network model can be selected to respectively construct an original first classification model and a second classification model; using a text sample set pre-labeled with classification results, the original segment semantic extraction model, the text semantic extraction model, the first classification model, and the second classification model are jointly trained end-to-end in a supervised learning manner.
[0129] Among them, the BERT model and ALBERT model are pre-trained models with rich prior experience.
[0130] In this embodiment, text classification is implemented in combination with the BERT model or the ALBERT model. Training can be performed based on the pre-trained model in combination with the text classification scenario of this specification. Good use effects can be achieved through fine-tuning, with fewer iterations and high training efficiency.
[0131] In addition, based on the fact that the BERT model and the ALBERT model can accurately extract semantics, and the advantage of the pre-trained model having a large amount of prior experience, combining the BERT model or the ALBERT model in this embodiment can improve the accuracy of text classification.
[0132] Both the BERT model and the ALBERT model include an embedding layer and several encoder layers connected in series.
[0133] In actual implementation, the embedding layer of the BERT model can be used as the embedding layer in the original fragment semantic extraction model, and several serially connected encoder layers of the BERT model can be used as several serially connected fragment semantic extraction layers in the original fragment semantic extraction model.
[0134] Alternatively, the embedding layer of the ALBERT model can be used as the embedding layer in the original fragment semantic extraction model, and several serially connected encoder layers of the ALBERT model can be used as several serially connected fragment semantic extraction layers in the original fragment semantic extraction model.
[0135] Similarly, several serial encoder layers of the BERT model can be used as several serial text semantic extraction layers in the original text semantic extraction model;
[0136] Alternatively, several serially connected encoder layers of the ALBERT model may be used as several serially connected text semantic extraction layers in the original text semantic extraction model.
[0137] The fragment semantic extraction model and the text semantic extraction model can both be constructed using the BERT model or the ALBERT model, or one of them can be constructed using the BERT model and the other can be constructed using the ALBERT model.
[0138] The last encoder layer used in the fragment semantic extraction model and the first encoder layer used in the text semantic extraction model can be two adjacent layers or two non-adjacent layers in the original BERT model or ALBERT model; at the same time, the number of encoder layers used in the fragment semantic extraction model and the text semantic extraction model can be equal or unequal.
[0139] For example, you can initialize the 1st to 6th encoder layers of the BERT model to build the original fragment semantic extraction model, and initialize the 7th to 12th encoder layers of the BERT model to build the original text semantic extraction model; you can also initialize the 1st to 6th encoder layers of the BERT model to build the original fragment semantic extraction model, and initialize the 4th to 8th encoder layers of the BERT model to build the original text semantic extraction model.
[0140] like Figure 7 As shown, it is a flow chart of a method for jointly performing end-to-end training of the segment semantic extraction model, the text semantic extraction model, the first classification model and the second classification model, which includes the following specific steps:
[0141] Step 702, segment the text sample marked with the classification result to obtain a number of text segment samples.
[0142] Step 704: for each text segment sample, input the text segment sample as an input parameter into a segment semantic extraction model to perform semantic extraction on the text segment sample, and obtain a segment semantic vector corresponding to the text segment sample.
[0143] Step 706: Input the segment semantic vector as an input parameter into a first classification model to classify the text segment sample, and obtain a classification result of the text segment sample.
[0144] Step 708: Input the plurality of segment semantic vectors corresponding to the plurality of text segment samples into a text semantic extraction model to perform semantic extraction on the text samples to obtain text semantic vectors corresponding to the text samples.
[0145] Step 710: Input the text semantic vector into a second classification model to classify the text sample to obtain a classification result of the text sample.
[0146] The specific implementation process of segmenting text samples, extracting segment semantic vectors, classifying text segment samples, extracting text semantic vectors and classifying text samples described in the above steps 702 to 710 can be found in steps 102 to 112 described above and will not be repeated here.
[0147] Step 712, calculating the fusion loss according to the difference between the classification result of each text segment sample and the marked classification result, and the difference between the classification result of the text sample and the marked classification result.
[0148] In this embodiment, a supervised learning training method is adopted to perform end-to-end joint training on each model based on the labeled classification results of text samples.
[0149] When designing the loss function of the entire model, the differences between the classification results output by the first classification model for each text fragment sample and the labeled classification results, as well as the differences between the classification results output by the second classification model for the text sample and the labeled classification results, are integrated to enable the entire model to take into account the accuracy of the first and second text classifications, that is, to take into account the accuracy of long and short text classification at the same time.
[0150] For example, suppose the difference between the classification results output by the first classification model for each text segment sample and the labeled classification results is loss 11 、loss 12 、loss 13 、loss 14 and loss 15 , the difference between the classification result output by the second classification model for the text sample and the labeled classification result is loss 2 , in an example, the overall loss function of the model can be designed as:
[0151] Loss = w 0 *loss 2 +w 1 *loss 11 +w 2 *loss 12 +w 3 *loss 13 +w 4 *loss 14 +w 5 *loss 15
[0152] Among them, the weight value w 0 tow 5 It can be set manually by a technician or after computer learning, and the weight values can be equal or unequal.
[0153] In another example, the overall loss function of the model can also be designed as:
[0154] Loss = w 0 *loss 2 +w 1 *Maxpool(loss 11 ,loss 12 ,loss 13 ,loss 14 ,loss 15 )
[0155] Among them, Maxpool(loss 11 ,loss 12,loss 13 ,loss 14 ,loss 15 ) is used to pool the losses of several classification results output by the first classification model for several text fragment samples, and the designed fusion loss is the weighted addition of the loss of the first classification model and the loss of the second classification model after the pooling process; the weight value w 0 、w 1 It can be set manually by a technician or after computer learning, and the two can be equal or different.
[0156] During end-to-end training, based on the designed loss function, the fusion loss is calculated jointly according to the difference between the classification results of the current first classification model for each text fragment sample and the labeled classification results, and the difference between the classification results of the current second classification model for the text sample and the labeled classification results.
[0157] Step 714: Based on the fusion loss, determine whether each model has converged. If not, continue to train each model.
[0158] In this embodiment, based on the fusion loss calculated in step 712, it is determined whether each current model has reached a preset convergence condition, or whether the current number of iterations has reached a preset iteration threshold.
[0159] If not, the models are further trained using methods such as back propagation and stochastic gradient descent. For details, please refer to the relevant technology and will not be described in detail.
[0160] If so, it is determined that the model training is completed, and a trained segment semantic extraction model, a text semantic extraction model, a first classification model, and a second classification model are obtained.
[0161] In this embodiment, the fragment semantic extraction model, the text semantic extraction model, the first classification model and the second classification model are jointly trained in an end-to-end training manner, and the model training is performed with a fusion loss based on the difference between the classification results output by each of the first classification model and the second classification model and the labeled classification results. During the training process, the classification effects of long and short texts are taken into account, thereby improving the accuracy of the text classification model in this embodiment in classifying long and short texts.
[0162] Corresponding to the above-mentioned embodiments of the method for text classification, this specification also provides embodiments of an apparatus for text classification.
[0163] The embodiments of the text classification device provided in this specification can be applied to electronic devices. The device embodiments can be implemented by software, hardware, or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of the electronic device in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. From the hardware level, if Figure 8 The figure is a hardware structure diagram of an electronic device in which the text classification device provided in this specification is located, except Figure 8 In addition to the processor, memory, network interface, and non-volatile memory shown, the electronic device in which the device in the embodiment is located may also include other hardware according to the actual function of the electronic device, which will not be described in detail.
[0164] Fig. 9 It is a block diagram of a text classification device shown in an exemplary embodiment of this specification.
[0165] Please refer to Fig. 9 The text classification device 800 can be applied in the aforementioned Figure 8 The electronic device shown includes a text segmentation unit 810, a segment semantic extraction unit 820, a segment classification unit 830, a text extraction unit 840, and a text classification unit 850:
[0166] The text segmentation unit 810 segments the text to be classified into a plurality of text segments;
[0167] A segment semantic extraction unit 820, for each text segment, inputs the text segment as an input parameter into a trained segment semantic extraction model to perform semantic extraction on the text segment, and obtains a segment semantic vector corresponding to the text segment;
[0168] A segment classification unit 830, inputting the segment semantic vector as an input parameter into a trained first classification model to classify the text segment, and obtaining a classification result of the text segment;
[0169] If the classification result of any text segment meets the preset confidence requirement, then determine the text category to which the to-be-classified text belongs according to the classification result that meets the preset confidence threshold;
[0170] The text extraction unit 840, if the classification results of all the text segments do not meet the confidence requirement, inputs the multiple segment semantic vectors corresponding to the multiple text segments as input parameters into the trained text semantic extraction model to perform semantic extraction on the text to be classified, and obtains the text semantic vector corresponding to the text to be classified;
[0171] The text classification unit 850 inputs the text semantic vector as an input parameter into a trained second classification model to classify the text to be classified, and determines the text category to which the text to be classified belongs.
[0172] Optionally, the fragment semantic extraction model includes an embedding layer and a plurality of serially connected fragment semantic extraction layers;
[0173] The embedding layer is used to convert the text fragment into a corresponding number of embedding vectors;
[0174] Each of the segment semantic extraction layers is used to perform semantic extraction based on the vector output by the previous layer, and output the intermediate segment semantic vector extracted by the current layer;
[0175] The segment semantic extraction model is used to determine the segment semantic vector corresponding to the text segment based on the intermediate segment semantic vector output by the last segment semantic extraction layer.
[0176] Optionally, the text semantic extraction model includes a plurality of text semantic extraction layers connected in series;
[0177] Each of the text semantic extraction layers is used to perform semantic extraction based on the vector input into this layer, and output the intermediate text semantic vector obtained by the extraction of this layer;
[0178] The text semantic extraction model is used to determine the text semantic vector corresponding to the text to be classified based on the intermediate text semantic vector output by the last text semantic extraction layer.
[0179] Optionally, the fragment semantic extraction model is a transformer-based bidirectional encoder representation model BERT, or a lightweight transformer-based bidirectional encoder representation model ALBERT;
[0180] The text semantic extraction model is a transformer-based bidirectional encoder representation model BERT, or a lightweight transformer-based bidirectional encoder representation model ALBERT.
[0181] Optionally, the text segmentation unit 810 segments the text to be classified by using a sliding window according to a preset window length and the number of text segments;
[0182] If the number of text segments obtained after segmentation does not reach the number of text segments, the segmentation result is supplemented with the preset text segments;
[0183] If the number of text segments obtained after segmentation exceeds the number of text segments, the text segments exceeding the number of text segments are discarded.
[0184] Optionally, the segment semantic extraction model, the text semantic extraction model, the first classification model, and the second classification model are jointly trained in an end-to-end training manner, and the training process includes:
[0185] Segment the text samples marked with classification results to obtain several text segment samples;
[0186] For each text segment sample, input the text segment sample into a segment semantic extraction model to perform semantic extraction on the text segment sample, and obtain a segment semantic vector corresponding to the text segment sample;
[0187] Inputting the segment semantic vector into a first classification model to classify the text segment sample, and obtaining a classification result of the text segment sample;
[0188] Inputting the plurality of segment semantic vectors corresponding to the plurality of text segment samples into a text semantic extraction model to perform semantic extraction on the text samples, thereby obtaining text semantic vectors corresponding to the text samples;
[0189] Inputting the text semantic vector into a second classification model to classify the text sample, and obtaining a classification result of the text sample;
[0190] Calculate the fusion loss based on the difference between the classification results of each text fragment sample and the labeled classification results, and the difference between the classification results of the text sample and the labeled classification results;
[0191] Based on the fusion loss, it is determined whether each model has converged. If not, the training of each model continues.
[0192] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0193] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can refer to the partial description of the method embodiments. The device embodiments described above are only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this specification. Ordinary technicians in this field can understand and implement it without paying creative work.
[0194] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver, a game console, a tablet computer, a wearable device or a combination of any of these devices.
[0195] Corresponding to the embodiment of the above-mentioned method for text classification, this specification also provides an electronic device, which includes: a processor and a memory for storing machine executable instructions. The processor and the memory are usually connected to each other via an internal bus. In other possible implementations, the device may also include an external interface to enable communication with other devices or components.
[0196] In this embodiment, the processor is prompted to implement the steps of the method described in any of the above embodiments by reading and executing machine executable instructions corresponding to the logic of text classification stored in the memory.
[0197] Corresponding to the embodiments of the aforementioned text classification method, this specification also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any of the aforementioned embodiments are implemented.
[0198] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0199] The above description is only a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this specification should be included in the scope of protection of this specification.
Claims
1. A method for text classification, the method comprising: Segment the text to be classified to obtain several text fragments; For each text segment, the text segment is input as an input parameter into a trained segment semantic extraction model to perform semantic extraction on the text segment, and a segment semantic vector corresponding to the text segment is obtained; Inputting the segment semantic vector as an input parameter into a trained first classification model to classify the text segment, thereby obtaining a classification result of the text segment; If the classification result of any text segment meets the preset confidence requirement, then determine the text category to which the to-be-classified text belongs according to the classification result that meets the preset confidence requirement; If the classification results of all text segments do not meet the confidence requirement, the segment semantic vectors corresponding to the segment semantic vectors are input as input parameters into the trained text semantic extraction model to perform semantic extraction on the text to be classified, so as to obtain the text semantic vectors corresponding to the text to be classified; Inputting the text semantic vector as an input parameter into a trained second classification model to classify the text to be classified, and determining the text category to which the text to be classified belongs; Among them, the fragment semantic extraction model includes an embedding layer and several serially connected fragment semantic extraction layers; the embedding layer is used to convert the text fragment into several corresponding embedding vectors; each of the fragment semantic extraction layers is used to perform semantic extraction based on the vector output by the previous layer, and output the intermediate fragment semantic vector obtained by extraction in this layer; the fragment semantic extraction model is used to determine the fragment semantic vector corresponding to the text fragment based on the intermediate fragment semantic vector output by the last layer of the fragment semantic extraction layer; the text semantic extraction model includes several serially connected text semantic extraction layers; each of the text semantic extraction layers is used to perform semantic extraction based on the vector input to this layer, and output the intermediate text semantic vector obtained by extraction in this layer; the text semantic extraction model is used to determine the text semantic vector corresponding to the text to be classified based on the intermediate text semantic vector output by the last layer of the text semantic extraction layer.
2. The method according to claim 1, The fragment semantic extraction model is a transformer-based bidirectional encoder representation model BERT, or a lightweight transformer-based bidirectional encoder representation model ALBERT; The text semantic extraction model is a transformer-based bidirectional encoder representation model BERT, or a lightweight transformer-based bidirectional encoder representation model ALBERT.
3. According to the method of claim 1, segmenting the text to be classified comprises: According to the preset window length and the number of text segments, the text to be classified is segmented by a sliding window method; If the number of text segments obtained after segmentation does not reach the number of text segments, the segmentation result is supplemented with the preset text segments; If the number of text segments obtained after segmentation exceeds the number of text segments, the text segments exceeding the number of text segments are discarded in a semantic order.
4. According to the method of claim 1, the segment semantic extraction model, the text semantic extraction model, the first classification model and the second classification model are jointly trained in an end-to-end training manner, and the training process includes: Segment the text samples marked with classification results to obtain several text segment samples; For each text segment sample, input the text segment sample into a segment semantic extraction model to perform semantic extraction on the text segment sample, and obtain a segment semantic vector corresponding to the text segment sample; Inputting the segment semantic vector into a first classification model to classify the text segment sample, and obtaining a classification result of the text segment sample; Inputting the plurality of segment semantic vectors corresponding to the plurality of text segment samples into a text semantic extraction model to perform semantic extraction on the text samples, thereby obtaining text semantic vectors corresponding to the text samples; Inputting the text semantic vector into a second classification model to classify the text sample, and obtaining a classification result of the text sample; Calculate the fusion loss based on the difference between the classification results of each text fragment sample and the labeled classification results, and the difference between the classification results of the text sample and the labeled classification results; Based on the fusion loss, it is determined whether each model has converged. If not, the training of each model continues.
5. A device for text classification, the device comprising: A text segmentation unit segments the text to be classified into several text segments; A segment semantic extraction unit, for each text segment, inputs the text segment as an input parameter into a trained segment semantic extraction model to perform semantic extraction on the text segment, and obtains a segment semantic vector corresponding to the text segment; A segment classification unit, inputting the segment semantic vector as an input parameter into a trained first classification model to classify the text segment, and obtaining a classification result of the text segment; If the classification result of any text segment meets the preset confidence requirement, then determine the text category to which the to-be-classified text belongs according to the classification result that meets the preset confidence requirement; A text extraction unit, if the classification results of all text segments do not meet the confidence requirement, inputs the multiple segment semantic vectors corresponding to the multiple text segments as input parameters into a trained text semantic extraction model to perform semantic extraction on the text to be classified, and obtains the text semantic vector corresponding to the text to be classified; A text classification unit, inputting the text semantic vector as an input parameter into a trained second classification model to classify the text to be classified, and determining the text category to which the text to be classified belongs; Among them, the fragment semantic extraction model includes an embedding layer and several serially connected fragment semantic extraction layers; the embedding layer is used to convert the text fragment into several corresponding embedding vectors; each of the fragment semantic extraction layers is used to perform semantic extraction based on the vector output by the previous layer, and output the intermediate fragment semantic vector obtained by extraction in this layer; the fragment semantic extraction model is used to determine the fragment semantic vector corresponding to the text fragment based on the intermediate fragment semantic vector output by the last layer of the fragment semantic extraction layer; the text semantic extraction model includes several serially connected text semantic extraction layers; each of the text semantic extraction layers is used to perform semantic extraction based on the vector input to this layer, and output the intermediate text semantic vector obtained by extraction in this layer; the text semantic extraction model is used to determine the text semantic vector corresponding to the text to be classified based on the intermediate text semantic vector output by the last layer of the text semantic extraction layer.
6. The device according to claim 5, The fragment semantic extraction model is a transformer-based bidirectional encoder representation model BERT, or a lightweight transformer-based bidirectional encoder representation model ALBERT; The text semantic extraction model is a transformer-based bidirectional encoder representation model BERT, or a lightweight transformer-based bidirectional encoder representation model ALBERT.
7. The device according to claim 5, The text segmentation unit segments the text to be classified by using a sliding window method according to a preset window length and the number of text segments; If the number of text segments obtained after segmentation does not reach the number of text segments, the segmentation result is supplemented with the preset text segments; If the number of text segments obtained after segmentation exceeds the number of text segments, the text segments exceeding the number of text segments are discarded in a semantic order.
8. The device according to claim 5, wherein the segment semantic extraction model, the text semantic extraction model, the first classification model and the second classification model are jointly trained in an end-to-end training manner, and the training process includes: Segment the text samples marked with classification results to obtain several text segment samples; For each text segment sample, input the text segment sample into a segment semantic extraction model to perform semantic extraction on the text segment sample, and obtain a segment semantic vector corresponding to the text segment sample; Inputting the segment semantic vector into a first classification model to classify the text segment sample, and obtaining a classification result of the text segment sample; Inputting the plurality of segment semantic vectors corresponding to the plurality of text segment samples into a text semantic extraction model to perform semantic extraction on the text samples, thereby obtaining text semantic vectors corresponding to the text samples; Inputting the text semantic vector into a second classification model to classify the text sample, and obtaining a classification result of the text sample; Calculate the fusion loss based on the difference between the classification results of each text fragment sample and the labeled classification results, and the difference between the classification results of the text sample and the labeled classification results; Based on the fusion loss, it is determined whether each model has converged. If not, the training of each model continues.
9. An electronic device, comprising: processor; memory for storing machine-executable instructions; Wherein, the processor implements the steps of the method as described in any one of claims 1 to 4 by reading and executing machine executable instructions corresponding to the logic of text classification stored in the memory.
10. A computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Neural network sentiment analysis method in combination with user and product information
CN106383815A
Text enhancement system
EP2391105A1