Text classification method and device, equipment, medium and product
By using a BERT encoder and a bidirectional recurrent neural network unit to segment text information and combining multiple classifiers, the problem of insufficient accuracy in existing text classification schemes is solved, achieving more efficient and accurate text information classification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2026-03-13
AI Technical Summary
Existing text classification schemes are insufficient in accuracy, especially when identifying the content type and evaluation type of text information, as they are computationally expensive and lack robustness.
The system employs a bidirectional language representation transformation model, BERT encoder and attention layer, combined with a gated bidirectional recurrent neural network unit. By segmenting text information and using multiple classifiers for feature representation and classification, multiple predicted probabilities are output to determine the final classification result.
It improves the accuracy of text information content type and evaluation type classification, reduces the amount of computation, improves computational efficiency, and ensures the accuracy of results.
Smart Images

Figure CN121658656A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more particularly to text classification methods, apparatus, devices, storage media, and products. Background Technology
[0002] With the rapid development of information technology and the popularization of mobile internet, users can publish various text information through internet platforms, such as comments on objects in different fields, such as e-commerce, catering, cultural tourism, or media, and publish product reviews, restaurant reviews, cultural tourism reviews, or media reviews.
[0003] Text information typically contains rich content, and in some application scenarios, content recognition is required to classify the text information. However, current classification schemes lack accuracy and need improvement. Summary of the Invention
[0004] This disclosure provides text classification methods, apparatus, devices, storage media, and products that can optimize existing text classification schemes.
[0005] In a first aspect, embodiments of this disclosure provide a text classification method, including:
[0006] Obtain the target text information to be processed;
[0007] Based on the target text information, a first category recognition result and a second category recognition result corresponding to the target text information are determined; wherein, the first category recognition result is determined based on the target text information, and the first category recognition result represents the content type information and evaluation type information corresponding to the target text information; the second category recognition result is determined based on the content category features and evaluation category features of the target text information, and the second category recognition result is used to represent the probability of the content type corresponding to the target text information and the evaluation type corresponding to the content type.
[0008] Based on the first category recognition result and the second category recognition result corresponding to the target text information, the content type of the target text information and the evaluation type corresponding to the content type are determined.
[0009] Secondly, embodiments of this disclosure also provide a text classification apparatus, including:
[0010] The text information acquisition module is used to acquire the target text information to be processed;
[0011] The recognition result determination module is used to determine a first category recognition result and a second category recognition result corresponding to the target text information based on the target text information; wherein, the first category recognition result is determined based on the target text information, and the first category recognition result represents the content type information and evaluation type information corresponding to the target text information; the second category recognition result is determined based on the content category features and evaluation category features of the target text information, and the second category recognition result is used to represent the probability of the content type corresponding to the target text information and the evaluation type corresponding to the content type.
[0012] The type determination module is used to determine the content type of the target text information and the evaluation type corresponding to the content type based on the first category recognition result and the second category recognition result corresponding to the target text information.
[0013] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0014] One or more processors;
[0015] Storage device for storing one or more programs.
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the methods provided in the embodiments of this disclosure.
[0017] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the methods provided in embodiments of this disclosure.
[0018] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the method provided in embodiments of this disclosure.
[0019] The text classification scheme provided in this disclosure acquires target text information to be processed. Based on the target text information, it determines a first category recognition result and a second category recognition result corresponding to the target text information. The first category recognition result is determined based on the target text information and represents the content type and evaluation type information corresponding to the target text information. The second category recognition result is determined based on the content type and evaluation type features of the target text information and represents the probability of the content type and the corresponding evaluation type. Based on the first and second category recognition results corresponding to the target text information, the content type and the corresponding evaluation type of the target text information are determined. By adopting the above technical solution, category recognition results for content type and evaluation type identification of target text information are obtained from different perspectives. Combining the two category recognition results to determine the final classification result can effectively improve the accuracy of content type and evaluation type classification for text information. Attached Figure Description
[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0021] Figure 1 A flowchart illustrating a text classification method provided in an embodiment of this disclosure;
[0022] Figure 2 This is a flowchart illustrating another text classification method provided in an embodiment of the present disclosure;
[0023] Figure 3 A flowchart illustrating yet another text classification method provided in this disclosure embodiment;
[0024] Figure 4 This is a schematic diagram of a model training and inference process provided in an embodiment of the present disclosure;
[0025] Figure 5 This is a schematic diagram of a data augmentation process provided in an embodiment of the present disclosure;
[0026] Figure 6 This is a schematic diagram of an encoding process provided in an embodiment of the present disclosure;
[0027] Figure 7 This is a schematic diagram of a classification and detection process provided in an embodiment of the present disclosure;
[0028] Figure 8This is a schematic diagram of the structure of a text classification device provided in an embodiment of the present disclosure;
[0029] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0031] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0032] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0034] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0035] Figure 1 This is a flowchart illustrating a text classification method provided in an embodiment of the present disclosure. This embodiment is applicable to situations where text information such as comment information is classified by content type and evaluation type. The method can be executed by a model training device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, such as a personal computer (PC) or a server.
[0036] like Figure 1As shown, the method includes:
[0037] Step 101: Obtain the target text information to be processed.
[0038] In this embodiment, the text information can be comments on objects within a preset domain, such as comment text. The text information can be published by users through an internet platform. The preset domain may include, for example, e-commerce, catering, cultural tourism, or media. Specifically, the text information can be product reviews, restaurant reviews, cultural tourism reviews, or media work reviews. For ease of explanation, the following description uses cultural tourism reviews (such as tourist attractions) as an example.
[0039] For example, the target text information can be any text information that needs to be categorized. The purpose of categorization is to determine the content type of the target text information and the corresponding evaluation type. Specifically, it can be to determine which content types the target text information involves and the corresponding evaluation types for those content types.
[0040] For example, a set of preset content types and a set of preset evaluation types are set for text information in a preset domain. The preset content types in the preset content type set can be categories of text aspects (such as comments) that the text information may involve, also known as aspect categories. The preset evaluation types in the preset evaluation type set can be categories used to represent the evaluation tendency or evaluation polarity of a single preset content type, such as positive, neutral, and negative, or good, medium, and poor.
[0041] For example, in the cultural and tourism sector, the content type definition can be based on the national A-level scenic spot evaluation indicators and actual text data to define the content type and evaluation type used for analyzing cultural and tourism reviews.
[0042] For example, based on the national A-level scenic area evaluation indicators and actual text data, nine preset content types are defined: "tourism transportation", "scenic area tour", "tourism safety", "scenic area hygiene", "service quality", "tourism shopping", "business planning", "resources and environment" and "overall evaluation". The set of preset content types is obtained, and the specific explanation of each preset content type is shown in Table 1.
[0043] Table 1 Content Type Description
[0044]
[0045]
[0046] For example, let the preset content type set be... |C| represents the set size, which is 9. i It is one of the nine preset content types mentioned above.
[0047] Step 102: Based on the target text information, determine the first category recognition result and the second category recognition result corresponding to the target text information; wherein, the first category recognition result is determined based on the target text information, and the first category recognition result represents the content type information and evaluation type information corresponding to the target text information; the second category recognition result is determined based on the content category features and evaluation category features of the target text information, and the second category recognition result is used to represent the probability of the content type and the evaluation type corresponding to the content type of the target text information.
[0048] For example, a machine learning model can be used to determine the first category recognition result and the second category recognition result corresponding to the target text information based on the target text information.
[0049] The first category of identification results represents the content type information and evaluation type information corresponding to the target text information. The content type information includes whether the target text information involves a certain preset content type, and the evaluation type information includes whether the target text information involves a certain preset evaluation type. The first category of identification results can represent the target text information as a whole whether it involves a certain preset content type and a certain preset evaluation type involving that preset content type. For example, it can be used to represent whether the target cultural tourism review information involves a positive evaluation type of tourism transportation.
[0050] The second category identification result is used to characterize the probability of the content type corresponding to the target text information and the evaluation type corresponding to the content type. Specifically, it can be the probability of each evaluation type of the target text information involving a certain preset content type when the target text information involves a certain preset content type. For example, when the target cultural and tourism review information involves tourism and transportation, it can be the probability of positive evaluation type, neutral positive evaluation type and negative positive evaluation type respectively.
[0051] Step 103: Based on the first category recognition result and the second category recognition result corresponding to the target text information, determine the content type of the target text information and the evaluation type corresponding to the content type.
[0052] For example, the content type of the target text information and the corresponding evaluation type can be determined based on the consistency between the first category identification result and the second category identification result. For instance, if both the first and second category identification results indicate that the target text information involves a certain preset content type, then the content type of the target text information can be determined to include that preset content type. If only one category identification result indicates that the target text information involves a certain preset content type, then the content type of the target text information can be determined to exclude that preset content type, thus improving accuracy. After determining that the content type of the target text information includes a certain preset content type, the specific preset evaluation type involved in that preset content type can be determined by comprehensively considering the first and second category identification results. For example, if the first and second category identification results indicate the same preset evaluation type involved in that preset content type, then that same preset evaluation type is determined as the evaluation type corresponding to the content type of the target text information. If no identical preset evaluation type exists, it indicates that the evaluation types are inconsistent; for example, one is a positive evaluation type and the other is a negative evaluation type, creating a contradiction. In this case, the confidence levels of the two category identification results can be determined, and the category identification result with the higher confidence level is taken as the standard.
[0053] The text classification method provided in this disclosure acquires target text information to be processed, and determines a first category recognition result and a second category recognition result corresponding to the target text information based on the target text information. The first category recognition result is determined based on the target text information and represents the content type and evaluation type information corresponding to the target text information. The second category recognition result is determined based on the content type and evaluation type features of the target text information and represents the probability of the content type and the corresponding evaluation type of the content type. Based on the first and second category recognition results corresponding to the target text information, the content type and the corresponding evaluation type of the target text information are determined. By adopting the above technical solution, category recognition results for content type and evaluation type identification of target text information are obtained from different perspectives. Combining the two category recognition results to determine the final classification result can effectively improve the accuracy of content type and evaluation type classification for text information.
[0054] Figure 2 This is a flowchart illustrating another text classification method provided in this disclosure embodiment. This disclosure embodiment optimizes upon the various optional solutions described above, such as... Figure 2 As shown, the method includes:
[0055] Step 201: Obtain the target text information to be processed.
[0056] Step 202: For each preset content type in the preset content type set, input the target text information, preset auxiliary statements, and preset evaluation type set into the target text representation and classification model to obtain the prediction probability output by the target text representation and classification model. Based on the prediction probability output by the target text representation and classification model, determine the first category recognition result and the second category recognition result of the target text information for the current preset content type.
[0057] The preset auxiliary statements are constructed based on the current preset content type and preset template; the content category features and evaluation category features of the target text information are extracted by the target text representation and classification model from the target text information, the preset auxiliary statements, and the preset evaluation type set.
[0058] In this embodiment of the disclosure, the target text representation and classification model can be a model obtained by training a preset text representation and classification model using a preset training sample set. The preset training samples in the preset training sample set include sample text information, and each preset training sample is associated with a sample label. The sample label includes the sample content type involved in the sample text information and the sample evaluation type corresponding to the sample content type involved in the sample text information. The sample content type is a preset content type in the preset content type set, and the sample evaluation type is a preset evaluation type in the preset evaluation type set.
[0059] For example, the sample text information is the text information used for model training. It can be real text information or text information obtained through data augmentation or other methods. Different sample text information can serve as different preset training samples. The preset training sample set includes multiple preset training samples, and the specific number is not limited. After obtaining the sample text information, sample labels can be associated with the sample text information through manual annotation or machine annotation, that is, sample labels can be associated with the preset training samples. Based on the preset content type set and the preset evaluation type set, the sample text information is labeled. If the sample text information involves the preset content type c... i If the corresponding preset evaluation type is specified, then it should be marked; otherwise, it should be marked as "None" or not marked. i Related information; if there is a conflict in the evaluation type, this sample text information can be deleted. Specific annotation examples are shown in Table 2.
[0060] Table 2 Annotation Examples
[0061]
[0062]
[0063] To facilitate understanding of the technical solutions of the embodiments of this disclosure, the relevant technologies are described below. Currently, the main schemes for identifying content type and evaluation type of text information include Cartesian product-based schemes and hierarchical classification (pipeline) schemes. The Cartesian product-based scheme enumerates all combinations of content type and evaluation type through Cartesian product operations, then inputs each combination along with the text information into an encoder for encoding, and finally uses a binary classifier to determine whether the combination is valid. However, this scheme greatly expands the amount of data during training and inference, significantly increasing computational costs, resulting in low efficiency. Furthermore, it may generate multiple different evaluation types for a single content type, leading to inaccurate calculation results. The hierarchical classification scheme performs content type judgment first, followed by evaluation type judgment. If a lower level does not detect a content type, subsequent judgments will not be performed, resulting in a one-way dependency from content type to evaluation type. If the content type detection is incorrect, the subsequent correct evaluation type judgment is meaningless, leading to inaccurate calculation results and insufficient robustness.
[0064] In this embodiment of the disclosure, for each preset content type, the target text information, preset auxiliary statements, and preset evaluation type set can be input into the target text representation and classification model, and the two category recognition results corresponding to the current preset content type can be quickly and accurately determined according to the model output, thereby obtaining an accurate classification result.
[0065] In some embodiments, the target text representation and classification model includes a bidirectional language representation transformation model (BERT) encoder and an attention layer; the BERT encoder is used to encode the preset auxiliary statement into a preset auxiliary statement representation, encode the target text information into an initial text information representation, and encode each preset evaluation type in the preset evaluation type set into a preset evaluation type representation; the attention layer is used to determine the content type representation corresponding to the preset auxiliary statement representation, the text information representation corresponding to the initial text information representation, and the evaluation type representation corresponding to the preset evaluation type representation; the content type representation, the text information representation, and the evaluation type representation are used to determine the prediction probability.
[0066] In this embodiment, the target text representation and classification model includes a Bidirectional Encoder Representations from Transformers (BERT) encoder and an attention layer. The input data of the target text representation and classification model includes target text information, preset auxiliary statements, and a preset set of evaluation types. The input data is sequentially encoded by the BERT encoder and the attention layer to obtain content type representation, text information representation, and evaluation type representation. This enables the target text representation and classification model to learn relevant information about text information, content type, and evaluation type based on richer, more granular encoded representations, and simultaneously outputs prediction information about preset content type and preset evaluation type. Compared with the aforementioned existing solutions, this reduces computational load, improves computational efficiency, and ensures the accuracy of the calculation results.
[0067] For example, the target text representation and classification model in this disclosure aims to simultaneously detect the preset content type involved in the text information and the preset evaluation type of the preset content type involved, that is, to extract the aspect evaluation pair (c,s) corresponding to the text information, where c represents the preset content type and s represents the preset evaluation type. To simplify the task, given a preset content type c (a single preset content type, for example, the current single preset content type can be denoted as the target preset content type), it can be determined whether the text information involves the given preset content type c and its corresponding preset evaluation type. Based on this simplification, for target text information, each preset content type c in the preset aspect category set C can be detected sequentially. i By determining whether the information is involved in the target text and identifying the corresponding preset evaluation type, the final classification result can be obtained.
[0068] For example, a preset auxiliary statement can be understood as a statement used to assist the target text representation and classification model in learning target text information. The preset template includes preset text and target characters. When constructing the preset auxiliary statement, the target characters can be replaced with the current preset content type. For example, the preset template could be "Evaluation type of content type c?", where "content type" and "evaluation type?" are preset text, and "c" is the target character. As in the example above, if the current preset content type (target preset content type) is "tourism and transportation", then the constructed preset auxiliary statement could be "Evaluation type of content type tourism and transportation?".
[0069] For example, after tokenizing the target text information, preset auxiliary statements, and each preset evaluation type in the preset evaluation type set, the input data for the target text representation and classification model is obtained, and the input data is fed into the target text representation and classification model. Specifically, the tokenization process can be as follows: add a classification token (CLS) to the beginning of each preset evaluation type in the target text information, preset auxiliary statements, and preset evaluation type set, and add a separator token (SEP) to the end of each preset evaluation type. The classification token is usually the first token in the input sequence.
[0070] For example, the input data of the target text representation and classification model is first encoded by the BERT encoder. The hidden state representation corresponding to the marker symbol [CLS] added to the beginning of the target text information by the BERT encoder is determined as the initial representation of the text information. The hidden state representation corresponding to the marker symbol [CLS] added to the beginning of the preset auxiliary statement by the BERT encoder is determined as the preset auxiliary statement representation. The hidden state representation corresponding to the marker symbol [CLS] added to the beginning of each preset evaluation type by the BERT encoder is determined as the preset evaluation type representation corresponding to each preset evaluation type.
[0071] For example, the preset auxiliary statement representation, the initial text information representation, and the evaluation type representation can be directly or indirectly input into the attention layer. The attention layer extracts the important information from them respectively, and then obtains the content type representation corresponding to the preset auxiliary statement representation, the text information representation corresponding to the initial text information representation, and the evaluation type representation corresponding to the preset evaluation type representation.
[0072] For example, the target text representation and classification model may also include a classifier, into which the content type representation, text information representation and evaluation type representation are input, and the probability value output by the classifier is used as the predicted probability output by the target representation and classification model.
[0073] Step 203: Based on the first category recognition result and the second category recognition result corresponding to the target text information, determine the content type of the target text information and the evaluation type corresponding to the content type.
[0074] For example, after obtaining the first category recognition result and the second category recognition result corresponding to each preset content type of the target text information, the classification result corresponding to each preset content type is determined, and then the content type of the target text information and the evaluation type corresponding to the content type are determined according to the classification result corresponding to each preset content type in the preset content type set.
[0075] The text classification method provided in this disclosure uses a target text representation and classification model to perform feature representation and classification on target text information, preset auxiliary statements constructed based on preset content types and preset templates, and preset evaluation type sets, and outputs preset probabilities to determine the recognition results of two categories, which can further improve the efficiency and accuracy of text classification.
[0076] In some embodiments, the preset text representation and classification model includes multiple classifiers, each corresponding to a classification task. After obtaining the content type representation, text information representation, and evaluation type representation, a multi-task approach can be used to output the corresponding predicted probabilities, so that the final classification result of the text information can be obtained based on the inference results of multiple classification tasks. The input to each classifier can be any one or a joint representation of at least two of the content type representation, text information representation, and evaluation type representation; no specific limitation is imposed.
[0077] In some embodiments, the target text representation and classification model further includes a first classifier, a second classifier, and a third classifier; wherein, the first classifier is used to determine a first predicted probability that the target text information involves a current preset content type based on the content type representation; the second classifier is used to determine several second predicted probabilities of a preset evaluation type corresponding to the current preset content type involved in the target text information based on the joint representation of the content type representation and the text information representation, wherein each second predicted probability corresponds to a different preset evaluation type; the third classifier is used to determine several third predicted probabilities that a preset evaluation type in the preset evaluation type set and the current preset content type are both involved in the target text information based on the joint representation of the content type representation and different evaluation type representations, wherein each third predicted probability corresponds to a different preset evaluation type; the predicted probabilities output by the target text representation and classification model include the first predicted probability, the second predicted probability, and the third predicted probability. The step of determining the first category recognition result and the second category recognition result corresponding to the current preset content type of the target text information based on the predicted probability output by the target text representation and classification model includes: determining the first category recognition result corresponding to the current preset content type of the target text information based on the third predicted probability; and determining the second category recognition result corresponding to the current preset content type of the target text information based on the first predicted probability and the second predicted probability.
[0078] Figure 3This is a flowchart illustrating another text classification method provided in this disclosure. This disclosure optimizes the various optional schemes in the above embodiments. The target text representation and classification model also includes a bidirectional recurrent neural network unit based on a gating mechanism, enabling the target text representation and classification model to better handle long texts. Figure 3 As shown, the method includes:
[0079] Step 301: Obtain the target text information to be processed.
[0080] Step 302: Segment the target text information to obtain multiple text information fragments.
[0081] For example, the target text information is segmented according to a target length value to obtain multiple text information fragments. Optionally, the target length value is a random value within a preset length range.
[0082] Step 303: For each preset content type in the preset content type set, input multiple text information fragments, preset auxiliary statements, and preset evaluation type sets into the target text representation and classification model to obtain the first prediction probability, second prediction probability, and third prediction probability output by the target text representation and classification model.
[0083] The target text representation and classification model includes a BERT encoder, a gating mechanism-based bidirectional recurrent neural network unit, an attention layer, a first classifier, a second classifier, and a third classifier.
[0084] For example, for each preset content type in the preset content type set, the input data corresponding to the target text information for the target text representation and classification model is determined, and the input data is input into the target text representation and classification model to obtain the corresponding output. As exemplified above, taking cultural and tourism text information as an example, if the preset content type set includes 9 preset content types, then 9 sets of input data need to be determined, and each set of input data is sequentially input into the target text representation and classification model to obtain 9 outputs from the target text representation and classification model.
[0085] For example, the BERT encoder can obtain the contextual representation of text, but as the text length increases, this contextual representation gradually weakens. In other words, there is a distance limit to the contextual interaction, which means that the BERT encoding may not be able to fully capture the complete contextual information, resulting in incomplete information acquisition and thus certain limitations. Furthermore, text information varies in length, with some text being very long, making it difficult for the BERT encoder to obtain an accurate representation. In this embodiment, the text information is segmented to effectively control the length of the text information input to the BERT encoder, enabling the BERT encoder to encode the text information in segments. Subsequently, a bidirectional recurrent neural network unit based on a gating mechanism is used to obtain the contextual interaction representation corresponding to the initial representation of the segmented text information output by the BERT encoder. This allows the attention layer to determine the text information representation corresponding to the target text information based on the contextual interaction representations corresponding to multiple initial segment representations, serving as the overall representation of the target text information.
[0086] For example, a bidirectional recurrent neural network unit based on a gating mechanism may include a bidirectional gated recurrent unit (Bi-GRU) or a bidirectional long short-term memory (Bi-LSTM) network unit.
[0087] Step 304: Determine the first category recognition result of the target text information for the current preset content type based on the third prediction probability, and determine the second category recognition result of the target text information for the current preset content type based on the first prediction probability and the second prediction probability.
[0088] For example, during the inference process, for a single preset content type c, the first prediction probability, the second prediction probability, and the third prediction probability are obtained through three classifiers, denoted as follows: and First, the second category recognition result can be determined using the following expression: Here, I() is an indicator function with a value of 1 indicating that the current preset content type c is involved, and other values indicate that it is not involved. That is, if the first prediction probability is greater than 0.5, then the function value of I() is 1. return The coordinates of the maximum value are 1, 2, and 3, representing positive, neutral, and negative directions, respectively. Next, the first category recognition result can be determined using the following expression: The `max()` function returns the maximum value, and `I()` is an indicator function whose value of 1 indicates that the current preset content type `c` is involved. That is, if the maximum value among the three probability values in the third prediction probability is greater than 0.5, then the value of `I()` is 1. return The coordinates of the maximum value are 1, 2, and 3, which represent positive, neutral, and negative directions, respectively.
[0089] Step 305: For each preset content type in the preset content type set, if both the first category recognition result and the second category recognition result corresponding to the current preset content type include target text information involving the current preset content type, then it is determined that the target text information involves the current preset content type, and the evaluation type contained in the one with the greater confidence in the first category recognition result and the second category recognition result is determined as the evaluation type corresponding to the current preset content type involved by the target text information, thus obtaining the classification result corresponding to the current preset content type.
[0090] For example, for two category recognition results, if both preset content types c are involved, the confidence levels of the first category recognition result and the second category recognition result can be determined separately, with the confidence level of the first category recognition result being... The confidence level of the second category identification result is The category identification result with higher confidence (either the first category identification result or the second category identification result) is selected as the final result. If one category identification result indicates that the target text information does not involve the preset content type c, the category identification result that involves the preset content type c is selected as the final result. If both category identification results indicate that the target text information does not involve the preset content type c, then the target text information does not involve the preset content type c.
[0091] For example, assuming a preset content type "tourism and transportation", the first predicted probability is obtained. The second preset probability is 0.91. The third preset probability is [0.005, 0.01, 0.985]. If the probability is [0.03, 0.05, 0.92], then the first category identification result is (1, 3), which involves tourism transportation and the preset evaluation type is negative. The second category identification result is also (1, 3), which involves tourism transportation and the preset evaluation type is negative. The two category identification results are consistent. If the obtained first predicted probability... The second preset probability is 0.91. The third preset probability is [0.005, 0.01, 0.985]. If the values are [0.03, 0.92, 0.05], then the first category identification result is (1, 3), which involves tourism transportation, and the preset evaluation type is negative. The second category identification result is (1, 2), which also involves tourism transportation, and the preset evaluation type is neutral. Both category identification results involve tourism transportation, but the corresponding preset evaluation types are different. The confidence level of the first category identification result is the product of 0.91 and 0.985, which is approximately equal to 0.897, while the confidence level of the second category identification result is 0.92. Therefore, the second category identification result shall prevail.
[0092] Step 306: Determine the content type of the target text information and the evaluation type corresponding to the content type based on the classification results corresponding to each preset content type in the preset content type set.
[0093] The text classification method provided in this embodiment of the present disclosure segments the target text information into multiple text information fragments. These fragments, along with preset auxiliary statements and a preset set of evaluation types, are input into a target text representation and classification model to obtain three predicted probabilities output by the model. The target text representation and classification model includes a BERT encoder, a bidirectional GRU, an attention layer, and three classifiers. The multiple text information fragments are sequentially encoded by the BERT encoder, the bidirectional GRU, and the attention layer to obtain a text information representation that fully captures complete contextual information. The preset auxiliary statements and the preset set of evaluation types are sequentially encoded by the BERT encoder and the attention layer to obtain content type representation and evaluation type representation. By combining different representations, multiple classifiers are used to output multiple predicted probabilities in a multi-task manner. Based on these multiple predicted probabilities, two category recognition results are determined. Combining these two category recognition results allows for more accurate classification of the text information.
[0094] In some embodiments, the target text representation and classification model is trained in the following manner: a preset training sample set is obtained; for each preset content type in the preset content type set, the sample text information, preset auxiliary statements, and the preset evaluation type set in the preset training sample set are input into the preset text representation and classification model to obtain the sample prediction probability output by the preset text representation and classification model; a target loss relationship is determined based on the sample prediction probability output by the preset text representation and classification model and the sample label; and the preset text representation and classification model is trained based on the target loss relationship to obtain the target text representation and classification model.
[0095] For example, Figure 4 This is a schematic diagram of a model training and inference process provided in an embodiment of the present disclosure, such as... Figure 4As shown, the content type and evaluation type are defined first. As mentioned above, taking cultural and tourism texts as an example, nine preset content types are defined to obtain a preset content type set. Then, the preset evaluation type set is determined. For example, “pos”, “neg”, and “neu” can represent three preset evaluation types: “positive”, “negative”, and “neutral”, respectively. Then, data annotation is performed.
[0096] For example, in the cultural and tourism field, there is a limited amount of text available for data annotation. If the model is trained based on a training set with insufficient or imbalanced samples, the model may be inaccurate. Therefore, data augmentation can be used to expand the training set.
[0097] Optionally, the preset training sample set is determined in the following way: Input data is constructed based on the text information of the sample to be enhanced, wherein the input data includes masked text information corresponding to the text information of the sample to be enhanced, a preset content type involved in the text information of the sample to be enhanced, and a preset evaluation type corresponding to the preset content type involved in the text information of the sample to be enhanced; the input data is input into a target pre-trained language model to obtain the output of the target pre-trained language model; the enhanced sample text information is determined based on the output of the target pre-trained language model; and the preset training sample set is determined based on the text information of the sample to be enhanced and the enhanced sample text information. The preset content type of the text information to be enhanced is the same as the preset content type of the text information to be enhanced, and the preset evaluation type corresponding to the preset content type of the text information to be enhanced is the same as the preset evaluation type corresponding to the preset content type of the text information to be enhanced. The mask text information corresponding to the text information to be enhanced is obtained by replacing each character in the text information to be enhanced with a preset mask symbol at a preset probability. The preset probability is positively correlated with the length of the text information to be enhanced. Therefore, the constructed input data includes not only mask text information but also preset content type and preset evaluation type, i.e., labels are introduced into the input, ensuring sufficient label information in the input. This allows the target pre-trained language model to better learn the features of the text information to be enhanced, thereby improving the accuracy of the generated enhanced text information.
[0098] The text information to be enhanced can be real text information. The preset content type involved in the text information to be enhanced and the preset evaluation type corresponding to the preset content type involved in the text information to be enhanced are determined by data annotation. Figure 5This is a schematic diagram of a data augmentation process provided in an embodiment of the present disclosure. The preset content type involved in the sample text information to be augmented, and the preset evaluation type corresponding to the preset content type (the aspect evaluation information to be augmented), are converted into natural language text according to a preset conversion template "evaluation type s of content type c is". All conversion results are then concatenated using [SEP] to obtain sequence X. pair For example, X is obtained from the data in serial number 1 of Table 2. pair The evaluation type for "Content Type Tourism Transportation is negative [SEP]", "Content Type Service Quality is negative [SEP]", "Content Type Business Management is negative [SEP]", and "Content Type Overall Evaluation is negative".
[0099] For example, for each character of the sample text information X (the relevant text that needs to be enhanced) with p mask The probability of (X) is replaced by the symbol [M] to obtain the sequence X. mask (Mask text information). p mask The specific representation of (X) is as follows:
[0100]
[0101] Where a, b, c, d, and e are preset values, and the specific values can be set according to actual needs. Optional values include a = 0.3, b = 0.6, c = 64, d = 192, and e = 256. |X| represents the length of the sample text information X to be enhanced. For example, the text length of serial number 1 in Table 1 is 71, then p mask The value of (X) is 0.311, and the masked text information is "Just entered the area [M][M] felt [M][M], I booked the network [M][M], [M] took [M], and was preparing to enter the network [M] entrance [M][M] as [M][M] said I [M] ticket [M][M] to enter. It seems that [M] does not encourage following the rules. I would like to remind everyone [M] to exchange the network ticket [M][M], and you can [M][M][M] enter."
[0102] For example, the sequence X is concatenated using the [SEP] symbol. pair and X mask Get input X data (Input data) is fed into the target pre-trained language model (PLMs), and the output is denoted as Y. data Y can be data The data is identified as sample text information to be enhanced. For example, the data in sequence 1 of Table 1 is used as the sample text information to be enhanced, resulting in X. dataThe evaluation type for the content type "Tourism Transportation" is negative [SEP], the evaluation type for the content type "Service Quality" is negative [SEP], the evaluation type for the content type "Operation Management" is negative [SEP], and the evaluation type for the content type "Overall Evaluation" is negative [SEP]. [M] Entering the area [M], feeling stuck in traffic [M][M][M][M][M], online ticket [M], I got my ticket and was about to use the online ticket [M][M][M], but the staff [M][M] said [M] exchanging tickets wouldn't let me in [M]. It seems that [M] here encourages following the rules, reminding everyone [M][M][M][M][M][M] to exchange [M] so they can enter directly. This inputs into PLMs to get the new text Y. data The text reads, "It's incredibly congested as soon as I entered. I just booked an online ticket, and I also bought a physical ticket and received it. I was about to enter through the online ticket entrance, but the staff told me to switch to a physical ticket. Seeing that everyone is not following the rules, they reminded everyone not to buy online tickets, but to buy physical tickets directly." Because the newly generated text is inconsistent with the original text (the relevant text that needs enhancement), it can be used as a piece of augmented data, i.e., augmented sample text information, and its corresponding aspect evaluation information is consistent with the aspect evaluation information that needs enhancement.
[0103] For example, after generating a certain number of enhanced sample text information in the above manner, the enhanced sample text information and the sample text information to be enhanced are used together as preset training samples in the preset training sample set to realize sample expansion.
[0104] For example, the labeled text information (text in the cultural tourism review dataset) can be determined as the initial text information, and then the target training sample set can be determined. The target training sample set includes multiple initial text information. The pre-trained language model is fine-tuned using the target training sample set to obtain the target pre-trained language model. The initial text information in the target training sample set can be used to construct sample input data in the manner described above for the text information to be enhanced. That is, the sample input data includes the masked text information corresponding to the initial text information, the preset content type involved in the initial text information, and the preset evaluation type corresponding to the preset content type involved in the initial text information. Using the initial text information as the target sequence, an input-output pair (sample input data, target sequence) is obtained, denoted as (X). data Y data The encoder-decoder model is used to perform sequence-to-sequence (Seq2Seq) modeling on the input-output pairs. First, the encoder models the sample input data as a context representation H. enc Then the decoder will generate the target sequence Y in an autoregressive manner. data At the t-th time step, the decoder will determine the context features H. encCalculate the decoder hidden state using the previously decoded tags.
[0105]
[0106] Among them, y [1:t-1] For the previous decoding markers; next utilize Calculate the label y t Conditional probability:
[0107]
[0108] Where W is the weight matrix, the model parameters are initialized using the pre-trained parameter weights of the pre-trained language model mt5-base with an Encoder-Decoder structure, and the training objective is to minimize the maximum log-likelihood estimate:
[0109]
[0110] Where m is the decoding length of the target sequence.
[0111] like Figure 4 As shown, the cultural and tourism reviews in the cultural and tourism review dataset obtained after data annotation can be used as the text information of the sample to be enhanced. After processing with random masking and label conversion, the data enhancement input-output pair is obtained, and the pre-trained language model is fine-tuned to obtain the target pre-trained language model. Then, the target pre-trained language model is used to generate the text information of the sample to be enhanced in order to enhance the dataset. The resulting model training dataset is the preset training sample set.
[0112] Optionally, the sample text information can be segmented according to a target length value to obtain multiple sample text information fragments, wherein the target length value is a random value within a preset length range.
[0113] For example, the sample text information X is evenly divided into T segments of length L, where The length L is a random value within the interval [32, 64]. For each sample text information, a value is randomly selected as the length L in each round. Simultaneously, the length of the first T-1 segments is L, and the length of the Tth segment is |X|%L, where % represents the modulo operation. The resulting multiple sample text information segments can be denoted as {x1, x2, ..., x...}. n}
[0114] Figure 6 This is a schematic diagram of an encoding process provided by an embodiment of the present disclosure. For example, let each sample text information fragment be... Where i∈[1,T], L i For X iThe length; [CLS] and [SEP] symbols are inserted at the beginning and end of each sample text information segment to obtain the sequence. The sequence is then input into the BERT encoding layer, and the hidden state representation of the first marker [CLS] is used as the representation r of each sample text information segment. i That is, the initial representation of the sample fragment.
[0115] The initial representation of each sample segment is encoded using a bidirectional GRU network to obtain the contextual interaction representation h of each text segment. i :
[0116]
[0117] Utilizing an attention mechanism and introducing a fragment-level context vector v s And using this vector to measure the importance of the fragment, we obtain the overall representation of the sample text information X (sample text information representation) d:
[0118] u i =tanh(W1h) i +b1),
[0119]
[0120] Where W1 is the weight matrix and b1 is the bias matrix. The fragment-level context vector v s It can be randomly initialized and jointly learned during the training process of the preset text representation and classification model.
[0121] For example, a preset auxiliary statement g is constructed for a single preset content type c according to the template "Evaluation type of content type c?". c ,and Where l is the length of the preset auxiliary statement. The [CLS] and [SEP] tokens are inserted at the beginning and end of the string for tokenization, and then input into the BERT encoder. The hidden state representation of the first token [CLS] is taken as the auxiliary sentence g. c The representation of (preset auxiliary statement representation) e c .
[0122] To extract fragment information important to the preset content type c, an attention mechanism is used to extract the representation (content type representation d) related to the preset content type. c :
[0123]
[0124] Where W2 and W3 are weight matrices, b2 is a bias matrix, and z1 is a column vector that can be randomly initialized and jointly learned during training.
[0125] The [CLS] and [SEP] markers are inserted at the beginning and end of the three preset evaluation types: "positive," "negative," and "neutral." These markers are then input into the BERT encoder, and the hidden state of the first marker [CLS] is used as the representation of the three preset evaluation types (preset evaluation type representation). pos e neg and e neu For each evaluation type, an attention mechanism is used to extract information related to the three preset evaluation types from the sample text, and the representations of the three preset evaluation types in the sample text (evaluation type representations) d are obtained. pos d neg and d neu :
[0126]
[0127] Among them, W4, W5, W6, W7, W8 and W9 are weight matrices, b3, b4 and b5 are bias matrices, and z2, z3 and z4 are column vectors, which can be randomly initialized and jointly learned during training.
[0128] In this embodiment of the disclosure, the first classifier is used to determine a first sample prediction probability that the sample text information involves the current preset content type based on the content type representation; the second classifier is used to determine a plurality of second sample prediction probabilities of a preset evaluation type corresponding to the current preset content type involved in the sample text information based on the joint representation of the content type representation and the sample text information representation, wherein each second prediction probability corresponds to a different preset evaluation type; the third classifier is used to determine a plurality of third sample prediction probabilities that the preset evaluation type in the preset evaluation type set and the current preset content type are both involved in the sample text information based on the joint representation of the content type representation and different evaluation type representations, wherein each third prediction probability corresponds to a different preset evaluation type; the sample prediction probabilities output by the preset text representation and classification model include the first sample prediction probability, the second sample prediction probability, and the third sample prediction probability.
[0129] Figure 7This is a schematic diagram of a classification detection process provided in an embodiment of the present disclosure. The first classifier corresponds to the Aspect Category Detection (ACD) task (which can be understood as a task used to detect whether text information involves a certain preset content type); the second classifier corresponds to the aspect category type classification task (which can be understood as a task used to classify a preset evaluation type corresponding to a certain preset content type involved in the text information); and the third classifier corresponds to the aspect category type detection task (which can be understood as a task used to detect whether the text information involves a certain preset content type and whether the corresponding evaluation type is a certain preset evaluation type).
[0130] For example, the first classifier can be a sigmoid function-based classifier. For the content type detection task, content type detection is performed on a preset content type c, and the content type representation d... c Based on this, the first classifier is applied for classification:
[0131]
[0132] Among them W task1 Let b be the weight matrix. task1 For bias, The probability that the sample text information X predicted for the ACD task involves a preset content type c is called the first prediction probability (or the first sample prediction probability).
[0133] For example, the second classifier can be a classifier based on the softmax function. For aspect category type classification tasks, first d c Concatenating d with the sample text information X yields a joint representation d of the preset content type c and the sample text information X. c,X =[d c ;d]; Next, d c,X The probability of different preset evaluation types corresponding to preset content types is input into the second classifier, which is also known as the second predicted probability (or the second sample predicted probability):
[0134]
[0135] Among them W task2 Let b be the weight matrix. task2 The bias matrix, It is a three-dimensional probability distribution, which is the probability distribution of the preset evaluation type of the preset content type c.
[0136] For example, the third classifier could be a sigmoid function-based classifier. For aspect category type detection tasks, concatenating d... c The representation of the three preset evaluation types d pos dneg and d neu The joint representation d of the preset content type and the three preset evaluation types is obtained. c,sp =[d c ;d sp ], where sp∈{neg,neu,pos}. A third classifier is applied to the above three joint representations to determine the relationship between the preset content type and the three preset evaluation types, obtaining the probability that for each preset evaluation type, the current preset evaluation type and a single preset content type are simultaneously involved in the sample text information, i.e., the third prediction probability (or the third sample prediction probability):
[0137]
[0138] Among them, W task3 Let b be the weight matrix. task3 The bias matrix, To evaluate the probability of the existence of (c,sp), another
[0139] For example, a first true probability corresponding to the predicted probability of the first sample is determined based on the sample label, and a first loss relationship is determined based on the predicted probability of the first sample and the first true probability.
[0140] For example, the first loss relationship can be determined based on the first loss function, which can be expressed by the following expression:
[0141]
[0142] in, denoted by p, where p represents the true probability, and c represents the predicted probability.
[0143] For example, the second true probability corresponding to the second sample prediction probability is determined based on the sample label, and the second loss relationship is determined based on the second sample prediction probability and the second true probability.
[0144] For example, the second loss relationship can be determined based on a second loss function, which can be expressed by the following expression:
[0145]
[0146] in, denoted by p, where p represents the true probability, and c represents the predicted probability.
[0147] For example, the third true probability corresponding to the third sample prediction probability is determined based on the sample label, and the third loss relationship is determined based on the third sample prediction probability and the third true probability.
[0148] For example, the third loss relationship can be determined based on the third loss function, which can be expressed by the following expression:
[0149]
[0150] in, denoted by p, where p represents the true probability, and c represents the predicted probability.
[0151] For example, a target loss relationship is determined based on the first loss relationship, the second loss relationship, and the third loss relationship.
[0152] For example, the target loss relation can be determined based on the target loss function, for the sample text information X, and the preset set of content types. The overall loss function (target loss function) can be expressed by the following expression:
[0153]
[0154] Among them, c i This indicates the preset content type. λ1 and λ2 are weighting coefficients with a value range of (0,1), and λ1+λ2<1. Here, we take λ1, and the value of λ2 can be set according to the actual situation, for example, 0.3 and 0.4 respectively.
[0155] For example, a preset text classification model is trained based on a target loss relation. During training, the goal is to minimize the target loss relation, and training methods such as backpropagation are used to continuously optimize the weight parameter values in the preset text classification model until a preset training cutoff condition is met. The specific training cutoff condition can be set according to actual needs, and this embodiment of the disclosure does not limit it. For example, it can be set based on the number of iterations, the degree of convergence of the loss value, or the model accuracy.
[0156] Figure 8 This is a schematic diagram of the structure of a text classification device provided in an embodiment of the present disclosure, as shown below. Figure 8 As shown, the device includes:
[0157] The text information acquisition module 801 is used to acquire the target text information to be processed.
[0158] The recognition result determination module 802 is used to determine a first category recognition result and a second category recognition result corresponding to the target text information based on the target text information; wherein, the first category recognition result is determined based on the target text information, and the first category recognition result represents the content type information and evaluation type information corresponding to the target text information; the second category recognition result is determined based on the content category features and evaluation category features of the target text information, and the second category recognition result is used to represent the probability of the content type corresponding to the target text information and the evaluation type corresponding to the content type.
[0159] The type determination module 803 is used to determine the content type of the target text information and the evaluation type corresponding to the content type based on the first category recognition result and the second category recognition result corresponding to the target text information.
[0160] The model training apparatus provided in this embodiment of the present disclosure obtains category recognition results for content type and evaluation type identification of target text information from different angles, and then combines the two category recognition results to determine the final classification result, which can effectively improve the accuracy of content type and evaluation type classification of text information.
[0161] Optionally, the recognition result determination module is used to: for each preset content type in the preset content type set, input the target text information, preset auxiliary statements, and preset evaluation type set into the target text representation and classification model to obtain the predicted probability output by the target text representation and classification model; and determine the first category recognition result and the second category recognition result of the target text information for the current preset content type based on the predicted probability output by the target text representation and classification model; wherein, the preset auxiliary statements are constructed based on the current preset content type and preset template; and the content category features and evaluation category features of the target text information are extracted by the target text representation and classification model from the target text information, the preset auxiliary statements, and the preset evaluation type set.
[0162] Optionally, the target text representation and classification model includes a bidirectional language representation transformation model (BERT) encoder and an attention layer; the BERT encoder is used to encode the preset auxiliary statement into a preset auxiliary statement representation, encode the target text information into an initial text information representation, and encode each preset evaluation type in the preset evaluation type set into an evaluation type representation; the attention layer is used to determine the content type representation corresponding to the preset auxiliary statement representation, the text information representation corresponding to the initial text information representation, and the evaluation type representation corresponding to the evaluation type representation; the content type representation, the text information representation, and the evaluation type representation are used to determine the prediction probability.
[0163] Optionally, the target text representation and classification model further includes a bidirectional recurrent neural network unit based on a gating mechanism; wherein, the device further includes: a segmentation module, used to segment the target text information into multiple text information fragments before inputting the target text information, preset auxiliary statements, and preset evaluation type set into the target text representation and classification model; wherein, a BERT encoder is used to encode the multiple text information fragments into initial text information representations, wherein the initial text information representations include fragment initial representations corresponding to the multiple text information fragments respectively; the bidirectional recurrent neural network unit is used to encode each fragment initial representation to obtain a context interaction representation corresponding to each fragment initial representation; an attention layer is used to determine the text information representation corresponding to the target text information based on the context interaction representations corresponding to the multiple fragment initial representations.
[0164] Optionally, the target text representation and classification model further includes a first classifier, a second classifier, and a third classifier; wherein, the first classifier is used to determine a first predicted probability that the target text information involves the current preset content type based on the content type representation; the second classifier is used to determine several second predicted probabilities of a preset evaluation type corresponding to the current preset content type involved by the target text information based on the joint representation of the content type representation and the text information representation, wherein each second predicted probability corresponds to a different preset evaluation type; the third classifier is used to determine whether the preset evaluation type in the preset evaluation type set is the same as the current preset content type based on the joint representation of the content type representation and different evaluation type representations. The method involves several third prediction probabilities related to the target text information, each of which corresponds to a different preset evaluation type. The prediction probabilities output by the target text representation and classification model include the first prediction probability, the second prediction probability, and the third prediction probability. The step of determining the first category recognition result and the second category recognition result corresponding to the current preset content type of the target text information based on the prediction probabilities output by the target text representation and classification model includes: determining the first category recognition result corresponding to the current preset content type of the target text information based on the third prediction probability; and determining the second category recognition result corresponding to the current preset content type of the target text information based on the first prediction probability and the second prediction probability.
[0165] Optional, the type determination module includes:
[0166] The classification result determination unit is used to determine, for each preset content type in the preset content type set, if both the first category identification result and the second category identification result corresponding to the current preset content type include the target text information involving the current preset content type, then determine that the target text information involves the current preset content type, and determine the evaluation type contained in the one with the greater confidence between the first category identification result and the second category identification result as the evaluation type corresponding to the current preset content type involving the target text information, thereby obtaining the classification result corresponding to the current preset content type;
[0167] The type determination unit is used to determine the content type of the target text information and the evaluation type corresponding to the content type based on the classification results corresponding to each preset content type in the preset content type set.
[0168] Optionally, the target text representation and classification model is trained in the following manner: A preset training sample set is obtained, wherein the preset training samples in the preset training sample set include sample text information, the preset training samples are associated with sample labels, the sample labels include the sample content type involved in the sample text information, and the sample evaluation type corresponding to the sample content type involved in the sample text information, the sample content type is a preset content type in the preset content type set, and the sample evaluation type is a preset evaluation type in the preset evaluation type set; for each preset content type in the preset content type set, the sample text information, preset auxiliary statements, and the preset evaluation type set in the preset training sample set are input into the preset text representation and classification model to obtain the sample prediction probability output by the preset text representation and classification model; a target loss relationship is determined based on the sample prediction probability output by the preset text representation and classification model and the sample labels, and the preset text representation and classification model is trained based on the target loss relationship to obtain the target text representation and classification model.
[0169] Optionally, the preset text representation and classification model further includes a bidirectional recurrent neural network unit based on a gating mechanism; before inputting the sample text information, preset auxiliary statements, and preset evaluation type set from the preset training sample set into the preset text representation and classification model, the model further includes: segmenting the sample text information according to a target length value to obtain multiple sample text information fragments, wherein the target length value is a random value within a preset length range; wherein inputting the sample text information from the preset training sample set into the preset text representation and classification model includes: inputting multiple sample text information fragments corresponding to the sample text information from the preset training sample set into the preset text representation and classification model.
[0170] Optionally, the preset text representation and classification model further includes a first classifier, a second classifier, and a third classifier; the first classifier is used to determine a first sample prediction probability that the sample text information involves the current preset content type based on the content type representation; the second classifier is used to determine several second sample prediction probabilities of a preset evaluation type corresponding to the current preset content type involved in the sample text information based on the joint representation of the content type representation and the sample text information representation, wherein each second prediction probability corresponds to a different preset evaluation type; the third classifier is used to determine several third sample prediction probabilities that the preset evaluation type in the preset evaluation type set and the current preset content type are simultaneously involved in the sample text information based on the joint representation of the content type representation and different evaluation type representations, wherein each third prediction probability corresponds to a different preset evaluation type; the sample prediction probabilities output by the preset text representation and classification model include the first sample prediction probability, the second sample prediction probability, and the third sample prediction probability.
[0171] The step of determining the target loss relationship based on the predicted probability output by the preset text representation and classification model and the sample label includes: determining a first true probability corresponding to the first sample predicted probability based on the sample label; determining a first loss relationship based on the first sample predicted probability and the first true probability; determining a second true probability corresponding to the second sample predicted probability based on the sample label; determining a second loss relationship based on the second sample predicted probability and the second true probability; determining a third true probability corresponding to the third sample predicted probability based on the sample label; determining a third loss relationship based on the third sample predicted probability and the third true probability; and determining a target loss relationship based on the first loss relationship, the second loss relationship, and the third loss relationship.
[0172] Optionally, the preset training sample set is determined in the following way: Input data is constructed based on the text information of the sample to be enhanced, wherein the input data includes masked text information corresponding to the text information of the sample to be enhanced, a preset content type involved in the text information of the sample to be enhanced, and a preset evaluation type corresponding to the preset content type involved in the text information of the sample to be enhanced; the input data is input into a target pre-trained language model to obtain the output of the target pre-trained language model; the enhanced sample text information is determined based on the output of the target pre-trained language model; and the preset training sample set is determined based on the text information of the sample to be enhanced and the enhanced sample text information. Wherein, the preset content type involved in the sample text information to be enhanced is the same as the preset content type involved in the enhanced sample text information, and the preset evaluation type corresponding to the preset content type involved in the sample text information to be enhanced is the same as the preset evaluation type corresponding to the preset content type involved in the enhanced sample text information; wherein, the mask text information corresponding to the sample text information to be enhanced is obtained by replacing each character in the sample text information to be enhanced with a preset mask symbol with a preset probability, thereby obtaining the mask text information corresponding to the sample text information to be enhanced; wherein, the preset probability is positively correlated with the length of the sample text information to be enhanced.
[0173] The information processing apparatus provided in this disclosure can execute the text classification method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.
[0174] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0175] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 9 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 9 The diagram below shows the structure of the terminal device or server 900. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0176] like Figure 9 As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An edit / output (I / O) interface 905 is also connected to the bus 904.
[0177] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0178] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.
[0179] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0180] The electronic device provided in this embodiment and the text classification method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0181] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the text classification method provided in the above embodiments.
[0182] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the text classification method provided in the above embodiments.
[0183] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0184] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0185] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to implement any text classification method of the present disclosure embodiments.
[0186] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0187] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0188] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a module does not necessarily limit the module itself; for example, a text information acquisition module can also be described as "a module for acquiring target text information to be processed".
[0189] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0190] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0191] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0192] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0193] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A text classification method, characterized in that, include: Obtain the target text information to be processed; Based on the target text information, a first category recognition result and a second category recognition result corresponding to the target text information are determined; wherein, the first category recognition result is determined based on the target text information, and the first category recognition result represents the content type information and evaluation type information corresponding to the target text information; the second category recognition result is determined based on the content category features and evaluation category features of the target text information, and the second category recognition result is used to represent the probability of the content type corresponding to the target text information and the evaluation type corresponding to the content type. Based on the first category recognition result and the second category recognition result corresponding to the target text information, the content type of the target text information and the evaluation type corresponding to the content type are determined.
2. The method according to claim 1, characterized in that, Based on the target text information, determine the first category recognition result and the second category recognition result corresponding to the target text information, including: For each preset content type in the preset content type set, the target text information, preset auxiliary statements, and preset evaluation type set are input into the target text representation and classification model to obtain the predicted probability output by the target text representation and classification model. Based on the predicted probability output by the target text representation and classification model, the first category recognition result and the second category recognition result of the target text information for the current preset content type are determined. The preset auxiliary statements are constructed based on the current preset content type and a preset template. The content category features and evaluation category features of the target text information are extracted by the target text representation and classification model from the target text information, the preset auxiliary statements, and the preset evaluation type set.
3. The method according to claim 1, characterized in that, The target text representation and classification model includes a bidirectional language representation transformation model (BERT) encoder and an attention layer. The BERT encoder is used to encode the preset auxiliary statement into a preset auxiliary statement representation, encode the target text information into an initial text information representation, and encode each preset evaluation type in the preset evaluation type set into a preset evaluation type representation. The attention layer is used to determine the content type representation corresponding to the preset auxiliary statement representation, the text information representation corresponding to the initial text information representation, and the evaluation type representation corresponding to the preset evaluation type representation. The content type representation, the text information representation, and the evaluation type representation are used to determine the prediction probability.
4. The method according to claim 3, characterized in that, The target text representation and classification model also includes a bidirectional recurrent neural network unit based on a gating mechanism; Before inputting the target text information, preset auxiliary statements, and preset evaluation type set into the target text representation and classification model, the following steps are also included: The target text information is segmented to obtain multiple text information fragments; The input of the target text information into the target text representation and classification model includes: The multiple text information fragments are input into the target text representation and classification model; The BERT encoder is used to encode the target text information into an initial representation of the text information, including: the BERT encoder is used to encode the plurality of text information segments into an initial representation of the text information, wherein the initial representation of the text information includes the initial representation of the segments corresponding to the plurality of text information segments respectively; The bidirectional recurrent neural network unit is used to encode the initial representation of each segment to obtain the context interaction representation corresponding to the initial representation of each segment; The attention layer is used to determine the text information representation corresponding to the initial representation of the text information, including: the attention layer is used to determine the text information representation corresponding to the target text information based on the context interaction representations corresponding to the initial representations of multiple segments.
5. The method according to claim 3, characterized in that, The target text representation and classification model also includes a first classifier, a second classifier, and a third classifier; Wherein, the first classifier is used to determine a first predicted probability that the target text information involves the current preset content type based on the content type representation; the second classifier is used to determine several second predicted probabilities of a preset evaluation type corresponding to the current preset content type involved by the target text information based on the joint representation of the content type representation and the text information representation, wherein each second predicted probability corresponds to a different preset evaluation type; the third classifier is used to determine several third predicted probabilities that a preset evaluation type in the preset evaluation type set and the current preset content type are both involved in the target text information based on the joint representation of the content type representation and different evaluation type representations, wherein each third predicted probability corresponds to a different preset evaluation type; the predicted probabilities output by the target text representation and classification model include the first predicted probability, the second predicted probability, and the third predicted probability. The step of determining the first category recognition result and the second category recognition result of the target text information for the current preset content type based on the predicted probability output by the target text representation and classification model includes: The target text information is identified as the first category corresponding to the current preset content type based on the third predicted probability. Based on the first predicted probability and the second predicted probability, the second category recognition result of the target text information corresponding to the current preset content type is determined.
6. The method according to claim 1, characterized in that, The step of determining the content type of the target text information and the evaluation type corresponding to the content type based on the first category recognition result and the second category recognition result corresponding to the target text information includes: For each preset content type in the preset content type set, if both the first category identification result and the second category identification result corresponding to the current preset content type include the target text information involving the current preset content type, then it is determined that the target text information involves the current preset content type, and the evaluation type contained in the one with the greater confidence in the first category identification result and the second category identification result is determined as the evaluation type corresponding to the current preset content type involving the target text information, thereby obtaining the classification result corresponding to the current preset content type; Based on the classification results corresponding to each preset content type in the preset content type set, the content type of the target text information and the evaluation type corresponding to the content type are determined.
7. The method according to claim 2, characterized in that, The target text representation and classification model is trained in the following way: Obtain a preset training sample set, wherein the preset training samples in the preset training sample set include sample text information, the preset training samples are associated with sample tags, the sample tags include the sample content type involved in the sample text information, and the sample evaluation type corresponding to the sample content type involved in the sample text information, the sample content type is a preset content type in the preset content type set, and the sample evaluation type is a preset evaluation type in the preset evaluation type set. For each preset content type in the preset content type set, the sample text information, preset auxiliary statements, and preset evaluation type set in the preset training sample set are input into the preset text representation and classification model to obtain the sample prediction probability output by the preset text representation and classification model. The target loss relationship is determined based on the sample prediction probability output by the preset text representation and classification model and the sample label. The preset text representation and classification model is then trained based on the target loss relationship to obtain the target text representation and classification model.
8. The method according to claim 7, characterized in that, The preset text representation and classification model further includes a bidirectional recurrent neural network unit based on a gating mechanism; before inputting the sample text information from the preset training sample set, the preset auxiliary statements, and the preset evaluation type set into the preset text representation and classification model, it also includes: The sample text information is segmented according to the target length value to obtain multiple sample text information fragments, wherein the target length value is a random value within a preset length range; The process of inputting sample text information from the preset training sample set into the preset text representation and classification model includes: Multiple sample text information fragments corresponding to the sample text information in the preset training sample set are input into the preset text representation and classification model.
9. The method according to claim 7, characterized in that, The preset text representation and classification model further includes a first classifier, a second classifier, and a third classifier. The first classifier is used to determine a first sample prediction probability that the sample text information involves the current preset content type based on the content type representation. The second classifier is used to determine several second sample prediction probabilities of a preset evaluation type corresponding to the current preset content type involved in the sample text information based on the joint representation of the content type representation and the sample text information representation, wherein each second prediction probability corresponds to a different preset evaluation type. The third classifier is used to determine several third sample prediction probabilities that the preset evaluation type in the preset evaluation type set and the current preset content type are both involved in the sample text information based on the joint representation of the content type representation and different evaluation type representations, wherein each third prediction probability corresponds to a different preset evaluation type. The sample prediction probabilities output by the preset text representation and classification model include the first sample prediction probability, the second sample prediction probability, and the third sample prediction probability. The step of determining the target loss relationship based on the preset text representation and the predicted probability output by the classification model and the sample label includes: Based on the sample label, determine the first true probability corresponding to the first sample prediction probability, and determine the first loss relationship based on the first sample prediction probability and the first true probability. Determine the second true probability corresponding to the second sample prediction probability based on the sample label, and determine the second loss relationship based on the second sample prediction probability and the second true probability. The third true probability corresponding to the third sample prediction probability is determined based on the sample label, and the third loss relationship is determined based on the third sample prediction probability and the third true probability. The target loss relationship is determined based on the first loss relationship, the second loss relationship, and the third loss relationship.
10. The method according to claim 7, characterized in that, The preset training sample set is determined in the following way: Input data is constructed based on the text information of the sample to be enhanced, wherein the input data includes the mask text information corresponding to the text information of the sample to be enhanced, the preset content type involved in the text information of the sample to be enhanced, and the preset evaluation type corresponding to the preset content type involved in the text information of the sample to be enhanced. The input data is fed into the target pre-trained language model to obtain the output of the target pre-trained language model; Based on the output of the target pre-trained language model, the enhanced sample text information is determined; A preset training sample set is determined based on the text information of the sample to be enhanced and the text information of the enhanced sample, wherein the preset content type involved in the text information of the sample to be enhanced is the same as the preset content type involved in the text information of the enhanced sample, and the preset evaluation type corresponding to the preset content type involved in the text information of the sample to be enhanced is the same as the preset evaluation type corresponding to the preset content type involved in the text information of the enhanced sample. The mask text information corresponding to the text information to be enhanced is obtained in the following way: Each character in the text information to be enhanced is replaced with a preset mask symbol with a preset probability to obtain the mask text information corresponding to the text information to be enhanced; wherein, the preset probability is positively correlated with the length of the text information to be enhanced.
11. A text classification device, characterized in that, include: The text information acquisition module is used to acquire the target text information to be processed; The recognition result determination module is used to determine a first category recognition result and a second category recognition result corresponding to the target text information based on the target text information; wherein, the first category recognition result is determined based on the target text information, and the first category recognition result represents the content type information and evaluation type information corresponding to the target text information; the second category recognition result is determined based on the content category features and evaluation category features of the target text information, and the second category recognition result is used to represent the probability of the content type corresponding to the target text information and the evaluation type corresponding to the content type. The type determination module is used to determine the content type of the target text information and the evaluation type corresponding to the content type based on the first category recognition result and the second category recognition result corresponding to the target text information.
12. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-10.
13. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the method as described in any one of claims 1-10.
14. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-10.