A text classification method, device, electronic device and storage medium
By fusing the vector representation of text to be classified and external entities, the problem of insufficient utilization of text external information in the prior art is solved, and the accuracy of text classification is improved.
Patent Information
- Application Number
- CN202210467956.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-04-29
AI Technical Summary
The prior art lacks the utilization of relevant information outside the text in text classification tasks, resulting in insufficient accuracy of classification results.
By obtaining the vector representation of the text to be classified and the external entity associated with it, the correlation between the two is calculated, and the vector representation of the external entity is processed with this as the weight to obtain the weighted vector representation. Then, the vector representation of the text to be classified is fused with the weighted vector representation of the external entity, and a fusion vector representation is generated to determine the category to which the text to be classified belongs.
Using information from external entities for text classification improves the accuracy of classification results and can more effectively utilize the correlation between external entities and text to be classified.
Smart Images

Figure CN114840669B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a text classification method, apparatus, electronic device, and storage medium. Background Art
[0002] The text classification task is a core subtask in the field of natural language processing and has important significance in aspects such as content category classification, spam filtering, news classification, etc.
[0003] In the related art, the text classification task only utilizes the information contained in the text itself and lacks the utilization of relevant information outside the text. Therefore, the accuracy of the results of the text classification task obtained by the related art still needs to be improved. Summary of the Invention
[0004] In view of the above problems, embodiments of the present invention provide a text classification method, apparatus, electronic device, and storage medium to overcome or at least partially solve the above problems.
[0005] In a first aspect of the embodiments of the present invention, a text classification method is provided, and the method includes:
[0006] Obtain a vector representation of a text to be classified and a vector representation of an external entity associated with the text to be classified;
[0007] Determine the relevance between the vector representation of the external entity and the vector representation of the text to be classified;
[0008] Process the vector representation of the external entity with the relevance as a weight to obtain a weighted vector representation of the external entity;
[0009] Fuse the vector representation of the text to be classified and the weighted vector representation of the external entity to obtain a fused vector representation of the text to be classified;
[0010] Determine the category to which the text to be classified belongs according to the fused vector representation.
[0011] Optionally, obtaining the vector representation of the text to be classified includes:
[0012] Determine the vector representation of each phrase in the text to be classified and the importance of each phrase in the text to be classified;
[0013] Process the vector representation of each phrase with the importance as a weight to obtain a weighted vector representation of the text to be classified.
[0014] Optionally, determining the category to which the text to be classified belongs according to the fused vector representation includes:
[0015] Predict the probabilities of the text to be classified belonging to each category according to the fusion vector representation;
[0016] Determine the category to which the text to be classified belongs from the various categories according to the magnitude relationship between the probabilities of the text to be classified belonging to each category.
[0017] Optionally, determining the category to which the text to be classified belongs according to the fusion vector representation includes:
[0018] In the case where the fusion vector representation is an N-segment vector representation, for each of the N segments, perform the following steps:
[0019] Obtain the prior category of this segment, and predict the probability that the text to be classified belongs to the prior category of this segment according to the fusion vector representation;
[0020] Predict the probability that this segment belongs to the corresponding prior category according to the vector elements of this segment in the N-segment vector representation;
[0021] Obtain the final probability that this segment belongs to the corresponding prior category according to the probability that this segment belongs to the corresponding prior category and the probability that the text to be classified belongs to the prior category of this segment;
[0022] Determine the category to which the text to be classified belongs according to the final probabilities that the N segments belong to their respective corresponding prior categories.
[0023] Optionally, determining the category to which the text to be classified belongs according to the final probabilities that the N segments belong to their respective corresponding prior categories includes:
[0024] Add up the final probabilities corresponding to multiple segments corresponding to the same prior category to obtain the probability of this prior category;
[0025] Determine the category to which the text to be classified belongs according to the magnitude relationship between the probabilities of the various prior categories.
[0026] Optionally, after determining the category to which the text to be classified belongs, the method further includes:
[0027] Mark the class label of the category to which the text to be classified belongs;
[0028] Add the text to be classified to the information stream corresponding to the class label.
[0029] In the second aspect of the embodiments of the present invention, a text classification device is provided, and the device includes:
[0030] A vector representation acquisition module, configured to acquire the vector representation of the text to be classified and the vector representation of an external entity associated with the text to be classified;
[0031] A relevance determination module, configured to determine the relevance between the vector representation of the external entity and the vector representation of the text to be classified;
[0032] An external entity weighting module, configured to process the vector representation of the external entity with the relevance as a weight to obtain a weighted vector representation of the external entity;
[0033] A vector representation fusion module, configured to fuse the vector representation of the text to be classified and the weighted vector representation of the external entity to obtain a fused vector representation of the text to be classified;
[0034] A category determination module, configured to determine the category to which the text to be classified belongs according to the fused vector representation.
[0035] Optionally, the vector representation acquisition module includes:
[0036] An importance unit, configured to determine the vector representation of each phrase in the text to be classified and the importance of each phrase in the text to be classified;
[0037] A phrase weighting unit, configured to process the vector representation of each phrase with the importance as a weight to obtain a weighted vector representation of the text to be classified.
[0038] In a third aspect of the embodiments of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the text classification method disclosed in the embodiments of the present application is implemented.
[0039] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the text classification method disclosed in the embodiments of the present application is implemented.
[0040] The embodiments of the present invention have the following advantages:
[0041] In the embodiments of the present invention, the weighted vector representation of an external entity and the vector representation of the text to be classified are fused to obtain the fused vector representation of the text to be classified, and then the category to which the text to be classified belongs is determined. By using the information of entities outside the text to be classified, it helps to improve the accuracy of determining the category to which the text to be classified belongs. Through the correlation between the vector representation of the external entity and the vector representation of the text to be classified, the weighted vector representation of the external entity associated with the text to be classified is obtained by processing the vector representation of the external entity. Considering the relevance of the external entity to the text to be classified, it can more effectively utilize the external entity. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0043] Figure 1 is a flowchart of the steps of a text classification method in the embodiments of the present invention;
[0044] Figure 2 is a schematic flowchart of determining the category to which the text to be classified belongs according to the fused vector in the embodiments of the present invention;
[0045] Figure 3 is a schematic structural diagram of a text classification device in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0047] To solve the technical problem in the related art that the text classification task lacks the utilization of relevant information outside the text, the applicant proposes to use the information of external entities associated with the text to be classified to assist in the classification of the text to be classified.
[0048] Refer to Figure 1 as shown, which shows a flowchart of the steps of a text classification method in the embodiments of the present invention. As Figure 1 shown, the text classification method may specifically include the following steps:
[0049] Step S11: Obtain the vector representation of the text to be classified and the vector representation of the external entity associated with the text to be classified.
[0050] An external entity associated with the text to be classified refers to an entity that is related to the text to be classified but belongs outside the text to be classified. For example, if the text to be classified is a review of a certain coffee shop, the external entity can be the geographical location of the coffee shop, the rating of the coffee shop, etc. that are not included in the review; the external entity can also be an entity extracted from the accompanying picture corresponding to the text to be classified, etc.
[0051] The method for obtaining the vector representations of the text to be classified and the external entity respectively can refer to related technologies, such as using the BERT model (Bidirectional Encoder Representation from Transformers, a language representation model), etc., which will not be elaborated here.
[0052] Step S12: Determine the relevance between the vector representation of the external entity and the vector representation of the text to be classified.
[0053] The relevance between the vector representation of the external entity and the vector representation of the text to be classified can be determined by calculating the cosine distance or Euclidean distance, etc. between the vector representation of the external entity and the vector representation of the text to be classified.
[0054] Step S13: Process the vector representation of the external entity with the relevance as the weight to obtain the weighted vector representation of the external entity.
[0055] Using the relevance between the vector representation of each external entity and the vector representation of the text to be classified as the weight, perform weighted processing on the vector representation of the corresponding external entity to obtain the weighted vector representation of each external entity.
[0056] Step S14: Fuse the vector representation of the text to be classified and the weighted vector representation of the external entity to obtain the fused vector representation of the text to be classified.
[0057] Fusing the vector representation of the text to be classified and the weighted vector representation of the external entity to obtain the fused vector representation of the text to be classified includes: concatenating the weighted vector representation of the external entity after the vector representation of the text to be classified to obtain the fused vector representation of the text to be classified.
[0058] Step S15: Determine the category to which the text to be classified belongs according to the fused vector representation.
[0059] Determine the category to which the text to be classified belongs according to the information contained in the fused vector representation of the text to be classified.
[0060] By adopting the technical solution of the embodiment of the present application, the weighted vector representation of an external entity and the vector representation of the text to be classified are fused to obtain the fused vector representation of the text to be classified, and then the category to which the text to be classified belongs is determined. The information of entities outside the text to be classified is utilized, which helps to improve the accuracy of the determined category of the text to be classified. Through the correlation between the vector representation of the external entity and the vector representation of the text to be classified, the weighted vector representation of the external entity associated with the text to be classified is obtained by processing the vector representation of the external entity. Considering the correlation of the external entity with respect to the text to be classified, it can more effectively utilize the external entity.
[0061] Optionally, the text classification method can be implemented through a text classification model. The text to be classified and the external entity are input into the text classification model. Each module inside the text classification model obtains the respective vector representations of the text to be classified and the external entity, calculates the correlation between the vector representation of the external entity and the vector representation of the text to be classified, and then obtains the weighted representation of the external entity. Further, the fused vector representation of the text to be classified is obtained. Finally, according to the fused vector representation of the text to be classified, the category to which the text to be classified belongs is determined.
[0062] Among them, the text classification model can be obtained by training a preset model through the following steps: obtaining a text sample, an external entity sample associated with the text sample, and the true category to which the text sample belongs, and inputting the text sample and the external entity sample associated with the text sample into the preset model; the preset model obtains the vector representation of the text sample and the vector representation of the external entity sample associated with the text sample; the preset model determines the correlation between the vector representation of the external entity sample and the vector representation of the text sample, and uses this correlation as a weight to process the vector representation of the external entity sample to obtain the weighted vector representation of the external entity sample; the preset model fuses the vector representation of the text sample and the weighted vector representation of the external entity sample to obtain the fused vector representation of the text sample; and according to the fused vector of the text sample, the predicted category of the text sample is obtained. According to the difference between the predicted category of the text sample and the true category, a loss function is established to train the preset model to obtain the text classification model.
[0063] In this way, determining the category to which the text to be classified belongs through the text classification model is faster and more convenient.
[0064] Optionally, on the basis of the above technical solution, the vector representation of the text to be classified can be the weighted vector representation of the text to be classified obtained by weighting the vector representations of each phrase in the text to be classified.
[0065] Each phrase of the text to be classified can be obtained by dividing according to part of speech. For example, the subject, predicate, and object in a sentence are divided into different phrases; or the word segmentation algorithm in related technologies is used to obtain each phrase in the text to be classified.
[0066] The self-attention between each phrase can be calculated to determine the importance of each phrase in the text to be classified; or an encoder can be used to determine the importance of each phrase in the text to be classified.
[0067] Taking the importance of each phrase in the text to be classified as the weight of the corresponding phrase, the vector representation of each phrase is weighted to obtain the weighted vector representation of each phrase. According to the weighted vector representation of each phrase, the weighted vector representation of the text to be classified can be obtained.
[0068] Optionally, when the text to be classified is a long text, the importance of each sentence in the text to be classified can be calculated, and each sentence is weighted correspondingly to obtain the weighted vector representation of the text to be classified.
[0069] Taking the weighted vector representation of the text to be classified as the vector representation of the text to be classified, the correlation between the vector representation of the external entity and the vector representation of the text to be classified is determined and fused with the weighted vector representation of the external entity to obtain the fused vector representation of the text to be classified.
[0070] By adopting the technical solution of the embodiment of the present application, the vector representation of each phrase of the text to be classified can be weighted according to the importance of each phrase of the text to be classified, and then the weighted vector representation of the text to be classified can be obtained; when the vector representation of the text to be classified is processed subsequently, more attention will be paid to the phrases with high importance in the text to be classified, so the determined category to which the text to be classified belongs is more accurate.
[0071] Optionally, on the basis of the above technical solution, determining the category to which the text to be classified belongs according to the fused vector representation includes: predicting the probability that the text to be classified belongs to each category according to the fused vector representation; determining the category to which the text to be classified belongs from the various categories according to the magnitude relationship between the probabilities that the text to be classified belongs to each category.
[0072] Optionally, the fused vector representation can be input into a text classification model to obtain the probabilities that the text to be classified belongs to each category determined by the text classification model, and according to the magnitude relationship between the probabilities that the text to be classified belongs to each category, the category with the highest probability is determined as the category to which the text to be classified belongs. Among them, each category can be preset.
[0073] Optionally, based on the above technical solution, determining the category to which the text to be classified belongs according to the fusion vector representation includes: when the fusion vector representation is an N-segment vector representation, for each of the N segments, performing the following steps: obtaining the prior category of this segment, and predicting the probability that the text to be classified belongs to the prior category of this segment according to the fusion vector representation; predicting the probability that this segment belongs to the corresponding prior category according to the vector elements of this segment in the N-segment vector representation; obtaining the final probability that this segment belongs to the corresponding prior category according to the probability that this segment belongs to the corresponding prior category and the probability that the text to be classified belongs to the prior category of this segment; determining the category to which the text to be classified belongs according to the final probabilities that the N segments belong to their respective corresponding prior categories.
[0074] When the text to be classified is a collection of multiple contents, the relevance between each part of the text to be classified can be calculated, and the text to be classified is divided into N segments, and each segment of text corresponds to a sub-case. For example, the content of the text to be classified is a collection of store visits of multiple stores, and the multiple stores include a dine-in bar, a dine-in Chinese restaurant, a takeout milk tea shop, and a takeout steamed bun shop; then the text to be classified can be divided into four segments.
[0075] Correspondingly, the fusion vector of the text to be classified can also be divided into N segments, and each segment of vector representation characterizes a sub-case. For each of the N segments, obtain the prior category of this segment, where the prior category of each segment refers to the category corresponding to the sub-case characterized by this segment, and the categories that the text to be classified may belong to include the prior categories of each segment. For example, to determine whether the text to be classified belongs to the dine-in or takeout category, determine the prior category corresponding to a segment of vector representation. Assume that the sub-case characterized by this segment of vector representation is a certain bar. According to prior knowledge, it can be known that the bar corresponds to dine-in rather than takeout. Therefore, the prior category of this segment of vector representation is dine-in.
[0076] Predict the probability that the text to be classified belongs to the prior category of each segment according to the fusion vector representation of the text to be classified. Since the categories that the text to be classified may belong to include the prior categories of each segment, the probability that the text to be classified belongs to each possible category can be directly predicted according to the fusion vector representation of the text to be classified.
[0077] Predict the probability that this segment belongs to the corresponding prior category according to the vector elements of this segment in each segment of vector representation. For each segment of vector representation, obtain the final probability that this segment belongs to the corresponding prior category according to the product of the predicted probability that this segment belongs to the corresponding prior category and the probability that the text to be classified belongs to the prior category of this segment.
[0078] For the final probability that each segment of vector representation belongs to its corresponding prior category, the category to which the text to be classified belongs can be determined. Among them, the prior category corresponding to the vector representation with the highest final probability can be determined as the category to which the text to be classified belongs.
[0079] For example: the prior category corresponding to a segment of vector representation is category A. According to the fused vector representation of the text to be classified, the probability that the text to be classified belongs to category A is predicted to be 0.3, and according to this segment of vector representation, the probability that this segment of vector representation belongs to category A is predicted to be 0.6. Then the final probability of this segment of vector representation corresponding to category A is 0.3×0.6 = 0.18; the prior category corresponding to a segment of vector representation is category B. According to the fused vector representation of the text to be classified, the probability that the text to be classified belongs to category B is predicted to be 0.7, and according to this segment of vector representation, the probability that this segment of vector representation belongs to category B is predicted to be 0.4. Then the final probability of this segment of vector representation corresponding to category B is 0.7×0.4 = 0.28.
[0080] Optionally, as an embodiment, the final probabilities corresponding to multiple segments belonging to the same prior category can be added to obtain the probability of this prior category. According to the magnitude relationship between the probabilities of each prior category, the prior category with the highest probability is determined as the category to which the text to be classified belongs.
[0081] Figure 2 A flowchart showing the process of determining the category to which the text to be classified belongs according to the fused vector is shown. The possible categories to which the text to be classified may belong include category A and category B; the bar chart represents each segment of vector representation. The white bar chart represents the vector representation with the prior category of category A, and the gray bar chart represents the vector representation with the prior category of category B. The height of the bar chart represents the probability of each segment of vector representation, including the probability of belonging to the prior category and the final probability of the corresponding prior category; arranging the bar charts can obtain the probability distribution. According to the fused vector representation of the text to be classified, the probability P A that the text to be classified belongs to category A, and the probability P B that it belongs to category B, P A +P B = 1. According to the prior category of each segment of vector representation, the probability that this segment of vector representation belongs to the prior category is predicted. Multiply the probability that this segment of vector belongs to the prior category by the probability that the text to be classified belongs to this prior category to obtain the final probability of the prior category corresponding to this segment of vector. According to the final probabilities of the prior categories corresponding to each segment of vector, the final distribution can be obtained.
[0082] Adopting the technical solution of the embodiment of the present application, considering that different categories have different attributes, if the same weight is used for all categories, the differences between categories will be lost. Therefore, through the final probability of each segment of vector representation, the probability that the predicted text to be classified belongs to each possible category is fine-tuned, making the determined category to which the text to be classified belongs more accurate.
[0083] It is understandable that the probabilities of the predicted text to be classified belonging to each possible category can be fine-tuned by the final probabilities represented by each segment of vectors, and this can also be achieved by a model. The model can be trained using a supervised training method, which will not be elaborated here.
[0084] Optionally, based on the above technical solution, after determining the category to which the text to be classified belongs, a category label of the category to which it belongs can be marked for the text to be classified, and the text to be classified can be added to the information stream corresponding to the category label. In this way, when the information stream of the category label is sent down, the text to be classified can also be sent down.
[0085] Optionally, after determining the category to which the text to be classified belongs, in addition to adding the text to be classified to the information stream corresponding to the category label, there can be multiple uses. For example, when a new piece of content is in cold start, the understanding of the text to be classified can be enhanced through the category label.
[0086] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0087] Figure 3 is a schematic structural diagram of a text classification device according to an embodiment of the present invention, as Figure 3 shown, a text classification device includes a vector representation acquisition module, a relevance determination module, an external entity weighting module, a vector representation fusion module, and a category determination module, where:
[0088] The vector representation acquisition module is configured to acquire the vector representation of the text to be classified and the vector representation of an external entity associated with the text to be classified;
[0089] The relevance determination module is configured to determine the relevance between the vector representation of the external entity and the vector representation of the text to be classified;
[0090] The external entity weighting module is configured to process the vector representation of the external entity with the relevance as a weight to obtain a weighted vector representation of the external entity;
[0091] The vector representation fusion module is configured to fuse the vector representation of the text to be classified with the weighted vector representation of the external entity to obtain a fused vector representation of the text to be classified;
[0092] A category determination module, configured to determine the category to which the text to be classified belongs according to the fused vector representation.
[0093] Optionally, as an embodiment, the vector representation acquisition module includes:
[0094] An importance unit, configured to determine the vector representation of each phrase in the text to be classified, and the importance of each phrase in the text to be classified;
[0095] A phrase weighting unit, configured to process the vector representation of each phrase with the importance as the weight to obtain the weighted vector representation of the text to be classified.
[0096] Optionally, as an embodiment, the category determination module includes:
[0097] A probability determination unit, configured to predict the probability that the text to be classified belongs to each category according to the fused vector representation;
[0098] A category determination unit, configured to determine the category to which the text to be classified belongs from the categories according to the magnitude relationship between the probabilities that the text to be classified belongs to each category.
[0099] Optionally, as an embodiment, the category determination module includes:
[0100] A step execution unit, configured to, when the fused vector representation is an N-segment vector representation, for each of the N segments, execute the following steps:
[0101] Obtain the prior category of this segment, and predict the probability that the text to be classified belongs to the prior category of this segment according to the fused vector representation;
[0102] Predict the probability that this segment belongs to the corresponding prior category according to the vector elements of this segment in the N-segment vector representation;
[0103] Obtain the final probability that this segment belongs to the corresponding prior category according to the probability that this segment belongs to the corresponding prior category and the probability that the text to be classified belongs to the prior category of this segment;
[0104] Determine the category to which the text to be classified belongs according to the final probabilities that the N segments belong to their respective corresponding prior categories.
[0105] Optionally, as an embodiment, determining the category to which the text to be classified belongs according to the final probabilities that the N segments belong to their respective corresponding prior categories includes:
[0106] Add up the final probabilities corresponding to multiple segments corresponding to the same prior category to obtain the probability of this prior category;
[0107] Determine the category to which the text to be classified belongs according to the magnitude relationship among the probabilities of each prior category.
[0108] Optionally, as an embodiment, after determining the category to which the text to be classified belongs, the apparatus further includes:
[0109] A label marking module, configured to mark the category label of the category to which the text to be classified belongs;
[0110] An adding module, configured to add the text to be classified to the information stream corresponding to the category label.
[0111] It should be noted that the apparatus embodiment is similar to the method embodiment, so the description is relatively simple. For related parts, please refer to the method embodiment.
[0112] An embodiment of the present invention further provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the text classification method disclosed in the embodiments of the present application is implemented.
[0113] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the text classification method disclosed in the embodiments of the present application is implemented.
[0114] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, please refer to each other.
[0115] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0116] Embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0117] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0119] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
[0120] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the said element.
[0121] The above provides a detailed introduction to a text classification method, apparatus, electronic device and storage medium provided by the present application. Specific examples are used in this text to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A text classification method, characterized in that, The method includes: Obtaining a vector representation of the text to be classified and a vector representation of an external entity associated with the text to be classified, where the external entity associated with the text to be classified refers to an entity that is related to the text to be classified but is outside the text to be classified; Determining the relevance between the vector representation of the external entity and the vector representation of the text to be classified; Processing the vector representation of the external entity with the relevance as a weight to obtain a weighted vector representation of the external entity; Fusing the vector representation of the text to be classified and the weighted vector representation of the external entity to obtain a fused vector representation of the text to be classified; Determining the category to which the text to be classified belongs according to the fused vector representation.
2. The method according to claim 1, characterized in that, Obtaining the vector representation of the text to be classified includes: Determining the vector representation of each phrase in the text to be classified and the importance of each phrase in the text to be classified; Processing the vector representation of each phrase with the importance as a weight to obtain a weighted vector representation of the text to be classified.
3. The method according to claim 1, characterized in that, Determining the category to which the text to be classified belongs according to the fused vector representation includes: Predicting the probability that the text to be classified belongs to each category according to the fused vector representation; Determining the category to which the text to be classified belongs from the categories according to the magnitude relationship between the probabilities that the text to be classified belongs to each category.
4. The method according to claim 1, characterized in that, Determining the category to which the text to be classified belongs according to the fused vector representation includes: In the case where the fused vector representation is an N-segment vector representation, for each of the N segments, perform the following steps: Obtaining the prior category of this segment and predicting the probability that the text to be classified belongs to the prior category of this segment according to the fused vector representation; Predicting the probability that this segment belongs to the corresponding prior category according to the vector elements of this segment in the N-segment vector representation; Obtaining the final probability that this segment belongs to the corresponding prior category according to the probability that this segment belongs to the corresponding prior category and the probability that the text to be classified belongs to the prior category of this segment; Determining the category to which the text to be classified belongs according to the final probabilities that the N segments belong to their respective corresponding prior categories.
5. The method according to claim 4, characterized in that, Determining the category to which the text to be classified belongs according to the final probabilities that the N segments belong to their respective corresponding prior categories includes: Adding the final probabilities corresponding to multiple segments corresponding to the same prior category to obtain the probability of this prior category; Determining the category to which the text to be classified belongs according to the magnitude relationship between the probabilities of each prior category.
6. The method according to any one of claims 1-5, characterized in that, After determining the category to which the text to be classified belongs, the method further includes: Marking the category label of the category to which the text to be classified belongs; Adding the text to be classified to the information stream corresponding to the category label.
7. A text classification device, characterized in that, The apparatus includes: A vector representation acquisition module, configured to obtain a vector representation of the text to be classified and a vector representation of an external entity associated with the text to be classified, where the external entity associated with the text to be classified refers to an entity that is related to the text to be classified but is outside the text to be classified; A relevance determination module, configured to determine the relevance between the vector representation of the external entity and the vector representation of the text to be classified; An external entity weighting module, configured to process the vector representation of the external entity with the relevance as a weight to obtain a weighted vector representation of the external entity; A vector representation fusion module, configured to fuse the vector representation of the text to be classified with the weighted vector representation of the external entity to obtain a fused vector representation of the text to be classified; A category determination module, configured to determine the category to which the text to be classified belongs according to the fused vector representation.
8. The device according to claim 7, characterized in that, The vector representation acquisition module includes: An importance unit, configured to determine the vector representation of each phrase in the text to be classified and the importance of each phrase in the text to be classified; A phrase weighting unit, configured to process the vector representation of each phrase with the importance as a weight to obtain a weighted vector representation of the text to be classified.
9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the text classification method according to any one of claims 1 to 6.
10. A computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the text classification method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Text recognition method based on position information and related device
CN113590832A
Text classification method and device based on artificial intelligence, equipment and medium
CN113868419A