Text content classification method, device, electronic device and storage medium

By extracting and correcting industry characteristic information in Internet text content, forming a database, and combining manual annotation and classification model prediction, the accuracy and recall of Internet text content classification are solved, achieving more efficient industry classification results.

CN114722205BActive Publication Date: 2025-05-09SHIQU INTERACTIVE (BEIJING) TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210416313.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-20
Publication Date
2025-05-09
Estimated Expiration
2042-04-20

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately and effectively classify Internet text content in industry, especially when faced with colloquialization, arbitrary expression, diverse coverage scenarios and industry categories, and unconstrained by grammatical rules.

Method used

By obtaining content data from various industries, representative feature information is extracted, including feature words and word combinations, and the feature information is repaired and retained through manual labeling to form an industry feature database. Then, based on the database matching feature information in the text content, masking unrelated content fragments, using industry classification models to predict, and using the processed text content to train the model.

Benefits of technology

The accuracy and recall rate of text content classification are improved, ensuring that both accuracy and recall rate are taken into account in actual use scenarios, and the representative characteristics are corrected through manual labeling processing process, which improves the classification effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114722205B_ABST
    Figure CN114722205B_ABST
Patent Text Reader

Abstract

The present invention provides a text content classification method, device, electronic device and storage medium, which splits the two logics of "capturing industry features in content" and "classifying content by industry" and independently optimizes and organically integrates the strategy to ensure that the accuracy and recall rate of content industry classification are taken into account in the actual use scenario. At the same time, between the above two functions, a certain manual annotation processing process is added to correct the representative features of each industry identified from the content, which effectively improves the accuracy of text content classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a text content classification method, device, electronic device and storage medium. Background Art

[0002] With the continuous rise of social media and self-media platforms, there are articles posted by various users every day on social media such as Weibo and WeChat, video websites such as Douyin and Bilibili, and social e-commerce platforms such as Xiaohongshu. For example, ordinary users’ daily experience and sharing of products in various industries, marketing promotion copywriting of brand owners in various industries, and endorsement and promotion of products by big V anchors. In many scenarios, it is necessary to extract and classify various updated text information to determine the industries involved in the relevant content and the brands mentioned. The characteristics of these online texts are that the content posted by ordinary users is mostly colloquial, casually expressed, covering many scenes and industry categories, and not constrained by grammatical rules; while some expression routines are often used in official marketing copywriting, such as a large section of artistic conception in the early stage of the content, pointing east and hitting west, and finally turning around to point out the brand to be promoted; or in a content of big V and anchors promoting products, it usually covers many industries and categories. These bring great difficulties to accurately judge the content.

[0003] Therefore, how to accurately and effectively classify Internet text content according to industry is an urgent problem that needs to be solved. Summary of the invention

[0004] In order to improve the above problems, the present invention provides a text content classification method, device, electronic device and storage medium.

[0005] According to a first aspect of an embodiment of the present invention, a text content classification method is provided, the method comprising:

[0006] Acquire content data of various industries, and extract representative feature information of each industry through feature mining, wherein the feature information includes feature words and word combinations;

[0007] Receiving externally input manual annotation information, repairing and retaining the extracted characteristic words and word combinations, and removing incorrectly mined characteristic information;

[0008] The processed characteristic information of each industry is saved in the industry characteristic database;

[0009] Acquire the text content to be classified, and match and extract the feature information in the text content according to the feature information stored in the industry feature database;

[0010] Masking other irrelevant content fragments in the text content according to a preset probability;

[0011] According to the industry classification model, the processed text content is predicted to obtain the industry classification result of the text content;

[0012] Use the processed text content to train the industry classification model.

[0013] Optionally, the step of matching and extracting the feature information in the text content according to the feature information stored in the industry feature database specifically includes:

[0014] Loading characteristic information of each industry from the industry characteristic database;

[0015] Traverse the currently acquired text content character by character and extract the matched feature information.

[0016] Optionally, the step of masking other irrelevant content segments in the text content according to a preset probability specifically includes:

[0017] For characters that are not matched in the text content, they are randomly masked according to the probability of the following formula:

[0018] Pm=(1-1.0 / (a*log(len(Text) / (len(F)+1))+1))*(1-b*(Wl*Rl+Wr*Rr) / 2)

[0019] Among them, Pm represents the probability that the current character is masked, Text represents the input text content, len(Text) is the character length of the input text, F represents the concatenation of the character string of extracted feature information, len(F) is the cumulative length of these extracted characters, Wl represents the classification weight of the nearest feature word to the left of the current traversal position, Wr represents the classification weight of the nearest feature word to the right of the current traversal position, Rl and Rr represent the distance ratio of the currently traversed unmatched character position relative to the feature words on the left and right sides, a and b are adjustment factors, the value range of a is a non-negative integer from 0 to positive infinity, and the value range of b is [0,1].

[0020] Optionally, the method further comprises:

[0021] The predicted text content is used in combination with the predicted industry classification results to update the feature information in the industry feature database.

[0022] A second aspect of an embodiment of the present invention provides a text content classification device, the device comprising:

[0023] A data acquisition unit, used to acquire content data of various industries, and extract representative feature information of various industries through feature mining, wherein the feature information includes feature words and word combinations;

[0024] A tag repair unit receives externally input manual tag information, repairs and retains the extracted feature words and word combinations, and removes incorrectly mined feature information;

[0025] A database management unit, used for saving the processed characteristic information of each industry into an industry characteristic database;

[0026] A feature matching unit, used to obtain the text content to be classified, and to match and extract the feature information in the text content according to the feature information stored in the industry feature database;

[0027] A random masking unit, used to mask out other irrelevant content segments in the text content according to a preset probability;

[0028] A classification prediction unit is used to predict the processed text content according to the industry classification model to obtain the industry classification result of the text content;

[0029] The model training unit is used to train the industry classification model using the processed text content.

[0030] Optionally, the feature matching unit is specifically used to:

[0031] Loading characteristic information of each industry from the industry characteristic database;

[0032] Traverse the currently acquired text content character by character and extract the matched feature information.

[0033] Optionally, the random masking unit is specifically used to:

[0034] For characters that are not matched in the text content, they are randomly masked according to the probability of the following formula:

[0035] Pm=(1-1.0 / (a*log(len(Text) / (len(F)+1))+1))*(1-b*(Wl*Rl+Wr*Rr) / 2)

[0036] Among them, Pm represents the probability that the current character is masked, Text represents the input text content, len(Text) is the character length of the input text, F represents the concatenation of the character string of extracted feature information, len(F) is the cumulative length of these extracted characters, Wl represents the classification weight of the nearest feature word to the left of the current traversal position, Wr represents the classification weight of the nearest feature word to the right of the current traversal position, Rl and Rr represent the distance ratio of the currently traversed unmatched character position relative to the feature words on the left and right sides, a and b are adjustment factors, the value range of a is a non-negative integer from 0 to positive infinity, and the value range of b is [0,1].

[0037] Optionally, the device further comprises:

[0038] The information feedback unit is used to update the feature information in the industry feature database using the predicted text content in combination with the predicted industry classification result.

[0039] According to a third aspect of an embodiment of the present invention, there is provided an electronic device, characterized in that it includes:

[0040] One or more processors; a memory; one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method as described in the first aspect.

[0041] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein a program code is stored in the computer-readable storage medium, and the program code can be called by a processor to execute the method described in the first aspect.

[0042] In summary, the present invention provides a text content classification method, device, electronic device and storage medium, which splits the two logics of "capturing industry features in content" and "classifying content by industry" and independently optimizes and organically integrates the strategy to ensure that the accuracy and recall rate of content industry classification are taken into account in actual use scenarios. At the same time, between the above two functions, a certain manual annotation processing process is added to correct the representative features of each industry identified from the content, which effectively improves the accuracy of text content classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0044] Figure 1 A schematic diagram of an application scenario of the text content classification method and device according to an embodiment of the present invention;

[0045] Figure 2 A method flow chart of a text content classification method according to an embodiment of the present invention;

[0046] Figure 3 A method flow chart of a text content classification method according to another embodiment of the present invention;

[0047] Figure 4 A functional module block diagram of a text content classification device according to an embodiment of the present invention;

[0048] Figure 5 The present invention is a structural block diagram of an electronic device for executing the text content classification method according to an embodiment of the present application.

[0049] Figure 6 It is a structural block diagram of a computer-readable storage medium for storing or carrying program codes for implementing the text content classification method according to an embodiment of the present application.

[0050] icon:

[0051] Cloud server 100; client 200; data acquisition unit 110; annotation repair unit 120; database management unit 130; feature matching unit 140; random masking unit 150; classification prediction unit 160; model training unit 170; information feedback unit 180; electronic device 300; processor 310; memory 320; computer readable storage medium 400; program code 410. DETAILED DESCRIPTION

[0052] With the continuous rise of social media and self-media platforms, there are articles posted by various users every day on social media such as Weibo and WeChat, video websites such as Douyin and Bilibili, and social e-commerce platforms such as Xiaohongshu. For example, ordinary users’ daily experience and sharing of products in various industries, marketing promotion copywriting of brand owners in various industries, and endorsement and promotion of products by big V anchors. In many scenarios, it is necessary to extract and classify various updated text information to determine the industries involved in the relevant content and the brands mentioned. The characteristics of these online texts are that the content posted by ordinary users is mostly colloquial, casually expressed, covering many scenes and industry categories, and not constrained by grammatical rules; while some expression routines are often used in official marketing copywriting, such as a large section of artistic conception in the early stage of the content, pointing east and hitting west, and finally turning around to point out the brand to be promoted; or in a content of big V and anchors promoting products, it usually covers many industries and categories. These bring great difficulties to accurately judge the content. Therefore, how to accurately and effectively classify Internet text content according to industry is an urgent problem that needs to be solved.

[0053] The commonly used text classification scheme at present first annotates the corresponding category labels of the text content, captures the important feature fragments in the content through deep learning model training such as Bert, and builds a classification model to determine the industry to which the content belongs.

[0054] This method has many problems that affect the effect in the scenario of judging the industry and brand mentioned in the network content, especially in the scenario of network text content. Including:

[0055] Since there are many industries and brands that need to be classified (dozens of industries, tens of thousands of brands), a large amount of high-quality labeled data is usually required, and the labeling cost is relatively high;.

[0056] Since the same online text content often contains multiple topics, involves multiple industries and brands, and contains a large amount of interfering content, it is difficult to completely rely on training a black box deep model to lock in the industry features in the content and perform accurate classification;

[0057] In addition, from the perspective of business needs, it is also necessary to ensure the recall rate of the industry classification strategy to identify and extract as many industries and brands mentioned in the content as possible for use in subsequent business analysis processes.

[0058] In view of this, the designers of the present invention have designed a text content classification method, device, electronic device and storage medium, which splits the two logics of "capturing industry features in content" and "classifying content by industry" and independently optimizes and organically integrates the strategy to ensure that the accuracy and recall rate of content industry classification are taken into account in the actual use scenario. At the same time, between the above two functions, a certain manual annotation processing flow is added to correct the representative features of each industry identified from the content, which effectively improves the accuracy of text content classification.

[0059] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0060] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0061] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0062] In the description of the present invention, it should be noted that the terms "top", "bottom", "inside", "outside", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the invented product is usually placed when in use, which is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0063] In the description of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0064] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0065] Example

[0066] See also Figure 1 , a schematic diagram of an application scenario of a text content classification method and device provided in this embodiment.

[0067] like Figure 1 As shown, a text content classification method and device provided by the present invention can be applied to a cloud server 100, and the cloud server 100 is connected to a client 200 via the Internet or other means. When performing text content classification, the cloud server 100 obtains content data via the Internet or the client 200, and obtains manually annotated information via the client 200. The industry feature database for storing feature information can be set in the cloud server 100, or in other devices or terminals connected to the cloud server 100.

[0068] It should be noted that the text content classification method and device provided by the present invention can also be applied to local terminals other than the cloud server 100, such as PCs, smart phones, tablet computers, or other devices with data processing and data interaction functions.

[0069] On the basis of the above, if Figure 2 As shown, a text content classification method provided by an embodiment of the present invention includes:

[0070] Step S101, obtaining content data of various industries, and extracting representative feature information of various industries through feature mining, wherein the feature information includes feature words and word combinations.

[0071] In the cold start phase, a relatively small amount of content data of various industries can be accumulated first, and then the representative feature information of each industry can be extracted through feature mining. The feature mining method can refer to the patent application number 202010664165.4, or any other feature mining algorithm currently disclosed can be used.

[0072] Step S102, receiving externally input manual annotation information, repairing and retaining the extracted feature words and word combinations, and removing incorrectly mined feature information.

[0073] Considering that the scenarios usually used when the method is executed are relatively complex, feature mining algorithms alone cannot achieve good results, so manual labeling is introduced in the early stage. Analysts involved in manual labeling judge whether to inject the feature words marked by the algorithm into the library based on the relevance of the feature words to the industry and brand (such as whether it is an ingredient, efficacy, user needs and pain points, public opinion feedback, etc.), and manually label the mined feature information to indicate the subsequent processing method. Specific processing methods include: repair, retention, and removal of error information. Among them, repair means that the feature information needs to be modified before it can be stored in the library, retention means that the feature information can be directly stored in the library, and removal of error information means that the feature information cannot be stored in the library due to errors. After the manually labeled information is sent to the server through the client, the server receives the corresponding manually labeled information, and performs corresponding operations based on the labeled information to process the mined feature information.

[0074] Step S103, saving the processed characteristic information of each industry into the industry characteristic database.

[0075] The processed feature information is saved in the industry feature database as the feature attributes of the current industry and brand. As the method is continuously executed, the content data of each industry and the data in the database are continuously accumulated and updated, and the feature information saved becomes more complete and accurate in representing the corresponding industry.

[0076] Step S104, obtaining text content to be classified, and matching and extracting feature information in the text content according to feature information stored in the industry feature database.

[0077] After the database is established, the text content on the network begins to be classified. The first action is to match the feature information stored in the database with the text content to be classified and extract the feature information that matches successfully.

[0078] Step S105: mask out other irrelevant content segments in the text content according to a preset probability.

[0079] The unmatched content fragments in the text to be classified are masked and then handed over to the subsequent classification model for processing, so as to reduce the impact of irrelevant content in the text.

[0080] Step S106, predicting the processed text content according to the industry classification model to obtain the industry classification result of the text content.

[0081] It should be noted that the type of industry classification model used here is not the main innovation of the present invention, so the corresponding model that has been publicly disclosed, such as the currently popular open source deep model such as Bert, can be used, and a smaller network and parameter scale can be used to speed up the calculation efficiency and avoid overfitting. Through the prediction of the industry classification model, the industry to which the text content belongs can be output. Since the object predicted by the industry classification model is the text content that has been optimized for content segment masking, the model classification effect will be better.

[0082] Step S107, using the processed text content to train the industry classification model.

[0083] In addition to prediction, the processed text content can also be used as data to further train the industry classification model. The next time this method is executed to classify text content, the trained and optimized model can be used. With the above method, as the amount of text content processed gradually increases, the classification model is continuously iterated and optimized, so that the accuracy and recall rate of content classification are improved accordingly.

[0084] The text content classification method provided in this embodiment separates the two logics of "capturing industry features in content" and "classifying content by industry" and independently optimizes and organically integrates them to ensure that the accuracy and recall rate of content industry classification are taken into account in actual use scenarios. At the same time, a certain manual annotation processing process is added between the above two functions to correct the representative features of each industry identified from the content, which effectively improves the accuracy of text content classification.

[0085] like Figure 3 As shown, a text content classification method according to another embodiment of the present invention comprises:

[0086] Step S201, obtain content data of various industries, and extract representative feature information of each industry through feature mining, wherein the feature information includes feature words and word combinations.

[0087] Step S202, receiving external input manual annotation information, repairing and retaining the extracted feature words and word combinations, and removing incorrect feature information.

[0088] Step S203, save the processed characteristic information of each industry to the industry characteristic database.

[0089] Step S204, obtaining the text content to be classified, and loading the characteristic information of each industry from the industry characteristic database.

[0090] In this embodiment, when loading various industry features from the industry feature library, the AC automaton algorithm may be used to load various feature words so as to efficiently query and match from the content.

[0091] Step S205, traverse the currently acquired text content character by character, and extract the matched feature information.

[0092] In this embodiment, the currently acquired text content is traversed character by character, and industry feature words are matched and extracted. A maximum matching strategy can be used to avoid position overlap and included feature extraction results.

[0093] Step S206: For characters that are not matched in the text content, they are randomly masked according to the probability of the formula.

[0094] In this embodiment, the formula is:

[0095] Pm=(1-1.0 / (a*log(len(Text) / (len(F)+1))+1))*(1-b*(Wl*Rl+Wr*Rr) / 2)

[0096] Among them, Pm represents the probability of the current character being masked, Text represents the input text content, len(Text) is the character length of the input text, F represents the concatenation of the character string of the extracted feature information, and len(F) is the cumulative length of these extracted characters. The meaning of the first half of the formula is that the smaller the proportion of industry features extracted from the content, the greater the probability of hiding industry-irrelevant characters in the content to reduce interference with industry classification. The second half of the formula, as an option, can add the importance of feature words in the content to industry classification into the calculation results.

[0097] Wl represents the classification weight of the nearest feature word to the left of the current traversal position, and Wr represents the classification weight of the nearest feature word to the right of the current traversal position. The classification weights Wl and Wr here are calculated in the feature mining step, and can be calculated using information gain, chi-square distribution and other calculation methods. In the case of multiple categories, the average value can be taken and normalized to [0,1]. The closer to 1, the more important it is to the classification.

[0098] Rl and Rr represent the distance ratio of the currently traversed unmatched character position to the feature words on the left and right sides, respectively. For example, Rl = 1-ll / l, Rr = ll / l, where ll represents the character distance between the currently traversed invalid character position and the feature word on the left, l represents the character distance between the feature words on the left and right sides of the current position, and the value range of Rl and Rr is [0,1], Rl+Rr=1. If there is no feature word on one side, the feature weight and distance ratio of the corresponding side can be set to 0. It can be considered that the probability of whether the currently traversed invalid character needs to be retained also refers to the classification importance of the feature words on the left and right sides and the character distance. The higher the classification importance of the feature words on both sides, the higher the probability that the current invalid character is covered.

[0099] a and b are adjustment factors. The value range of a is a non-negative integer from 0 to positive infinity, and the value is usually 2; the value range of b is [0,1], which is used to control the influence of this part on the Pm result, and the value is usually 0.3.

[0100] For characters that are not in the feature extraction results, they are randomly masked according to the probability of the above formula, which can reduce the interference of irrelevant content on industry classification.

[0101] The following is a specific example to illustrate:

[0102] Taking the common situation of microblog content as an example, len(Text) is 200, len(F) is 5, a is 2, then Pm=0.88; that is, when traversing to the current irrelevant character, the calculated random floating point number between [0,1] needs to be greater than 0.88 to be retained and retained in the text in the original character form; otherwise, the current character is overwritten by the invalid character [UNK].

[0103] Add the second half of the formula, if Wl=0.8, Wr=0.5, Rl=0.3, Rr=1-Rl=0.7, b is 0.3, then the second half is calculated to be 0.91, and finally Pm=0.88*0.91=0.8. In addition, if there is no industry feature word in the current text, then the values ​​of len(F), Wl, Wr, Rl, and Rr in the above formula are all 0, then Pm=0.91.

[0104] Step S207: predict the processed text content according to the industry classification model to obtain the industry classification result of the text content.

[0105] Step S208: Use the processed text content to train the industry classification model.

[0106] Step S209, using the predicted text content in combination with the predicted industry classification result, the feature information in the industry feature database is updated.

[0107] The optimized industry classification results are output to subsequent data insight applications of various industries on the one hand, and can be fed back to the feature extraction process of step S201 on the other hand to iteratively optimize the mining effect of industry feature information. After the processing of step S202, the industry feature rule library is further enriched.

[0108] In order to address the problem that deep models find it difficult to accurately and fully lock on to the features of the required classification categories in scenarios with cluttered content, the method implemented in this paper first extracts representative features and combinations of industries from the content through the accumulated representative features and combinations of various industries through a rule matching strategy, and masks out other irrelevant content fragments according to a set probability, and then hands them over to subsequent classification models for processing, including the training and prediction stages of the model. The same feature extraction and random masking strategies are used to train industry classification models for the representative features and combinations in the content.

[0109] In summary, the text content classification method provided in this embodiment splits the two logics of "capturing industry features in content" and "classifying content by industry", and independently optimizes and organically integrates strategies to ensure that in actual use scenarios, both the accuracy and recall of content industry classification are taken into account. At the same time, between the above two functions, a certain manual annotation processing flow is added to correct the representative features of each industry identified from the content, which effectively improves the accuracy of text content classification. In view of the characteristics of network text content, a weight enhancement strategy for features that play an important role in classification is also added, and the impact of interfering content on classification results is reduced by masking by probability, while also improving the accuracy of data classification results and business interpretability.

[0110] like Figure 4 As shown, the text content classification device provided by the present invention comprises:

[0111] The data acquisition unit 110 is used to acquire content data of various industries and extract representative feature information of various industries through feature mining, wherein the feature information includes feature words and word combinations;

[0112] The annotation repair unit 120 receives the manual annotation information input from the outside, repairs and retains the extracted characteristic words and word combinations, and removes the characteristic information mined incorrectly;

[0113] The database management unit 130 is used to save the processed characteristic information of each industry into the industry characteristic database;

[0114] The feature matching unit 140 is used to obtain the text content to be classified, and match and extract the feature information in the text content according to the feature information stored in the industry feature database;

[0115] A random masking unit 150, used to mask out other irrelevant content segments in the text content according to a preset probability;

[0116] The classification prediction unit 160 is used to predict the processed text content according to the industry classification model to obtain the industry classification result of the text content;

[0117] The model training unit 170 is used to train the industry classification model using the processed text content.

[0118] As a preferred implementation of this embodiment, the feature matching unit 140 is specifically used for:

[0119] Loading characteristic information of each industry from the industry characteristic database;

[0120] Traverse the currently acquired text content character by character and extract the matched feature information.

[0121] As a preferred implementation of this embodiment, the random masking unit 150 is specifically used for:

[0122] For characters that are not matched in the text content, they are randomly masked according to the probability of the following formula:

[0123] Pm=(1-1.0 / (a*log(len(Text) / (len(F)+1))+1))*(1-b*(Wl*Rl+Wr*Rr) / 2)

[0124] Among them, Pm represents the probability that the current character is masked, Text represents the input text content, len(Text) is the character length of the input text, F represents the concatenation of the character string of extracted feature information, len(F) is the cumulative length of these extracted characters, Wl represents the classification weight of the nearest feature word to the left of the current traversal position, Wr represents the classification weight of the nearest feature word to the right of the current traversal position, Rl and Rr represent the distance ratio of the currently traversed unmatched character position relative to the feature words on the left and right sides, a and b are adjustment factors, the value range of a is a non-negative integer from 0 to positive infinity, and the value range of b is [0,1].

[0125] As a preferred implementation of this embodiment, the device also includes:

[0126] The information feedback unit 180 is used to update the feature information in the industry feature database using the predicted text content in combination with the predicted industry classification result.

[0127] The text content classification device provided in the embodiment of the present invention is used to implement the above-mentioned text content classification method, so the specific implementation method is the same as the above-mentioned method and will not be repeated here.

[0128] like Figure 5 As shown, a structural block diagram of an electronic device 300 provided in an embodiment of the present invention. The electronic device 300 may be an electronic device 300 capable of running applications, such as a smart phone, a tablet computer, an e-book, etc. The electronic device 300 in the present application may include one or more of the following components: a processor 310, a memory 320, and one or more applications, wherein the one or more applications may be stored in the memory 320 and configured to be executed by one or more processors 310, and the one or more programs are configured to execute the method described in the aforementioned method embodiment.

[0129] The processor 310 may include one or more processing cores. The processor 310 uses various interfaces and lines to connect various parts of the entire electronic device 300, and executes various functions and processes data of the electronic device 300 by running or executing instructions, programs, code sets or instruction sets stored in the memory 320, and calling data stored in the memory 320. Optionally, the processor 310 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 310 can integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 310, but may be implemented separately through a communication chip.

[0130] The memory 320 may include a random access memory (RAM) or a read-only memory (ROM). The memory 320 may be used to store instructions, programs, codes, code sets or instruction sets. The memory 320 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created by the terminal during use (such as a phone book, audio and video data, chat record data), etc.

[0131] like Figure 6 FIG. 4 is a block diagram of a computer-readable storage medium 400 provided in an embodiment of the present invention. The computer-readable storage medium stores program code 410, which can be called by a processor to execute the method described in the above method embodiment.

[0132] The computer readable storage medium 400 may be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk, or a ROM. Optionally, the computer readable storage medium 400 includes a non-transitory computer-readable storage medium. The computer readable storage medium 400 has storage space for program codes 410 that perform any of the method steps in the above method. These program codes 410 may be read from or written to one or more computer program products. The program codes 410 may be compressed, for example, in an appropriate form.

[0133] In summary, the present invention provides a text content classification method, device, electronic device and storage medium, which splits the two logics of "capturing industry features in content" and "classifying content by industry" and independently optimizes and organically integrates strategies to ensure that in actual use scenarios, the accuracy and recall rate of content industry classification are taken into account. At the same time, between the above two functions, a certain manual annotation processing flow is added to correct the representative features of each industry identified from the content, which effectively improves the accuracy of text content classification. In view of the characteristics of network text content, a weight enhancement strategy for features that play an important role in classification is also added, and the impact of interfering content on classification results is reduced by masking by probability, while also improving the accuracy of data classification results and business interpretability.

[0134] In several embodiments disclosed in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and a module, a program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0135] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0136] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. A text content classification method, characterized in that: The method comprises: Acquire content data of various industries, and extract representative feature information of each industry through feature mining, wherein the feature information includes feature words and word combinations; Receiving externally input manual annotation information, repairing and retaining the extracted characteristic words and word combinations, and removing incorrectly mined characteristic information; The processed characteristic information of each industry is saved in the industry characteristic database; Acquire the text content to be classified, and match and extract the feature information in the text content according to the feature information stored in the industry feature database; Masking other irrelevant content fragments in the text content according to a preset probability; According to the industry classification model, the processed text content is predicted to obtain the industry classification result of the text content; Use the processed text content to train the industry classification model.

2. The text content classification method according to claim 1, characterized in that: The step of matching and extracting the feature information in the text content according to the feature information stored in the industry feature database specifically includes: Loading characteristic information of each industry from the industry characteristic database; Traverse the currently acquired text content character by character and extract the matched feature information.

3. The text content classification method according to claim 2, characterized in that: The step of masking other irrelevant content segments in the text content according to a preset probability specifically includes: For characters that are not matched in the text content, they are randomly masked according to the probability of the following formula: Pm=(1-1.0 / (a*log(len(Text) / (len(F)+1))+1))*(1-b*(Wl*Rl+Wr*Rr) / 2) Among them, Pm represents the probability that the current character is masked, Text represents the input text content, len(Text) is the character length of the input text, F represents the concatenation of the character string of extracted feature information, len(F) is the cumulative length of these extracted characters, Wl represents the classification weight of the nearest feature word to the left of the current traversal position, Wr represents the classification weight of the nearest feature word to the right of the current traversal position, Rl and Rr represent the distance ratio of the currently traversed unmatched character position relative to the feature words on the left and right sides, a and b are adjustment factors, the value range of a is a non-negative integer from 0 to positive infinity, and the value range of b is [0,1].

4. The text content classification method according to claim 3, characterized in that: The method further comprises: The predicted text content is used in combination with the predicted industry classification results to update the feature information in the industry feature database.

5. A text content classification device, characterized in that: The device comprises: A data acquisition unit, used to acquire content data of various industries, and extract representative feature information of various industries through feature mining, wherein the feature information includes feature words and word combinations; A tag repair unit receives externally input manual tag information, repairs and retains the extracted feature words and word combinations, and removes incorrectly mined feature information; A database management unit, used for saving the processed characteristic information of each industry into an industry characteristic database; A feature matching unit, used to obtain the text content to be classified, and to match and extract the feature information in the text content according to the feature information stored in the industry feature database; A random masking unit, used to mask out other irrelevant content segments in the text content according to a preset probability; A classification prediction unit is used to predict the processed text content according to the industry classification model to obtain the industry classification result of the text content; The model training unit is used to train the industry classification model using the processed text content.

6. The text content classification device according to claim 5, characterized in that: The feature matching unit is specifically used for: Loading characteristic information of each industry from the industry characteristic database; Traverse the currently acquired text content character by character and extract the matched feature information.

7. The text content classification device according to claim 6, characterized in that: The random masking unit is specifically used for: For characters that are not matched in the text content, they are randomly masked according to the probability of the following formula: Pm=(1-1.0 / (a*log(len(Text) / (len(F)+1))+1))*(1-b*(Wl*Rl+Wr*Rr) / 2) Among them, Pm represents the probability that the current character is masked, Text represents the input text content, len(Text) is the character length of the input text, F represents the concatenation of the character string of extracted feature information, len(F) is the cumulative length of these extracted characters, Wl represents the classification weight of the nearest feature word to the left of the current traversal position, Wr represents the classification weight of the nearest feature word to the right of the current traversal position, Rl and Rr represent the distance ratio of the currently traversed unmatched character position relative to the feature words on the left and right sides, a and b are adjustment factors, the value range of a is a non-negative integer from 0 to positive infinity, and the value range of b is [0,1].

8. The text content classification device according to claim 7, characterized in that: The device also includes: The information feedback unit is used to update the feature information in the industry feature database using the predicted text content in combination with the predicted industry classification result.

9. An electronic device, characterized in that: include: one or more processors; Memory; One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • New keyword mining method and device and electronic equipment

    CN111898010A

  • Industry categorizing method and system of text, computer equipment and storage medium

    CN108520041A

  • Text classification method, system, computer equipment and storage medium

    CN108536800A