Text processing method, device, equipment and storage medium

By extracting and splicing text features and theme features, using BERT and LDA models to predict and classify text topics, the problem of low manual review efficiency is solved and efficient text classification is achieved.

CN114020921BActive Publication Date: 2025-09-05CHENGDU SHULIANYUNSUAN TECH CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111550733.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-17
Publication Date
2025-09-05
Estimated Expiration
2041-12-17

AI Technical Summary

Technical Problem

The existing text review methods rely on manual processing, resulting in high labor costs and low efficiency, making it difficult to efficiently classify text.

Method used

By extracting the text features and theme features of the text to be classified, performing splicing processing, and using the BERT model and the LDA theme model for text theme prediction and classification.

Benefits of technology

Improve the accuracy of text topic prediction and text classification accuracy, reduce labor costs, and improve efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114020921B_ABST
    Figure CN114020921B_ABST
Patent Text Reader

Abstract

The present application discloses a text processing method, apparatus, device, and storage medium, wherein the method includes: extracting text features of a text to be classified to obtain text features corresponding to the text to be classified; extracting text topic features of the text to be classified based on the text features to obtain topic features corresponding to the text to be classified; splicing the text features and topic features, and predicting the text topic of the text to be classified based on the splicing results to obtain the text topic corresponding to the text to be classified; and classifying the text to be classified based on the text topic to determine the category to which the text to be classified belongs. The present application simplifies the text review process and improves the accuracy of text classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a text processing method, apparatus, device, and storage medium. Background Art

[0002] In the context of scarce natural resources, the traditional energy industry faces many challenges, such as how to save costs, how to increase production capacity, and how to prevent potential problems. To prevent potential problems, it is usually necessary to review some business data to promptly discover and resolve them. Taking oilfield safety generation and daily management as an example, it is usually necessary to review the relevant texts of safety generation and daily management. Current text review is still at the stage of manually reviewing and classifying issues. Due to the large volume, complex types, and diverse sources of the problem texts to be reviewed, purely manual processing methods have problems such as high labor costs and efficiency. Therefore, in the field of text review, how to perform efficient text classification has become one of the hot issues in current research. Summary of the Invention

[0003] The embodiments of the present application provide a text processing method, apparatus, device and storage medium to improve the accuracy of text classification.

[0004] In one aspect, an embodiment of the present application provides a text processing method, comprising:

[0005] Perform text feature extraction on the text to be classified to obtain text features corresponding to the text to be classified;

[0006] Extracting text theme features from the text to be classified based on the text features to obtain theme features corresponding to the text to be classified;

[0007] The text features and the topic features are spliced ​​together, and a text topic prediction is performed on the text to be classified based on the splicing result to obtain a text topic corresponding to the text to be classified;

[0008] The text to be classified is classified based on the text theme to determine the category to which the text to be classified belongs.

[0009] In one aspect, an embodiment of the present application further provides a text processing device, comprising:

[0010] An extraction unit, configured to extract text features from the text to be classified, and obtain text features corresponding to the text to be classified;

[0011] The extraction unit is further configured to extract text topic features from the text to be classified based on the text features to obtain topic features corresponding to the text to be classified;

[0012] A splicing unit, configured to splice the text features and the topic features;

[0013] A prediction unit, configured to perform text topic prediction on the text to be classified based on the splicing processing result, to obtain a text topic corresponding to the text to be classified;

[0014] The determining unit is configured to perform text classification on the text to be classified based on the text theme, and determine the category to which the text to be classified belongs.

[0015] In one aspect, an embodiment of the present application provides a text processing device, comprising: a processor adapted to execute one or more computer programs; and a computer storage medium storing one or more computer programs, wherein the one or more computer programs are adapted to be loaded and executed by the processor.

[0016] Performing text feature extraction on the text to be classified to obtain text features corresponding to the text to be classified; performing text topic feature extraction on the text to be classified based on the text features to obtain topic features corresponding to the text to be classified; performing splicing processing on the text features and the topic features, and performing text topic prediction on the text to be classified based on the splicing processing result to obtain a text topic corresponding to the text to be classified; performing text classification on the text to be classified based on the text topic to determine the category to which the text to be classified belongs.

[0017] In one aspect, an embodiment of the present application provides a computer storage medium, wherein the computer storage medium stores a computer program, and when the computer program is executed by a processor, is configured to perform:

[0018] Performing text feature extraction on the text to be classified to obtain text features corresponding to the text to be classified; performing text topic feature extraction on the text to be classified based on the text features to obtain topic features corresponding to the text to be classified; performing splicing processing on the text features and the topic features, and performing text topic prediction on the text to be classified based on the splicing processing result to obtain a text topic corresponding to the text to be classified; performing text classification on the text to be classified based on the text topic to determine the category to which the text to be classified belongs.

[0019] In one aspect, an embodiment of the present application provides a computer program product or a computer program. The computer program product includes a computer program stored in a computer storage medium. A processor of a text processing device reads the computer program from the computer storage medium and executes the computer program, causing the text processing device to perform:

[0020] Performing text feature extraction on the text to be classified to obtain text features corresponding to the text to be classified; performing text topic feature extraction on the text to be classified based on the text features to obtain topic features corresponding to the text to be classified; performing splicing processing on the text features and the topic features, and performing text topic prediction on the text to be classified based on the splicing processing result to obtain a text topic corresponding to the text to be classified; performing text classification on the text to be classified based on the text topic to determine the category to which the text to be classified belongs.

[0021] In an embodiment of the present application, when classifying a text to be classified, the text features of the text to be classified are first extracted, and then the text theme features of the text to be classified are extracted based on the text features to obtain the theme features corresponding to the text to be classified. Finally, the text features and the theme features are spliced ​​together, and a theme prediction is performed on the text to be classified based on the splicing results, thereby determining the text theme corresponding to the text to be classified based on the predicted text theme. It should be understood that text features can be used to reflect the semantic features of the text to be classified, and splicing and fusing text features and theme features effectively combines the semantic information of the text and obtains a better theme representation. When performing theme prediction based on the splicing results, the accuracy of theme prediction can be improved, thereby also providing the accuracy of text classification based on text themes. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0023] Figure 1 This is a structural diagram of a text processing system provided in an embodiment of the present application;

[0024] Figure 2 This is a flowchart of a text processing method provided by an embodiment of the present application;

[0025] Figure 3 This is a flowchart of another text processing method provided in an embodiment of the present application;

[0026] Figure 4 This is a schematic diagram of an LDA topic model generated text provided by the present application;

[0027] Figure 5 This is a structural diagram of a text classification model provided in an embodiment of the present application;

[0028] Figure 6This is a flowchart of a text processing device provided by an embodiment of the present application;

[0029] Figure 7 It is a schematic diagram of the structure of a text processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.

[0031] An embodiment of the present application provides a text processing solution that can be used to classify text to be classified. In a specific implementation, the text features of the text to be classified are first extracted, and then the topic features of the text to be classified are extracted based on the text features. Furthermore, the text features and the topic features are spliced ​​together, and text topic prediction is performed based on the splicing results to determine the topic corresponding to the text to be classified; finally, the text to be classified is classified based on the topic corresponding to the text to be classified to determine the type of the text to be classified.

[0032] The text processing solution can be executed by a text processing device, which can be a terminal, such as

[0033] Smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, car terminals, smart home appliances, smart voice interaction devices, etc.; or, the text processing device can also be a server, such as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0034] Alternatively, the text processing solution can also be jointly executed by a text processing device and a text management server, such as a text processing device obtaining a text to be classified, and transmitting the text to be classified to a text management server, and the text management server performs the steps of text feature extraction, topic feature extraction, text topic prediction, and determining the category to which the text to be classified belongs. For another example, a text processing device obtains a text to be classified, and after transmitting the text to be classified to a text management server, the text management server performs text feature extraction, topic feature extraction, and text topic prediction, and returns the predicted text topic to the text processing device, and the text processing device performs classification processing on the text to be classified based on the predicted text topic. It should be understood that the above are only two feasible implementation methods listed in the embodiments of the present application. In specific applications, depending on different products and different application scenarios, it can be set which device performs the text processing solution.

[0035] In one embodiment, if the text processing method is performed by a text processing device and a text management server, then the embodiment of the present application can provide a text processing system, see Figure 1 , is a structural diagram of a text processing system provided in an embodiment of the present application. Figure 1 The text processing system may include a text processing device 101 and a text management server 102 , and the text processing device 101 and the text management server 102 are connected via a wired or wireless manner.

[0036] The text processing device 101 obtains the text to be classified and transmits the text to be classified to the text management server 102. After receiving the text to be classified, the text management server 102 performs text feature extraction on the text to be classified to obtain text features corresponding to the text to be classified. Furthermore, based on the text features, the text topic features of the text to be classified are extracted to obtain topic features corresponding to the topic of the text to be classified.

[0037] The text management server 102 then combines the text features and the topic features and, based on the combined results, predicts the text topic of the text to be classified, obtaining the text topic corresponding to the text to be classified. The text management server 102 can return the text topic corresponding to the text to be classified to the text processing device 101. The text processing device 101 then classifies the text to be classified based on the text topic and determines the category to which the text to be classified belongs.

[0038] Alternatively, the text management server 102 classifies the text to be classified based on the text's subject matter, determines the category to which the text to be classified belongs, and then notifies the text processing device 101 of the category to which the text to be classified belongs. Text features can be used to reflect the semantic features of the text to be classified. By combining and fusing text features with subject features, this effectively combines the semantic information of the text and obtains a more optimal subject representation. Furthermore, when subject prediction is performed based on the combined processing results, the accuracy of subject prediction can be improved, thereby also improving the accuracy of text classification based on text subject matter.

[0039] Based on the above text processing solution and text processing system, the present application embodiment provides a text processing method, see Figure 2 , which is a flow chart of a text processing method provided in an embodiment of the present application. Figure 2 The text processing method can be executed by a text processing device, specifically by a processor of the text processing device. Figure 2 The text processing method may include the following steps:

[0040] Step S201: extract text features from the text to be classified to obtain text features corresponding to the text to be classified.

[0041] Among them, the text to be classified usually refers to that related to a certain regulation or specification. In other words, the text to be classified is used to record some content that violates or complies with a certain regulation or specification. For example, the text to be classified can be some short texts generated in petroleum engineering. These short texts can be used to record texts for safe generation and daily management in petroleum engineering. Such short texts may be related to a certain regulation or specification in petroleum engineering. For example, the content recorded in a short text violates a certain specification or regulation; or, the text to be classified can also be any type of text, for example, the text to be classified can be text used to record the operation log of a certain application, etc.

[0042] The text features corresponding to the text to be classified can be used to reflect the semantic information of the text to be classified. The semantic information of a text serves as the basic reference for subsequent processing of the text. Text feature extraction for the text to be classified can be performed by calling a text feature processing model. Currently, most text feature processing models are based on recurrent neural networks (RNNs). However, simple RNNs suffer from the vanishing / exploding gradient problem, making them unable to model long-term contextual dependencies. Compared to traditional RNNs, LSTMs incorporate a gating mechanism that enables more efficient feature extraction and training, addressing the vanishing and exploding gradient issues. LSTMs are a special type of RNN. The difference between them is that a single recurrent structure in a conventional RNN has only one state, while a single recurrent structure (also known as a cell) in an LSTM has four states. Compared to traditional RNNs, LSTMs maintain a persistent cell state that is continuously passed down between recurrent structures to determine which information to forget or pass on. However, because LSTMs rely on the previous computational results during sequential processing, they suffer from low parallel computing efficiency and slow model execution.

[0043] The text feature processing models (also called language models) that are currently flourishing in various natural language processing businesses, such as the GPT model and the Transformer bidirectional encoder (Bidirectional Encoder Representations from Transformers, BERT) model, are all based on the Transformer model. Among them, the BERT model is different from other language representation models. BERT aims to pre-train deep bidirectional representations by jointly adjusting the context in all layers. Therefore, the pre-trained BERT model representation can be fine-tuned through an additional output layer, which is suitable for the construction of state-of-the-art models for a wide range of tasks. Based on the advantages of the BERT model, the text feature processing model used for text extraction in the embodiments of the present application can use the BERT model.

[0044] Step S202: extracting text topic features from the text to be classified based on the text features to obtain topic features corresponding to the text to be classified.

[0045] In order to accurately classify the text to be classified, it is necessary not only to base the classification on the semantic information of the text to be classified, but also to mine the subject features of the text to be classified. In the embodiment of the present application, the subject features corresponding to the text to be classified can be obtained by feature extraction based on text features. In a specific implementation, the subject features corresponding to the text to be classified are executed by calling a text subject processing model, and the text features extracted in step S201 are input into the text subject processing model so that the text subject processing model extracts the text subject through the text features, thereby obtaining the subject features corresponding to the text to be classified.

[0046] A text topic processing model can be a Latent Dirichlet Allocation (LDA) topic model. The LDA topic model is a document generation model that assumes that a text or article has multiple topics, each corresponding to different words. During the construction of a text or article, a topic is first selected with a certain probability, and then a word is selected within this topic with a certain probability, thus generating the first word of the text or article. This process is repeated continuously to generate a text or article. The use of the LDA topic model is the reverse process of this document generation process. It searches for a topic within a text or article and finds the words corresponding to these topics.

[0047] The LDA topic model is an unsupervised machine learning model that uses the bag-of-words model. A trained LDA topic model can be used to extract topic features based on the text features of the text to be classified, obtaining the topic features corresponding to the text to be classified.

[0048] Step S203: splicing the text features and the topic features, and performing text topic prediction on the text to be classified based on the splicing result to obtain the text topic corresponding to the text to be classified.

[0049] Optionally, step S203 can be performed by calling a text classification model. In a specific implementation, the text classification model may include a fully connected layer and a classification layer. The fully connected layer is called to concatenate text features and topic features to obtain a concatenated result of a target length. Furthermore, the classification layer is called to perform classification prediction based on the concatenated result using a classification function to obtain a text topic distribution corresponding to the text to be classified, and the text topic corresponding to the text to be classified is determined based on the text topic distribution.

[0050] The target length is the same as the preset number of text topics. For example, if the preset number of text topics is 5, then the target length is 5. The output text topic distribution can include the probability of each text topic in multiple text topics. Here, the number of text topics is the same as the target length, assuming it is expressed as N. The target length and the number of text topics are both N, where N is a positive integer greater than or equal to 1. The text topic distribution includes the probability corresponding to the i-th text topic, where i is a positive integer greater than or equal to 1 and less than or equal to N.

[0051] In one embodiment, determining the text topic corresponding to the text to be classified based on the text topic distribution may include: determining the text topic with the highest probability in the text topic distribution as the text topic corresponding to the text to be classified. For example, if the text topic distribution includes a first text topic and a second text topic, and the probability of the first text topic corresponding to the text to be classified is greater than the probability of the second text topic corresponding to the text topic, then the first text topic is determined as the text topic corresponding to the text to be classified.

[0052] In another embodiment, determining the text topic corresponding to the text to be classified based on the text topic distribution may include: determining the text topic whose probability meets the probability threshold in the text topic distribution as the text topic corresponding to the text to be classified. The probability threshold may be pre-set, the probability threshold may be a specific value, and a certain probability meeting the probability threshold may refer to a certain probability being greater than or equal to a certain probability threshold; or, the probability threshold may also be a range, and a certain probability meeting the probability threshold may refer to a certain probability falling within the range required by the probability threshold. Optionally, if there are multiple probabilities that meet the probability threshold, the text topic with the larger probability may be determined as the text topic corresponding to the text to be classified. Alternatively, any text topic whose probability meets the probability threshold may be randomly determined as the text topic corresponding to the text to be classified. Here, the embodiment of the present application only lists two feasible implementation methods. In specific applications, due to different product forms and different application scenarios, the implementation method of determining the text topic corresponding to the text to be classified based on the text topic distribution can be determined based on the product form and application scenario, and the embodiment of the present application does not make specific limitations.

[0053] Step S204: classify the text to be classified based on the text theme to determine the category to which the text to be classified belongs.

[0054] After determining the text theme corresponding to the text to be classified, the text to be classified can be further classified based on the text theme in step S204 to determine the category to which the text to be classified belongs. In a specific implementation, a correspondence between text themes and audit specifications can be pre-set. After determining the text theme corresponding to the text to be classified, the correspondence between the pre-set text themes and audit specifications is obtained, and based on the correspondence, a target audit specification that matches the text theme corresponding to the text to be classified is determined; and the text to be classified is classified as a text that violates the target audit specification.

[0055] For example, if the text to be classified is a text related to petroleum engineering, the category to which the text belongs can be understood as determining which audit specification the text violates. There is a direct relationship between the corresponding audit specification and the text subject. For example, if the audit specification "Petroleum Enterprise Site Safety Inspection Specification: Downhole Operations" corresponds to the text subject "Contractors and / or Suppliers," then if the text subject to be classified is determined to be contractors and / or suppliers, then based on this correspondence, the category to which the text to be classified belongs can be determined to be a violation of the audit specification "Petroleum Enterprise Site Safety Inspection Specification: Downhole Operations."

[0056] In an embodiment of the present application, when classifying a text to be classified, the text features of the text to be classified are first extracted, and then the text theme features of the text to be classified are extracted based on the text features to obtain the theme features corresponding to the text to be classified. Finally, the text features and the theme features are spliced ​​together, and a theme prediction is performed on the text to be classified based on the splicing results, thereby determining the text theme corresponding to the text to be classified based on the predicted text theme. It should be understood that text features can be used to reflect the semantic features of the text to be classified, and splicing and fusing text features and theme features effectively combines the semantic information of the text and obtains a better theme representation. When performing theme prediction based on the splicing results, the accuracy of theme prediction can be improved, thereby also providing the accuracy of text classification based on text themes.

[0057] Based on the above text processing method embodiment, the present application embodiment provides another text processing method. Figure 3 , which is a flow chart of another text processing method provided in an embodiment of the present application. Figure 3 The text processing method can be executed by a text processing device, specifically by a processor of the text processing device. Figure 3 The text processing method may include the following steps:

[0058] Step S301: pre-process the text to be classified, and call a text feature processing model to extract features of the pre-processed text to be classified to obtain text features corresponding to the text to be classified.

[0059] Optionally, preprocessing the text to be classified may include word segmentation and stop word removal. In a specific implementation, it includes: performing word segmentation on the text to be classified to obtain a word set corresponding to the text to be classified, where the word set includes one or more characters or words; performing stop word removal on the word set to obtain the feature words included in the text to be classified. Among them, the purpose of word segmentation is to split the Chinese character sequence in the text to be classified into individual independent words. In the embodiments of the present application, Jieba can be used as a word segmentation tool. For example, if the text to be classified is expressed as "Fighting across the seas for a victory today, I will not be defeated again", after using the Jieba word segmentation tool to segment it, the obtained word set is ("Fighting", "across the seas", "only", "for", "today", "a victory", ",", "I", "will not", "be defeated again", ".").

[0060] Stop words generally refer to words with a very high frequency of occurrence but little actual meaning. For example, common stop words may include "of", "in", "and", etc. Optionally, in the embodiments of the present application, a stop word list can be preset according to experience. The stop word list includes multiple stop words. Performing stop word removal on the word set to obtain the feature words included in the text to be classified may include: removing the words that appear in the stop word list from the word set, and the remaining words are used as the feature words included in the text to be classified.

[0061] Because stop words generally have little actual meaning, if stop words are also used as feature words of the text to be classified for analysis, it will increase the time for text feature extraction, thereby reducing the efficiency of text feature extraction. After stop word removal from the text to be classified, the number of feature words included in the text to be classified is significantly reduced, which can reduce the time for text feature extraction, significantly reduce the operation time of the text feature processing model, and thus improve the efficiency of text feature extraction.

[0062] In one embodiment, feature extraction is performed on the preprocessed text to be classified to obtain the text features corresponding to the text to be classified, which may include: S1: performing vector embedding processing on the text to be classified to obtain the word vectors corresponding to each feature word in the text to be classified; S2: performing linear transformation processing on the word vectors corresponding to each feature word, and performing attention weight calculation on each word vector after the linear transformation processing to obtain the text features of the text to be classified.

[0063] In the specific implementation, feature extraction of the pre-processed text to be classified is performed by calling the text feature processing model. Optionally, the text feature processing model may include a vector embedding layer and a Transformer encoder layer. The above s1 is performed by calling the embedding layer in the text feature processing model. Specifically, the word vector of each feature word contains the word vector of each feature word, the text vector of the text to be classified, and the position vector of each word in the text to be classified. For any feature word, assuming that the word vector is represented as , the text vector of the text to be classified is represented as , and the position vector of any feature word in the text to be classified is represented as , then after the embedding layer performs vector embedding processing on the classified text, the word vector mapped to any feature word is represented as .

[0064] After obtaining the word vector corresponding to each feature word, the word vector corresponding to each feature word is linearly transformed through the above step s2, and the attention weight operation is performed on each word vector after the linear transformation to obtain the text features of the text to be classified. In a specific implementation, step s2 can be executed by calling the Transformer encoder layer in the text feature processing model. Transformer can be a bidirectional encoder. In order to enable the text feature processing model to learn more information, the Transformer encoder layer connects the multi-head mechanism and the feedforward layer through the residual network result. The multi-head mechanism performs multiple linear transformations on the input vector to obtain different linear values, and then calculates the attention weight. The attention weight calculation formula is as follows:

[0065] MultiHead ( Q,K,V ) = Concat ( head 1 , head 2 ,…, head h )W o (1)

[0066] head f = Attention (2)

[0067] In formula (1), Q, K, and V are the input word vector matrices, Attention represents the attention weight operation, and Concat represents the concatenation process. is the fth hyperparameter head, f is an integer greater than or equal to 1 and less than or equal to h, It can be expressed as , Expressed as a weight matrix, , as well as are the weight matrices corresponding to the fth hyperparameter head. From the above formulas (1) and (2), it can be seen that the principle of performing linear transformation on the word vector corresponding to each feature word and performing attention weight operation on each word vector after linear transformation is: Q, K, and V are mapped through the corresponding weight matrix and then the Attention operation is performed. After repeating h times, the calculation results are concatenated to obtain the text features of the text to be classified. The text features of the text to be classified reflect the semantic relationship and grammatical structure of the text to be classified.

[0068] Step S302: Call the text topic processing model to extract text topic features from the text to be classified based on the text features, and obtain topic features corresponding to the text to be classified.

[0069] As can be seen from the above, the text topic processing model can be an LDA topic model. Before introducing step S302, we first introduce the relevant knowledge of the LDA topic model. The LDA topic model is a three-layer Bayesian probability generation model of "document-topic-word". It models the text as a probability distribution on a mixed topic through the generation process of the model text. The modeling process of the LDA topic model can be as follows: Figure 4 As shown. Figure 4 In the example, suppose a text set includes m texts, and the text set is represented as ,in, Represents the i-th text in the text set, where i is an integer greater than or equal to 1 and less than or equal to m; assuming that the i-th text contains n words, then the i-th text can be represented as , assuming Each component in the text represents the topic word vector corresponding to each text in the text set, specifically Indicates the topic corresponding to each word in the i-th text, Given hyperparameters, the process of generating the i-th text by the LDA topic model is:

[0070] 1) For a given i-th text ,according to The obedience parameter is Dirichlet distribution Determine a topic distribution , in simple terms, it is to get the topic distribution of the i-th text;

[0071] 2) For the nth word in the i-th text , according to z obey Multinomial distribution of ,for Determine a topic number ;

[0072] 3) According to The obedience parameter is Dirichlet distribution , determine a topic-word distribution matrix and word distribution ;

[0073] 4) According to the word obey Multinomial distribution of Generator ;

[0074] 5) Traverse all words in the i-th text and repeat steps 2)-4) to generate the i-th text ;

[0075] 6) Traverse all texts and repeat steps 1)-5) to generate the entire text set D.

[0076] When training the LDA topic model, the above text set can be a sample text set for training. When training the LDA topic model, after the entire text set is generated through the above steps 1)-6), the text set is then Perform Gibbs Sampling to estimate the above parameters and , the specific training process can be as follows:

[0077] 1) Random initialization, that is, randomly assigning a topic to each word in each text;

[0078] 2) Repeatedly scan the entire text set and re-extract the topic of each feature word using Gibbs sampling;

[0079] 3) Repeat the above sampling process until Gibbs sampling converges;

[0080] 4) Calculate the topic-word co-occurrence frequency matrix of the entire text set;

[0081] After training, the topic-word distribution is obtained according to the topic-word co-occurrence frequency matrix , count the frequency distribution of the topics contained in each text, and get the topic distribution of the text set .

[0082] After the above steps, the LDA topic model is trained. and is stable.

[0083] It should be understood that the above is only a method for training an LDA topic model proposed in an embodiment of the present application. In practical applications, in addition to using Gibbs sampling to calculate the topic distribution, the word vector output by the BERT model can also be used as the input of the LDA model, and the BERT model can be used instead of the Gibbs sampling algorithm to calculate the topic distribution.

[0084] In an embodiment of the present application, a trained LDA topic model is called to extract text topic features from the text to be classified based on text features, thereby obtaining topic features corresponding to the text to be classified. In a specific implementation, characteristic words included in the text features are obtained, and a text topic number is assigned to each characteristic word; Gibbs sampling is performed on the text topic number of each characteristic word. When Gibbs sampling converges, the text topic distribution in the text to be classified is counted, and the text topic distribution is determined as the topic feature of the text to be classified.

[0085] Step S303: calling the fully connected layer in the text classification model to perform splicing processing on the text features and the topic features to obtain the splicing processing result; and calling the classification layer in the text classification model to use the classification function to perform classification prediction based on the splicing processing result to obtain the text topic distribution corresponding to the text to be classified.

[0086] In one embodiment, the text topic processing model and the text feature processing model can both be deployed in the text classification model, and the text feature processing model is connected to the text topic processing model. In this way, the text features extracted by the text feature processing model can be transmitted to the text topic processing model. Optionally, the text features extracted by the text feature processing model (which can be represented as Feature tensor) can pass through an embedding layer before being transmitted to the text topic processing model. The embedding layer can be mapped to a low-dimensional word vector through one-hot encoding. As can be seen from the above, the text classification model also includes a fully connected layer and a classification layer. Based on this, the embodiment of the present application provides a text classification model, see Figure 5 , is a structural diagram of a text classification model provided in an embodiment of the present application. Figure 5 501 represents a text feature processing model, which can be a BERT model, and can specifically include an embedding layer and a TransformerIncoder encoder layer. Figure 5As described in the article, the text features output by the BERT model are processed into low-dimensional word vectors through an embedding layer and then used as the input of the LDA topic model. The LDA topic extracts text topic features based on the input and outputs the topic features of the text to be classified (which can also be expressed as a Feature tensor). Finally, the text feature tensor output by the BERT model and the topic feature tensor output by the LDA topic model are concatenated and fused through a fully connected layer. The concatenated result is then classified and predicted through a SoftMax layer. Finally, the text topic of the text to be classified is output, that is, the topic classification of the text to be classified is output.

[0087] Outputting the text topic of the text to be classified may specifically include: first determining the topic distribution, and then determining the text topic of the text to be classified according to the topic distribution in step S304. Figure 2 The description of step S203 in the embodiment will not be repeated here.

[0088] In one embodiment, the text classification model can be pre-trained based on a corpus, which can include multiple training texts and topic tags corresponding to each training text. During training, the corpus can be input into the text classification model, and the initial learning rate LR, the batch size of each training text input, the dropout rate and the number of epoch training times can be set. The Adam optimizer is used to dynamically adjust the learning rate to accelerate model convergence. During training, the confusion matrix is ​​obtained based on the model prediction results and the F1-score and AUC indicators are calculated as evaluation indicators for testing the model effect. Among them, F1-score can also be called F1 score, which is an indicator used in statistics to measure the accuracy of a binary classification model. It takes into account both the precision and recall of the classification model. The F1 score can be regarded as a harmonic average of the model precision and recall rate, with a maximum value of 1 and a minimum value of 0. The F1 score can be expressed by the following formula (3):

[0089] (3)

[0090] Among them, precision represents the accuracy rate, and its calculation formula is: , recall represents the recall rate, which is calculated as follows TP (true positive) means that the positive example is correctly classified as a positive example, FN (false negative) means that the positive example is incorrectly classified as a negative example, TN (true negative) means that the negative example is correctly classified as a negative example, and FP (false positive) means that the negative example is incorrectly classified as a positive example. In the embodiment of the present application, correctly classifying a positive example as a positive example and correctly classifying a negative example as a negative example can mean that: for a training text, the predicted text topic is the same as the topic label corresponding to the training text; incorrectly classifying a positive example as a negative example and incorrectly classifying a negative example as a positive example can mean that: for a training text, the predicted text topic is different from the topic label corresponding to the training text. The larger the value of the F1 score, the higher the precision and recall rate of the model, and the more accurate the text classification model is when performing text classification. AUC (Area under Curve) is an indicator to measure the quality of a learner. It refers to the area under the ROC curve and has a value between 0 and 1. Similarly, the larger the value of AUC, the better the training of the text classification model.

[0091] Since the cross entropy loss function makes up for the defect that the derivative form of the sigmoid function is prone to saturation, and at the same time introduces Softmax as the prediction result, it has a good performance in the classification task, so the cross entropy loss function cross-entropy error is used as the loss function of the training text classification model in the embodiment of the present application. The principle of using the cross entropy loss function to train the text classification model can be: calling the text feature processing model in the text classification model to extract text features from the training text to obtain text features, and then calling the text topic processing model to extract text topic features based on the text features to obtain topic features; splicing the text features and the topic features, and predicting the predicted topic of the training text based on the splicing processing result; calculating the value of the cross entropy loss function based on the predicted topic of the training text and the topic label of the training text; adjusting the model parameters of the text classification model in the direction of reducing the value of the cross entropy loss function until the text classification model reaches convergence.

[0092] Alternatively, assuming that the number of training texts is m, the value of the cross entropy loss function is calculated based on the predicted topics of the training texts and the topic labels of the training texts, which can be expressed by the following formula (4):

[0093] (4)

[0094] In formula (4), Indicates the topic label corresponding to the jth training text, and the value of j ranges from 1 to m. Represents the predicted topic corresponding to the j-th training text.

[0095] After training the text classification model, the fully connected layer in the text classification model is called to concatenate the text features and the topic features to obtain a concatenated result. Furthermore, the classification layer in the text classification model is called to perform classification prediction based on the concatenated result using a classification function to obtain the text topic distribution corresponding to the text to be classified. The text features extracted by the text feature processing model and the topic features extracted by the text topic processing model have the same dimension. Therefore, concatenating the text features and the topic features can refer to directly adding the text features and the topic features.

[0096] Step S304: determining the text topic corresponding to the text to be classified according to the text topic distribution, and classifying the text to be classified based on the text topic to determine the category to which the text to be classified belongs.

[0097] In one embodiment, some feasible implementations included in step S304 are already Figure 2 Step S203 is described in detail in the embodiment and will not be repeated here.

[0098] In the embodiments of this application, the text features proposed by the text feature processing model are integrated with the topic information of the text topic processing model, effectively combining the semantic information of the text and obtaining a more optimized topic vector. Furthermore, to eliminate the impact of noise words on accuracy, the text to be classified is processed to remove stop words. Furthermore, using the cross-entropy loss function to train the text classification model can improve the accuracy of the text classification model.

[0099] Based on the above text processing method embodiment, the present application embodiment provides a text processing device, see Figure 6 , is a structural diagram of a text processing device provided in an embodiment of the present application. Figure 6 The text processing device can run the following units:

[0100] An extraction unit 601 is used to extract text features from the text to be classified to obtain text features corresponding to the text to be classified;

[0101] The extraction unit 601 is further configured to extract text topic features from the text to be classified based on the text features to obtain topic features corresponding to the text to be classified;

[0102] A splicing unit 602 is used to splice the text features and the topic features;

[0103] The prediction unit 603 is configured to perform text topic prediction on the text to be classified based on the splicing processing result to obtain a text topic corresponding to the text to be classified;

[0104] The determining unit 604 is configured to perform text classification on the text to be classified based on the text theme, and determine the category to which the text to be classified belongs.

[0105] In one embodiment, when extracting features from the text to be classified, the extraction unit 601 performs the following steps: preprocessing the text to be classified, and extracting features from the preprocessed text to be classified;

[0106] The text to be classified is preprocessed, including: performing word segmentation processing on the text to be classified to obtain a word set corresponding to the text to be classified, the word set including one or more characters or words; and removing stop words from the word set to obtain characteristic words included in the text to be classified.

[0107] In one embodiment, when the extraction unit 601 extracts features from the text to be classified and obtains text features corresponding to the text to be classified, the extraction unit 601 performs the following steps:

[0108] Performing vector embedding processing on the text to be classified to obtain a word vector corresponding to each characteristic word in the text to be classified;

[0109] The word vector corresponding to each feature word is linearly transformed, and an attention weight operation is performed on each word vector after the linear transformation to obtain the text features of the text to be classified.

[0110] In one embodiment, when extracting text topic features from the text to be classified based on the text features, the extraction unit 601 performs the following steps:

[0111] Acquire characteristic words included in the text characteristics, and assign a text topic number to each characteristic word;

[0112] The text topic number of each characteristic word is subjected to Gibbs sampling. When Gibbs sampling converges, the text topic distribution in the to-be-classified text is counted, and the text topic distribution is determined as the topic feature of the to-be-classified text.

[0113] In one embodiment, the feature extraction of the text to be classified is performed by calling a text feature processing model, and the text topic feature extraction of the text to be classified based on the text features is performed by calling a text topic processing model; the text feature processing model and the text topic processing model are both deployed in a text classification model, and the text feature processing model is connected to the text topic processing model; the text classification model further includes a fully connected layer and a classification layer; the splicing unit 602 performs the following steps when splicing the text features and the topic features:

[0114] Calling the fully connected layer to splice the text features and the topic features to obtain a splicing processing result of a target length;

[0115] When the prediction unit 603 performs text topic prediction on the text to be classified based on the splicing processing result and obtains the text topic corresponding to the text to be classified, the prediction unit 603 performs the following steps:

[0116] The classification layer is called to adopt a classification function to perform classification prediction based on the splicing processing result to obtain a text topic distribution corresponding to the text to be classified, and the text topic corresponding to the text to be classified is determined according to the text topic distribution.

[0117] In one embodiment, the text topic processing model is obtained by training based on a training text set, wherein the training text set includes multiple training texts, and each training text includes one or more feature words; the text processing device also includes a training unit 605, and the training unit 605 is used to perform: randomly assigning a topic to the feature words in each training text; scanning the training text set, and repeatedly re-extracting a topic for each feature word in each training text according to Gibbs C sampling until Gibbs sampling converges; counting the co-occurrence frequency matrix of the topics and words in the training text set, and determining the text topic processing model based on the co-occurrence matrix of the topics and words.

[0118] In one embodiment, when the determining unit 604 classifies the text to be classified based on the text theme and determines the category to which the text to be classified belongs, the determining unit 604 performs the following steps:

[0119] Obtaining the correspondence between text topics and audit specifications; determining a target audit specification that matches the text topic corresponding to the text to be classified based on the correspondence between the text topics and the audit specifications; and classifying the text to be classified as a text that violates the target audit specification.

[0120] According to one embodiment of the present application, Figure 2 as well as Figure 3 The steps involved in the text processing method shown can be Figure 6 The text processing device shown in FIG. Figure 2 The steps S201 and S202 can be performed by Figure 6 The text processing device shown in FIG. 1 is used to extract the text. ... Figure 6 The text processing device is implemented by the splicing unit 602 and the prediction unit 603. Step S204 can be performed by Figure 6 The determination unit 604 in the text processing device is executed; for example, Figure 3Step S301 and step S302 can be obtained by Figure 6 The text processing device is executed by the extraction unit 601, and step S303 can be performed by Figure 6 The text processing device is implemented by the splicing unit 602 and the prediction unit 603; step S304 can be performed by Figure 6 The determination unit 604 in the text processing device is used for execution.

[0121] According to another embodiment of the present application, Figure 6 The various units in the text processing device shown can be individually or all combined into one or several other units to form a whole, or one (or some) of the units can be further divided into multiple functionally smaller units to form a whole, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, other units can also be included based on the text processing device. In actual applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.

[0122] According to another embodiment of the present application, the program can be executed by running on a general computing device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM) and other processing elements and storage elements. Figure 2 as well as Figure 3 A computer program (including program code) for each step involved in the corresponding method shown in FIG. Figure 7 The text processing device shown in and the text processing method of the embodiment of the present application are implemented. The computer program can be recorded on, for example, a computer-readable storage medium, and loaded into the text processing device through the computer-readable storage medium and run therein.

[0123] In an embodiment of the present application, when classifying a text to be classified, the text features of the text to be classified are first extracted, and then the text theme features of the text to be classified are extracted based on the text features to obtain the theme features corresponding to the text to be classified. Finally, the text features and the theme features are spliced ​​together, and a theme prediction is performed on the text to be classified based on the splicing results, thereby determining the text theme corresponding to the text to be classified based on the predicted text theme. It should be understood that text features can be used to reflect the semantic features of the text to be classified, and splicing and fusing text features and theme features effectively combines the semantic information of the text and obtains a better theme representation. When performing theme prediction based on the splicing results, the accuracy of theme prediction can be improved, thereby also providing the accuracy of text classification based on text themes.

[0124] Based on the above-mentioned text processing method embodiment and text processing device embodiment, the present application embodiment provides a text processing device, see Figure 7 , is a structural diagram of a text processing device provided in an embodiment of the present application. Figure 7 The text processing device may include a processor 701, an input interface 702, an output interface 703, and a computer storage medium 704. The processor 701, the input interface 702, the output interface 703, and the computer storage medium 704 may be connected via a bus or other means.

[0125] Computer storage medium 704 may be stored in the memory of the text processing device. The computer storage medium 704 is used to store computer programs, and the processor 701 is used to execute the computer programs stored in the computer storage medium 704. The processor 701 (also known as the CPU (Central Processing Unit)) is the computing and control core of the text processing device and is suitable for implementing one or more computer programs, specifically loading and executing:

[0126] Performing text feature extraction on the text to be classified to obtain text features corresponding to the text to be classified; performing text topic feature extraction on the text to be classified based on the text features to obtain topic features corresponding to the text to be classified; performing splicing processing on the text features and the topic features, and performing text topic prediction on the text to be classified based on the splicing processing result to obtain a text topic corresponding to the text to be classified; performing text classification on the text to be classified based on the text topic to determine the category to which the text to be classified belongs.

[0127] In an embodiment of the present application, when classifying a text to be classified, the text features of the text to be classified are first extracted, and then the text theme features of the text to be classified are extracted based on the text features to obtain the theme features corresponding to the text to be classified. Finally, the text features and the theme features are spliced ​​together, and a theme prediction is performed on the text to be classified based on the splicing results, thereby determining the text theme corresponding to the text to be classified based on the predicted text theme. It should be understood that text features can be used to reflect the semantic features of the text to be classified, and splicing and fusing text features and theme features effectively combines the semantic information of the text and obtains a better theme representation. When performing theme prediction based on the splicing results, the accuracy of theme prediction can be improved, thereby also providing the accuracy of text classification based on text themes.

[0128] The embodiment of the present application also provides a computer storage medium (Memory), which is a memory device of a text processing device and is used to store programs and data. It is understandable that the computer storage medium here can include both the built-in storage medium of the text processing device and, of course, the extended storage medium supported by the text processing device. The computer storage medium provides a storage space that stores the operating system of the text processing device. In addition, one or more computer programs suitable for being loaded and executed by the processor 701 are also stored in the storage space. It should be noted that the computer storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk storage; optionally, it can also be at least one computer storage medium located away from the aforementioned processor.

[0129] In one embodiment, the one or more computer programs stored in the computer storage medium may be loaded and executed by the processor 701:

[0130] Performing text feature extraction on the text to be classified to obtain text features corresponding to the text to be classified; performing text topic feature extraction on the text to be classified based on the text features to obtain topic features corresponding to the text to be classified; performing splicing processing on the text features and the topic features, and performing text topic prediction on the text to be classified based on the splicing processing result to obtain a text topic corresponding to the text to be classified; performing text classification on the text to be classified based on the text topic to determine the category to which the text to be classified belongs.

[0131] In one embodiment, when performing feature extraction on the text to be classified, the processor 701 performs the following steps: preprocessing the text to be classified, and performing feature extraction on the preprocessed text to be classified;

[0132] The preprocessing of the text to be classified includes: performing word segmentation processing on the text to be classified to obtain a word set corresponding to the text to be classified, and the word set includes one or more characters or words; performing stop word removal processing on the word set to obtain characteristic words included in the text to be classified.

[0133] In one embodiment, when the processor 701 extracts features from the text to be classified and obtains text features corresponding to the text to be classified, the processor 701 performs the following steps:

[0134] The text to be classified is subjected to vector embedding processing to obtain a word vector corresponding to each feature word in the text to be classified; the word vector corresponding to each feature word is subjected to linear transformation processing, and an attention weight operation is performed on each word vector after the linear transformation processing to obtain the text feature of the text to be classified.

[0135] In one embodiment, when extracting text topic features from the text to be classified based on the text features, the processor 701 performs the following steps:

[0136] Acquire the characteristic words included in the text features, and assign a text topic number to each of the characteristic words; perform Gibbs sampling on the text topic number of each characteristic word, and when Gibbs sampling converges, count the text topic distribution in the text to be classified, and determine the text topic distribution as the topic feature of the text to be classified.

[0137] In one embodiment, the feature extraction of the text to be classified is performed by calling a text feature processing model, and the text topic feature extraction of the text to be classified based on the text features is performed by calling a text topic processing model; the text feature processing model and the text topic processing model are both deployed in a text classification model, and the text feature processing model is connected to the text topic processing model; the text classification model also includes a fully connected layer and a classification layer;

[0138] When the processor 701 performs splicing processing on the text features and the topic features, and performs text topic prediction on the text to be classified based on the splicing processing result to obtain the text topic corresponding to the text to be classified, the processor 701 performs the following steps:

[0139] Calling the fully connected layer to splice the text features and the topic features to obtain a splicing processing result of a target length;

[0140] The classification layer is called to adopt a classification function to perform classification prediction based on the splicing processing result to obtain a text topic distribution corresponding to the text to be classified, and the text topic corresponding to the text to be classified is determined according to the text topic distribution.

[0141] In one embodiment, the text topic processing model is obtained based on a training text set, wherein the training text set includes a plurality of training texts, each training text includes one or more feature words; the processor 701 is further configured to:

[0142] Randomly assign a topic to a feature word in each training text; scan the training text set and repeatedly re-extract a topic for each feature word in each training text according to Gibbs C sampling until Gibbs sampling converges;

[0143] The co-occurrence frequency matrix of the topics and words of the training text set is counted, and a text topic processing model is determined according to the co-occurrence matrix of the topics and words.

[0144] In one embodiment, when the processor 701 classifies the text to be classified based on the text theme and determines the category to which the text to be classified belongs, the processor 701 performs the following steps:

[0145] Obtain the correspondence between text topics and review specifications;

[0146] Based on the correspondence between the text subject and the audit specification, determining a target audit specification that matches the text subject corresponding to the text to be classified;

[0147] The to-be-classified text is classified as a text that violates the target review specification.

[0148] In an embodiment of the present application, when classifying a text to be classified, the text features of the text to be classified are first extracted, and then the text theme features of the text to be classified are extracted based on the text features to obtain the theme features corresponding to the text to be classified. Finally, the text features and the theme features are spliced ​​together, and a theme prediction is performed on the text to be classified based on the splicing results, thereby determining the text theme corresponding to the text to be classified based on the predicted text theme. It should be understood that text features can be used to reflect the semantic features of the text to be classified, and splicing and fusing text features and theme features effectively combines the semantic information of the text and obtains a better theme representation. When performing theme prediction based on the splicing results, the accuracy of theme prediction can be improved, thereby also providing the accuracy of text classification based on text themes.

[0149] The present embodiment provides a computer program product or computer program, wherein the computer program product includes a computer program, and when the computer program is executed by the processor 701, the computer program is configured to load and execute:

[0150] Performing text feature extraction on the text to be classified to obtain text features corresponding to the text to be classified; performing text topic feature extraction on the text to be classified based on the text features to obtain topic features corresponding to the text to be classified; performing splicing processing on the text features and the topic features, and performing text topic prediction on the text to be classified based on the splicing processing result to obtain a text topic corresponding to the text to be classified; performing text classification on the text to be classified based on the text topic to determine the category to which the text to be classified belongs.

[0151] In an embodiment of the present application, when classifying a text to be classified, the text features of the text to be classified are first extracted, and then the text theme features of the text to be classified are extracted based on the text features to obtain the theme features corresponding to the text to be classified. Finally, the text features and the theme features are spliced ​​together, and a theme prediction is performed on the text to be classified based on the splicing results, thereby determining the text theme corresponding to the text to be classified based on the predicted text theme. It should be understood that text features can be used to reflect the semantic features of the text to be classified, and splicing and fusing text features and theme features effectively combines the semantic information of the text and obtains a better theme representation. When performing theme prediction based on the splicing results, the accuracy of theme prediction can be improved, thereby also providing the accuracy of text classification based on text themes.

Claims

1. A text processing method, characterized in that: include: Perform text feature extraction on the text to be classified to obtain text features corresponding to the text to be classified; The feature extraction of the text to be classified is performed by calling a text feature processing model; Based on the text features, text topic features are extracted from the text to be classified to obtain topic features corresponding to the text to be classified; the text topic features are extracted from the text to be classified based on the text features by calling a text topic processing model; the text feature processing model and the text topic processing model are both deployed in a text classification model, and the text feature processing model is connected to the text topic processing model; the text classification model also includes a fully connected layer and a classification layer; The text features and the topic features are spliced ​​together, and a text topic prediction is performed on the text to be classified based on the splicing processing result to obtain a text topic corresponding to the text to be classified; the text features and the topic features are spliced ​​together, and a text topic prediction is performed on the text to be classified based on the splicing processing result to obtain a text topic corresponding to the text to be classified, including: calling the fully connected layer to splice the text features and the topic features to obtain a splicing processing result of a target length; Calling the classification layer to use a classification function to perform classification prediction based on the splicing processing result, obtaining a text topic distribution corresponding to the text to be classified, and determining a text topic corresponding to the text to be classified according to the text topic distribution; The text to be classified is classified based on the text theme to determine the category to which the text to be classified belongs.

2. The method according to claim 1, wherein The feature extraction of the text to be classified includes: Preprocessing the text to be classified, and performing feature extraction on the preprocessed text to be classified; The preprocessing of the text to be classified includes: Performing word segmentation on the text to be classified to obtain a word set corresponding to the text to be classified, wherein the word set includes one or more characters or words; The word set is processed to remove stop words to obtain characteristic words included in the text to be classified.

3. The method according to claim 2, wherein The feature extraction of the text to be classified to obtain text features corresponding to the text to be classified includes: Performing vector embedding processing on the text to be classified to obtain a word vector corresponding to each characteristic word in the text to be classified; The word vector corresponding to each feature word is linearly transformed, and an attention weight operation is performed on each word vector after the linear transformation to obtain the text features of the text to be classified.

4. The method according to claim 1, wherein The extracting text topic features of the text to be classified based on the text features includes: Acquire characteristic words included in the text characteristics, and assign a text topic number to each characteristic word; The text topic number of each characteristic word is subjected to Gibbs sampling. When Gibbs sampling converges, the text topic distribution in the to-be-classified text is counted, and the text topic distribution is determined as the topic feature of the to-be-classified text.

5. The method according to claim 1, wherein The text topic processing model is obtained by training based on a training text set, wherein the training text set includes a plurality of training texts, each training text includes one or more feature words, and the method further includes: Randomly assign a topic to the feature words in each training text; Scanning the training text set, and repeatedly re-extracting a topic for each feature word in each training text according to Gibbs C sampling until Gibbs sampling converges; The co-occurrence frequency matrix of the topics and words of the training text set is counted, and a text topic processing model is determined according to the co-occurrence frequency matrix of the topics and words.

6. The method according to claim 1, wherein The step of classifying the text to be classified based on the text subject to determine the category to which the text to be classified belongs includes: Obtain the correspondence between text topics and review specifications; Based on the correspondence between the text subject and the audit specification, determining a target audit specification that matches the text subject corresponding to the text to be classified; The to-be-classified text is classified as a text that violates the target review specification.

7. A text processing device, characterized in that: include: An extraction unit, configured to extract text features from the text to be classified, and obtain text features corresponding to the text to be classified; The feature extraction of the text to be classified is performed by calling a text feature processing model; The extraction unit is used to extract text topic features from the text to be classified based on the text features to obtain topic features corresponding to the text to be classified; the text topic feature extraction from the text to be classified based on the text features is performed by calling a text topic processing model; the text feature processing model and the text topic processing model are both deployed in a text classification model, and the text feature processing model is connected to the text topic processing model; the text classification model also includes a fully connected layer and a classification layer; A splicing unit is used to splice the text features and the topic features, and predict the text topic of the text to be classified based on the splicing processing result to obtain the text topic corresponding to the text to be classified; the splicing of the text features and the topic features, and predicting the text topic of the text to be classified based on the splicing processing result to obtain the text topic corresponding to the text to be classified, including: calling the fully connected layer to splice the text features and the topic features to obtain a splicing processing result of a target length; calling the classification layer to use a classification function to perform classification prediction based on the splicing processing result to obtain the text topic distribution corresponding to the text to be classified, and determining the text topic corresponding to the text to be classified based on the text topic distribution; The determining unit is configured to perform text classification on the text to be classified based on the text theme, and determine the category to which the text to be classified belongs.

8. A text processing device, characterized in that: include: A processor adapted to implement one or more computer programs; a computer storage medium storing one or more computer programs, wherein the one or more computer programs are adapted to be loaded by the processor and executed by the text processing method according to any one of claims 1 to 6.

9. A computer storage medium, characterized in that The computer storage medium stores a computer program, and when the computer program is executed by a processor, the computer program is used to load and execute the text processing method according to any one of claims 1 to 6.

10. A computer program product or a computer program, wherein the computer program product comprises a computer program, wherein the computer program is stored in a computer storage medium, and when the computer program is executed by a processor, is used to load and execute the text processing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device to output information

    CN108121699A

  • Text auditing method and device, computer equipment and readable storage medium

    CN111274782A