Abnormal text detection model training method, abnormal text detection method and device
By replacing the time characters in the training text and determining the label based on the generation time, the text classification model is trained to generate an abnormal text detection model, which solves the problem of low accuracy in abnormal text detection in the prior art, and achieves more efficient abnormal text recognition.
Patent Information
- Application Number
- CN202210590192.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-05-26
AI Technical Summary
The existing abnormal text detection technology has low accuracy, mainly due to the limited number of pre-stored keywords, which makes it impossible to effectively identify abnormal content in text data.
By obtaining the training sample set, including the training text and classification labels, and replacing the characters representing time in the training text with preset characters, the classification labels are determined based on the generation time, and the text classification model is trained to generate an abnormal text detection model.
It improves the accuracy of abnormal text detection, and can detect abnormal text from the first time to the current time from the text data before the current time, with a wide range of applications.
Smart Images

Figure CN115269830B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to an abnormal text detection model training method, an abnormal text detection method and a device. Background Art
[0002] Natural Language Processing (NLP) is an important field in computer science and artificial intelligence. NLP is often used to classify text data. Text classification refers to the classification of text data according to its content and is widely used in abnormal text detection.
[0003] Abnormal text detection technology is to find objects in a text data set that are significantly different from other text data. In related technologies, the text content to be detected is identified by pre-stored keywords, and the text content matching the pre-stored keywords is determined as abnormal text content.
[0004] However, the number of pre-stored keywords is limited, the above method has certain limitations, and the accuracy of abnormal text detection is low. Summary of the invention
[0005] The present application provides an abnormal text detection model training method, an abnormal text detection method and a device, which can improve the accuracy of abnormal text detection.
[0006] In a first aspect, the present application provides a method for training an abnormal text detection model, comprising:
[0007] Acquire a training sample set, the training sample set comprising training text and classification labels of the training text, wherein characters used to indicate time in the training text are replaced with preset characters, wherein the classification label of the training text generated before a first time is a first classification label, and the classification label of the training text generated after the first time is a second classification label, and the first time is before the current time;
[0008] Performing text classification model training according to the training sample set, and in any training process, taking the training text as the input of the text classification model and outputting the classification prediction probability of the training text;
[0009] According to the classification prediction probability of the training text and the classification label of the training text, the parameters of the text classification model are adjusted until the training stop condition is met;
[0010] The text classification model determined by satisfying the training stop condition is output as an abnormal text detection model.
[0011] In a second aspect, the present application provides a method for detecting abnormal text, comprising:
[0012] After receiving the detection instruction, obtaining a text set to be detected, the text set to be detected includes texts generated before the current time, and the detection instruction carries a first time period;
[0013] Determine, according to the first time period, an abnormal text detection model corresponding to a first time, the abnormal text detection model being trained according to the method of the first aspect, the first time being a time point whose time interval from the current time is the first time period;
[0014] Input each text in the to-be-detected text set into the abnormal text detection model in turn to obtain a classification prediction probability of each text;
[0015] According to the classification prediction probability of each text, the target text whose classification label is the second classification label and whose classification prediction probability is greater than the first preset threshold is determined as an abnormal text within the first time period, and the abnormal text is a text that has not appeared in the text set to be detected before the first time, and the classification label of the text generated after the first time is the second classification label.
[0016] In a third aspect, the present application provides an abnormal text detection model training device, comprising:
[0017] An acquisition module is used to acquire a training sample set, wherein the training sample set includes a training text and a classification label of the training text, wherein characters used to indicate time in the training text are replaced with preset characters, wherein the classification label of the training text generated before a first time is a first classification label, and the classification label of the training text generated after the first time is a second classification label, and the first time is before the current time;
[0018] A processing module, used for training a text classification model according to the training sample set, and in any training process, taking the training text as the input of the text classification model and outputting the classification prediction probability of the training text;
[0019] An adjustment module, used for adjusting the parameters of the text classification model according to the classification prediction probability of the training text and the classification label of the training text until a training stop condition is met;
[0020] The output module is used to output the text classification model determined by satisfying the training stop condition as an abnormal text detection model.
[0021] In a fourth aspect, the present application provides an abnormal text detection device, comprising:
[0022] An acquisition module, configured to acquire a text set to be detected after receiving a detection instruction, wherein the text set to be detected includes texts generated before a current time, and the detection instruction carries a first time period;
[0023] A first determination module is used to determine an abnormal text detection model corresponding to a first time according to the first time period, the abnormal text detection model is trained according to the method described in the first aspect, and the first time is a time point whose time interval with the current time is the first time period;
[0024] A processing module, used for inputting each text in the to-be-detected text set into the abnormal text detection model in sequence to obtain a classification prediction probability of each text;
[0025] The second determination module is used to determine the target text whose classification label is the second classification label and whose classification prediction probability is greater than the first preset threshold as an abnormal text within the first time period according to the classification prediction probability of each text, wherein the abnormal text is the text that has not appeared in the text set to be detected before the first time, and the classification label of the text generated after the first time is the second classification label.
[0026] In a fifth aspect, the present application provides a computer device, comprising: a processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method of the first aspect or the second aspect.
[0027] In a sixth aspect, the present application provides a computer-readable storage medium comprising instructions, which, when executed on a computer program, causes the computer to execute the method of the first aspect or the second aspect.
[0028] In a seventh aspect, the present application provides a computer program product comprising instructions, which, when executed on a computer, causes the computer to execute the method of the first aspect or the second aspect.
[0029] In summary, in the present application, when obtaining a training sample set, the training sample set includes training text and the true classification label of the training text, and the characters used to represent time in the training text are replaced with preset characters, and the true classification label of the training text is determined through a time point, and the classification label of the training text generated before the first time is the first classification label, and the classification label of the training text generated after the first time is the second classification label. When training a text classification model based on the training sample set, in any training process, the training text is used as the input of the text classification model, and the classification prediction probability of the training text is output. According to the classification prediction probability of the training text and the classification label of the training text, the parameters of the text classification model are adjusted until the training stop condition is met, and the text classification model finally trained is an abnormal text detection model. Since the characters used to indicate time in the training text are replaced with preset characters, the training text does not contain time information. The real classification label of the training text is determined by generating time before and after the first time. Using the training text processed as above to train the text classification model, the trained text classification model can detect abnormal texts from the first time to the current time from the set of texts to be detected generated before the current time through the semantic information of the training text. The abnormal text here is the text that has not appeared in the set of texts to be detected before the first time. Compared with the abnormal text recognition using pre-stored keywords, the abnormal text detection model trained by the abnormal text detection model training method provided in the embodiment of the present application can improve the accuracy of abnormal text detection and has a wide range of applications.
[0030] Furthermore, in the present application, after receiving the detection instruction, a set of texts to be detected is obtained, the set of texts to be detected includes texts whose generation time is before the current time, an abnormal text detection model corresponding to the first time is determined according to the first time period carried by the detection instruction, each text in the set of texts to be detected is input into the abnormal text detection model in turn, and the classification prediction probability of each text is obtained. According to the classification prediction probability of each text, the target text whose classification label is the second classification label and whose classification prediction probability is greater than the first preset threshold is determined as an abnormal text in the first time period, and the abnormal text is the text that has not appeared in the set of texts to be detected before the first time. Thus, the accuracy of abnormal text detection is improved, and the scope of application is wide.
[0031] Furthermore, after detecting the abnormal text within the first time period from the text set to be detected, each text in the abnormal text within the first time period is segmented to obtain a segmentation set, and then based on the TF-IDF value of each segmentation in the segmentation set and the posterior probability of each segmentation, the keywords in the abnormal text within the first time period are screened out and output, so that the user can perform all-round control and analysis based on the keywords, for example, search for the corresponding abnormal text based on the keywords, reduce the number of abnormal texts to be viewed, and improve processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 A schematic diagram of an abnormal text detection model training method and an implementation scenario of the abnormal text detection method provided in an embodiment of the present application;
[0033] Figure 2 A flowchart of an abnormal text detection model training method provided in an embodiment of the present application;
[0034] Figure 3 A flowchart of an abnormal text detection model training method provided in an embodiment of the present application;
[0035] Figure 4 A flowchart of an abnormal text detection model training method provided in an embodiment of the present application;
[0036] Figure 5 A flowchart of an abnormal text detection method provided in an embodiment of the present application;
[0037] Figure 6 A flowchart of an abnormal text detection method provided in an embodiment of the present application;
[0038] Figure 7 A schematic diagram of the structure of an abnormal text detection model training device provided in an embodiment of the present application;
[0039] Figure 8 A schematic diagram of the structure of an abnormal text detection device provided in an embodiment of the present application;
[0040] Fig. 9 It is a schematic block diagram of a computer device 700 provided in an embodiment of the present application. DETAILED DESCRIPTION
[0041] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0042] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0043] Before introducing the technical solution of this application, the following is an introduction to the relevant knowledge of this application:
[0044] Artificial Intelligence (AI): It is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0045] Machine Learning (ML): It is a multi-disciplinary interdisciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0046] Deep Learning (DL): It is a branch of machine learning and an algorithm that attempts to use multiple processing layers containing complex structures or composed of multiple nonlinear transformations to perform high-level abstraction of data. Deep learning is to learn the inherent laws and representation levels of training sample data. The information obtained in the learning process is of great help in the interpretation of data such as text, images and sounds. The ultimate goal of deep learning is to enable machines to have analytical learning capabilities like humans and to be able to recognize data such as text, images and sounds. Deep learning is a complex machine learning algorithm that has achieved results in speech and image recognition that far exceed previous related technologies.
[0047] Neural Network (NN): A deep learning model in the field of machine learning and cognitive science that imitates the structure and function of biological neural networks.
[0048] Pre-training: A process of training a neural network model using a large dataset so that the neural network model learns the common features in the dataset. The purpose of pre-training is to provide high-quality model parameters for subsequent training of the neural network model on a specific dataset. In the embodiments of the present application, pre-training may refer to the process of training the BERT model using unlabeled training text.
[0049] Fine-tuning: A process of further training a pre-trained neural network model using a specific data set. Typically, the amount of data in the data set used in the fine-tuning phase is smaller than that in the pre-training phase, and the training samples in the data set used in the fine-tuning phase contain annotation information. Fine-tuning in the embodiments of the present application refers to the process of training a BERT model (pre-trained BERT model) using training text containing classification labels.
[0050] In the related art, the accuracy of abnormal text detection is low. To solve this problem, the present application trains an abnormal text detection model. When training the model, a training sample set is obtained. The training sample set includes training text and the true classification label of the training text. The characters used to represent time in the training text are replaced with preset characters. The true classification label of the training text is determined by a time point. The classification label of the training text generated before the first time is the first classification label, and the classification label of the training text generated after the first time is the second classification label. When training the text classification model according to the training sample set, in any training process, the training text is used as the input of the text classification model, and the classification prediction probability of the training text is output. According to the classification prediction probability of the training text and the classification label of the training text, the parameters of the text classification model are adjusted until the training stop condition is met. The text classification model finally trained is the abnormal text detection model. Since the characters used to indicate time in the training text are replaced with preset characters, the training text does not contain time information. The real classification label of the training text is determined by generating time before and after the first time. Using the training text processed as above to train the text classification model, the trained text classification model can detect abnormal texts from the first time to the current time from the set of texts to be detected generated before the current time through the semantic information of the training text. The abnormal text here is the text that has not appeared in the set of texts to be detected before the first time. Compared with the abnormal text recognition using pre-stored keywords, the abnormal text detection model trained by the abnormal text detection model training method provided in the embodiment of the present application can improve the accuracy of abnormal text detection and has a wide range of applications.
[0051] Furthermore, after the abnormal text detection model trained by the abnormal text detection model training method provided in the embodiment of the present application detects the abnormal text within the first time period from the text set to be detected, each text in the abnormal text within the first time period is segmented to obtain a segmentation set, and then based on the TF-IDF value of each segmentation in the segmentation set and the posterior probability of each segmentation, the keywords in the abnormal text within the first time period are screened out and output, so that the user can perform all-round control and analysis according to the keywords, for example, search for the corresponding abnormal text according to the keywords, reduce the number of abnormal texts to be viewed, and improve processing efficiency.
[0052] The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving and the like.
[0053] The following is a brief introduction to the application scenarios to which the technical solutions of the embodiments of the present application can be applied. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present application and are not limited. In specific implementation, the technical solutions provided by the embodiments of the present application can be flexibly applied according to actual needs.
[0054] The abnormal text detection model training method and the abnormal text detection method provided in the embodiments of the present application can be applied to text information classification scenarios, and are described below in conjunction with several application scenarios.
[0055] 1. In the electronic payment application scenario, it is necessary to detect abnormal complaints that have occurred recently.
[0056] In the electronic payment application scenario, the abnormal text detection model training method and the abnormal text detection method provided in the embodiment of the present application can be applied to the electronic payment server. In order to detect abnormal complaints (i.e. abnormal text) that have appeared recently, the electronic payment server first obtains a training sample set based on the historical text data before the current time, and the training sample set includes training text and classification labels of the training text. The process of obtaining the training sample set can be specifically: determining the first time, obtaining the historical text data before the current time, the historical text data includes text information and the generation time of the text information, replacing the characters used to represent the time in the text information with preset characters, obtaining the training text, and determining the classification label of the training text according to the generation time of the training text and the first time, wherein the classification label of the training text generated before the first time is the first classification label, and the classification label of the training text generated after the first time is the second classification label. Then, the electronic payment server performs text classification model training based on the training sample set to train an abnormal text detection model. In the model application stage, after receiving the detection instruction, the electronic payment server obtains the text set to be detected, which includes texts generated before the current time. The detection instruction carries the first time period, and each text in the text set to be detected is input into the text classification model in turn to obtain the classification prediction probability of each text. According to the classification prediction probability of each text, the target text with the classification label as the second classification label and the classification prediction probability greater than the first preset threshold is determined as an abnormal text in the first time period. The abnormal text is the text that has not appeared in the text set to be detected before the first time. Thus, abnormal complaints (i.e., abnormal texts) that have appeared recently can be detected.
[0057] 2. Text content review scenario, it is necessary to review the abnormal text content that has appeared recently
[0058] Text content audit scenarios include but are not limited to social platform release content audit, comment content audit, social information audit, multimedia file description information audit, etc. Taking the social platform release content audit as an example, the abnormal text detection model training method and abnormal text detection method provided in the embodiment of the present application can be applied to the server. The server needs to audit the abnormal text content released recently. First, a training sample set is obtained based on the historical text data before the current time. The training sample set includes training text and classification labels of the training text. The process of obtaining the training sample set can be specifically as follows: determine the first time, obtain the historical text data before the current time, the historical text data includes text information and the generation time of the text information (i.e., the release time), replace the characters used to represent the time in the text information with preset characters, obtain the training text, and determine the classification label of the training text according to the generation time of the training text and the first time, wherein the classification label of the training text generated before the first time is the first classification label, and the classification label of the training text generated after the first time is the second classification label. Then, the server performs text classification model training according to the training sample set to train an abnormal text detection model. In the model application stage, after receiving the detection instruction, the server obtains the text set to be detected, which includes texts generated before the current time. The detection instruction carries the first time period, and each text in the text set to be detected is input into the text classification model in turn to obtain the classification prediction probability of each text. According to the classification prediction probability of each text, the target text with the classification label as the second classification label and the classification prediction probability greater than the first preset threshold is determined as the abnormal text in the first time period. The abnormal text is the text that has not appeared in the text set to be detected before the first time. In this way, the abnormal text content published recently can be detected.
[0059] The above is only a schematic illustration of several common application scenarios. The method provided in the embodiments of the present application can also be applied to other scenarios that require abnormal detection or classification of text content. The embodiments of the present application do not limit the actual application scenarios.
[0060] For example, Figure 1 A schematic diagram of an abnormal text detection model training method and an implementation scenario of the abnormal text detection method provided in the embodiment of the present application is shown in FIG. Figure 1 As shown, the implementation scenario of the embodiment of the present application involves a server 1 and a user terminal 2, and the user terminal 2 can communicate data with the server 1 through a communication network.
[0061] Among them, in some possible implementations, the user terminal 2 refers to a type of device that has rich human-computer interaction methods, has the ability to access the Internet, is usually equipped with various operating systems, and has strong processing capabilities. The user terminal can be a user terminal such as a smart phone, a tablet computer, a portable laptop, a desktop computer, or a phone watch, etc., but is not limited thereto. Optionally, in the embodiment of the present application, the user terminal 2 is installed with an electronic payment application or an application with an electronic payment function.
[0062] Among them, in some possible implementations, the user terminal 2 includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, car terminals, etc.
[0063] Figure 1 The server 1 in the embodiment may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. This application does not limit this. In the embodiment of this application, the server 1 may be a background server of an electronic payment application in the user terminal 2.
[0064] In some possible implementations, Figure 1 One user terminal and one server are shown as an example, but other numbers of user terminals and servers may actually be included, and this application does not impose any limitation on this.
[0065] Exemplarily, a target application with an electronic payment function can be installed and run on the user terminal 2. The user can operate the target application installed on the user terminal 2 to implement the electronic payment service. The server 1 can process the electronic payment request sent by the user terminal 2 and store the corresponding payment information. Optionally, it can also receive the user's complaint information against the merchant sent by the user terminal 2. The complaint information may include the name of the merchant complained and the reason for the complaint. The server 1 can receive and store the complaint information. It is understandable that the data volume of the complaint information is large. Furthermore, the server 1 can detect the abnormal complaints that have appeared recently in these complaint information, so that the electronic payment developers or operation and maintenance personnel can handle the abnormal complaints accordingly.
[0066] The technical solution of this application will be described in detail below:
[0067] Figure 2 A flowchart of a method for training an abnormal text detection model provided in an embodiment of the present application, the execution subject of the method may be a server, such as Figure 2 As shown, the method may include:
[0068] S101. Obtain a training sample set, where the training sample set includes training text and classification labels of the training text, where characters used to indicate time in the training text are replaced with preset characters, where the classification label of the training text generated before a first time is a first classification label, and the classification label of the training text generated after the first time is a second classification label, and the first time is before the current time.
[0069] Specifically, the first time is before the current time, and the first time is, for example, a time period from the current time to a preset time period, and the preset time period is, for example, one week, half a month, one month or other time, etc. Optionally, the first time can be determined according to the preset time period, and can also be determined according to the first time period carried in the model training instruction.
[0070] As an implementable manner, obtaining a training sample set in S101 may specifically include:
[0071] S1011. Determine the first time.
[0072] Optionally, determining the first time may specifically include: receiving a model training instruction, the model training instruction carrying a first time period; or determining a preset time period as the first time period. A time point when the time interval and the current time are the first time period is determined as the first time. For example, if the first time period is one month, the first time is a time point one month from the current time.
[0073] S1012: Acquire historical text data before the current time, where the historical text data includes text information and the time when the text information is generated.
[0074] Specifically, the historical text data may be all text information generated between a preset time node and the current time, where the preset time node may be, for example, one year ago, two years ago, and so on.
[0075] S1013: Replace characters used to indicate time in the text information with preset characters to obtain training text.
[0076] Specifically, for example, replace all the characters representing time in the text information with the preset characters "Eight" or "8". As an implementable method, S1013 can specifically be to replace the uppercase and lowercase numbers in the text information with single characters or symbols respectively, and combine the numbers. For example, a piece of text information is "Don't know what's going on. There will be an automatic deduction of 648 yuan on October 1st". If the preset character "8" is used to replace the uppercase and lowercase numbers in this text information respectively and the numbers are combined, the training text obtained is: "Don't know what's going on. There will be an automatic deduction of 8 yuan on August 8th". If the preset character "Eight" is used to replace the uppercase and lowercase numbers in this text information respectively and the numbers are combined, the training text obtained is: "Don't know what's going on. There will be an automatic deduction of Eight yuan on the Eighth of August".
[0077] S1014. Determine the classification label of the training text according to the generation time of the training text and the first time.
[0078] Specifically, after the first time is determined and the training text is also determined, the classification label of the training text can be determined according to the generation time of the training text and the first time. For example, determine the classification label of the training text whose generation time is before the first time as the first classification label, and determine the classification label of the training text whose generation time is after the first time as the second classification label. Optionally, the first classification label is 0 and the second classification label is 1.
[0079] S102. Train the text classification model according to the training sample set. During any training process, use the training text as the input of the text classification model and output the classification prediction probability of the training text.
[0080] Optionally, the text classification model in this embodiment is used to: extract the semantic information of the training text, convert the semantic information into a vector, convert the vector into a scalar using a fully connected layer, and use an activation function to convert the scalar to obtain the classification prediction probability of the training text.
[0081] Among them, the text classification model can be a text convolutional neural network (textCNN) model or a Bidirectional Encoder Representations from Transformer (BERT) model. Optionally, the text classification model can be a pre-trained BERT model.
[0082] Among them, the classification prediction probability of the training text can be the prediction probability that the classification label of the training text is the first classification label or the second classification label, specifically a value between 0 and 1.
[0083] S103: According to the classification prediction probability of the training text and the classification label of the training text, the parameters of the text classification model are adjusted until the training stop condition is met.
[0084] Optionally, in S103, the parameters of the text classification model are adjusted according to the classification prediction probability of the training text and the classification label of the training text, which may be specifically:
[0085] S1031. Construct a loss function based on the classification prediction probability of the training text and the classification label of the training text, and adjust the parameters of the text classification model by back propagation based on the loss function.
[0086] Optionally, the loss function may be a cross entropy loss function. Taking the cross entropy loss function as an example, y i is the classification label of the training sample, that is, the actual classification label, y i ' is the classification prediction probability of the training text. The loss function constructed based on the classification prediction probability of the training text and the classification label of the training text can be shown in the following formula (1):
[0087]
[0088] Among them, CE(y i ,y i ')=-y i logy i '-(1-y i )log((1-y i ')), batch_size is the number of training samples used in any training.
[0089] The training stop condition may be a preset training stop condition.
[0090] As an implementable method, when training a text classification model based on a training sample set, the training sample set can be first divided into a training set and a validation set. The division method can be random, for example, the training set accounts for 70% and the validation set accounts for 30%. Use the training set to train the text classification model, and during the training iteration process, calculate the accuracy (AUC) of the text classification model on the training set and the validation set at a preset number of times. Stop training when the AUC on the training set exceeds the preset value of the AUC of the validation set (for example, 2%) and the AUC of the validation set no longer increases, and record the AUC of the validation set at this time as C. AUC Then, the full data (i.e., the unsplit training sample set) is used to retrain the text classification model. When the AUC of the trained text classification model is greater than C AUCThe training is stopped when the AUC of the trained text classification model is greater than C AUC If the preset multiple is reached, the training will be stopped.
[0091] S104: Outputting the text classification model determined by satisfying the training stop condition as an abnormal text detection model.
[0092] The abnormal text detection model training method provided in this embodiment is as follows: when obtaining a training sample set, the training sample set includes training text and the real classification label of the training text, the characters used to indicate time in the training text are replaced with preset characters, the real classification label of the training text is determined by a time point, the classification label of the training text generated before the first time is the first classification label, and the classification label of the training text generated after the first time is the second classification label. When training the text classification model according to the training sample set, in any training process, the training text is used as the input of the text classification model, the classification prediction probability of the training text is output, and the parameters of the text classification model are adjusted according to the classification prediction probability of the training text and the classification label of the training text until the training stop condition is met, and the text classification model finally trained is the abnormal text detection model. Since the characters used to indicate time in the training text are replaced with preset characters, the training text does not contain time information. The real classification label of the training text is determined by generating time before and after the first time. Using the training text processed as above to train the text classification model, the trained text classification model can detect abnormal texts from the first time to the current time from the set of texts to be detected generated before the current time through the semantic information of the training text. The abnormal text here is the text that has not appeared in the set of texts to be detected before the first time. Compared with the abnormal text recognition using pre-stored keywords, the abnormal text detection model trained by the abnormal text detection model training method provided in the embodiment of the present application can improve the accuracy of abnormal text detection and has a wide range of applications.
[0093] Combine the following Figure 3 Taking the text classification model as a pre-trained BERT model as an example, fine-tuning is performed when training the pre-trained BERT model, and the training process of the abnormal text detection model is explained in detail. Figure 3 A flowchart of a method for training an abnormal text detection model provided in an embodiment of the present application, the execution subject of the method may be a server, such as Figure 3 As shown, the method may include:
[0094] S201. Obtain a training sample set, where the training sample set includes training text and classification labels of the training text, where characters used to indicate time in the training text are replaced with preset characters, where the classification label of the training text generated before a first time is a first classification label, and the classification label of the training text generated after the first time is a second classification label, and the first time is before the current time.
[0095] Among them, the process of S201 is the same as that of S101. The detailed process can be found in the description of S101, which will not be repeated here.
[0096] S202. Perform BERT model training based on the training sample set. In any training process, the training text is used as the input of the BERT model, and the classification prediction probability of the training text is output.
[0097] Figure 4 A flowchart of a method for training an abnormal text detection model provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, in any training process, the training text is input into the BERT model, and the BERT model extracts the semantic information of the training text and converts the semantic information into a vector. Optionally, to prevent the model from overfitting, a random dropout layer can be added. The vector passes through the dropout layer (i.e., the vector is randomly disabled with a preset probability (such as 15%) during training), and then the vector is converted into a scalar using a fully connected layer (MLP layer), and then the scalar is converted using an activation function (sigmoid function) to obtain the classification prediction probability of the training text. By adding a dropout layer, the model is prevented from overfitting.
[0098] For example, the classification prediction probability y of the training text i 'for:
[0099] y i '=sigmoid(MLP(dropout(textclassifier(x i ),0.15)))
[0100] Among them, 0.15 is the preset probability.
[0101] S203, constructing a loss function according to the classification prediction probability of the training text and the classification label of the training text, and adjusting the parameters of the text classification model by back propagation according to the loss function until the training stop condition is met.
[0102] Specifically, in one practicable manner, in combination Figure 4, according to the classification prediction probability of the training text and the classification label of the training text, the cross entropy of the classification prediction probability of each training text and the classification label of each training text is calculated, and the result is multiplied by the weight of the training sample. Finally, the sum of the weighted cross entropies of all training samples is calculated, which is the loss function represented by the above formula (1).
[0103] The stopping condition for training can be the one in the above embodiment: the AUC of the trained text classification model is greater than C AUC The training is stopped when the preset multiple is reached. For details, please refer to the above embodiment and will not be described again here.
[0104] S204: Outputting the BERT model determined by satisfying the training stop condition as an abnormal text detection model.
[0105] Figure 5 A flowchart of an abnormal text detection method provided in an embodiment of the present application, the execution subject of the method may be a server, such as Figure 5 As shown, the method may include:
[0106] S301: After receiving a detection instruction, obtain a text set to be detected, where the text set to be detected includes texts generated before the current time, and the detection instruction carries a first time period.
[0107] Among them, the text set to be detected includes texts generated before the current time, specifically all text information generated between a preset time node and the current time, where the preset time node is, for example, half a year ago, one year ago, two years ago, etc.
[0108] The first time period is, for example, one week, one month, etc.
[0109] S302. Determine an abnormal text detection model corresponding to a first time according to a first time period, where the first time is a time point whose time interval with the current time is the first time period.
[0110] Among them, the abnormal text detection model is based on Figure 2 or Figure 3 The method shown is trained.
[0111] As an implementable manner, in S302, determining the abnormal text detection model corresponding to the first time according to the first time period may specifically be:
[0112] S3021. Determine a time point where the time interval and the current time are a first time period as a first time.
[0113] S3022. Perform training of an abnormal text detection model as soon as possible.
[0114] S3023. Determine the trained abnormal text detection model as the abnormal text detection model corresponding to the first time.
[0115] The above implementation method is an online training method, that is, after receiving the detection instruction, the abnormal text detection model is trained first. The specific training process can be found in Figure 2 The illustrated embodiment describes the process.
[0116] As another practicable manner, in S302, determining the abnormal text detection model corresponding to the first time according to the first time period may specifically be:
[0117] S3021′, determine the time point at which the time interval and the current time are the first time period as the first time.
[0118] S3022', determining an abnormal text detection model corresponding to the first time from a set of pre-trained abnormal text detection models.
[0119] Specifically, for example, different abnormal text detection models corresponding to first times can be pre-trained and stored in a normal text detection model set. After receiving a detection instruction, the first time is first determined based on the first time period carried by the detection instruction, and then the abnormal text detection model corresponding to the first time is determined from the pre-stored abnormal text detection model set.
[0120] S3023', determining the determined abnormal text detection model as the abnormal text detection model corresponding to the first time.
[0121] This embodiment method can be called an offline acquisition method of an abnormal text detection model. Compared with the two methods, the online training method is more accurate because the training text used is closer to the detection text, and the offline acquisition method is less complex and takes less time.
[0122] S303: Input each text in the to-be-detected text set into the abnormal text detection model in sequence to obtain the classification prediction probability of each text.
[0123] S304. According to the classification prediction probability of each text, the target text whose classification label is the second classification label and whose classification prediction probability is greater than the first preset threshold is determined as an abnormal text within the first time period. The abnormal text is the text that has not appeared in the text collection to be detected before the first time, and the classification label of the text generated after the first time is the second classification label.
[0124] Specifically, optionally, the first preset threshold is, for example, 0.8, then based on the classification prediction probability of each text, the target text whose classification label is the second classification label and whose classification prediction probability is greater than 0.8 is determined as abnormal text within the first time period, wherein the classification label of the text generated after the first time is the second classification label.
[0125] As an implementable manner, in S304, according to the classification prediction probability of each text, the target text whose classification label is the second classification label and whose classification prediction probability is greater than the first preset threshold is determined as an abnormal text in the first time period, which can be specifically:
[0126] S3041. Divide the to-be-detected text set into a first text set, a second text set, and a third text set according to the classification prediction probability of each text.
[0127] Among them, the classification prediction probability of the text in the first text set is greater than 0 and less than or equal to the first threshold, the classification prediction probability of the text in the second text set is greater than the first threshold and less than or equal to the second threshold, the classification prediction probability of the text in the third text set is greater than the second threshold, and the first threshold is less than the second threshold.
[0128] For example, the first threshold is a value between 0-0.2, and the second threshold is a value between 0.5-0.8.
[0129] S3042: Determine the proportion of texts with the second classification label in the first text set, the proportion of texts with the second classification label in the second text set, and the proportion of texts with the second classification label in the third text set.
[0130] S3043: Determine the text in the text set whose proportion in the first text set, the second text set, and the third text set is greater than a first preset threshold as the target text.
[0131] Furthermore, in an implementable manner, the method of this embodiment may also include:
[0132] S305: Segment each text in the abnormal text within the first time period to obtain a segmentation set.
[0133] Specifically, a word segmentation tool may be used to segment each text in the abnormal text within the first time period to obtain a word segmentation set.
[0134] S306: Calculate the TF-IDF value of each word in the word set and the posterior probability of each word.
[0135] Optionally, the TF-IDF value of each word can be calculated as follows:
[0136]
[0137] TF=the number of abnormal texts containing the segmented word in the abnormal texts in the first time period.
[0138] TF-IDF value = IDF*TF. Multiply the above two values to get the TF-IDF value.
[0139] S307 . Select the first N segmented words from the segmented word set in descending order of TF-IDF values to form a candidate keyword set, where N is a preset positive integer.
[0140] S308: Determine the candidate keywords in the candidate keyword set whose posterior probabilities are greater than a first preset threshold as keywords in the abnormal text within the first time period.
[0141] Among them, the TF-IDF value is to sort each word from the perspective of information content. However, this still results in a large number of "important" but not "special" words. Therefore, in this embodiment, the posterior probability (Bayesian formula) is used to screen the special words.
[0142] For a word segmentation, its posterior probability is:
[0143]
[0144] A is a single text in which the specified word x appears in the abnormal text within the first time period, B is any text in which any word appears in the abnormal text within the first time period, and C is a character x in which x appears in any text in the abnormal text within the first time period. j , D is a character x in the specified word x j Appears in any text of the abnormal text in the first time period.
[0145] S309: Output keywords.
[0146] In this embodiment, after detecting abnormal texts within the first time period from the text set to be detected, each text in the abnormal texts within the first time period is segmented to obtain a segmentation set, and then based on the TF-IDF value of each segmentation in the segmentation set and the posterior probability of each segmentation, keywords in the abnormal texts within the first time period are screened out and output, so that users can perform all-round control and analysis based on keywords, for example, search for corresponding abnormal texts based on keywords, reduce the number of abnormal texts to be viewed, and improve processing efficiency.
[0147] The abnormal text detection method provided in this embodiment obtains a set of texts to be detected after receiving a detection instruction, the set of texts to be detected includes texts whose generation time is before the current time, determines an abnormal text detection model corresponding to the first time according to the first time period carried by the detection instruction, inputs each text in the set of texts to be detected into the abnormal text detection model in turn, obtains the classification prediction probability of each text, and determines the target text whose classification label is the second classification label and whose classification prediction probability is greater than the first preset threshold as an abnormal text in the first time period according to the classification prediction probability of each text, and the abnormal text is the text that has not appeared in the set of texts to be detected before the first time. Thus, the accuracy of abnormal text detection is improved, and the scope of application is wide.
[0148] Combine the following Figure 6 , a specific embodiment is used to describe in detail the abnormal text detection method provided in the embodiment of the present application. In this embodiment, the online training of the abnormal text detection model is used as an example for description.
[0149] Figure 6 A flowchart of an abnormal text detection method provided in an embodiment of the present application, the execution subject of the method may be a server, such as Figure 6 As shown, the method may include:
[0150] S401: Receive a detection instruction, where the detection instruction carries a first time period.
[0151] S402: Determine a time point where the time interval and the current time are a first time period as a first time.
[0152] S403: Acquire historical text data before the current time, where the historical text data includes text information and the time when the text information is generated.
[0153] S404: Replace characters used to indicate time in the text information with preset characters to obtain training text.
[0154] S405 , determining the classification label of the training text according to the generation time of the training text and the first time, and obtaining a training sample set, wherein the training sample set includes the training text and the classification label of the training text.
[0155] Specifically, after the first time is determined, the training text is also determined, and the classification label of the training text can be determined according to the generation time of the training text and the first time. Figure 6 As shown, the training text generated before the first time is determined as a negative sample, and the classification label of the negative sample is the first classification label, and the classification label of the training text generated after the first time is determined as a positive sample, and the classification label of the positive sample is the second classification label. Among them, the first classification label is 0, and the second classification label is 1.
[0156] S406: Perform text classification model training based on the training sample set. In any training process, the training text is used as the input of the text classification model, and the classification prediction probability of the training text is output.
[0157] S407: According to the classification prediction probability of the training text and the classification label of the training text, the parameters of the text classification model are adjusted until the training stop condition is met.
[0158] S408: Determine the text classification model that meets the training stop condition as an abnormal text detection model.
[0159] S409: Obtain a text set to be detected, where the text set to be detected includes texts generated before the current time.
[0160] S410 , input each text in the to-be-detected text set into the abnormal text detection model in sequence to obtain the classification prediction probability of each text.
[0161] Specifically, each text in the text set to be detected is sequentially input into the abnormal text detection model trained through S401-S408.
[0162] S411 . Divide the to-be-detected text set into a first text set, a second text set, and a third text set according to the classification prediction probability of each text.
[0163] Among them, the classification prediction probability of the text in the first text set is greater than 0 and less than or equal to the first threshold, the classification prediction probability of the text in the second text set is greater than the first threshold and less than or equal to the second threshold, the classification prediction probability of the text in the third text set is greater than the second threshold, and the first threshold is less than the second threshold.
[0164] For example, the first threshold is a value between 0 and 0.2, and the second threshold is a value between 0.5 and 0.8. Figure 8 As shown, the classification prediction probability of the text in the first text set is greater than 0 and less than or equal to 0.15, the classification prediction probability of the text in the second text set is greater than 0.15 and less than or equal to 0.75, and the classification prediction probability of the text in the third text set is greater than 0.75.
[0165] S412: Determine the proportion of texts with the second classification label in the first text set, the proportion of texts with the second classification label in the second text set, and the proportion of texts with the second classification label in the third text set.
[0166] S413: Determine the text in the text set whose proportion in the first text set, the second text set, and the third text set is greater than a first preset threshold as the target text.
[0167] S414: Segment each text in the abnormal text within the first time period to obtain a segmentation set.
[0168] S415: Calculate the TF-IDF value of each word in the word set and the posterior probability of each word.
[0169] The calculation of the TF-IDF value of each word and the posterior probability of each word can be found in Figure 5 The description in the illustrated embodiment will not be repeated here.
[0170] S416. Select the first N word segments from the word set in descending order of TF-IDF values to form a candidate keyword set, where N is a preset positive integer.
[0171] S417: Determine the candidate keywords in the candidate keyword set whose posterior probabilities are greater than a first preset threshold as keywords in the abnormal text within the first time period.
[0172] S418. Output keywords.
[0173] After inputting the keywords, operation and maintenance analysts can conduct analysis or verify the situation based on the keywords.
[0174] For example, the following table 1 is an example of keywords in an abnormal text within a first time period:
[0175] Table 1 Keywords in abnormal texts during the first period
[0176]
[0177]
[0178] As shown in Table 1 above, the current time is October 1st, and September 1st is taken as the first time, and the keywords for abnormal text detection are performed on the text set to be detected (including texts generated before the current time). The keywords are sorted in order from large to small according to the TF-IDF value. The bold ones are the final keywords filtered after the posterior probability, and the keywords screened are those with a posterior probability greater than 0.8. It can be seen that the TF-IDF value alone is not enough to filter common words, and it is necessary to filter in combination with the delayed probability. Among the final keywords, XX6, XX13, XX14, XX15, XX18 and XX22 are the keywords corresponding to the abnormal texts added in September. The method of the embodiment of the present application can accurately detect recent abnormal complaints without human intervention, and output keywords.
[0179] Figure 7 A structural diagram of an abnormal text detection model training device provided in an embodiment of the present application is shown in FIG. Figure 7As shown, the device may include: an acquisition module 11, a processing module 12, an adjustment module 13 and an output module 14, wherein:
[0180] The acquisition module 11 is used to obtain a training sample set, which includes training text and classification labels of the training text. The characters used to represent time in the training text are replaced with preset characters, wherein the classification label of the training text generated before a first time is a first classification label, and the classification label of the training text generated after the first time is a second classification label, and the first time is before the current time.
[0181] The processing module 12 is used to train the text classification model according to the training sample set. In any training process, the training text is used as the input of the text classification model, and the classification prediction probability of the training text is output.
[0182] The adjustment module 13 is used to adjust the parameters of the text classification model according to the classification prediction probability of the training text and the classification label of the training text until the training stop condition is met.
[0183] The output module 14 is used to output the text classification model determined by satisfying the training stop condition as an abnormal text detection model.
[0184] Optionally, the acquisition module 11 is used for:
[0185] Determine the first time;
[0186] Obtaining historical text data before the current time, the historical text data including text information and the time when the text information was generated;
[0187] The characters used to indicate time in the text information are replaced with preset characters to obtain training text;
[0188] The classification label of the training text is determined according to the generation time of the training text and the first time.
[0189] Optionally, the acquisition module 11 is specifically used for:
[0190] Receiving a model training instruction, where the model training instruction carries a first time period; or determining a preset time period as the first time period;
[0191] A time point at which the time interval and the current time are a first time period is determined as a first time.
[0192] Optionally, the adjustment module 13 is used for:
[0193] Construct a loss function based on the classification prediction probability of the training text and the classification label of the training text;
[0194] Based on the loss function, back propagation adjusts the parameters of the text classification model.
[0195] Optionally, the text classification model is used to extract semantic information of the training text, convert the semantic information into a vector, convert the vector into a scalar using a fully connected layer, convert the scalar using an activation function, and obtain the classification prediction probability of the training text.
[0196] Optionally, the text classification model is a pre-trained bidirectional encoder representation BERT model from the transformer.
[0197] Figure 8 A schematic diagram of the structure of an abnormal text detection device provided in an embodiment of the present application is shown in FIG. Figure 8 As shown, the device may include: an acquisition module 21, a first determination module 22, a processing module 23, and a second determination module 14, wherein:
[0198] The acquisition module 21 is used to acquire a text set to be detected after receiving a detection instruction, where the text set to be detected includes texts generated before the current time, and the detection instruction carries a first time period.
[0199] The first determination module 22 is used to determine an abnormal text detection model corresponding to the first time according to the first time period. Figure 2 The method training of the illustrated embodiment obtains that the first time is a time point whose time interval with the current time is a first time period.
[0200] The processing module 23 is used to input each text in the to-be-detected text set into the abnormal text detection model in sequence to obtain the classification prediction probability of each text.
[0201] The second determination module 24 is used to determine the target text whose classification label is the second classification label and whose classification prediction probability is greater than the first preset threshold as abnormal text within the first time period according to the classification prediction probability of each text. The abnormal text is the text that has not appeared in the text set to be detected before the first time, and the classification label of the text generated after the first time is the second classification label.
[0202] Optionally, the first determining module 22 is used to:
[0203] Determine the time point where the time interval and the current time are the first time period as the first time;
[0204] According to the first time to train the abnormal text detection model;
[0205] The trained abnormal text detection model is determined as the abnormal text detection model corresponding to the first time.
[0206] Optionally, the first determining module 22 is used to:
[0207] Determining, from a set of pre-trained abnormal text detection models, an abnormal text detection model corresponding to the first time;
[0208] The determined abnormal text detection model is determined as the abnormal text detection model corresponding to the first time.
[0209] Optionally, the second determining module 24 is used to:
[0210] According to the classification prediction probability of each text, the text set to be detected is divided into a first text set, a second text set and a third text set;
[0211] The classification prediction probability of the text in the first text set is greater than 0 and less than or equal to the first threshold, the classification prediction probability of the text in the second text set is greater than the first threshold and less than or equal to the second threshold, the classification prediction probability of the text in the third text set is greater than the second threshold, and the first threshold is less than the second threshold;
[0212] Determine the proportion of texts whose classification label is the second classification label in the first text set, the proportion of texts whose classification label is the second classification label in the second text set, and the proportion of texts whose classification label is the second classification label in the third text set;
[0213] The texts in the text sets whose proportions among the first text set, the second text set and the third text set are greater than a first preset threshold are determined as target texts.
[0214] Optionally, the processing module 23 is further used to: filter keywords in the abnormal text within the first time period according to a preset method, and output the keywords.
[0215] Optionally, the processing module 23 is used to: segment each text in the abnormal text within the first time period to obtain a segmentation set;
[0216] Calculate the TF-IDF value of each word in the word set and the posterior probability of each word;
[0217] In descending order of TF-IDF values, select the first N words from the word set to form a candidate keyword set, where N is a preset positive integer;
[0218] The candidate keywords in the candidate keyword set whose posterior probabilities are greater than a first preset threshold are determined as keywords in the abnormal text within the first time period.
[0219] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Figure 7 The abnormal text detection model training device shown or Figure 8The abnormal text detection device shown can execute the method embodiment corresponding to the computer device, and the aforementioned and other operations and / or functions of each module in the device are respectively for implementing the method embodiment corresponding to the computer device, which will not be repeated here for the sake of brevity.
[0220] In the above text, the abnormal text detection model training device and the abnormal text detection device of the embodiment of the present application are described from the perspective of functional modules in combination with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to execute, or a combination of hardware and software modules in the decoding processor to execute. Optionally, the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory, and completes the steps in the above method embodiment in combination with its hardware.
[0221] Fig. 9 It is a schematic block diagram of a computer device 700 provided in an embodiment of the present application.
[0222] like Fig. 9 As shown, the computer device 700 may include:
[0223] The memory 710 and the processor 720, the memory 710 is used to store the computer program and transmit the program code to the processor 720. In other words, the processor 720 can call and run the computer program from the memory 710 to implement the method in the embodiment of the present application.
[0224] For example, the processor 720 may be configured to execute the above method embodiments according to instructions in the computer program.
[0225] In some embodiments of the present application, the processor 720 may include but is not limited to:
[0226] General-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.
[0227] In some embodiments of the present application, the memory 710 includes but is not limited to:
[0228] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).
[0229] In some embodiments of the present application, the computer program may be divided into one or more modules, which are stored in the memory 710 and executed by the processor 720 to complete the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.
[0230] like Fig. 9 As shown, the computer device may also include:
[0231] The transceiver 730 may be connected to the processor 720 or the memory 710 .
[0232] The processor 720 may control the transceiver 730 to communicate with other devices, specifically, to send information or data to other devices, or to receive information or data sent by other devices. The transceiver 730 may include a transmitter and a receiver. The transceiver 730 may further include an antenna, and the number of antennas may be one or more.
[0233] It should be understood that the various components in the electronic device are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.
[0234] The present application also provides a computer storage medium on which a computer program is stored, and when the computer program is executed by a computer, the computer can perform the method of the above method embodiment. In other words, the present application embodiment also provides a computer program product containing instructions, and when the instructions are executed by a computer, the computer can perform the method of the above method embodiment.
[0235] When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integration. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (digital video disc, DVD)), or a semiconductor medium (e.g., a solid state drive (solid state disk, SSD)), etc.
[0236] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0237] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the module is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0238] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. For example, each functional module in each embodiment of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0239] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for training an abnormal text detection model, characterized in that: include: Acquire a training sample set, the training sample set comprising training text and classification labels of the training text, wherein characters used to indicate time in the training text are replaced with preset characters, wherein the classification label of the training text generated before a first time is a first classification label, and the classification label of the training text generated after the first time is a second classification label, and the first time is before the current time; Wherein, obtaining a training sample set includes: determining the first time; Acquire historical text data before the current time, the historical text data including text information and the time when the text information was generated; Replacing the characters used to indicate time in the text information with the preset characters so that the training text does not contain time information, thereby obtaining the training text; Determining a classification label of the training text according to the generation time of the training text and the first time; Performing text classification model training according to the training sample set, and in any training process, taking the training text as the input of the text classification model and outputting the classification prediction probability of the training text; According to the classification prediction probability of the training text and the classification label of the training text, the parameters of the text classification model are adjusted until the training stop condition is met; Outputting the text classification model determined by satisfying the training stop condition as an abnormal text detection model, the abnormal text detection model is used to detect abnormal texts between the first time and the current time from the to-be-detected text set generated before the current time by using semantic information of the training text, the abnormal texts being texts that have not appeared in the to-be-detected text set before the first time; Wherein, the text classification model is used for: Extracting semantic information of the training text; Converting the semantic information into a vector; Convert the vector into a scalar using a fully connected layer; The scalar is transformed using an activation function to obtain the classification prediction probability of the training text.
2. The method according to claim 1, characterized in that The determining the first time includes: receiving a model training instruction, wherein the model training instruction carries a first time period; or determining a preset time period as the first time period; The time point at which the time interval and the current time are the first time period is determined as the first time.
3. The method according to claim 1, characterized in that: The step of adjusting the parameters of the text classification model according to the classification prediction probability of the training text and the classification label of the training text includes: Constructing a loss function according to the classification prediction probability of the training text and the classification label of the training text; According to the loss function, back propagation adjusts the parameters of the text classification model.
4. A method for detecting abnormal text, characterized in that: include: After receiving the detection instruction, obtaining a text set to be detected, the text set to be detected includes texts generated before the current time, and the detection instruction carries a first time period; Determine, according to the first time period, an abnormal text detection model corresponding to the first time, wherein the abnormal text detection model is trained according to the method according to any one of claims 1 to 3, and the first time is a time point whose time interval with the current time is the first time period; Input each text in the to-be-detected text set into the abnormal text detection model in turn to obtain a classification prediction probability of each text; According to the classification prediction probability of each text, the text set to be detected is divided into a first text set, a second text set and a third text set; The classification prediction probability of the text in the first text set is greater than 0 and less than or equal to a first threshold, the classification prediction probability of the text in the second text set is greater than the first threshold and less than or equal to a second threshold, the classification prediction probability of the text in the third text set is greater than the second threshold, and the first threshold is less than the second threshold; Determine the proportion of texts with the second classification label in the first text set, the proportion of texts with the second classification label in the second text set, and the proportion of texts with the second classification label in the third text set; Determine the text in the text set whose proportion in the first text set, the second text set and the third text set is greater than a first preset threshold as the target text; According to the classification prediction probability of each text, the target text whose classification label is the second classification label and whose classification prediction probability is greater than the first preset threshold is determined as an abnormal text within the first time period. The abnormal text is a text that has not appeared in the text set to be detected before the first time, and the classification label of the text generated after the first time is the second classification label.
5. The method according to claim 4, characterized in that The determining, according to the first time period, an abnormal text detection model corresponding to the first time period includes: Determine the time point where the time interval and the current time are the first time period as the first time; Training an abnormal text detection model according to the first time; The trained abnormal text detection model is determined as the abnormal text detection model corresponding to the first time.
6. The method according to claim 4, characterized in that The determining, according to the first time period, an abnormal text detection model corresponding to the first time period includes: Determine the time point where the time interval and the current time are the first time period as the first time; Determining, from a set of pre-trained abnormal text detection models, an abnormal text detection model corresponding to the first time; The determined abnormal text detection model is determined as the abnormal text detection model corresponding to the first time.
7. The method according to claim 4, characterized in that The method further comprises: Performing word segmentation on each text in the abnormal text within the first time period to obtain a word segmentation set; Calculate the TF-IDF value of each word in the word set and the posterior probability of each word; In descending order of TF-IDF values, the first N segmented words are selected from the segmented word set to form a candidate keyword set, where N is a preset positive integer; Determine the candidate keywords in the candidate keyword set whose posterior probabilities are greater than a first preset threshold as keywords in the abnormal text within the first time period; The keyword is output.
8. An abnormal text detection model training device, characterized in that: include: An acquisition module is used to acquire a training sample set, wherein the training sample set includes a training text and a classification label of the training text, wherein characters used to indicate time in the training text are replaced with preset characters, wherein the classification label of the training text generated before a first time is a first classification label, and the classification label of the training text generated after the first time is a second classification label, and the first time is before the current time; Wherein, obtaining a training sample set includes: determining the first time; Acquire historical text data before the current time, the historical text data including text information and the time when the text information was generated; Replacing the characters used to indicate time in the text information with the preset characters so that the training text does not contain time information, thereby obtaining the training text; Determining a classification label of the training text according to the generation time of the training text and the first time; A processing module, used for training a text classification model according to the training sample set, and in any training process, taking the training text as the input of the text classification model and outputting the classification prediction probability of the training text; An adjustment module, used for adjusting the parameters of the text classification model according to the classification prediction probability of the training text and the classification label of the training text until a training stop condition is met; An output module, used for outputting the text classification model determined by satisfying the training stop condition as an abnormal text detection model, wherein the abnormal text detection model is used for detecting abnormal texts between the first time and the current time from the text set to be detected generated before the current time by using semantic information of the training text, wherein the abnormal texts are texts that have not appeared in the text set to be detected before the first time; Wherein, the text classification model is used for: Extracting semantic information of the training text; Converting the semantic information into a vector; Convert the vector into a scalar using a fully connected layer; The scalar is transformed using an activation function to obtain the classification prediction probability of the training text.
9. An abnormal text detection device, characterized in that: include: An acquisition module, configured to acquire a text set to be detected after receiving a detection instruction, wherein the text set to be detected includes texts generated before a current time, and the detection instruction carries a first time period; A first determination module is used to determine an abnormal text detection model corresponding to a first time according to the first time period, wherein the abnormal text detection model is trained according to the method according to any one of claims 1 to 3, and the first time is a time point whose time interval with the current time is the first time period; A processing module, used for inputting each text in the to-be-detected text set into the abnormal text detection model in sequence to obtain a classification prediction probability of each text; A second determination module is used to divide the to-be-detected text set into a first text set, a second text set and a third text set according to the classification prediction probability of each text; The classification prediction probability of the text in the first text set is greater than 0 and less than or equal to a first threshold, the classification prediction probability of the text in the second text set is greater than the first threshold and less than or equal to a second threshold, the classification prediction probability of the text in the third text set is greater than the second threshold, and the first threshold is less than the second threshold; Determine the proportion of texts with the second classification label in the first text set, the proportion of texts with the second classification label in the second text set, and the proportion of texts with the second classification label in the third text set; Determine the text in the text set whose proportion in the first text set, the second text set and the third text set is greater than a first preset threshold as the target text; According to the classification prediction probability of each text, the target text whose classification label is the second classification label and whose classification prediction probability is greater than the first preset threshold is determined as an abnormal text within the first time period. The abnormal text is a text that has not appeared in the text set to be detected before the first time, and the classification label of the text generated after the first time is the second classification label.
10. A computer device, characterized in that: include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 3 or 4 to 7.
11. A computer-readable storage medium, characterized in that: Comprising instructions which, when run on a computer program, cause the computer to perform the method as claimed in any one of claims 1 to 3 or 4 to 7.
12. A computer program product comprising instructions, characterized in that When the instructions are executed on a computer, the computer is caused to perform the method according to any one of claims 1 to 3 or 4 to 7.
Citation Information
Patent Citations
Abnormal text detection method and device
CN109582833A
Fault detection method, device and equipment
CN111061581A