Method, apparatus, computer-readable medium, and electronic device for detecting abnormal text
By performing feature extraction and multi-model mapping processing on the text, the method of determining abnormal fragments in the text is solved, and the problem of low positioning accuracy in the prior art is improved.
Patent Information
- Application Number
- CN202210073277.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-01-21
AI Technical Summary
In the prior art, the positioning accuracy of abnormal fragments in text is low, resulting in low detection accuracy.
By obtaining the feature sequence of the text to be detected, and mapping the feature sequences using multiple preset models, obtaining the abnormal probability of the feature fragment, and determining the abnormal fragment in the text based on these probabilities.
The accuracy and accuracy of abnormal text detection are improved, and the accuracy and accuracy of detection results are enhanced by considering the local features in the text to be detected and model detection of multiple granularities.
Smart Images

Figure CN114490935B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical fields of computers and artificial intelligence, and particularly relates to a method, device, computer-readable medium, and electronic device for detecting abnormal text. Background Art
[0002] In many cases, it is necessary to proofread text data to solve abnormal problems such as typos, semantic and grammatical errors in the text. Currently, in natural language processing, the commonly used detection method is to identify abnormal positions in the text through a sequence labeling model. Sequence labeling means given a text sequence, analyzing each element in the text sequence to determine the abnormal probability of each element, and finally identifying the element with a relatively large abnormal probability as an abnormal element. However, the sequence labeling method usually has a positioning offset problem, that is, there is a certain gap between the identified abnormal element and the true abnormal element. Therefore, the detection accuracy of this detection method is not high and needs to be improved.
[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of this application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The purpose of this application is to provide a method, device, computer-readable medium, and electronic device for detecting abnormal text, so as to solve the problem of low positioning accuracy of abnormal fragments in the text in the related art.
[0005] Other features and advantages of this application will become apparent through the following detailed description, or will be partially learned through the practice of this application.
[0006] According to one aspect of the embodiments of this application, a method for detecting abnormal text is provided, including:
[0007] Obtain a text to be detected composed of multiple characters;
[0008] Extract features from the text to be detected to obtain a feature sequence of the text to be detected, where the feature sequence includes context features corresponding to multiple characters in the text to be detected;
[0009] Perform mapping processing on the feature sequence through multiple preset models respectively to obtain processing results corresponding to each preset model; wherein, the processing result of the preset model includes the abnormal probability of a feature segment in the feature sequence, and the feature segment includes context features of at least one character; the lengths of the feature segments corresponding to the processing results of different preset models are different;
[0010] Determine the abnormal segments in the text to be detected according to the abnormal probabilities of the feature segments indicated by the processing results of each preset model.
[0011] According to one aspect of the embodiments of the present application, there is provided a detection device for abnormal text, including:
[0012] A text acquisition module, configured to acquire the text to be detected composed of multiple characters;
[0013] A feature extraction module, configured to extract features from the text to be detected to obtain a feature sequence of the text to be detected, where the feature sequence includes context features corresponding to multiple characters in the text to be detected;
[0014] A mapping processing module, configured to perform mapping processing on the feature sequence through multiple preset models respectively to obtain processing results corresponding to each preset model; wherein, the processing result of the preset model includes the abnormal probability of different feature segments in the feature sequence, and the feature segment includes context features of at least one character; the lengths of the feature segments corresponding to the processing results of different preset models are different;
[0015] An abnormal segment determination module, configured to determine the abnormal segments in the text to be detected according to the abnormal probabilities of the feature segments indicated by the processing results of each preset model.
[0016] In an embodiment of the present application, the device further includes:
[0017] A sample data acquisition module, configured to acquire sample data composed of multiple characters, where the characters in the sample data have a first label indicating an abnormal state;
[0018] A second label generation module, configured to determine multiple sample segments in the sample data according to each preset segment length based on multiple preset segment lengths, and assign a second label indicating an abnormal state to the sample segments according to the first label corresponding to the sample segments;
[0019] A model training module, configured to use the sample data with the second label corresponding to each preset segment length as training samples to train a neural network model through the training samples to obtain a preset model corresponding to each preset segment length.
[0020] In an embodiment of the present application, the second label generation module is specifically configured to:
[0021] Set a window with the preset segment length as the window width, and use all the characters included in the window in the sample data as sample segments, where the window slides from the start position to the end position of the sample data according to a set step size.
[0022] In one embodiment of the present application, the first label includes a normal label and an abnormal label; the second label generation module is further configured to:
[0023] Generate a second label for the sample segment according to the total number of abnormal labels within the window and the window width.
[0024] In one embodiment of the present application, during the training process of the neural network model, the cross entropy between the predicted value of the neural network model for the training sample and the second label of the training sample is used as the loss function, and the model parameters of the neural network model are updated based on the loss function.
[0025] In one embodiment of the present application, the feature extraction module includes:
[0026] A character segmentation unit, configured to perform character segmentation on the text to be detected, obtain a plurality of characters arranged in sequence, and convert each character in the plurality of characters arranged in sequence into a corresponding character label according to a preset dictionary, so as to obtain a character sequence of the text to be detected;
[0027] A feature extraction unit, configured to perform context feature extraction on the character sequence to obtain a feature sequence of the text to be detected.
[0028] In one embodiment of the present application, the feature extraction unit is specifically configured to:
[0029] Determine a semantic vector and a position vector corresponding to the character label according to the character label in the character sequence;
[0030] Generate a vector to be feature-extracted according to the character label in the character sequence and the semantic vector and position vector corresponding to the character label;
[0031] Perform context feature extraction on the vector to be feature-extracted to obtain a feature sequence of the text to be detected.
[0032] In one embodiment of the present application, the mapping processing module includes:
[0033] A feature segment determination unit, configured to determine a feature segment in the feature sequence by means of a sliding window method according to a preset segment length corresponding to the preset model;
[0034] An abnormal feature representation acquisition unit, configured to acquire an abnormal feature representation of the feature segment through a convolutional layer of the preset model;
[0035] An abnormal probability acquisition unit, configured to perform mapping processing on the abnormal feature representation through a fully connected layer of the preset model to obtain an abnormal probability of the feature segment.
[0036] In one embodiment of the present application, the abnormal feature representation obtaining unit is specifically configured to:
[0037] Fuse all context features in the feature segment according to the model parameters of the convolutional layer of the preset model to obtain the abnormal feature representation of the feature segment.
[0038] In one embodiment of the present application, the model parameters of the convolutional layer include a first weight parameter and a first base value parameter; the abnormal feature representation obtaining unit is further configured to:
[0039] Perform weighted summation on all context features in the feature segment through the first weight parameter to obtain a weighted feature;
[0040] Superimpose the weighted feature and the first base value parameter to obtain the abnormal feature representation of the feature segment.
[0041] In one embodiment of the present application, the fully connected layer of the preset model includes a second weight parameter and a second base value parameter; the abnormal probability obtaining unit is specifically configured to:
[0042] Multiply the abnormal feature representation by the second weight parameter and then add the second base value parameter to obtain a feature to be activated;
[0043] Process the feature to be activated through a preset activation function to obtain the abnormal probability of the feature segment.
[0044] In one embodiment of the present application, the abnormal segment determining module is specifically configured to:
[0045] Determine the maximum abnormal probability among the abnormal probabilities of the feature segments indicated by the processing results of each preset model;
[0046] Use the multiple characters indicated by the feature segment corresponding to the maximum abnormal probability as the abnormal segment in the text to be detected.
[0047] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the method for detecting abnormal text in the above technical solution.
[0048] According to one aspect of the embodiments of the present application, there is provided an electronic device, which includes: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the method for detecting abnormal text in the above technical solution by executing the executable instructions.
[0049] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for detecting abnormal text in the above technical solution.
[0050] In the technical solution provided by the embodiments of the present application, multiple preset models are respectively used to process a feature sequence to obtain a processing result. Among them, the processing result includes the abnormal probability of a feature segment, that is, the text to be detected is divided into multiple segments for abnormal detection, rather than directly detecting the entire sentence of the text to be detected, so that the local features in the text to be detected are fully considered in the abnormal detection process, improving the detection accuracy; at the same time, since the lengths of the feature segments corresponding to the processing results of different preset models are different, it is equivalent to detecting the text to be detected by models with multiple granularities respectively, further improving the accuracy and precision of the detection result.
[0051] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0053] Figure 1 Schematically shown is an exemplary system architecture block diagram applying the technical solution of the present application.
[0054] Figure 2 Schematically shown is a flowchart of a method for detecting abnormal text provided by an embodiment of the present application.
[0055] Figure 3 Schematically shown is a flowchart of feature extraction for text to be detected provided by an embodiment of the present application.
[0056] Figure 4 Schematically shown is a schematic diagram for determining a preset segment length provided by an embodiment of the present application.
[0057] Figure 5 Schematically shown is a flowchart of a method for constructing a preset model provided by an embodiment of the present application.
[0058] Figure 6 The model structure diagram applying the technical solution of the present application is schematically shown.
[0059] Figure 7 The application flow chart of the technical solution of the present application in a scenario is schematically shown.
[0060] Figure 8 The structural block diagram of the detection device for abnormal text provided by the embodiment of the present application is schematically shown.
[0061] Figure 9 The structural block diagram of the computer system of the electronic device suitable for implementing the embodiment of the present application is schematically shown. Detailed implementation manners
[0062] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0063] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will recognize that the technical solution of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be employed. In other instances, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0064] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0065] The flow charts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0066] Figure 1 The exemplary system architecture block diagram applying the technical solution of the present application is schematically shown.
[0067] Such as Figure 1As shown, the system architecture 100 may include a terminal device 110, a network 120, and a server 130. The terminal device 110 may include a smart phone, a tablet computer, a laptop computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, and so on. The server 130 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The network 120 may be a communication medium of various connection types capable of providing a communication link between the terminal device 110 and the server 130. For example, it may be a wired communication link or a wireless communication link.
[0068] According to the implementation requirements, the system architecture in the embodiments of the present application may have any number of terminal devices, networks, and servers. For example, the server 130 may be a server group composed of multiple server devices. In addition, the technical solutions provided in the embodiments of the present application may be applied to the terminal device 110, or may be applied to the server 130, or may be jointly implemented by the terminal device 110 and the server 130. The present application does not make special limitations on this.
[0069] The method for detecting abnormal text provided in the embodiments of the present application is executed by the server 130. Correspondingly, the device for detecting abnormal text is set in the server 130. However, it is easy for those skilled in the art to understand that the method for detecting abnormal text provided in the embodiments of the present application may also be executed by the terminal device 110. Correspondingly, the device for detecting abnormal text may also be set in the terminal device 110. No special limitations are made on this in this exemplary embodiment.
[0070] For example, the server 130 obtains a text to be detected composed of multiple characters, and then extracts features from the text to be detected to obtain a feature sequence of the text to be detected. The feature sequence includes the context features corresponding to multiple characters in the text to be detected. Next, the server 130 performs mapping processing on the feature sequence through multiple preset models respectively to obtain the processing results corresponding to each preset model; among them, the lengths of the feature segments corresponding to the processing results of different preset models are different. And, in the processing result of a preset model, it includes the abnormal probabilities of multiple feature segments in the feature sequence, and the feature segment includes the context features of at least one character. Finally, the server 130 determines the abnormal segment in the text to be detected according to the abnormal probabilities of the feature segments indicated by the processing results of each preset model.
[0071] In one embodiment of the present application, after the server 130 determines the abnormal segment in the text to be detected, it can return the abnormal segment in the text to be detected to the terminal device 110 through the network 120. The terminal device 110 can mark the abnormal segment in the text to be detected on the display interface. For example, the abnormal segment in the text to be detected can be highlighted, so that the abnormal segment in the text to be detected can be quickly and conveniently known through the display interface of the terminal device 110.
[0072] The technical solution provided by the embodiments of the present application can be implemented through artificial intelligence technology, such as generating a preset model through artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making.
[0073] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, machine learning / deep learning, autonomous driving, and intelligent transportation.
[0074] The following will make a detailed description of the method for detecting abnormal text provided by the present application in combination with specific embodiments.
[0075] Figure 2 Schematically shows a flowchart of the method for detecting abnormal text provided by an embodiment of the present application. This method can be implemented by a terminal device, such as Figure 1 the terminal device 110 shown; this method can also be implemented by a server, such as Figure 1 the server 130 shown. As Figure 2 shown, the method for detecting abnormal text provided by the embodiments of the present application includes steps 210 to 240, specifically as follows:
[0076] Step 210: Obtain the text to be detected composed of multiple characters.
[0077] Specifically, the text to be detected is composed of multiple characters, which can be a sentence or multiple sentences. The text to be detected is text data that has been determined to be abnormal but the abnormal location is unclear. The abnormal situation of the text data includes typos, grammatical errors, semantic errors, etc. in the text data. The text to be detected can be content obtained from text data, for example, a title or text sentence confirmed to be abnormal in an article. The text to be detected can also be abnormal text data obtained after performing speech recognition on voice data, or abnormal text data obtained by performing text recognition on image data, which is not limited by the embodiments of the present application. Among them, confirming that the text data is abnormal text can be achieved by a trained text detection model.
[0078] Step 220 , extract features from the text to be detected to obtain a feature sequence of the text to be detected, where the feature sequence includes context features corresponding to multiple characters in the text to be detected.
[0079] Specifically, feature extraction of the text to be detected is to extract contextual semantic features of the text to be detected, convert the text to be detected from text to feature vectors, obtain contextual features corresponding to multiple characters in the text to be detected, and form a feature sequence. The contextual features of a character contain the semantic information of the character in the text to be detected, that is, contain abnormal information of the character in the text to be detected.
[0080] In one embodiment of the present application, Figure 3 As shown, the process of extracting features of the text to be detected includes steps 310 to 320, specifically:
[0081] Step 310 , the text to be detected is processed by word segmentation to obtain a plurality of characters arranged in sequence, and each of the plurality of characters arranged in sequence is converted into a corresponding character label according to a preset dictionary to obtain a character sequence of the text to be detected.
[0082] Specifically, word segmentation processing is to segment the text to be detected to obtain multiple words arranged in order. The order of these multiple words is the order of the words in the text to be detected. Since it is not possible to process the text directly, after segmenting to obtain multiple words, each word needs to be converted into a corresponding word label, so as to obtain multiple word labels arranged in order, that is, the word sequence of the text to be detected. The word label is a kind of identification of the word, which is equivalent to the ID of the word and is recorded as Token.
[0083] In one embodiment of the present application, a character can be converted into a corresponding character tag according to a preset dictionary. The preset dictionary contains a large number of characters and character tags corresponding to the characters. A plurality of characters arranged in order are traversed, and for each character, a character identical to the character is searched in the preset dictionary, and the character tag corresponding to the identical character is used as the character tag of the character.
[0084] In one embodiment of the present application, when the text to be detected includes multiple sentences, in order to identify the sentences, a sentence start identifier [CLS] can be set at the beginning of the text to be detected, and a sentence end identifier [SEP] can be set at the end of each sentence. Generally speaking, the beginning of the text to be detected is before the first character of the text to be detected. There will be punctuation marks at the end of the sentence, so the punctuation marks in the sentence can be identified first, and then the sentence end identifier can be set at the punctuation marks. In one case, when there is no punctuation mark or other character after a character, it can be considered that the character is the end of the sentence, and the sentence end identifier can be set after the character. Of course, when the text to be detected has only one sentence, the sentence end identifier can be set only after the last character tag. Then, the final obtained character sequence is composed of the sentence start identifier, character tags, and sentence end identifiers.
[0085] Exemplarily, the text to be detected is "I am Chinese, I love China". The beginning of the text to be detected is before "I", so the sentence start identifier [CLS] is set before the character tag of "I". The "," in the text to be detected is regarded as the end of the sentence, and the "China" in "I love China" is regarded as the end of the sentence, so two sentence end identifiers [SEP] need to be set. After conversion, the obtained character sequence is: [CLS]Token1 Token2Token3 Token4Token5[SEP]Token6 Token7Token8 Token9[SEP], where the Token numbers only represent the sorting numbers of the corresponding characters in the text to be detected.
[0086] Step 320: Extract context features from the character sequence to obtain the feature sequence of the text to be detected.
[0087] Specifically, the context feature is a feature that can represent information such as the context and semantics of a word or phrase, and naturally can also reflect the abnormal information of the position where the word or phrase is located.
[0088] In one embodiment of the present application, the context features of the character sequence can be extracted by a pre-trained language model, and the pre-trained language model can be a BERT (Bidirectional Encoder Representation from Transformers) model.
[0089] The pre-trained language model is a model in natural language processing. Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies.
[0090] In one embodiment of the present application, the process of extracting context features from a character sequence specifically includes: determining a semantic vector and a position vector corresponding to a character label according to the character label in the character sequence; generating a vector to be feature-extracted according to the character label in the character sequence and the corresponding semantic vector and position vector; and performing context feature extraction on the vector to be feature-extracted to obtain a feature sequence of the text to be detected.
[0091] Specifically, the semantic vector of a character represents the information after fusing the global semantic information of the text to be detected and the semantic information of the character. For example, the semantic vector of a character can indicate the sentence in which the character is located (for example, if the text to be detected includes sentence A and sentence B, the semantic vector of the character can indicate whether the character is in sentence A or sentence B), the type of the sentence in which the character is located (such as a title or the main text), and the semantic vector is determined by the pre-trained language model according to the character label and the text to be detected. The position vector of a character represents the position information of the character in the text to be detected. Since the semantic information carried by characters in different positions in the text to be detected will be different, the position vector of the character is added so that the position information of the character is considered in the process of context feature extraction, thereby making the context feature extraction more accurate. The position vector is determined by the pre-trained language model according to the character label and the text to be detected.
[0092] After determining the semantic vectors and position vectors corresponding to the word tags in the word sequence, the word tags are superimposed with the corresponding semantic vectors and position vectors to obtain the vector to be feature-extracted. Exemplarily, for the word sequence: [CLS]Token1Token2 Token3 Token4Token5[SEP], the semantic vectors corresponding to each word tag (sorted according to the word sequence) are: EC E1 E2 E3 E4E5 ES, and the position vectors corresponding to each word tag (sorted according to the word sequence) are: PC P1P2 P3 P4P5 PS. Then the vector to be feature-extracted generated is: [CLS]+EC+PC Token1+E1+P1 Token2+E2+P2Token3+E3+P3 Token4+E4+P4Token5+E5+P5[SEP]+E6+P6.
[0093] Finally, context features are extracted from the vector to be feature-extracted to obtain a feature sequence. The feature sequence includes the context features of each word. Denote the context feature of the i-th word as h i , then the feature sequence corresponding to the text to be detected with n words is h1h2h3…h i …h n .
[0094] Continue to refer to Figure 2 , step 230, map the feature sequence through multiple preset models respectively to obtain the processing results corresponding to each preset model; wherein, the processing result of the preset model includes the anomaly probability of the feature segment in the feature sequence, and the feature segment includes the context features of at least one word; the lengths of the feature segments corresponding to the processing results of different preset models are different.
[0095] Specifically, the preset model is used to predict the anomaly situation of the feature sequence. The processing result obtained by the preset model mapping the feature sequence includes the anomaly probabilities of multiple feature segments. That is to say, the preset model divides the feature sequence into multiple feature segments, and then predicts the anomaly probability of each feature segment. A feature segment is equivalent to a part of the feature sequence, so a feature segment includes the context features of at least one word.
[0096] In an embodiment of the present application, the feature sequence is respectively mapped by multiple preset models. Among the processing results obtained by each preset model, the lengths of the feature segments are different. The length of the feature segment refers to the number of context features that make up the feature segment (equivalent to the number of characters corresponding to the feature segment). Exemplarily, in the present application, 5 preset models are used to respectively map the feature sequence. The length of the feature segment corresponding to the 1st preset model is 1, the length of the feature segment corresponding to the 2nd preset model is 2, the length of the feature segment corresponding to the 3rd preset model is 3, the length of the feature segment corresponding to the 4th preset model is 4, and the length of the feature segment corresponding to the 5th preset model is 5.
[0097] In an embodiment of the present application, the process of the preset model for processing the feature sequence is as follows: according to the preset segment length corresponding to the preset model, the feature segments in the feature sequence are determined by the sliding window method, and the feature segments are mapped to obtain the anomaly probability of the feature segments.
[0098] Specifically, the length of the feature segment that the preset model needs to divide the feature sequence into is preset in advance, that is, the preset segment length. Taking the preset segment length as the width of a window, and then sliding this window along the feature sequence. During the sliding of the window, the segment of the feature sequence within the window is the feature segment. Thus, by the sliding window method, the feature sequence is divided into multiple feature segments. After the feature segments are divided, the feature segments are mapped to obtain the anomaly probability of the feature segments.
[0099] During the sliding of the window, the window slides with a set step length. Taking the head of the window as the calculation starting point, the distance between the current window head and the previous window head is the set step length. Generally, the set step length is 1, that is to say, the window moves backward by the distance of one character each time, and the window head slides from the starting position to the ending position of the feature sequence. Then, when dividing the feature sequence, when the preset segment length is k (k>1), let the feature sequence be h1h2h3…h i …h n ,then the feature segment is [h i :h i+k , indicating that the feature segment is from the context feature h i of the i-th character to the context feature h i+k of the (i + k)-th character, where the value of i ranges from 1 to n. It can be seen that when i = n - k, h i+k is h n , when i increases further, i + k will be greater than n, and at this time h i+kIt can be replaced by 0. When the preset segment length is 1, in fact, the context features in the feature sequence are segmented one by one, that is, the text to be detected is segmented into single characters, and the obtained feature segment is the context feature corresponding to a single character in the text to be detected. Exemplarily, as Figure 4 shown, taking the preset segment length of 2 as an example, the feature sequence is h1h2h3h4h5, and the set step size for window sliding is 1, and the obtained feature segments are: [h1:h2], [h2:h3], [h3:h4], [h4:h5], [h5:0].
[0100] In an embodiment of the present application, the process of mapping the feature segment is as follows: obtaining the abnormal feature representation of the feature segment through the convolutional layer of the preset model; performing mapping processing on the abnormal feature representation through the fully connected layer of the preset model to obtain the abnormal probability of the feature segment.
[0101] Specifically, the preset model has a convolutional layer and a fully connected layer. The convolutional layer is used to extract the abnormal feature representation of the feature segment, and the fully connected layer is used to calculate the abnormal probability according to the abnormal feature representation. Specifically, the convolutional layer fuses all the context features in the feature segment through the model parameters to obtain the abnormal feature representation. The model parameters of the convolutional layer include the first weight parameter and the first base value parameter. The first weight parameter and the first base value parameter are the model parameters obtained during the model training process. Different preset models have different model parameters for their convolutional layers. When performing the fusion processing, first, the first weight parameter is used to perform weighted summation on all the context features in the feature segment to obtain the weighted feature; then, the weighted feature is superimposed with the first base value parameter to obtain the abnormal feature representation of the feature segment. Denote the first weight parameter of the convolutional layer of the preset model corresponding to the preset segment length k as W 1k , and the first base value parameter as b 1k , denote all the context features in the i-th feature segment as [h i :h i+k , then the abnormal feature representation r ki of the i-th feature segment is as follows:
[0102] r ki =(W 1k [h i :h i+k +b 1k )
[0103] where k represents the preset segment length, r ki is the abnormal feature representation of the i-th feature segment extracted under the preset segment length k, [h i :h i+k represents the i-th feature segment, W 1k , b 1kThey are the model parameters of the convolutional layer in the preset model corresponding to the preset segment length k. The preset segment length is also equivalent to the width of the convolutional kernel of the convolutional layer. In fact, in this application, the feature sequences are processed by convolutional models with multiple granularities respectively.
[0104] After obtaining the abnormal feature representation, it is mapped through a fully connected layer to obtain the abnormal probability of the feature segment. Specifically, the fully connected layer has a second weight parameter and a second bias parameter. First, the abnormal feature representation is multiplied by the second weight parameter, and then added to the second bias parameter to obtain the feature to be activated; finally, the preset activation function set by the fully connected layer is used to process the feature to be activated to obtain the abnormal probability of the feature segment. The preset activation function can be a ReLU function, a Sigmoid function, a Softmax function, a Linear function, etc., and can be selected according to actual needs when setting. Exemplarily, the second weight parameter of the fully connected layer of the preset model corresponding to the preset segment length k is denoted as W 2k and the second bias parameter is denoted as b 2k . If the Softmax function is used as the preset activation function, the abnormal probability of the feature segment is shown in the following formula:
[0105] p ki = softmax(W 2k r ki + b 2k )
[0106] where p ki is the abnormal probability of the i-th feature segment extracted under the preset segment length k, r ki is the abnormal feature representation of the i-th feature segment extracted under the preset segment length k, and W 2k , b 2k are the model parameters of the fully connected layer in the preset model corresponding to the preset segment length k.
[0107] In an embodiment of this application, before mapping the feature sequence through the preset model, it further includes the construction process of the preset model, as Figure 5 shown. This process includes steps 510 to 530, specifically:
[0108] Step 510, obtain sample data composed of multiple words, and the words in the sample data have a first label indicating the abnormal state.
[0109] Specifically, the sample data is abnormal text data with abnormal information annotation. The abnormal text data is also composed of multiple characters, and each character has a first label indicating the abnormal state. The abnormal state of a character refers to whether the character is in an abnormal state, and this abnormal state can be reflected by different annotation information of the first label. For example, if the first label is annotated as 0, it means that the character is not in an abnormal state (i.e., the character is in a normal state), so this type of first label can be recorded as a normal label. If the first label is annotated as 1, it means that the character is in an abnormal state, so this type of first label can be recorded as an abnormal label.
[0110] In an embodiment of the present application, for the same sample data, different annotation methods may result in different first labels for each character in the sample data. For example, for the sample data "The status of pets is higher than that of people", annotator 1 may think that "dou ren" is ungrammatical (i.e., abnormal), so both the characters "dou" and "ren" are marked as 1, and the rest of the characters are marked as 0; annotator 2 may think that the character "ren" is redundant, so the character "ren" is marked as 1, and the rest of the characters are marked as 0; annotator 3 may think that the character "dou" is ungrammatical, so the character "dou" is marked as 1, and the rest of the characters are marked as 0. Thus, for the sample data "The status of pets is higher than that of people", the following three annotation situations can be obtained as shown in the table below:
[0111] Table 1
[0112] Sample data Pet Animal 's Location Position All People Tall Annotator 1 0 0 0 0 0 1 1 0 Annotator 2 0 0 0 0 0 0 1 0 Annotator 3 0 0 0 0 0 1 0 0
[0113] Step 520: Based on multiple preset segment lengths, determine multiple sample segments in the sample data according to each preset segment length, and assign a second label indicating the abnormal state to the sample segment according to the first label corresponding to the sample segment.
[0114] Specifically, set multiple preset segment lengths, divide the sample data for each preset segment length to obtain multiple sample segments corresponding to the sample data, and assign a second label to the sample segment according to the first label of each character in the sample segment. The second label is calculated based on the first label and is used to indicate the abnormal condition of the sample segment.
[0115] The process of obtaining multiple sample segments from sample data according to the preset segment length is as follows: Set a window with the preset segment length as the window width, and take all the characters within the window in the sample data as a sample segment. Herein, the window slides from the start position to the end position of the sample data according to the set step size. That is to say, set a window with the preset segment length as the window width, and then slide the head of the window from the start position to the end position of the sample data according to the set step size. Each time it slides, all the characters within the window in the sample data form a sample segment, so as to obtain multiple sample segments. When the number of characters in the window is insufficient, it is padded with 0. Generally, the set step size is 1. The process of obtaining sample segments can refer to the process of obtaining feature segments in the previous text, and the two processes are similar. Exemplarily, taking the sample data in Table 1 above with a preset segment length of 2 as an example, the sample segments can be obtained: pet, of the, of the, status, position, all people, people are tall, tall 0.
[0116] Since there are only two cases, 0 and 1, for the first label of each character in the sample data, when directly using the first label for model training, there will be a large error in the loss function calculated by the predicted value of the model and the first label. For example, taking the annotation of the person 2 in Table 1 above as an example, the first label of the character "person" is 1, but the model will also predict a relatively high probability for the character "all", which will result in a large loss for the model. Then when the model updates the parameters through gradient backpropagation and retrains, it will mislead the model to make the probability prediction of the character "all" very small. This will cause the model to be confused, thereby reducing the prediction accuracy of the model.
[0117] Considering the above situation, this application reassigns a second label to the sample segments. During the window movement process, the second label of the sample segment is generated according to the total number of abnormal labels within the window and the window width. The second label of the sample segment includes two parts: a normal identifier and an abnormal identifier. Among them, the normal identifier is used to represent the probability that the sample segment is in a normal state, and the abnormal identifier is used to represent the probability that the sample segment is in an abnormal state. The sum of the normal identifier and the abnormal identifier is 1. Then, once the abnormal identifier is determined, the normal identifier is also determined.
[0118] In an embodiment of the present application, during the window movement, the ratio of the total number of abnormal tags in the window to the window width is used as the abnormal identifier in the second tag of the sample segment, and the normal identifier of the sample segment is 1 minus the abnormal identifier. When the window width is 1 (i.e., the preset segment length is 1), at this time the sample segment is a single character in the sample data, so the second tag of the sample segment is the same as the first tag of each character. When the window width is 2, taking the annotation of annotator 2 in Table 1 above as an example, the sample segments and their corresponding second tags (expressed in the form of (abnormal identifier, normal identifier)) can be obtained as follows: pet: (0,1), of thing: (0,1), of the: (0,1), status: (0,1), position capital: (0,1), capital people: (0.5,0.5), people tall: (0.5,0.5), tall 0: (0,1). Among them, pet: (0,1) means that the abnormal identifier of the sample segment "pet" is 0 and the normal identifier is 1. When the window width is 3, taking the annotation of annotator 2 in Table 1 above as an example, the sample segments and their corresponding second tags can be obtained as follows: pet of: (0,1), of thing the: (0,1), of the status: (0,1), status capital: (0,1), position capital people: (0.33,0.67), capital people tall: (0.33,0.67), people tall 0: (0.33,0.67), tall 00: (0,1). It can be seen that the data in the second tag no longer has only two values, and the data in the second tag is smoother, which can effectively alleviate the influence of the error in the first tag on the model.
[0119] Step 530: Use the sample data with second tags corresponding to each preset segment length as training samples, and train the neural network model with the training samples to obtain preset models corresponding to each preset segment length.
[0120] Specifically, after assigning second tags to the sample segments in the sample data, the sample data can be used as training samples to train the neural network model. Through the processing of the foregoing steps, a kind of sample data with second tags can be obtained for each preset segment length, that is, each preset segment length corresponds to a training sample. During the training process, the training sample corresponding to the preset segment length is used to train the neural network model of the preset segment length, and the preset model corresponding to the preset segment length is obtained.
[0121] In an embodiment of the present application, during the training process of the neural network model, the cross entropy between the predicted value of the neural network model for the training sample and the second tag of the training sample is used as the loss function, and the model parameters of the neural network model are updated based on the loss function. The model parameters are W 1k , b 1k , W 2k , b 2k and other parameters.
[0122] Specifically, the loss function Loss k is calculated as follows:
[0123]
[0124] where y 0ki represents the normal label in the i-th sample segment under the preset segment length k, and y 1ki represents the abnormal label in the i-th sample segment under the preset segment length k, and y 0ki + y 1ki = 1; p 0ki represents the probability that the i-th sample segment predicted by the neural network model with the preset segment length k is normal, and p 1ki represents the probability that the i-th sample segment predicted by the neural network model with the preset segment length k is abnormal.
[0125] In an embodiment of the present application, when training the neural network models corresponding to each preset segment length, an iterative training method can be adopted, that is, training each model in ascending order of the preset segment length.
[0126] In an embodiment of the present application, other suitable machine learning models can also be used to train the preset model. Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0127] Continue to refer to Figure 2 , step 240: Determine the abnormal segments in the text to be detected according to the abnormal probabilities of the feature segments indicated by the processing results of each preset model.
[0128] Specifically, the processing result of a preset model includes the abnormal probabilities of multiple feature segments of the same length, and the processing results of multiple preset models include the abnormal probabilities of multiple feature segments of multiple lengths. Determine the maximum abnormal probability among the multiple abnormal probabilities, and use the multiple characters indicated by the feature segment corresponding to the maximum abnormal probability as the abnormal segment in the text to be detected, thereby determining the length and position of the abnormal segment in the text to be detected. For example, if the feature segment corresponding to the maximum abnormal probability is the i-th feature segment with a preset segment length of k, then it is determined that the abnormal segment in the text to be detected is the segment composed of the i-th character to the (i + k)-th character.
[0129] In the technical solution provided by the embodiments of the present application, the feature sequence is processed by multiple preset models respectively to obtain the processing results. Among them, the processing results include the abnormal probabilities of the feature segments, that is, the text to be detected is divided into multiple segments for abnormal detection, rather than directly detecting the entire sentence of the text to be detected, so that the local features in the text to be detected are fully considered during the abnormal detection process, improving the detection accuracy; at the same time, since the lengths of the feature segments corresponding to the processing results of different preset models are different, it is equivalent to detecting the text to be detected by models with multiple granularities respectively, further improving the accuracy and precision of the detection results.
[0130] Figure 6 Schematically shows a model structure applying the technical solution of the present application. As Figure 6 shown, the model structure includes:
[0131] A text embedding module (TokenEMBEDDING) 610, which is used to perform word segmentation on the text to be detected, so as to convert the text to be detected 611 into a word sequence composed of word tags (tokens). Specifically, reference can be made to the relevant description of the foregoing step 310, which will not be elaborated here.
[0132] A vector superposition module (TASKEMBEDDING) 620, which is used to superpose the word tags in the word sequence with the corresponding semantic vectors and position vectors to generate a vector to be feature-extracted 621. Specifically, reference can be made to the relevant description of the foregoing step 320, which will not be elaborated here.
[0133] A BERT model (BERT MODEL) 630, where the BERT model is a pre-trained language model, which is used to perform context feature extraction on the vector to be feature-extracted and output a feature sequence 631.
[0134] Multi-granularity Convolution Module 640, which includes convolution models of 5 granularities (grams). Here, the granularity is the size of the convolution kernel, that is, the preset segment length. In the embodiments of this application, the granularities of the 5 convolution models are respectively: 1, 2, 3, 4, and 5. Each convolution model of a granularity performs mapping processing on the output feature sequence 631 to obtain the abnormal probability of the feature segments in the feature sequence. The length of the feature segment is the same as the granularity of the corresponding convolution model. The abnormal probability of the feature segment is equivalent to the prediction score (score) of the convolution model for this feature segment. Finally, the maximum value (MAXscore) is selected from each score to determine the abnormal segment in the text to be detected 611.
[0135] Figure 7 Schematically shows the application flowchart of the technical solution of this application in a scenario. As Figure 7 shown, this process includes:
[0136] S710. Obtain unsmooth text. The unsmooth text is the abnormal text, and the abnormal segment in the unsmooth text can be determined through the technical solution of this application.
[0137] S720. Input the unsmooth text into the unsmooth segment detection model. The unsmooth segment detection model is the model implementing the technical solution of this application, that is, the unsmooth segment detection model extracts features from the unsmooth segment to obtain a feature sequence; then multiple preset models respectively perform mapping processing on the feature sequence to obtain the processing results corresponding to each preset model. Among them, the processing result of the preset model includes the abnormal probability of the feature segment in the feature sequence, and the feature segment includes the context features of at least one character; the lengths of the feature segments corresponding to the processing results of different preset models are different. Finally, according to the maximum value of the abnormal probability of the feature segment indicated by the processing results of each preset model, the abnormal segment in the text to be detected is determined.
[0138] S730. Locate the unsmooth segment according to the model prediction result. Determine the specific position of the unsmooth segment according to the output result of the unsmooth segment detection model.
[0139] S740. Machine review system. Input the positioning result into the machine review system, and this system can perform manual review.
[0140] S750. Highlight the unsmooth segment. Through highlighting, the object can more quickly determine the unsmooth segment in the unsmooth text.
[0141] It should be noted that although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0142] The following introduces the device embodiments of this application, which can be used to execute the abnormal text detection method in the above embodiments of this application. Figure 8 The structural block diagram of the abnormal text detection device provided by the embodiments of this application is schematically shown. As Figure 8 shown, the abnormal text detection device provided by the embodiments of this application includes:
[0143] A text acquisition module 810, configured to acquire a text to be detected composed of multiple characters;
[0144] A feature extraction module 820, configured to perform feature extraction on the text to be detected to obtain a feature sequence of the text to be detected, where the feature sequence includes context features corresponding to multiple characters in the text to be detected;
[0145] A mapping processing module 830, configured to perform mapping processing on the feature sequence through multiple preset models respectively to obtain processing results corresponding to each preset model; wherein, the processing result of the preset model includes the abnormal probability of different feature segments in the feature sequence, and the feature segment includes context features of at least one character; the lengths of the feature segments corresponding to the processing results of different preset models are different;
[0146] An abnormal segment determination module 840, configured to determine an abnormal segment in the text to be detected according to the abnormal probability of the feature segment indicated by the processing result of each preset model.
[0147] In an embodiment of this application, the device further includes:
[0148] A sample data acquisition module, configured to acquire sample data composed of multiple characters, and the characters in the sample data have a first label indicating an abnormal state;
[0149] A second label generation module, configured to determine multiple sample segments in the sample data according to each preset segment length based on multiple preset segment lengths, and assign a second label indicating an abnormal state to the sample segment according to the first label corresponding to the sample segment;
[0150] A model training module, configured to use the sample data with a second label corresponding to each preset segment length as training samples, and train a neural network model with the training samples to obtain a preset model corresponding to each preset segment length.
[0151] In an embodiment of the present application, the second label generation module is specifically configured to:
[0152] Set a window with the preset segment length as the window width, and use all the characters included in the window in the sample data as sample segments, where the window slides from the start position to the end position of the sample data according to a set step size.
[0153] In an embodiment of the present application, the first label includes a normal label and an abnormal label; the second label generation module is further configured to:
[0154] Generate a second label for the sample segment according to the total number of abnormal labels in the window and the window width.
[0155] In an embodiment of the present application, during the training process of the neural network model, the cross entropy between the predicted value of the neural network model for the training sample and the second label of the training sample is used as a loss function, and the model parameters of the neural network model are updated based on the loss function.
[0156] In an embodiment of the present application, the feature extraction module 820 includes:
[0157] A character splitting unit, configured to split the text to be detected into characters to obtain a plurality of characters arranged in sequence, and convert each character in the plurality of characters arranged in sequence into a corresponding character label according to a preset dictionary to obtain a character sequence of the text to be detected;
[0158] A feature extraction unit, configured to perform context feature extraction on the character sequence to obtain a feature sequence of the text to be detected.
[0159] In an embodiment of the present application, the feature extraction unit is specifically configured to:
[0160] Determine a semantic vector and a position vector corresponding to the character label according to the character label in the character sequence;
[0161] Generate a vector to be feature-extracted according to the character label in the character sequence and the semantic vector and position vector corresponding to the character label;
[0162] Perform context feature extraction on the vector to be feature-extracted to obtain a feature sequence of the text to be detected.
[0163] In an embodiment of the present application, the mapping processing module 830 includes:
[0164] A feature segment determination unit, configured to determine a feature segment in the feature sequence by means of a sliding window method according to a preset segment length corresponding to the preset model;
[0165] An abnormal feature representation acquisition unit, configured to acquire an abnormal feature representation of the feature segment through a convolutional layer of the preset model;
[0166] An abnormal probability acquisition unit, configured to perform a mapping process on the abnormal feature representation through a fully connected layer of the preset model to obtain an abnormal probability of the feature segment.
[0167] In an embodiment of the present application, the abnormal feature representation acquisition unit is specifically configured to:
[0168] Fuse all context features in the feature segment according to model parameters of the convolutional layer of the preset model to obtain an abnormal feature representation of the feature segment.
[0169] In an embodiment of the present application, the model parameters of the convolutional layer include first weight parameters and first base value parameters; the abnormal feature representation acquisition unit is further configured to:
[0170] Perform a weighted sum on all context features in the feature segment through the first weight parameters to obtain weighted features;
[0171] Superimpose the weighted features and the first base value parameters to obtain an abnormal feature representation of the feature segment.
[0172] In an embodiment of the present application, the fully connected layer of the preset model includes second weight parameters and second base value parameters; the abnormal probability acquisition unit is specifically configured to:
[0173] Multiply the abnormal feature representation by the second weight parameters and then add the second base value parameters to obtain a feature to be activated;
[0174] Process the feature to be activated through a preset activation function to obtain an abnormal probability of the feature segment.
[0175] In an embodiment of the present application, the abnormal segment determination module 840 is specifically configured to:
[0176] Determine the maximum abnormal probability among the abnormal probabilities of the feature segments indicated by the processing results of each preset model;
[0177] Use multiple characters indicated by the feature segment corresponding to the maximum abnormal probability as the abnormal segment in the text to be detected.
[0178] Details of the detection device for abnormal text provided in the embodiments of the present application have been described in detail in the corresponding method embodiments and will not be elaborated here.
[0179] Figure 9 Schematically shows a block diagram of a computer system of an electronic device for implementing the embodiments of the present application.
[0180] It should be noted that Figure 9 The computer system 900 of the shown electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0181] As Figure 9 shown, the computer system 900 includes a central processing unit 901 (Central Processing Unit, CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 902 (Read-Only Memory, ROM) or the program loaded from the storage section 908 into the random access memory 903 (Random Access Memory, RAM). In the random access memory 903, various programs and data required for system operation are also stored. The central processing unit 901, the read-only memory 902, and the random access memory 903 are connected to each other via a bus 904. The input / output interface 905 (Input / Output interface, that is, I / O interface) is also connected to the bus 904.
[0182] The following components are connected to the input / output interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including such as a cathode ray tube (Cathode Ray Tube, CRT), a liquid crystal display (Liquid Crystal Display, LCD), etc. and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a local area network card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that the computer program read from it can be installed into the storage section 908 as needed.
[0183] In particular, according to embodiments of the present application, the processes described in each method flowchart can be implemented as computer software programs. For example, embodiments of the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by the central processing unit 901, various functions defined in the system of the present application are performed.
[0184] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0185] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the above-mentioned module, segment of a program, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0186] It should be noted that although several modules or units of devices for performing actions are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0187] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a portable hard drive, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0188] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0189] It should be understood that the present application is not limited to the exact structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A method for detecting abnormal text, characterized in that, Including: Obtain the text to be detected composed of multiple characters; Extract features from the text to be detected to obtain a feature sequence of the text to be detected, where the feature sequence includes context features corresponding to multiple characters in the text to be detected; Perform mapping processing on the feature sequence through multiple preset models respectively to obtain processing results corresponding to each preset model; wherein, the processing result of the preset model includes the abnormal probability of a feature segment in the feature sequence, and the feature segment includes context features of at least one character; the lengths of the feature segments corresponding to the processing results of different preset models are different; Determine the abnormal segment in the text to be detected according to the abnormal probability of the feature segment indicated by the processing results of each preset model; Among them, performing mapping processing on the feature sequence through multiple preset models respectively to obtain processing results corresponding to each preset model includes: Determine the feature segments in the feature sequence by the sliding window method according to the preset segment length corresponding to the preset model; Perform weighted summation on all context features in the feature segment through the first weight parameter of the convolutional layer of the preset model to obtain a weighted feature; superimpose the weighted feature and the first base value parameter of the convolutional layer to obtain an abnormal feature representation of the feature segment; Perform mapping processing on the abnormal feature representation through the fully connected layer of the preset model to obtain the abnormal probability of the feature segment.
2. The method for detecting abnormal text according to claim 1, characterized in that, Before performing mapping processing on the feature sequence through multiple preset models respectively to obtain processing results corresponding to each preset model, the method further includes: Obtain sample data composed of multiple characters, where the characters in the sample data have a first label indicating an abnormal state; Based on multiple preset segment lengths, determine multiple sample segments in the sample data according to each preset segment length, and assign a second label indicating an abnormal state to the sample segment according to the first label corresponding to the sample segment; Use the sample data with the second label corresponding to each preset segment length as a training sample, and train a neural network model through the training sample to obtain a preset model corresponding to each preset segment length.
3. The method for detecting abnormal text according to claim 2, characterized in that, Determining multiple sample segments in the sample data according to each preset segment length includes: Set a window with the preset segment length as the window width, and use all the characters included in the window in the sample data as a sample segment, where the window slides from the start position to the end position of the sample data according to a set step size.
4. The method for detecting abnormal text according to claim 3, characterized in that, The first label includes a normal label and an abnormal label; assigning a second label indicating an abnormal state to the sample segment according to the first label corresponding to the sample segment includes: Generate the second label of the sample segment according to the total amount of abnormal labels in the window and the window width.
5. The method for detecting abnormal text according to claim 3, characterized in that, During the training process of the neural network model, use the cross entropy between the predicted value of the neural network model for the training sample and the second label of the training sample as the loss function, and update the model parameters of the neural network model based on the loss function.
6. The method for detecting abnormal text according to claim 1, characterized in that, Feature extraction is performed on the text to be detected to obtain a feature sequence of the text to be detected, including: The text to be detected is segmented into characters to obtain a plurality of characters arranged in sequence, and each character in the plurality of characters arranged in sequence is converted into a corresponding character label according to a preset dictionary to obtain a character sequence of the text to be detected; Context feature extraction is performed on the character sequence to obtain a feature sequence of the text to be detected.
7. The method for detecting abnormal text according to claim 6, characterized in that, Context feature extraction is performed on the character sequence to obtain a feature sequence of the text to be detected, including: Determine the semantic vector and position vector corresponding to the character label according to the character label in the character sequence; Generate a vector to be feature-extracted according to the character label in the character sequence and the semantic vector and position vector corresponding to the character label; Context feature extraction is performed on the vector to be feature-extracted to obtain a feature sequence of the text to be detected.
8. The method for detecting abnormal text according to claim 1, characterized in that, The fully connected layer of the preset model includes a second weight parameter and a second base value parameter; mapping processing is performed on the abnormal feature representation through the fully connected layer of the preset model to obtain the abnormal probability of the feature segment, including: Multiply the abnormal feature representation by the second weight parameter and then add the second base value parameter to obtain a feature to be activated; Process the feature to be activated through a preset activation function to obtain the abnormal probability of the feature segment.
9. The method for detecting abnormal text according to claim 1, characterized in that, Determine the abnormal segment in the text to be detected according to the abnormal probability of the feature segment indicated by the processing results of each preset model, including: Determine the maximum abnormal probability among the abnormal probabilities of the feature segments indicated by the processing results of each preset model; Regard the plurality of characters indicated by the feature segment corresponding to the maximum abnormal probability as the abnormal segment in the text to be detected.
10. A detection device for abnormal text, characterized in that, Including: A text acquisition module for acquiring the text to be detected composed of a plurality of characters; A feature extraction module for performing feature extraction on the text to be detected to obtain a feature sequence of the text to be detected, and the feature sequence includes the context features corresponding to a plurality of characters in the text to be detected; A mapping processing module for performing mapping processing on the feature sequence through a plurality of preset models respectively to obtain the processing results corresponding to each preset model; wherein, the processing result of the preset model includes the abnormal probability of different feature segments in the feature sequence, and the feature segment includes the context features of at least one character; the lengths of the feature segments corresponding to the processing results of different preset models are different; An abnormal segment determination module for determining the abnormal segment in the text to be detected according to the abnormal probability of the feature segment indicated by the processing results of each preset model; Among them, the mapping processing module is specifically used for: Determine the feature segment in the feature sequence by means of a sliding window method according to the preset segment length corresponding to the preset model; Perform weighted summation on all context features in the feature segment through the first weight parameter of the convolution layer of the preset model to obtain a weighted feature; superimpose the weighted feature and the first base value parameter of the convolution layer to obtain the abnormal feature representation of the feature segment; The abnormal probability of the feature segment is obtained by performing a mapping process on the abnormal feature representation through the fully connected layer of the preset model.
11. A computer-readable medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the method for detecting abnormal text according to any one of claims 1 to 9.
12. An electronic device, characterized in that, It includes: A processor; And A memory for storing the executable instructions of the processor; Wherein, when the processor executes the executable instructions, the electronic device executes the method for detecting abnormal text according to any one of claims 1 to 9.
13. A computer program product, characterized in that, The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; The processor of the computer device reads and executes the computer instructions from the computer-readable storage medium, so that the computer device executes the method for detecting abnormal text according to any one of claims 1 to 9.
Citation Information
Patent Citations
Text error detection method and device based on artificial intelligence, and computer equipment
CN112434131A
BERT-based machine reading understanding method, apparatus and device, and storage medium
CN112464641A
Sensitive tendency expression detection method, apparatus and device, and storage medium
CN112732912A