Text Detection Method, Apparatus, Computer Device, and Readable Storage Medium
Through entity word segmentation and semantic model processing, combined with confusion degree and vector distance, the problem of inaccurate spam message recognition in the prior art is solved, and the high accuracy recognition of spam messages using text deformation is achieved.
Patent Information
- Application Number
- CN202110070007.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-19
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-01-19
AI Technical Summary
The prior art is difficult to accurately identify spam messages that use text deformation to bypass detection, resulting in inaccurate identification results.
The entity byte unit is obtained through entity word segmentation processing, and the local semantic model is used to predict the occurrence probability of entity byte unit and calculate the confusion degree; at the same time, the target global representation vector is obtained through the global semantic model, and the normal text vector distance is calculated, and the normal or anomalies of the text are determined based on the confusion degree and vector distance.
It improves the recognition accuracy of spam messages and can effectively identify text messages that use text deformation to convey spam messages.
Smart Images

Figure CN113569041B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a text detection method, apparatus, computer device, and readable storage medium. Background Art
[0002] With the rapid development of network technology, people's communication methods are constantly diversifying. Due to the rapidity and conciseness of text messages, they have become an important communication method in people's lives and work. However, because text messages cover a large number of users and have a low dissemination cost, they have become a medium for advertising or the transmission of some illegal information, resulting in the proliferation of spam messages. Users are frequently harassed by various spam messages. More seriously, some spam messages are phishing and fraudulent messages, which will cause economic losses and psychological harm to some teenagers and the elderly with less social experience.
[0003] In the prior art, a sensitive keyword-based method is usually used to identify spam messages. When it is detected that the text of a message contains sensitive keywords, the text of the message can be identified as a spam message, thereby realizing the automatic interception of spam messages. However, currently, in order to bypass the detection of the detection engine, there are endless ways to transform spam messages. For example, spam information is interspersed with normal text, homophonic transformation, substitution with pictographic characters, and substitution with pinyin are used to transmit spam information. Since the sensitive keywords in such message texts are separated or replaced, it is easy to bypass the detection, resulting in inaccurate spam message recognition results. Summary of the Invention
[0004] Embodiments of this application provide a text detection method, apparatus, computer device, and readable storage medium, which can screen out spam texts that bypass detection by text transformation and improve the recognition accuracy of spam texts.
[0005] On the one hand, embodiments of this application provide a text detection method, including:
[0006] Obtain a target text to be detected, perform entity word segmentation processing on the content of the target text to be detected, and obtain a number of entity byte units after entity word segmentation processing;
[0007] Perform entity prediction processing on a number of entity byte units to obtain the occurrence probabilities of the number of entity byte units in the target text to be detected, and determine the perplexity for measuring the semantic coherence of the target text to be detected according to the corresponding occurrence probabilities of the number of entity byte units;
[0008] Perform semantic feature extraction processing on the target text to be detected to obtain a target global representation vector of the target text to be detected;
[0009] Determine a normal text vector distance for measuring the normal semantics of the target text to be detected according to the target global representation vector and the reference global representation vector of the normal text type;
[0010] Determine whether the target text to be detected belongs to normal text or abnormal text according to the perplexity and the normal text vector distance.
[0011] An embodiment of the present application provides a text detection device on the one hand, including:
[0012] A target text processing module, configured to obtain the target text to be detected, perform entity word segmentation processing on the content of the target text to be detected, and obtain a number of entity byte units after entity word segmentation processing;
[0013] A perplexity determination module, configured to perform entity prediction processing on a number of entity byte units to obtain the occurrence probabilities of the number of entity byte units in the target text to be detected, and determine a perplexity for measuring the semantic coherence of the target text to be detected according to the occurrence probabilities corresponding to the number of entity byte units;
[0014] A vector acquisition module, configured to perform semantic feature extraction processing on the target text to be detected to obtain a target global representation vector of the target text to be detected;
[0015] A vector distance determination module, configured to determine a normal text vector distance for measuring the normal semantics of the target text to be detected according to the target global representation vector and the reference global representation vector of the normal text type;
[0016] A target text determination module, configured to determine whether the target text to be detected belongs to normal text or abnormal text according to the perplexity and the normal text vector distance.
[0017] The perplexity determination module includes:
[0018] A to-be-input word set generation unit, configured to obtain a start tag word and an end tag word, determine the start tag word, the end tag word, and a number of entity byte units as to-be-input words, and generate a to-be-input word set according to the to-be-input words;
[0019] An index determination unit, configured to obtain the word index corresponding to each to-be-input word in a word dictionary;
[0020] A local model prediction unit, configured to identify the sorting position of each to-be-input word in the to-be-input word set, call a local semantic model, and perform prediction on the word index corresponding to each to-be-input word in the local semantic model according to the sorting position to obtain a word prediction distribution corresponding to each entity byte unit; the word prediction distribution is used to represent the occurrence probability of the words in the word dictionary in the target text;
[0021] An occurrence probability determination unit for determining the occurrence probability of each entity byte unit in the target text to be detected according to the word prediction distribution corresponding to each entity byte unit respectively;
[0022] A perplexity determination unit for determining the perplexity for measuring the semantic coherence of the target text to be detected according to the occurrence probability of each entity byte unit in the target text to be detected.
[0023] Among them, the word prediction distribution corresponding to each entity byte unit respectively includes a forward prediction distribution and a backward prediction distribution; the local semantic model includes a first embedding layer, a bidirectional memory network layer, and a normalization layer;
[0024] The local model prediction unit includes:
[0025] A first embedding layer subunit for calling the first embedding layer to perform embedding feature processing on the word index corresponding to each input word respectively, and obtaining the word embedding vector corresponding to each input word respectively;
[0026] A first network layer subunit for identifying the sorting position of each input word in the input word set, and calling the bidirectional memory network layer to perform hidden layer feature processing on the word embedding vector corresponding to each input word respectively according to the sorting position, and obtaining the forward hidden layer representation vector and the backward hidden layer representation vector corresponding to each entity byte unit respectively;
[0027] A normalization layer subunit for calling the normalization layer to perform normalization processing on the forward hidden layer representation vector and the backward hidden layer representation vector corresponding to each entity byte unit respectively, and obtaining the forward prediction distribution and the backward prediction distribution corresponding to each entity byte unit respectively.
[0028] Among them, the first network layer subunit is specifically used for identifying the forward sorting position of each input word in the input word set, and calling the bidirectional memory network layer to perform hidden layer feature processing on the word embedding vector corresponding to each input word respectively according to the forward sorting position, and obtaining the forward hidden layer representation vector corresponding to each entity byte unit respectively;
[0029] The first network layer subunit is specifically further used for identifying the backward sorting position of each input word in the input word set, and calling the bidirectional memory network layer to perform hidden layer feature processing on the word embedding vector corresponding to each input word respectively according to the backward sorting position, and obtaining the backward hidden layer representation vector corresponding to each entity byte unit respectively.
[0030] Among them, the occurrence probability determination unit includes:
[0031] A word average prediction subunit, configured to generate an average prediction distribution corresponding to each entity byte unit according to the forward prediction distribution and the backward prediction distribution respectively corresponding to each entity byte unit; the plurality of entity byte units include an entity byte unit M.
[0032] A target word determination subunit, configured to find a word that matches the word index corresponding to the entity byte unit M in the average prediction distribution corresponding to the entity byte unit M as the target word.
[0033] An occurrence probability determination subunit, configured to determine the occurrence probability corresponding to the target word in the average prediction distribution corresponding to the entity byte unit M as the occurrence probability of the entity byte unit M in the target text to be detected.
[0034] A vector acquisition module, including:
[0035] A text word set generation unit, configured to obtain a head tag word and a tail tag word, determine the head tag word, the tail tag word, and the plurality of entity byte units as words to be processed, and generate a text word set according to the words to be processed.
[0036] A first set determination unit, configured to obtain the word index corresponding to each word to be processed in a word dictionary and generate a set of word indices to be processed.
[0037] A second set determination unit, configured to obtain the word position information of each word to be processed in the text word set and generate a set of word position information.
[0038] A third set determination unit, configured to obtain the sentence position information of each word to be processed in the text word set and generate a set of sentence position information.
[0039] A global model unit, configured to call a global semantic model to perform semantic feature extraction processing on the set of word indices to be processed, the set of word position information, and the set of sentence position information in the global semantic model, and obtain a target global representation vector of the target text to be detected.
[0040] Wherein, the global semantic model includes a second embedding layer and a self-attention network layer.
[0041] The global model unit includes:
[0042] A second embedding layer subunit, configured to call the second embedding layer to perform embedding feature processing on the set of word indices to be processed, the set of word position information, and the set of sentence position information, and obtain a word index embedding vector to be processed, a word position information embedding vector, and a sentence position information embedding vector respectively corresponding to each word to be processed.
[0043] A fusion vector determination subunit, configured to generate a fusion embedding vector corresponding to each word to be processed according to the word index embedding vector, word position information embedding vector, and sentence position information embedding vector corresponding to each word to be processed respectively; the fusion embedding vector corresponding to a word to be processed is generated by the word index embedding vector, word position information embedding vector, and sentence position information embedding vector corresponding to this word to be processed;
[0044] A second network layer subunit, which invokes a self-attention network layer to perform semantic representation processing on the fusion embedding vector corresponding to each word to be processed, and obtains a hidden layer vector corresponding to each word to be processed;
[0045] A global vector determination subunit, configured to determine, among the hidden layer vectors corresponding to each word to be processed respectively, the hidden layer vector corresponding to the head tag word as the target global representation vector of the text to be detected.
[0046] Wherein, the number of normal text types is at least two;
[0047] A vector distance determination module, including:
[0048] A reference distance determination unit, configured to respectively obtain the vector distance between the reference global representation vector of each normal text type and the target global representation vector as the reference distance;
[0049] A vector selection unit, configured to sequentially obtain Q reference global representation vectors from at least two reference global representation vectors according to the reference distance; Q is a positive integer, and Q is less than or equal to the number of normal text types;
[0050] A vector distance determination unit, configured to generate the average vector distance between the target global representation vector and the Q reference global representation vectors, and use the average vector distance as the normal text vector distance of the text to be detected.
[0051] Wherein, a target text determination module, including:
[0052] A harmonic unit, configured to harmonize the perplexity and the normal text vector distance according to the reference weights corresponding to the perplexity and the normal text vector distance respectively, and obtain the text semantic relevance of the text to be detected;
[0053] An abnormal text determination unit, configured to determine that the text to be detected belongs to abnormal text if the text semantic relevance is greater than the standard relevance threshold;
[0054] A normal text determination unit, configured to determine that the text to be detected belongs to normal text if the text semantic relevance is less than or equal to the standard relevance threshold.
[0055] Among them, the above text detection device further includes:
[0056] A reference normal text acquisition module, configured to acquire at least two reference normal texts;
[0057] A global vector acquisition module, configured to acquire global representation vectors corresponding to at least two reference normal texts respectively based on a global semantic model;
[0058] A clustering module, configured to perform clustering processing on the global representation vectors corresponding to at least two reference normal texts respectively to obtain N clustering clusters; each clustering cluster respectively represents a different normal text type;
[0059] A reference vector determination module, configured to use the global representation vectors located at the cluster centers in each clustering cluster as reference global representation vectors of the normal text types respectively.
[0060] Among them, the above text detection device further includes:
[0061] A first text processing module, configured to acquire a first normal sample text, perform entity word segmentation processing on the content of the first normal sample text to obtain a number of first sample entity byte units after entity word segmentation processing;
[0062] A predicted sample set generation module, configured to determine the start tag word, the end tag word, and a number of first sample entity byte units as to-be-input sample words, and generate a to-be-input sample word set according to the to-be-input sample words;
[0063] A sample index acquisition module, configured to acquire the word index corresponding to each to-be-input sample word in a word dictionary;
[0064] A sample prediction acquisition module, configured to identify the sample sorting position of each to-be-input sample word in the to-be-input sample word set, call a local semantic initial model, and perform prediction on the word index corresponding to each to-be-input sample word according to the sample sorting position in the local semantic initial model to obtain the sample word prediction distribution corresponding to each first sample entity byte unit;
[0065] A first training module, configured to train the local semantic initial model based on the word label distribution and the sample word prediction distribution corresponding to each first sample entity byte unit respectively to obtain a local semantic model.
[0066] Among them, the sample word prediction distribution corresponding to each first sample entity byte unit includes a sample forward prediction distribution and a sample reverse prediction distribution; the local semantic model includes a first initial embedding layer, an initial bidirectional memory network layer, and an initial normalization layer;
[0067] The sample prediction acquisition module includes:
[0068] A first initial embedding layer unit, which is used to call a first initial embedding layer to perform initial embedding feature processing on the word indices corresponding to each input sample word respectively, so as to obtain sample word embedding vectors corresponding to each input sample word respectively;
[0069] An initial network layer unit, which is used to identify the sample sorting positions of each input sample word in the input sample word set respectively, and call an initial bidirectional memory network layer to perform initial hidden layer feature processing on the word embedding vectors corresponding to each input sample word respectively, so as to obtain a sample forward hidden layer representation vector and a sample backward hidden layer representation vector corresponding to each first sample entity byte unit respectively;
[0070] An initial normalization layer unit, which is used to call an initial normalization layer to perform normalization processing on the sample forward hidden layer representation vector and the backward hidden layer representation vector corresponding to each first sample entity byte unit respectively, so as to obtain a sample forward prediction distribution and a sample backward prediction distribution corresponding to each first sample entity byte unit respectively;
[0071] Then the first training module includes:
[0072] A first loss determination unit, which is used to construct a cross-entropy loss function of a local semantic initial model according to the sample forward prediction distribution corresponding to each first sample entity byte unit, the sample backward prediction distribution corresponding to each first sample entity byte unit, and the word label distribution corresponding to each first sample entity byte unit;
[0073] A first training unit, which is used to adjust the parameters of the local semantic initial model according to the cross-entropy loss function; when the adjusted local semantic initial model meets the model convergence condition, the adjusted local semantic initial model is determined as the local semantic model.
[0074] Wherein, the above text detection device includes:
[0075] A second text processing module, which is used to obtain a second normal sample text, perform entity word segmentation processing on the content of the second normal sample text, so as to obtain a number of second sample entity byte units after entity word segmentation processing;
[0076] A sample word set generation module, which is used to determine both the head label word, the tail label word, and a number of second sample entity byte units as first sample words, and generate a sample text word set according to the first sample words;
[0077] A random replacement module, which is used to obtain N random label words, and replace N first sample words in the sample text word set with the N random label words respectively, so as to obtain a replaced sample text word set; N is a positive integer less than the total number of first sample words in the sample text word set;
[0078] A label position information determination module, configured to obtain label position information of N random label words in the set of sample text words after replacement, and generate a set of label position information according to the label position information;
[0079] A second training module, configured to train a global semantic initial model according to the set of sample text words after replacement, the set of label position information, and the random label words, and generate a global semantic model.
[0080] Among them, the second training module includes:
[0081] A random prediction determination unit, configured to call a global semantic processing model to perform initial semantic feature extraction processing on the set of sample text words after replacement, and obtain a sample word prediction distribution corresponding to each of the N random label words;
[0082] A target sample word determination unit, configured to obtain, according to the N label position information in the set of label position information, a first sample word at the N label position information in the set of sample text words as a target sample word;
[0083] A word label distribution acquisition unit, configured to acquire a word label distribution corresponding to each of the N target sample words;
[0084] A second loss determination unit, configured to constitute a cross-entropy loss function of the global semantic initial model according to the sample word prediction distribution corresponding to each of the N random label words and the word label distribution corresponding to each of the N target sample words;
[0085] A second training unit, configured to adjust parameters in the global semantic initial model according to the cross-entropy loss function, and when the adjusted global semantic initial model meets the model convergence condition, determine the adjusted global semantic initial model as the global semantic model.
[0086] An embodiment of the present application provides a computer device on the one hand, including: a processor, a memory, and a network interface;
[0087] The above-mentioned processor is connected to the above-mentioned memory and the above-mentioned network interface. Among them, the above-mentioned network interface is used to provide a data communication function, the above-mentioned memory is used to store a computer program, and the above-mentioned processor is used to call the above-mentioned computer program to execute the method in the embodiment of the present application.
[0088] An embodiment of the present application provides a computer-readable storage medium on the one hand. The above-mentioned computer-readable storage medium stores a computer program, and when the above-mentioned computer program is loaded and executed by a processor, it is used to execute the method in the embodiment of the present application.
[0089] One aspect of the embodiments of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, which are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method in the embodiments of the present application.
[0090] After obtaining the text to be detected in the embodiments of the present application, through entity word segmentation processing on the text to be detected, a number of entity byte units are obtained. Then, the occurrence probability of the entity byte units in the text to be detected is predicted, and the perplexity of the text to be detected is determined according to the corresponding occurrence probability of the entity byte units. Then, the target global representation vector of the text to be detected is obtained, and according to the target global representation vector and the reference global representation vector of the normal text type, the normal text vector distance of the text to be detected is determined. Finally, according to the perplexity of the target text and the normal text vector distance, it is determined whether the text to be detected belongs to normal text or abnormal text. Among them, predicting the occurrence probability of the entity byte units in the text to be detected can be implemented based on a local semantic model; obtaining the target global representation vector of the text to be detected can be implemented based on a global semantic model. Among them, both the local semantic model and the global semantic model are trained using normal sample texts. By using the method proposed in the embodiments of the present application, the text semantic relevance of the text to be detected is determined through the perplexity of the text to be detected and the normal text vector distance, and abnormal texts that use text transformation to transmit spam information can be screened out according to the text semantic relevance, improving the recognition accuracy of spam texts. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0092] Figure 1a is a schematic diagram of a network architecture provided by the embodiments of the present application;
[0093] Figure 1b is a schematic diagram of a scenario of a text detection method provided by the embodiments of the present application;
[0094] Figure 2 is a schematic flowchart of a text detection method provided by the embodiments of the present application;
[0095] Figure 3 is a schematic flowchart of a text detection method provided by the embodiments of the present application;
[0096] Figure 4a It is a schematic structural diagram of a text detection model provided by an embodiment of the present application;
[0097] Figure 4b It is a schematic structural diagram of another text detection model provided by an embodiment of the present application;
[0098] Figure 5 It is a schematic flow diagram of training a local semantic initial model provided by an embodiment of the present application;
[0099] Figure 6 It is a schematic flow diagram of training a global semantic initial model provided by an embodiment of the present application;
[0100] Figure 7 It is a flow chart of a target text recognition method provided by an embodiment of the present application;
[0101] Figure 8 It is a schematic structural diagram of a text detection device provided by an embodiment of the present invention;
[0102] Figure 9 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0103] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0104] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0105] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.
[0106] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0107] The solution provided in the embodiments of this application involves technologies such as natural language processing and machine learning in artificial intelligence. The specific description is as follows through the following embodiments. Please refer to Figure 1a , Figure 1a is a schematic diagram of a network architecture provided in the embodiments of this application. As Figure 1aAs shown in the figure, the system may include a service server 100 and a user terminal cluster. The user terminal cluster may include user terminals 10a, 10b, …, 10n. Among them, there may be communication connections between the user terminal clusters. For example, there is a communication connection between user terminal 10a and user terminal 10b, and there is a communication connection between user terminal 10b and user terminal 10n. Moreover, any user terminal in the user terminal cluster may have a communication connection with the service server 100. For example, there is a communication connection between user terminal 10a and the service server 100, and there is a communication connection between user terminal 10b and the service server 100.
[0108] The user terminal clusters (including the above-mentioned user terminals 10a, 10b, and 10n) can transmit SMS texts through the above-mentioned communication connections. However, due to black industry personnel using SMS to place advertisements or spread junk content, etc., the user experience is reduced, and the valuable time of users is wasted. More seriously, some junk SMS are phishing and fraudulent SMS, which will cause economic losses and psychological harm to some teenagers and the elderly with less social experience. Therefore, when a user terminal (i.e., any user terminal in the user terminal cluster) receives an SMS, it can first send a text recognition request to the service server 100. After the service server 100 responds to the text recognition request, it will first classify the SMS text to identify whether the SMS text is a normal text, and then determine whether to intercept the SMS text according to the recognition result. For the specific process, please refer to Figure 1b , Figure 1b which is a schematic diagram of the scenario of a text detection method provided by an embodiment of the present application. As Figure 1b shown, user terminal 10b receives a user terminal (which can be the above-mentioned Figure 1aAny terminal other than the user terminal 10b) text message, the content of the text message is "Dear Mr. Wang, you have passed our interview. There will be a dedicated person to contact you for the fake ID at 50 yuan each. For details, please consult xxxx". Obviously, this text message is an advertisement for fake IDs. The sender interspersed it among the text messages related to the normal interview, attempting to transmit the junk information to user A. If the user terminal 10b does not send a text recognition request to the service server 100 to recognize this text message and directly displays the text message to user A, it is very likely to cause discomfort or even disgust to user A, or user A may believe such junk information due to insufficient social experience and thus be deceived. Therefore, it is necessary to process the text message to determine whether the text message belongs to normal text or abnormal text. If the text message is normal text, the service server 100 will send a display instruction to the user terminal 10b so that the user terminal 10b can display the text message in the text message list and user A can view the text message normally; if the text message is abnormal text, the service server 100 will send an interception instruction to the user terminal 10b so that the user terminal 10b can place the text message in the trash can and user A can only view the text message in the trash can of the user terminal 10b.
[0109] Specifically, when the user terminal 10b receives a text message, it transmits a text recognition request to the service server 100. After the service server 100 responds to the text recognition request, it can regard the real-time text message as the text to be detected. After determining the text recognition result of the text to be detected, it will perform subsequent processing on the text to be detected. The service server 100 will perform entity word segmentation processing on the content of the text to be detected to obtain several entity byte units, obtain the occurrence probability of several entity byte units in the text to be detected through the local semantic model, obtain the target global representation vector of the text to be detected through the global semantic model, and then perform recognition processing on the text to be detected based on the occurrence probability and the target global representation vector. The specific process can be as follows:
[0110] The service server 100 performs entity word segmentation on the content of the text to be detected, obtaining a number of entity byte units. Then the service server 100 acquires the start tag word and the end tag word, takes the start tag word, the end tag word, and the number of entity byte units as words to be input, then obtains the word index corresponding to each word to be input in the word dictionary, takes the word index corresponding to each word to be input as the first text data, and then calls the local semantic model to perform prediction processing on the first text data, obtaining the word prediction distribution corresponding to each entity byte unit. Among them, the word dictionary includes words and the word indices corresponding to the words. Among them, the words in the word dictionary include entity byte units, some tag words, and so on. Among them, the word index can be a number used to refer to a word, and each word has a corresponding word index. Among them, the word prediction distribution is used to represent the occurrence probability of the words in the word dictionary in the text to be detected. According to the word prediction distribution corresponding to each entity byte unit, the occurrence probability of each entity byte unit in the text to be detected can be obtained, and then according to the occurrence probability corresponding to each entity byte unit, the perplexity of the text to be detected is determined. Among them, the perplexity can measure the semantic coherence degree of the text to be detected. While constructing the first text data, the service server 100 acquires the head tag word and the tail tag word, determines the head tag word, the tail tag word, and the number of entity byte units as words to be processed, generates a text word set according to the words to be processed, then the service server 100 acquires the word index corresponding to each word to be processed in the dictionary words, generates a set of word indices to be processed, then generates a set of word position information according to the position information of each word to be processed in the text word set, and further acquires the sentence position information of each word to be processed in the text word set, generating a set of sentence position information. The service server 100 can take the set of word indices, the set of word position information, and the set of sentence position information as the second text data, and then call the global semantic model to perform semantic feature extraction processing on the second text data, obtaining the target global representation vector of the text to be detected. Then the service server 100 acquires the reference global representation vector of the normal text type, and together with the target global representation vector, obtains the normal text vector distance of the text to be detected. After obtaining the perplexity and the normal text vector distance of the text to be detected, the service server 100 can determine whether the text to be detected belongs to normal text or abnormal text according to the perplexity and the normal text vector distance.
[0111] Such as Figure 1bAs shown in the figure, after the above processing, the service server 100 will determine that the SMS text sent by user A is an abnormal text. The service server 100 will intercept the SMS text, that is, send an interception instruction to the user terminal 10b. In response to the interception instruction, the user terminal 10b will place the SMS in the trash bin, and user A will not receive the SMS in the normal SMS list of the user terminal 10b, avoiding harassment to user A and meaningless time waste.
[0112] Optionally, if the user terminal locally stores the trained local semantic model and global semantic model, the user terminal can perform an SMS recognition task on the received SMS text locally to obtain the text recognition result of the SMS text, and then determine whether to display the SMS text in the normal SMS list of the user terminal based on the text recognition result of the SMS text. Among them, since training the local semantic initial model and the global semantic initial model involves a large amount of offline computing, the local semantic initial model and the global semantic initial model on the user terminal can be sent to the user terminal by the server 100 after training.
[0113] It can be understood that the method provided in the embodiments of the present application can be executed by a computer device. The computer device includes but is not limited to a terminal or a server. The server 100 in the embodiments of the present application can be a computer device, and the user terminals in the user terminal cluster can also be computer devices, which is not limited here. The above server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The above terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here.
[0114] Among them, Figure 1a the server 100, user terminal 10a, user terminal 10b, and user terminal 10n in can include mobile phones, tablet computers, notebook computers, palm computers, smart speakers, mobile internet devices (MIDs), point-of-sale (POS) machines, wearable devices (such as smart watches, smart bracelets, etc.).
[0115] It should be noted that in the embodiments of the present application, the target text to be detected is described by taking real-time SMS as an example. However, in actual application scenarios, the target text can also include emails, papers, works, novels, etc.
[0116] Further, please refer to Figure 2 , Figure 2 which is a schematic flowchart of a text detection method provided by an embodiment of the present application. This method is executed by the computer device described in Figure 1a , that is, it can be the server 100 in Figure 1a or the user terminal cluster in Figure 1a (which also includes user terminals 10a, 10b, and user terminal 10n). As shown in Figure 2 , the text detection process includes the following steps:
[0117] S101: Obtain the target text to be detected, perform entity word segmentation processing on the content of the target text to be detected, and obtain a number of entity byte units after entity word segmentation processing.
[0118] Specifically, the target text to be detected can be a sentence text. A sentence text can be understood as a sequence composed of one or more words, and each word is the basic unit that makes up the sentence text. For a sentence text, the semantic information of each word is very important. An entity byte unit can be an entity word, that is, the basic unit that makes up the sentence text. A number of entity byte units can be at least one entity word. Entity word segmentation processing is the process of recombining a continuous character sequence into a word sequence according to certain specifications, that is, converting the content included in the target text to be detected into the representation of entity byte units. To perform word segmentation on the target text to be detected to obtain entity byte units, a rule-based word segmentation method can be used to perform word segmentation on the target text to be detected. Mainly, a word library, also called a dictionary or word dictionary, is established in advance, and the sentence is divided by dictionary matching; or a word segmentation tool can be used to perform word segmentation on the target text to be detected, or other methods, which are not limited here. For example, if the target text to be detected is "Dear Mr. Wang! Have a nice life!", the entity byte units obtained after word segmentation can be ("Dear", "Wang", "Mr.", "!", "life", "nice", "!"). After obtaining the entity byte units, that is, the entity words, can the target text to be detected be further analyzed and processed.
[0119] S102: Perform entity prediction processing on the number of entity byte units to obtain the occurrence probability of the number of entity byte units in the target text to be detected, and determine the perplexity for measuring the semantic coherence of the target text to be detected according to the occurrence probability corresponding to the number of entity byte units.
[0120] Specifically, perplexity (PPL) is used to measure the degree of semantic coherence of a text. The higher the perplexity, the worse the semantic coherence. The perplexity can be jointly determined by the occurrence probabilities of each entity byte unit included in the text to be detected in the text to be detected, that is, it can be calculated by formula (1):
[0121]
[0122] where PPL is the perplexity of the text to be detected, q is the maximum sentence length of the text to be detected, is the occurrence probability of the t-th entity byte unit in the text to be detected in the text to be detected. By obtaining the perplexity of the text to be detected through formula (1), the local semantic coherence of the text to be detected can be measured.
[0123] Specifically, predicting the occurrence probability of an entity byte unit in the text to be detected can be completed through a local semantic model. Among them, the local semantic model can be trained based on normal sample texts and a Bi-LSTM model, and can predict the next word based on the known words. Among them, Bi-LSTM is short for Bi-directional Long-Short Term Memory. Bi-LSTM is a recurrent neural network that can effectively model the context-dependent information of the text. Entity prediction processing can be a calculation process of predicting the occurrence probability of an entity byte unit in the text to be detected through a local semantic model.
[0124] S103: Perform semantic feature extraction processing on the text to be detected to obtain the target global representation vector of the text to be detected.
[0125] Specifically, perform semantic feature extraction processing on the text to be detected to obtain the target global representation vector. Among them, semantic feature extraction processing refers to representing the text as a low-dimensional, dense, real-valued vector to facilitate subsequent characterization of semantic, syntactic, and other features of the text. The target global representation vector of the text to be detected can be obtained through a global semantic model. Among them, the global semantic model is obtained by training a BERT model based on normal sample texts. Among them, BERT is the Bidirectional Encoder Representations from Transformers, a pre-trained language model that can effectively learn word vectors with context semantics and dense sentence representations without labeled data.
[0126] S104: Determine a normal text vector distance for measuring the normal semantics of the target text to be detected according to the target global representation vector and the reference global representation vectors of normal text types.
[0127] Specifically, the larger the normal text vector distance is, the farther the target text to be detected deviates from the normal semantics. The number of normal text types is at least two. According to the target global representation vector and the reference global representation vector of one normal text type, a vector distance can be calculated as a reference distance. Suppose there are M normal text types, then there are M reference global representation vectors, and M reference distances can be obtained. The computer device will sequentially select Q reference global representation vectors corresponding to the Q reference distances from small to large among the M reference distances, and then calculate the average vector distance between the target global representation vector and the selected Q reference global representation vectors, and use this average vector distance as the normal text vector distance of the target text to be detected.
[0128] Specifically, for the generation of the reference global representation vectors of normal text types, the computer device can obtain at least two reference normal texts, and then obtain the global representation vectors corresponding to the reference normal texts based on the global semantic model; then the computer device will perform clustering processing on the global representation vectors corresponding to the reference normal texts to obtain N clustering clusters, and each clustering cluster can represent a normal text type. Among them, the clustering processing can adopt K-means clustering or other clustering methods, which are not limited here. Then, the computer device will use the global representation vector corresponding to the clustering center in each clustering cluster as the reference global representation vector of the normal text type.
[0129] Specifically, the average vector distance can be calculated through the following formula (2):
[0130]
[0131] where C is the target global representation vector of the target text to be detected, Γ is the index set of the Q clustering centers closest to the vector C, and q i is the vector representation of the i-th clustering center.
[0132] S105: Determine whether the target text to be detected belongs to normal text or abnormal text according to the perplexity and the normal text vector distance.
[0133] Specifically, when the computer device determines the perplexity and the normal text vector distance of the target text to be detected, it can calculate the text semantic relevance of the target text to be detected according to formula (3):
[0134] Score = λ * PPL + (1 - λ) * VD Formula (3)
[0135] Among them, Score is the text semantic relevance score, that is, the text semantic relatedness; PPL is the perplexity that measures local semantic coherence; VD is the normal text vector distance that measures global semantic normality; λ is a hyperparameter, that is, a parameter used for reconciliation, and the specific value can be adjusted according to different actual situations.
[0136] When Score is higher than the standard relevance threshold λ, it is determined that the text of the target to be detected belongs to abnormal text; when Score is less than or equal to θ, it is determined that the text of the target to be detected belongs to normal text. According to whether the text of the target to be detected is abnormal, the computer device can perform subsequent processing. For example, if the text of the target to be detected is a short message text, when it is determined that the text of the target to be detected is abnormal text, the computer device will determine that the short message is a spam message and intercept the short message or put it into the trash can to avoid disturbing the user.
[0137] Through the method provided by the embodiments of the present application, after the computer device tokenizes the text of the target to be detected, entity byte units are obtained, then the occurrence probability of the entity byte units in the text of the target to be detected is predicted, then the perplexity of the text of the target to be detected is determined according to the corresponding occurrence probability of the entity byte units, then the target global representation vector of the text of the target to be detected is obtained, and according to the target global representation vector and the reference global representation vector of the normal text type, the normal text vector distance of the text of the target to be detected is determined, and then the hyperparameter is used to combine the perplexity and the normal text vector distance to obtain the text semantic relevance score. When the text semantic relevance score is greater than the reference threshold, it is determined that the text of the target to be detected is abnormal text, and it can be intercepted. Among them, the perplexity is used to measure the coherence of the text of the target to be detected; the normal text vector distance is used to measure the global semantics of the text of the target to be detected. Based on the method provided by the embodiments of the present application, spam information transmitted by text transformation can be accurately screened out.
[0138] Further, please refer to Figure 3 , Figure 3 which is a schematic flowchart of a text detection method provided by the embodiments of the present application. This method is executed by the Figure 1a computer device described in Figure 1a , that is, it can be the Figure 1a server 100 in Figure 3 , or it can be the user terminal cluster (including user terminal 10a, user terminal 10b, and user terminal 10n) in
[0139] S201: Obtain the text of the target to be detected, perform entity tokenization processing on the content of the text of the target to be detected, and obtain a number of entity byte units after entity tokenization processing.
[0140] Specifically, for the implementation of step S201, reference can be made to the implementation of step S101 in the corresponding embodiment above, which will not be elaborated here. Figure 2
[0141] S202: Obtain a start tag word and an end tag word, determine both the start tag word, the end tag word, and the several entity byte units as words to be input, and generate a set of words to be input according to the words to be input.
[0142] Specifically, the start tag word can be the [SOS] tag word, and the end tag word can be the [EOS] tag word. In the set of words to be input, the sorting position of the entity byte units is the same as the sorting position of the entity byte units in the text to be detected. The start tag word [SOS] will be placed before the first entity byte unit, and the end tag word [EOS] will be placed after the last entity byte unit. For example, if the text to be detected is "The weather is nice today!", after performing word segmentation on the text to be detected to obtain entity byte units, and then together with the [SOS] word and the [EOS] word, the set of words to be input obtained is {[SOS], "The", "weather", "is", "nice", "today", "!", [EOS]}.
[0143] S203: Obtain the word index corresponding to each word to be input in the word dictionary.
[0144] Specifically, after obtaining the set of words to be input, a local semantic model can be used to predict the occurrence probability of the entity byte units in the text to be detected. For the convenience of calculation by the local semantic model, it is necessary to first perform data conversion on the words to be input. Here, the word index corresponding to each word to be input can be obtained in the word dictionary. Among them, the word index can also be called a word identifier, or word id (identifier), which is used to refer to the entity byte unit. Each entity byte unit corresponds to a word index, and the word index can be a string of numbers, such as 1, 12, 123. Among them, the word dictionary is a dictionary that includes entity byte units and the corresponding word indexes. For example, the word dictionary is (SOS-0, today-1, tomorrow-2, the day after tomorrow-3, weather-4, sky-5, nice-6, bad-7,!-8, EOS-9), and the set of words to be input is the above {"[SOS]", "today", "weather", "nice", "!", "[EOS]"}, then query the word index corresponding to each word to be input in the word dictionary, and the word index of this text to be detected is obtained as (0, 1, 4, 6, 8, 9). Then, the computer device will use the word index as the input of the local semantic model.
[0145] S204: Identify the sorting position of each word to be input in the set of words to be input, call the local semantic model, and predict the word indices corresponding to each word to be input according to the sorting position in the local semantic model, so as to obtain the word prediction distribution corresponding to each entity byte unit.
[0146] Specifically, the computer device will input the word indices corresponding to each word to be input into the local semantic model according to the sorting position of each entity byte unit in the set of words to be input. For example, in the set of words to be input corresponding to the above text to be detected "The weather is nice today!", the input order of the word indices corresponding to the words to be input is (0, 1, 4, 6, 8, 9). Then, the computer device outputs the word prediction distribution corresponding to each entity byte unit through the local semantic model. Among them, the word prediction distribution is used to represent the occurrence probability of the words in the word dictionary in the text to be detected, and the local semantic model will output the word prediction distribution corresponding to each entity byte unit. For example, after inputting the above (0, 1, 4, 6, 8, 9) into the local semantic model, the output will obtain the word prediction distribution corresponding to each entity byte unit. For example, the word prediction distribution corresponding to the entity byte unit "today" is [0, 0.5, 0.3, 0.2, 0, 0, 0, 0, 0, 0]. This prediction distribution indicates that in the word dictionary, the occurrence probability of "today" in the text to be detected is 0.5, the occurrence probability of "tomorrow" in the text to be detected is 0.3, the occurrence probability of "the day after tomorrow" in the text to be detected is 0.2, and the occurrence probabilities of the remaining words are 0. According to the word prediction distribution corresponding to each entity byte unit, the computer device can determine the occurrence probability of each entity byte unit in the above text to be detected.
[0147] Specifically, the local semantic model may include a first embedding layer, a bidirectional memory network layer, and a normalization layer. Then, identify the sorting position of each word to be input in the word set to be input, call the local semantic model, and predict the word index corresponding to each word to be input according to the sorting position in the local semantic model, so as to obtain the word prediction distribution corresponding to each entity byte unit. It can be that the first embedding layer is called to perform embedding feature processing on the word index corresponding to each word to be input, so as to obtain the word embedding vector corresponding to each word to be input. Then, identify the sorting position of each word to be input in the word set to be input, and call the bidirectional memory network layer to perform hidden layer feature processing on the word embedding vector corresponding to each word to be input according to the sorting position, so as to obtain the forward hidden layer representation vector and the backward hidden layer representation vector corresponding to each entity byte unit. Finally, call the normalization layer to perform normalization processing on the forward hidden layer representation vector and the backward hidden layer representation vector corresponding to each entity byte unit, so as to obtain the forward prediction distribution and the backward prediction distribution corresponding to each entity byte unit. In other words, the computer will first input the word index corresponding to each word to be input above into the first embedding layer, and output the word embedding vector corresponding to each word to be input through the first embedding layer; then, according to the sorting position of each word to be input in the word set to be input, input the word embedding vector corresponding to each word to be input into the bidirectional memory network layer, and output the forward hidden layer representation vector and the backward hidden layer representation vector corresponding to each entity byte unit through the bidirectional memory network layer; finally, input the forward hidden layer representation vector corresponding to each entity byte unit into the normalization layer to obtain the forward prediction distribution corresponding to each entity byte unit, and input the backward hidden layer representation vector corresponding to each entity byte unit into the normalization layer to obtain the backward prediction distribution corresponding to each entity byte unit. Among them, the embedding feature processing is to convert the word index corresponding to the word to be processed into a fixed-dimensional vector representation form, which is convenient for subsequent calculations. Among them, to obtain the forward hidden layer representation vector corresponding to each entity byte unit, it can be to identify the forward sorting position of each word to be input in the word set to be input, and call the bidirectional memory network layer to perform hidden layer feature processing on the word embedding vector corresponding to each word to be input according to the forward sorting position, so as to obtain the forward hidden layer representation vector corresponding to each entity byte unit. Among them, the hidden layer feature processing can be a process of sequentially inputting the word embedding vectors corresponding to the words to be input into the bidirectional memory network layer according to the sorting position of the words to be input in the word set to be input, and outputting the forward hidden layer representation vector corresponding to each entity byte unit through calculation by the bidirectional memory network layer.Obtain the reverse hidden layer representation vector corresponding to each entity byte unit, which can be to identify the reverse sorting position of each word to be input in the word set to be input. Call the bidirectional memory network layer to perform hidden layer feature processing on the word embedding vectors corresponding to each word to be input according to the reverse sorting position, and obtain the reverse hidden layer representation vector corresponding to each entity byte unit. Among them, the hidden layer feature processing can be a process of sequentially and reversely inputting the word embedding vectors corresponding to the words to be input into the bidirectional memory network layer according to the sorting position of the words to be input in the word set to be input, and outputting the reverse hidden layer representation vector corresponding to each entity byte unit after calculation by the bidirectional memory network layer. Among them, the forward prediction distribution is used to predict the probability of the next word to be processed appearing, and the reverse prediction distribution is used to predict the probability of the previous word to be processed appearing.
[0148] For ease of understanding, please also refer to Figure 4a , Figure 4a which is a schematic structural diagram of a text detection model provided by an embodiment of the present application. As Figure 4a shown, the text detection model can be the above-mentioned local semantic model, and the text detection model includes an Embedding layer, a Bi-LSTM network layer, and a SoftMax layer. Among them, the Embedding layer is the above-mentioned first embedding layer, the Bi-LSTM network layer is the above-mentioned bidirectional memory network layer, and the SoftMax layer is the above-mentioned normalization layer. After the computer device obtains the text, it performs word segmentation on it. The entity byte units obtained after word segmentation are (w1, w2, w3, w4,..., w r ). After adding the start tag word and the end tag word, the word set to be processed {[SOS], w1, w2, w3, w4,..., w r , [EOS]} is obtained. Any word in the word set to be processed can be called a word to be processed. Then, according to the word dictionary, the word data corresponding to each word to be processed, that is, the word index, is obtained as (v1, v2, v3, v4,..., v t ). Among them, v1 is the word index corresponding to the start tag word [SOS], and v t is the word index corresponding to the end tag word [EOS]. Then the computer device will input the word data into the embedding layer including the embedding matrix . Among them, the embedding matrix W e is a random matrix, m is the number of rows of the embedding matrix and also the dimension of the embedding vector; D is the size of the word dictionary of the training set and also the number of columns of the embedding matrix. Since v t represents the index of the word to be processed w t in the word dictionary, the mathematical representation of the embedding layer is as shown in formula (4):
[0149]
[0150] wherein is the word embedding vector of the word w to be processed t , and W e [v t represents the v e -th column of the embedding matrix W t . The embedding vectors at each position are collected to form the word embedding features Then, the computer device inputs X into the Bi-LSTM network moment by moment, and obtains the bidirectional hidden layer representations and wherein is obtained by the following formula:
[0151]
[0152] c t = f t ⊙ c t-1 + i t ⊙ g t Formula (6)
[0153]
[0154] wherein i t , f t and o t are the input gate, forget gate and output gate respectively, g t is the input memory cell, σ is the Sigmoid activation function, M is the learnable parameter matrix, is the forward hidden layer representation vector of the word w to be input t . Then is sent to the Softmax layer to be mapped to the vocabulary dimension and normalized in probability, and the forward prediction distribution corresponding to the word to be input generated by the forward information is obtained The mathematical representation is as follows:
[0155]
[0156] Similarly, by inputting the embedding features X into the Bi-LSTM network in reverse order moment by moment, the reverse prediction distribution corresponding to each word to be input generated by the reverse information can be obtained
[0157] It should be noted that by sequentially inputting the word index v t corresponding to the t-th word to be input moment by moment, the forward prediction distribution corresponding to the (t + 1)-th word to be input output by the local semantic modelHowever, since only the forward prediction distribution corresponding to the entity byte units in the words to be input is required for subsequent calculations, that is, after the word indices corresponding to q words to be input are sequentially input into the local semantic model, the finally truly output forward prediction distribution is Because among the q words to be input, starting from the second word to be input to the (q - 1)-th word to be input are the entity byte units. Similarly, when the word index v corresponding to the t-th word to be input is input into the local semantic model in reverse order at each moment, the output is the reverse prediction distribution corresponding to the (t - 1)-th word to be predicted. t When input into the local semantic model, the output is the reverse prediction distribution corresponding to the (t - 1)-th word to be predicted. Then, after the word indices corresponding to q words to be input are sequentially input into the local semantic model, the finally truly output reverse prediction distribution is
[0158] S205: Determine the occurrence probability of each entity byte unit in the target text to be detected according to the word prediction distribution corresponding to each entity byte unit; determine the perplexity of the target text to be detected according to the occurrence probability of each entity byte unit in the target text to be detected.
[0159] Specifically, after obtaining the forward prediction distribution and reverse prediction distribution corresponding to the entity byte unit, the computer device will generate the average prediction distribution corresponding to the entity byte unit according to the forward prediction distribution and reverse prediction distribution corresponding to the entity byte unit. Suppose the entity byte unit includes entity byte unit M. The computer device will search for the word that matches the word index corresponding to entity byte unit M in the average prediction distribution corresponding to entity byte unit M as the target word, and determine the occurrence probability corresponding to the target word in the average prediction distribution corresponding to entity byte unit M as the occurrence probability of entity byte unit M in the target text to be detected. For example, the word dictionary is {M - 1, N - 2, L - 3}, and the forward prediction distribution corresponding to the word position where entity byte unit M is located in the target text to be detected is [0.2, 0.3, 0.5], and the reverse prediction distribution is [0.4, 0.2, 0.4]. Add the forward prediction distribution and reverse prediction distribution corresponding to entity byte unit M, and then take the average to obtain the average prediction distribution [0.3, 0.25, 0.45] corresponding to entity byte unit M. Then, in the average prediction distribution, since the word index of the entity byte unit is 1, the computer device will obtain the occurrence probability at the first position as the occurrence probability of entity byte unit M in the target text to be detected, that is, 0.3. After obtaining the occurrence probability corresponding to each entity byte unit, the perplexity PPL of the target text to be detected can be calculated according to the above formula (1).
[0160] S206: Obtain the head tag word and the tail tag word, determine the head tag word, the tail tag word, and the several entity byte units as words to be processed, and generate a text word set according to the words to be processed.
[0161] Specifically, the number of several entity byte units in the target text to be detected is at least one. The head tag word can be represented by the [CLS] token. Usually, the computer device will insert [CLS] before all entity byte units, that is, at the first position of the text word set. The tail tag word can be represented by the [SEP] token. Usually, the computer device will insert [SEP] after the last entity byte unit in a sentence in the target text to be detected. For example, if the target text to be detected is "Xiaoming likes basketball. Xiaoming likes football.", after tokenizing the target text to be detected, the obtained entity byte units are {"Xiaoming", "likes", "basketball", "Xiaoming", "likes", "football"} in sequence. Then, insert the [CLS] token and the [SEP] token into the entity byte units to obtain {"[CLS]", "Xiaoming", "likes", "basketball", "[SEP]", "Xiaoming", "likes", "football", "[SEP]"}, that is, the text word set. Any word in the text word set can be used as a word to be processed.
[0162] S207: Obtain the word index corresponding to each word to be processed in the word dictionary, and generate a set of word indices to be processed; obtain the word position information of each word to be processed in the text word set respectively, and generate a set of word position information; obtain the sentence position information of each word to be processed in the text word set respectively, and generate a set of sentence position information.
[0163] Specifically, in addition to including entity byte units and the word indices corresponding to the entity byte units, the word dictionary also includes some special tokens (tokenizations) and the word indices corresponding to the special tokens. For example, the start tag word [SOS] and the end tag word [EOS] mentioned above, the [CLS] token and the [SEP] token, as well as the random mask [MASK] and the unknown character [UNK]. Among them, the random mask [MASK] is a token used to replace entity byte units during training, and the unknown character [UNK] is mainly used to replace entity byte units that are not in the word dictionary. That is, if when tokenizing the target text to be detected, the corresponding word cannot find a matching word in the word dictionary, the word index corresponding to [UNK] will be used to replace the word in the subsequent process.
[0164] Specifically, after the computer device obtains the text word set, it will obtain the word index corresponding to each word to be processed in the word dictionary and generate a set of word indices to be processed. Assuming the word dictionary is {CLS-0, SEP-1, like-2, basketball-3, football-4, UNK-5, MASK-6}, for the above text word set {"[CLS]", "Xiaowang", "like", "basketball", "[SEP]", "Xiaoming", "like", "football", "[SEP]"}, the corresponding set of word indices to be processed is {0, 5, 2, 3, 1, 5, 2, 4, 1}. Among them, neither "Xiaowang" nor "Xiaoming" is included in the word dictionary, and by default, the word index corresponding to UNK is used as their corresponding word index. Then the computer device will obtain the word position information of each word to be processed in the text word set respectively and generate a set of word position information. Among them, the word position information refers to which word to be processed in the text word set the word to be processed is. For example, the set of word position information corresponding to the above text word set is {0, 1, 2, 3, 4, 5, 6, 7, 8}. The computer device will also obtain the sentence position information of each word to be processed in the text word set respectively and generate a set of sentence position information. When inserting, the [SEP] token is placed at the end of each sentence of the text to be detected. Therefore, the words to be processed between the [SEP] tokens are of the same sentence information. For example, the set of sentence position information corresponding to the above text word set is {0, 0, 0, 0, 0, 1, 1, 1, 1}.
[0165] S208: Call the global semantic model, and perform semantic feature extraction processing on the set of word indices to be processed, the set of word position information, and the set of sentence position information in the global semantic model to obtain the target global representation vector of the text to be detected.
[0166] Specifically, the global semantic model may include a second embedding layer and a self-attention network layer. Among them, the second embedding layer can convert the index corresponding to the word to be processed into a fixed-dimensional vector representation. Therefore, the computer device will call the second embedding layer to perform embedding feature processing on the set of indices of the words to be processed, the set of word position information, and the set of sentence position information, and obtain the embedding vectors of the indices of the words to be processed, the embedding vectors of the word position information, and the embedding vectors of the sentence position information corresponding to each word to be processed respectively. Among them, the embedding feature processing may be a process in which the computer device inputs the set of indices of the words to be processed, the set of word position information, and the set of sentence position information into the second embedding layer, and outputs the embedding vectors of the indices of the words to be processed, the embedding vectors of the word position information, and the embedding vectors of the sentence position information corresponding to each word to be processed respectively after data processing by the second embedding layer. Then the computer device will generate the fused embedding vectors corresponding to each word to be processed respectively according to the embedding vectors of the indices of the words to be processed, the embedding vectors of the word position information, and the embedding vectors of the sentence position information corresponding to each word to be processed respectively. Among them, the fused embedding vector corresponding to a word to be processed is obtained by adding the embedding vector of the index of the word to be processed, the embedding vector of the word position information, and the embedding vector of the sentence position information corresponding to the word to be processed. Then the computer device calls the self-attention network layer to perform semantic representation processing on the fused embedding vectors corresponding to each word to be processed respectively, that is, inputs the fused embedding vectors corresponding to each word to be processed respectively into the self-attention network layer, and obtains the hidden layer vectors corresponding to each word to be processed respectively after L-layer operations by the self-attention network layer. Then the computer device will use the hidden layer vector corresponding to the [CLS] token in the last layer as the global representation vector of the text to be detected.
[0167] For ease of understanding, please also refer to Figure 4b , Figure 4b which is a schematic structural diagram of another text detection model provided by an embodiment of the present application. As Figure 4b shown, the text detection model may be the above-mentioned global semantic model. After the computer device tokenizes the text to be detected, it obtains multiple entity byte units, such as Tok1, Tokq, etc. Then, according to the position information of the entity byte units in the text to be detected, the computer device will insert the [CLS] and [SEP] words into the entity byte unit sequence to obtain a text word set. Each word in the text word set can be called a word to be processed. According to the word dictionary and the text word set, input_ids (character indices) can be constructed, that is, the set of indices of the words to be processed mentioned above. According to the end punctuation marks (period, exclamation mark, question mark) in the text word set and the text word set, seg_ids (sentence indices) can be constructed, that is, the set of sentence position information mentioned above. As Figure 4bAs shown, the sentence position information corresponding to the words in Sentence 1 can all be 1, and the sentence position information corresponding to the words in Sentence 2 can all be 2. Then, according to the position information of the word to be processed in the text word set, pos_ids (position index) can be constructed, that is, the set of word position information described above. As Figure 4b shown, [CLS] is the first word to be processed, so its corresponding word position information can be 0. Tok1 in Sentence 1 is the second word to be processed, so its corresponding word position information can be 1, and so on. The word position information corresponding to each word to be processed can be obtained. Then, input_ids (character index), seg_ids (sentence index), and pos_ids (position index) are used as the input of the global semantic model, and in the second embedding layer, that is, the three indexes are embedded. There are three randomly generated embedding matrices in the second embedding layer. Using the above formula (2), the character index embedding vector corresponding to the t-th word to be processed can be obtained Sentence index embedding vector and position index embedding vector The sum of these three representations corresponding to the t-th word to be processed is obtained as E t , and the representations at all times are collected to obtain the embedding matrix Then the computer device will input the embedding matrix E into the self-attention network layer, perform the operation of the L-layer transformer layer, and obtain the final hidden layer representation E L , and the mathematical representation is as follows:
[0168] E L = transformer(E L-1 ) Formula (9)
[0169] Among them, E L-1 is the vector matrix obtained by passing the embedding matrix E through the (L - 1)-th layer of the transformer layer.
[0170] As Figure 4b shown, Among them, the vector C is the hidden layer vector corresponding to the above-mentioned head label word [CLS], and it is used as the target global representation vector of the text to be detected, and is used to represent the global semantics of the text to be detected.
[0171] S209: Determine the normal text vector distance of the text to be detected according to the target global representation vector and the reference global representation vector of the normal text type.
[0172] S210: Determine whether the text to be detected belongs to normal text or abnormal text according to the perplexity and the normal text vector distance.
[0173] Specifically, for the implementation of steps S209 - S210, reference can be made to the implementation of steps S104 - S105 in the corresponding embodiment above, which will not be elaborated here. Figure 2
[0174] Using the method provided in the embodiment of the present application, after segmenting the text of the target to be detected into several entity byte units, the forward prediction distribution and the reverse prediction distribution corresponding to each entity byte unit in the text of the target to be detected can be predicted through a local semantic model, so as to determine the occurrence probability of each entity byte unit in the text of the target to be detected, and then determine the perplexity of the text of the target to be detected; meanwhile, obtain the head tag word and the tail tag word, generate a text word set together with the above-mentioned several entity byte units, obtain the input of the global semantic model based on the text word set, then output the target global representation vector of the text of the target to be detected through the global semantic model, and then determine the normal text vector distance of the text of the target to be detected through the target global representation vector and the reference global representation vector of the normal text type; finally, according to the perplexity and the normal text vector distance of the text of the target to be detected, identify whether the text of the target to be detected is a normal text or an abnormal text. Through the method provided in the embodiment of the present application, it is not necessary to manually annotate and perform manual feature engineering to determine the type of text transformation. Instead, the perplexity and the normal text vector distance are directly used to evaluate the semantic relevance of the text of the target to be detected, so as to screen out abnormal texts that use text transformation to transmit spam information, improve the recognition accuracy of spam texts, and save the labor cost brought by annotating the text transformation type at the same time.
[0175] Further, please refer to Figure 5 Figure 5 which is a schematic flowchart of a process for training a local semantic initial model provided by an embodiment of the present application. This process is executed by the computer device described in Figure 1a i.e., it can be the server 100 in Figure 1a . As shown in Figure 5 , this training process includes the following steps:
[0176] S301: Obtain the first normal sample text, perform entity word segmentation processing on the content of the first normal sample text, and obtain several first sample entity byte units after entity word segmentation processing.
[0177] S302: Determine both the start tag word, the end tag word, and the several first sample entity byte units as the sample words to be input, and generate a sample word set to be input according to the sample words to be input.
[0178] S303: Obtain the word index corresponding to each sample word to be input in the word dictionary.
[0179] Specifically, the difference between the first normal sample text and the above-mentioned text to be detected is that it is unknown whether the text to be detected belongs to normal text or abnormal text, while the first normal sample text is the collected normal text. If it is desired that the trained local semantic model be used to identify short message text, the first normal sample text can be the normal short message text obtained, that is, the short message without junk information. A number of first sample entity byte units can be, well, a number of first sample entity byte units. For the entity word segmentation process and index acquisition process of the first normal sample text in steps S301 - S303, reference can be made to Figure 3 The entity word segmentation process and index acquisition process of the text to be detected by the computer device in steps S201 - S203 in the corresponding embodiment.
[0180] S304: Identify the sorting position of each input sample word in the input sample word set, call the local semantic initial model, and predict the word index corresponding to each input sample word according to the sorting position in the local semantic initial model, so as to obtain the sample word prediction distribution corresponding to each first sample entity byte unit.
[0181] Specifically, the structure of the local semantic initial model can be referred to Figure 4a The structure of the text detection model shown. The local semantic initial model includes a first initial embedding layer, an initial bidirectional memory network layer, and an initial normalization layer. As Figure 4aThe embedding layer shown can be the first initial embedding layer, the Bi-LSTM layer can be the initial bidirectional memory network layer, and the SoftMax layer is the initial normalization layer. The computer device will call the first initial embedding layer to perform initial embedding feature processing on the word indices corresponding to each input sample word to obtain the sample word embedding vectors corresponding to each input sample word, and then identify the sample sorting positions of each input sample word in the input sample word set, and call the initial bidirectional memory network layer to perform initial hidden layer feature processing on the word embedding vectors corresponding to each input sample word to obtain the sample forward hidden layer representation vectors and sample backward hidden layer representation vectors corresponding to each first sample entity byte unit. Finally, the computer device will call the initial normalization layer to perform normalization processing on the sample forward hidden layer representation vectors and backward hidden layer representation vectors corresponding to each first sample entity byte unit to obtain the sample forward prediction distribution and sample backward prediction distribution corresponding to each first sample entity byte unit. In other words, the computer device will input the word indices corresponding to each input sample word into the first initial embedding layer, and output the sample word embedding vectors corresponding to each input sample word through the first initial embedding layer; then the computer device will, according to the sorting positions of each input sample word in the input sample word set, input the word embedding vectors corresponding to each input sample word into the initial bidirectional memory network layer, and output the sample forward hidden layer representation vectors and sample backward hidden layer representation vectors corresponding to each first sample entity byte unit through the initial bidirectional memory network layer; then the computer device will input the sample forward hidden layer representation vectors corresponding to each first sample entity byte unit into the initial normalization layer, and output the sample forward prediction distribution corresponding to each first sample entity byte unit through the initial normalization layer; the computer device will input the sample backward hidden layer representation vectors corresponding to each first sample entity byte unit into the initial normalization layer, and output the sample backward prediction distribution corresponding to each first sample entity byte unit through the initial normalization layer. The sample forward prediction distribution and sample backward prediction distribution obtained by the computer device through the local semantic initial model are the sample word prediction distributions corresponding to the first sample entity byte unit. For the specific implementation process of inputting the word index corresponding to the input sample word into the local semantic initial model and then obtaining the sample word prediction distribution of the first sample entity byte unit, reference can be made to the Figure 3 processing process of the computer device for the text to be detected through the local semantic model in step S204 in the corresponding embodiment.
[0182] S305: Train the local semantic initial model based on the word label distribution corresponding to each first sample entity byte unit and the sample word prediction distribution to obtain the local semantic model.
[0183] Specifically, each first-sample entity byte unit corresponds to a word tag distribution, which can also be called the true tag distribution of the sample. The true tag distribution also includes the occurrence probabilities of all words in the word dictionary. For the true tag distribution of the first-sample entity byte unit M, only the occurrence probability corresponding to the first-sample entity byte unit M is 1, and the rest are all 0. For example, if the word dictionary is (M-1, N-2, L-3), then the true tag distribution of the first-sample entity byte unit M is [1, 0, 0], and the true tag distribution of the first-sample entity byte unit L is [0, 0, 1]. According to the true tag distribution, the sample forward prediction distribution, and the sample backward prediction distribution of each first-sample entity byte unit, the cross-entropy loss function of the local semantic initial model can be constructed as shown in formula (10):
[0184]
[0185] where L1 is the loss value of the local semantic initial model, Ω is the training set, q is the maximum sentence length, p t is the true tag distribution of the t-th first-sample entity byte unit, is the sample backward prediction distribution of the t-th first-sample entity byte unit, is the sample forward prediction distribution of the t-th first-sample entity byte unit. Among them, the training set is the set including the first normal sample text. After the computer device determines the loss value according to formula (10), the local semantic initial model can use the ADAM optimization algorithm and the backpropagation algorithm to update parameters and learn. When the adjusted local semantic initial model meets the model convergence condition, the adjusted local semantic initial model is determined as the local semantic model. Among them, the convergence condition can be that the loss value of the local semantic initial model is basically stable.
[0186] Using the method provided in the embodiments of the present application to train the local semantic initial model based on the first normal sample text. After the trained local semantic initial model meets the model convergence condition, the local semantic model is obtained. Then, the local semantic model can be used to obtain the occurrence probability of each entity byte unit in the text to be detected, and the perplexity of the text to be detected can be determined based on the occurrence probability of each entity byte unit. Among them, the training of the local semantic initial model can be carried out offline, and the local semantic model can be directly used after obtaining the text to be detected, saving the time for determining the perplexity of the text to be detected.
[0187] Further, please refer to Figure 6 , Figure 6 which is a schematic flowchart of a process for training a global semantic initial model provided in the embodiments of the present application. This process is executed by the computer device described in Figure 1a i.e., it can beFigure 1a Server 100 in Figure 6 As shown, the training process includes the following steps:
[0188] S401: Obtain the second normal sample text, perform entity word segmentation on the content of the second normal sample text to obtain a number of second sample entity byte units after entity word segmentation.
[0189] S402: Determine both the head label word, the tail label word, and the number of second sample entity byte units as the first sample words, and generate a sample text word set according to the first sample words.
[0190] Specifically, the second normal sample text is also the collected normal text. It can be understood that the first normal sample text and the second normal sample text can be the same normal text. The number of second sample entity byte units can be the number of second sample entity byte units. For the process of word segmentation of the second normal sample text and generating the sample text word set, reference can be made to the relevant descriptions of steps S201 and steps S205 - S206 in the corresponding embodiments above. Figure 3 The relevant descriptions of steps S201 and steps S205 - S206 in the corresponding embodiments.
[0191] S403: Obtain N random label words, and use the N random label words to replace N first sample words in the sample text word set respectively to obtain a replaced sample text word set; N is a positive integer less than the total number of the first sample words in the sample text word set.
[0192] Specifically, the random label words are used to replace some of the second sample entity byte units, that is, the [MASK] words mentioned above. For example, the sample text word set can be {"[CLS]", "Xiaowang", "likes", "basketball", "[SEP]", "Xiaoming", "likes", "football", "[SEP]"}, and by randomly replacing some of the first sample words in the sample text word set, a replaced sample text word set can be obtained, such as {"[CLS]", "Xiaowang", "[MASK]", "basketball", "[SEP]", "Xiaoming", "likes", "[MASK]", "[SEP]"}. Among them, some of the first sample words are second sample entity byte units, that is to say, the [MASK] words will not replace the head label word or the tail label word. The number of second sample entity byte units to be replaced should be less than the total number of the first sample words in the sample text word set.
[0193] S404: Obtain the label position information of the N random label words in the replaced sample text word set, and generate a label position information set according to the label position information.
[0194] Specifically, after the computer device replaces the first sample word in the sample text word set with the [MASK] word, it will record the position information of the replaced first sample word as the label position information of the [MASK] word in the replaced sample text word set, and then obtain a set of label position information. For example, in the above sample text word set {"[CLS]", "Xiaowang", "likes", "basketball", "[SEP]", "Xiaoming", "likes", "football", "[SEP]"}, the replaced sample text word set is {"[CLS]", "Xiaowang", "[MASK]", "basketball", "[SEP]", "Xiaoming", "likes", "[MASK]", "[SEP]"}, and the first sample words to be replaced are "likes" and "football". There are two "likes" in the sample text word set. If only the first sample word replaced by the [MASK] word is recorded, the computer device cannot know which "likes" the [MASK] word replaces. Therefore, the computer device will obtain the position information 2 of the first "likes" as the label position information corresponding to the first [MASK] word, and the position information 7 of "football" as the label position information corresponding to the second [MASK] word.
[0195] S405: Train the global semantic initial model according to the replaced sample text word set, the set of label position information, and the random label word to generate the global semantic model.
[0196] Specifically, the computer device will call the global semantic processing model to perform initial semantic feature extraction processing on the replaced sample text word set to obtain the sample word prediction distributions corresponding to the N random label words. That is to say, the computer device will input the replaced sample text word set into the global semantic initial model, and then output the hidden layer vector corresponding to each first sample word in the replaced sample text word set through the global semantic initial model. Among them, the structure of the global semantic initial model can refer to the structure of the text detection model in the corresponding embodiment above. Figure 4b Then, the process of initial semantic feature extraction processing can refer to the process of the computer device performing semantic feature extraction processing in the corresponding embodiment above. Figure 3 Specifically, the computer device will obtain the hidden layer vector corresponding to each random label word from the hidden layer vector corresponding to the first sample word as the random hidden layer vector. Assume that among the random label words, the random hidden layer vector corresponding to the x-th random label word is
[0197] According to According to The sample word prediction distribution corresponding to the x-th random label word can be obtained The mathematical expression is as follows:
[0198]
[0199] The computer device obtains the sample word prediction distribution corresponding to each random label word through the random hidden layer vector of each random label word. Then, the computer device obtains the corresponding first sample word in the sample text word set as the target sample word according to the label position information in the label position information set. Then. The computer device obtains the word label distribution of the target sample word, and the obtaining process is the same as that of obtaining the word label distribution of the first sample entity byte unit in step S305 in the corresponding embodiment above, which will not be elaborated here. According to the obtained sample word prediction distribution corresponding to each random label word and the word label distribution corresponding to each target sample word, the cross-entropy loss function of the global semantic initial model can be obtained, that is, formula (12): Figure 3 where L2 is the loss value of the global semantic initial model, Ω is the training set, X is the label position information set, p
[0200]
[0201] is the word label distribution of the target sample word at the x-th position, x and is the sample word prediction distribution of the random label word at the x-th position. Among them, the training set is the set containing the second normal sample text. After the computer device determines the loss value of the global semantic initial model according to formula (12), the global semantic initial model can use the ADAM optimization algorithm and the backpropagation algorithm for parameter update and learning. When the adjusted global semantic initial model meets the model convergence condition, the adjusted global semantic initial model is determined as the global semantic model mentioned in the corresponding embodiment above Figure 3 where the convergence condition can be that the loss value of the global semantic initial model is basically stable.
[0202] By using the method provided in the embodiment of the present application, a global semantic model for obtaining the target global representation vector of the text to be detected can be obtained through the second normal sample text and the BERT model, and the text semantics of the text to be detected can be effectively learned without labeled data to obtain the target global representation vector.
[0203] Further, please refer to Figure 7 , Figure 7 which is a flowchart of a method for recognizing a text to be detected provided in the embodiment of the present application. This method is executed by the computer device described in Figure 1a , that is, it can be the server 100 in Figure 1a , or it can be Figure 1aThe user terminal cluster (including user terminal 10a, user terminal 10b, and user terminal 10n). Taking the scenario where this method is applied to SMS text recognition as an example, as Figure 7 shown, the computer device collects normal SMS samples authorized by the user and performs data preprocessing operations such as entity recognition, word segmentation, and dictionary construction to obtain a word dictionary. The computer device stores the processed normal SMS samples in the normal SMS sample library 71 as Figure 7 shown. When training data is needed, the computer device obtains normal sample texts from the normal SMS sample library 71, then performs word segmentation and entity recognition on them to obtain entity byte units. The computer device generates an AR sample set based on the entity byte units, start tag words, and end tag words, uses it as the training data required for local semantic coherence modeling, and performs language modeling using the Bi-LSTM model. After training is completed, the local semantic model 72 is obtained. The computer device can also generate an AE sample set based on the entity byte units, random tag words, head tag words, and tail tag words, use it as the training data required for global semantic coherence modeling, train the BERT model, and obtain the global semantic model 73 after training is completed. Then the computer device can input the normal sample text into the global semantic model, take the global representation vector corresponding to the [CLS] token as the global semantic representation of the SMS, and perform K-means clustering on the global semantic representations of all normal sample texts to obtain K clustering centers. The above training process can be carried out offline.
[0204] As Figure 7 shown, when the computer device obtains the SMS received by the terminal online, it uses it as a test sample and then performs data preprocessing on it. The computer device inputs each entity byte unit of the test sample into the local semantic model 72 moment by moment, predicts the distribution of the next entity byte unit, collects the probability p t of the real character, and calculates the perplexity PPL through the probabilities at all moments to measure the semantic coherence of the text. At the same time, after obtaining the sentence representation vector of the test sample according to the global semantic model 73, calculate the average distance to the Q nearest clustering centers as the vector distance score VD to measure the semantic normality of the text. Finally, fuse the above two scores through the hyperparameter λ to obtain the final semantic score, and distinguish spam SMS through the threshold parameter θ.
[0205] By using the method provided in the embodiments of the present application, on the one hand, Bi-LSTM is used for bidirectional language modeling (given historical words, predicting the next word), and the perplexity PPL of the entire sentence is calculated. The higher the perplexity, the worse the semantic coherence. On the other hand, the BERT model is trained with normal sample texts, and the average distance VD between the feature vector of the sample and the nearest Q cluster centers after clustering with the reference text is calculated. The larger the distance, the farther the deviation from the normal semantics. Then, by defining the hyperparameter λ, the above two heuristic semantic relevance scoring models are organically combined to obtain a semantic score, and new deformed spam texts are screened out by setting a threshold and the final semantic score. The embodiments of the present application do not require any manual annotation and feature engineering, can save a large amount of labor costs, propose a novel idea to screen out new deformed spam texts, improve the existing spam text detection method, and improve the accuracy of spam text recognition and the user experience.
[0206] Further, please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a text detection device provided by the embodiments of the present application. The above text detection device can be a computer program (including program code) running in a computer device. For example, the text detection device is an application software; the device can be used to execute the corresponding steps in the method provided by the embodiments of the present application. As Figure 8 shown, the text detection device 1 may include: a target text to be detected processing module 101, a perplexity determination module 102, a vector acquisition module 103, a vector distance determination module 104, and a target text determination module 105.
[0207] The target text processing module 101 is configured to obtain the target text to be detected, perform entity word segmentation processing on the content of the target text to be detected, and obtain a number of entity byte units after entity word segmentation processing;
[0208] The perplexity determination module 102 is configured to perform entity prediction processing on a number of entity byte units, obtain the occurrence probabilities of the number of entity byte units in the target text to be detected, and determine the perplexity for measuring the semantic coherence of the target text to be detected according to the corresponding occurrence probabilities of the number of entity byte units;
[0209] The vector acquisition module 103 is configured to perform semantic feature extraction processing on the target text to be detected, and obtain a target global representation vector of the target text to be detected;
[0210] The vector distance determination module 104 is configured to determine the normal text vector distance for measuring the normal semantics of the target text to be detected according to the target global representation vector and the reference global representation vector of the normal text type;
[0211] The target text determination module 105 is configured to determine whether the target text to be detected belongs to normal text or abnormal text according to the perplexity and the normal text vector distance.
[0212] Among them, for the specific implementation manners of the target text processing module 101, the perplexity determination module 102, the vector acquisition module 103, the vector distance determination module 104, and the target text determination module 105, reference may be made to the descriptions of steps S101 - S104 in the corresponding embodiments above. Figure 2 Details will not be described herein again.
[0213] Please refer to Figure 8 , the perplexity determination module 102 includes: an input word set generation unit 1021, an index determination unit 1022, a local model prediction unit 1023, an occurrence probability determination unit 1024, and a perplexity determination unit 1025.
[0214] The input word set generation unit 1021 is configured to obtain a start tag word and an end tag word, determine the start tag word, the end tag word, and a plurality of entity byte units as input words to be input, and generate an input word set according to the input words to be input.
[0215] The index determination unit 1022 is configured to obtain the word index corresponding to each input word to be input in a word dictionary.
[0216] The local model prediction unit 1023 is configured to identify the sorting position of each input word to be input in the input word set, call a local semantic model, and predict the word index corresponding to each input word to be input in the local semantic model according to the sorting position, so as to obtain the word prediction distribution corresponding to each entity byte unit; the word prediction distribution is used to represent the occurrence probability of the words in the word dictionary in the target text.
[0217] The occurrence probability determination unit 1024 is configured to determine the occurrence probability of each entity byte unit in the target text to be detected according to the word prediction distribution corresponding to each entity byte unit.
[0218] The perplexity determination unit 1025 is configured to determine the perplexity for measuring the semantic coherence of the target text to be detected according to the occurrence probability of each entity byte unit in the target text to be detected.
[0219] Among them, for the specific implementation manners of the input word set generation unit 1021, the index determination unit 1022, the local model prediction unit 1023, the occurrence probability determination unit 1024, and the perplexity determination unit 1025, reference may be made to the descriptions of steps S201 - S205 in the corresponding embodiments above. Figure 3 Details will not be described herein again.
[0220] Among them, the word prediction distribution corresponding to each entity byte unit respectively includes a forward prediction distribution and a backward prediction distribution; the local semantic model includes a first embedding layer, a bidirectional memory network layer, and a normalization layer;
[0221] Please refer to Figure 8 , the local model prediction unit 1023 includes: a first embedding layer subunit 10231, a first network layer subunit 10232, and a normalization layer subunit 10233.
[0222] The first embedding layer subunit 10231 is used to call the first embedding layer to perform embedding feature processing on the word index corresponding to each word to be input, and obtain the word embedding vector corresponding to each word to be input;
[0223] The first network layer subunit 10232 is used to identify the sorting position of each word to be input in the word set to be input, and call the bidirectional memory network layer to perform hidden layer feature processing on the word embedding vector corresponding to each word to be input according to the sorting position, and obtain the forward hidden layer representation vector and the backward hidden layer representation vector corresponding to each entity byte unit respectively;
[0224] The normalization layer subunit 10233 is used to call the normalization layer to perform normalization processing on the forward hidden layer representation vector and the backward hidden layer representation vector corresponding to each entity byte unit respectively, and obtain the forward prediction distribution and the backward prediction distribution corresponding to each entity byte unit respectively.
[0225] Among them, the first network layer subunit 10232 is specifically used to identify the forward sorting position of each word to be input in the word set to be input, and call the bidirectional memory network layer to perform hidden layer feature processing on the word embedding vector corresponding to each word to be input according to the forward sorting position, and obtain the forward hidden layer representation vector corresponding to each entity byte unit respectively;
[0226] The first network layer subunit 10232 is specifically further used to identify the backward sorting position of each word to be input in the word set to be input, and call the bidirectional memory network layer to perform hidden layer feature processing on the word embedding vector corresponding to each word to be input according to the backward sorting position, and obtain the backward hidden layer representation vector corresponding to each entity byte unit respectively.
[0227] Among them, for the specific implementation manners of the first embedding layer subunit 10231, the first network layer subunit 10232, and the normalization layer subunit 10233, reference can be made to the description of step S204 in the corresponding embodiment above, which will not be elaborated here. Figure 3 The description will not be repeated here.
[0228] Please refer to Figure 8, the occurrence probability determination unit 1024 includes: a word average prediction subunit 10241, a target word determination subunit 10242, and an occurrence probability determination subunit 10243.
[0229] The word average prediction subunit 10241 is configured to generate an average prediction distribution corresponding to each entity byte unit according to the forward prediction distribution and the backward prediction distribution respectively corresponding to each entity byte unit; a plurality of entity byte units include an entity byte unit M.
[0230] The target word determination subunit 10242 is configured to find a word that matches the word index corresponding to the entity byte unit M in the average prediction distribution corresponding to the entity byte unit M as the target word.
[0231] The occurrence probability determination subunit 10243 is configured to determine the occurrence probability corresponding to the target word in the average prediction distribution corresponding to the entity byte unit M as the occurrence probability of the entity byte unit M in the target text to be detected.
[0232] Among them, for the specific implementation manners of the word average prediction subunit 10241, the target word determination subunit 10242, and the occurrence probability determination subunit 10243, reference can be made to the description of step S205 in the corresponding embodiment above. Figure 3 Details will not be described herein again.
[0233] Please refer to Figure 8 , the vector acquisition module 103 includes: a text word set generation unit 1031, a first set determination unit 1032, a second set determination unit 1033, a third set determination unit 1034, and a global model unit 1035.
[0234] The text word set generation unit 1031 is configured to obtain a head tag word and a tail tag word, determine the head tag word, the tail tag word, and a plurality of entity byte units as words to be processed, and generate a text word set according to the words to be processed.
[0235] The first set determination unit 1032 is configured to obtain the word index corresponding to each word to be processed in the word dictionary and generate a set of word indices to be processed.
[0236] The second set determination unit 1033 is configured to obtain the word position information of each word to be processed in the text word set and generate a set of word position information.
[0237] The third set determination unit 1034 is configured to obtain the sentence position information of each word to be processed in the text word set and generate a set of sentence position information.
[0238] The global model unit 1035 is used to call the global semantic model to perform semantic feature extraction processing on the set of word indices to be processed, the set of word position information, and the set of sentence position information in the global semantic model, so as to obtain the target global representation vector of the text to be detected.
[0239] Among them, for the specific implementation manners of the text word set generation unit 1031, the first set determination unit 1032, the second set determination unit 1033, the third set determination unit 1034, and the global model unit 1035, reference can be made to the descriptions of steps S206 - S208 in the corresponding embodiments above. Figure 3 Details will not be elaborated here.
[0240] Among them, the global semantic model includes a second embedding layer and a self - attention network layer;
[0241] Please refer to Figure 8 , the global model unit 1035 includes: a second embedding layer sub - unit 10351, a fusion vector determination sub - unit 10352, a second network layer sub - unit 10353, and a global vector determination sub - unit 10354.
[0242] The second embedding layer sub - unit 10351 is used to call the second embedding layer to perform embedding feature processing on the set of word indices to be processed, the set of word position information, and the set of sentence position information, so as to obtain the word index embedding vector to be processed, the word position information embedding vector, and the sentence position information embedding vector corresponding to each word to be processed respectively;
[0243] The fusion vector determination sub - unit 10352 is used to generate the fusion embedding vector corresponding to each word to be processed according to the word index embedding vector to be processed, the word position information embedding vector, and the sentence position information embedding vector corresponding to each word to be processed respectively; the fusion embedding vector corresponding to a word to be processed is generated by the word index embedding vector to be processed, the word position information embedding vector, and the sentence position information embedding vector corresponding to this word to be processed;
[0244] The second network layer sub - unit 10353 is used to call the self - attention network layer to perform semantic representation processing on the fusion embedding vector corresponding to each word to be processed, so as to obtain the hidden layer vector corresponding to each word to be processed respectively;
[0245] The global vector determination sub - unit 10354 is used to determine, among the hidden layer vectors corresponding to each word to be processed respectively, the hidden layer vector corresponding to the head label word as the target global representation vector of the text to be detected.
[0246] Among them, for the specific implementation manners of the second embedding layer subunit 10351, the fusion vector determination subunit 10352, the second network layer subunit 10353, and the global vector determination subunit 10354, reference can be made to the description of step S208 in the corresponding embodiment above, which will not be elaborated here. Figure 3 The description of step S208 in the corresponding embodiment will not be elaborated here.
[0247] Among them, the number of normal text types is at least two;
[0248] Please refer to Figure 8 , the vector distance determination module 104 includes: a reference distance determination unit 1041, a vector selection unit 1042, and a vector distance determination unit 1043.
[0249] The reference distance determination unit 1041 is configured to respectively obtain the vector distance between the reference global representation vector of each normal text type and the target global representation vector as the reference distance;
[0250] The vector selection unit 1042 is configured to sequentially obtain Q reference global representation vectors from at least two reference global representation vectors according to the reference distance; Q is a positive integer, and Q is less than or equal to the number of normal text types;
[0251] The vector distance determination unit 1043 is configured to generate the average vector distance between the target global representation vector and the Q reference global representation vectors, and use the average vector distance as the normal text vector distance of the to-be-detected target text.
[0252] Among them, for the specific implementation manners of the reference distance determination unit 1041, the vector selection unit 1042, and the vector distance determination unit 1043, reference can be made to the description of step S104 in the corresponding embodiment above, which will not be elaborated here. Figure 2 The description of step S104 in the corresponding embodiment will not be elaborated here.
[0253] Please refer to Figure 8 , the target text determination module 105 includes: a harmonic unit 1051, an abnormal text determination unit 1052, and a normal text determination unit 1053.
[0254] The harmonic unit 1051 is configured to harmonize the perplexity and the normal text vector distance according to the reference weights corresponding to the perplexity and the normal text vector distance respectively, to obtain the text semantic relevance of the target text;
[0255] The abnormal text determination unit 1052 is configured to determine that the to-be-detected target text belongs to abnormal text if the text semantic relevance is greater than the standard relevance threshold;
[0256] The normal text determination unit 1053 is configured to determine that the to-be-detected target text belongs to normal text if the text semantic relevance is less than or equal to the standard relevance threshold.
[0257] Among them, for the specific implementation manners of the harmonic unit 1051, the abnormal text determination unit 1052, and the normal text determination unit 1053, reference can be made to the description of step S105 in the corresponding embodiment above, which will not be elaborated here. Figure 2 For the description of step S105 in the corresponding embodiment above, which will not be elaborated here.
[0258] Please refer to Figure 8 , the text detection device 1 further includes: a reference normal text acquisition module 106, a global vector acquisition module 107, a clustering module 108, and a reference vector determination module 109.
[0259] The reference normal text acquisition module 106 is configured to acquire at least two reference normal texts;
[0260] The global vector acquisition module 107 is configured to acquire global representation vectors corresponding to at least two reference normal texts respectively based on a global semantic model;
[0261] The clustering module 108 is configured to perform clustering processing on the global representation vectors corresponding to at least two reference normal texts respectively to obtain N clustering clusters; each clustering cluster respectively represents a different normal text type;
[0262] The reference vector determination module 109 is configured to use the global representation vectors located at the cluster centers in each clustering cluster as reference global representation vectors of the normal text types respectively.
[0263] Among them, for the specific implementation manners of the reference normal text acquisition module 106, the global vector acquisition module 107, the clustering module 108, and the reference vector determination module 109, reference can be made to the description of step S104 in the corresponding embodiment above, which will not be elaborated here. Figure 2 For the description of step S104 in the corresponding embodiment above, which will not be elaborated here.
[0264] Please refer to Figure 8 , the text detection device 1 further includes: a first text processing module 110, a first sample set generation module 111, a first index acquisition module 112, a sample prediction acquisition module 113, and a first training module 114.
[0265] The first text processing module 110 is configured to acquire a first normal sample text, perform entity word segmentation processing on the content of the first normal sample text to obtain a plurality of first sample entity byte units after entity word segmentation processing; the prediction sample set generation module 111 is configured to determine the start tag word, the end tag word, and a plurality of first sample entity byte units as to-be-input sample words, and generate a to-be-input sample word set according to the to-be-input sample words;
[0266] The sample index acquisition module 112 is configured to acquire the word index corresponding to each to-be-input sample word in a word dictionary;
[0267] A sample prediction acquisition module 113, configured to identify the sample sorting position of each to-be-input sample word in the to-be-input sample word set, call a local semantic initial model, and predict the word index corresponding to each to-be-input sample word in the local semantic initial model according to the sample sorting position, so as to obtain the sample word prediction distribution corresponding to each first sample entity byte unit;
[0268] A first training module 114, configured to train the local semantic initial model based on the word label distribution and the sample word prediction distribution corresponding to each first sample entity byte unit, so as to obtain a local semantic model.
[0269] Among them, for the specific implementation manners of the first text processing module 110, the first sample set generation module 111, the sample index acquisition module 112, the sample prediction acquisition module 113, and the first training module 114, reference may be made to the descriptions of steps S301 - S305 in the corresponding embodiments above. Details will not be elaborated here. Figure 5 The descriptions of steps S301 - S305 in the corresponding embodiments above will not be repeated here.
[0270] Among them, the sample word prediction distribution corresponding to each first sample entity byte unit includes a sample forward prediction distribution and a sample reverse prediction distribution; the local semantic model includes a first initial embedding layer, an initial bidirectional memory network layer, and an initial normalization layer;
[0271] Please refer to Figure 8 , the sample prediction acquisition module 113 includes: a first initial embedding layer unit 1131, an initial network layer unit 1132, and an initial normalization layer unit 1133.
[0272] The first initial embedding layer unit 1131 is configured to call the first initial embedding layer to perform initial embedding feature processing on the word index corresponding to each to-be-input sample word, so as to obtain the sample word embedding vector corresponding to each to-be-input sample word;
[0273] The initial network layer unit 1132 is configured to identify the sample sorting position of each to-be-input sample word in the to-be-input sample word set, and call the initial bidirectional memory network layer to perform initial hidden layer feature processing on the word embedding vector corresponding to each to-be-input sample word, so as to obtain the sample forward hidden layer representation vector and the sample reverse hidden layer representation vector corresponding to each first sample entity byte unit;
[0274] The initial normalization layer unit 1133 is configured to call the initial normalization layer to perform normalization processing on the sample forward hidden layer representation vector and the reverse hidden layer representation vector corresponding to each first sample entity byte unit, so as to obtain the sample forward prediction distribution and the sample reverse prediction distribution corresponding to each first sample entity byte unit.
[0275] Among them, for the specific implementation manners of the first initial embedding layer unit 1131, the initial network layer unit 1132, and the initial normalization layer unit 1133, reference may be made to the description of step S304 in the corresponding embodiment above, which will not be elaborated here. Figure 5 Please refer to
[0276] For example, the first training module 114 includes: a first loss determination unit 1141 and a first training unit 1142. Figure 8 The first loss determination unit 1141 is configured to construct a cross-entropy loss function of the local semantic initial model according to the sample forward prediction distribution corresponding to each first sample entity byte unit, the sample reverse prediction distribution corresponding to each first sample entity byte unit, and the word label distribution corresponding to each first sample entity byte unit;
[0277] The first training unit 1142 is configured to adjust the parameters of the local semantic initial model according to the cross-entropy loss function; when the adjusted local semantic initial model meets the model convergence condition, the adjusted local semantic initial model is determined as the local semantic model.
[0278] Among them, for the specific implementation manners of the first loss determination unit 1141 and the first training unit 1142, reference may be made to the description of step S305 in the corresponding embodiment above, which will not be elaborated here.
[0279] Please refer to Figure 5 For example, the text detection device 1 further includes: a second text processing module 115, a sample word set generation module 116, a random replacement module 117, a label position information determination module 118, and a second training module 119.
[0280] The second text processing module 115 is configured to obtain a second normal sample text, perform entity word segmentation processing on the content of the second normal sample text to obtain a plurality of second sample entity byte units after entity word segmentation processing; the sample word set generation module 116 is configured to determine both the head label word, the tail label word, and the plurality of second sample entity byte units as first sample words, and generate a sample text word set according to the first sample words; Figure 8 The random replacement module 117 is configured to obtain N random label words, and respectively replace N first sample words in the sample text word set with the N random label words to obtain a replaced sample text word set; N is a positive integer less than the total number of first sample words in the sample text word set;
[0281] The label position information determination module 118 is configured to determine the position information of the replaced first sample words in the replaced sample text word set;
[0282] The second training module 119 is configured to construct a cross-entropy loss function of the global semantic model according to the sample forward prediction distribution corresponding to each second sample entity byte unit, the sample reverse prediction distribution corresponding to each second sample entity byte unit, and the word label distribution corresponding to each second sample entity byte unit;
[0283] A label position information determination module 118, configured to obtain label position information of N random label words in the set of sample text words after replacement, and generate a set of label position information according to the label position information;
[0284] A second training module 119, configured to train a global semantic initial model according to the set of sample text words after replacement, the set of label position information, and the random label words, and generate a global semantic model.
[0285] Among them, for the specific implementation manners of the second text processing module 115, the sample word set generation module 116, the random replacement module 117, the label position information determination module 118, and the second training module 119, reference can be made to the descriptions of steps S401 - S405 in the corresponding embodiments above, which will not be elaborated here. Figure 6 For details, please refer to
[0286] Please refer to Figure 8 , the second training module 119 includes: a random prediction determination unit 1191, a target sample word determination unit 1192, a word label distribution acquisition unit 1193, a second loss determination unit 1194, and a second training unit 1195.
[0287] The random prediction determination unit 1191 is configured to call a global semantic processing model to perform initial semantic feature extraction processing on the set of sample text words after replacement, and obtain sample word prediction distributions corresponding to the N random label words respectively;
[0288] The target sample word determination unit 1192 is configured to obtain, according to the N label position information in the set of label position information, the first sample words in the set of sample text words at the N label position information as target sample words;
[0289] The word label distribution acquisition unit 1193 is configured to obtain word label distributions corresponding to the N target sample words respectively;
[0290] The second loss determination unit 1194 is configured to constitute a cross - entropy loss function of the global semantic initial model according to the sample word prediction distributions corresponding to the N random label words respectively and the word label distributions corresponding to the N target sample words respectively;
[0291] The second training unit 1195 is configured to adjust the parameters of the global semantic initial model according to the cross - entropy loss function, and when the adjusted global semantic initial model meets the model convergence condition, determine the adjusted global semantic initial model as the global semantic model.
[0292] The specific implementation of the random prediction determination unit 1191, the target sample word determination unit 1192, the word label distribution acquisition unit 1193, the second loss determination unit 1194 and the second training unit 1195 can be found in the above Figure 6 The description of step S404 in the corresponding embodiment will not be repeated here.
[0293] For further information, see Figure 9 , Figure 9 Schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 9 As shown, the computer device 2000 can be applied to a server, which can be the above-mentioned Figure 1a The business server 100 in the corresponding embodiment; the computer device 2000 can be applied to a terminal, which can be the above-mentioned Figure 1a The user terminal 10a, user terminal 10b, ..., user terminal 10n in the corresponding embodiment; the computer device 2000 may also be the above-mentioned Figure 2 The computer device in the corresponding embodiment. The computer device 2000 may include: a processor 2001, a network interface 2004 and a memory 2005. In addition, the above-mentioned computer device 2000 also includes: a user interface 2003, and at least one communication bus 2002. Among them, the communication bus 2002 is used to realize the connection and communication between these components. The network interface 2004 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface). The memory 2005 may be a high-speed RAM memory, or it may be a non-volatile memory (non-volatile memory), such as at least one disk storage. The memory 2005 may optionally also be at least one storage device located away from the aforementioned processor 2001. Figure 9 As shown, the memory 2005 as a computer-readable storage medium may include an operating system, a network communication module, a user interface module, and a device control application program.
[0294] exist Figure 9 In the computer device 2000 shown, the network interface 2004 can provide a network communication function; the user interface 2003 is mainly used to provide an input interface for the user; and the processor 2001 can be used to call the device control application stored in the memory 2005 to achieve:
[0295] Obtaining a target text to be detected, performing entity segmentation processing on the content of the target text to be detected, and obtaining a number of entity byte units processed by the entity segmentation;
[0296] Perform entity prediction processing on a number of entity byte units to obtain the occurrence probabilities of the number of entity byte units in the text to be detected, and determine the perplexity for measuring the semantic coherence of the text to be detected according to the corresponding occurrence probabilities of the number of entity byte units;
[0297] Perform semantic feature extraction processing on the text to be detected to obtain the target global representation vector of the text to be detected;
[0298] According to the target global representation vector and the reference global representation vector of the normal text type, determine the normal text vector distance for measuring the normal semantic property of the text to be detected;
[0299] According to the perplexity and the normal text vector distance, determine whether the text to be detected belongs to normal text or abnormal text.
[0300] It should be understood that the computer device 2000 described in the embodiments of the present application can execute the description of the corresponding embodiments in the foregoing Figure 2 and can also execute the description of the text detection device 1 in the corresponding embodiments in the foregoing Figure 8 which will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either.
[0301] In addition, it should be pointed out here that: The embodiments of the present application also provide a computer-readable storage medium, and the computer-readable storage medium stores the computer program executed by the foregoing text detection device 1. When the processor executes the computer program, it can execute the description of the text detection method in the corresponding embodiments in the foregoing Figure 2 which will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer storage medium involved in the present application, please refer to the description of the method embodiments of the present application.
[0302] The foregoing computer-readable storage medium may be the text detection device provided in any of the foregoing embodiments or the internal storage unit of the foregoing computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.
[0303] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A text detection method, characterized in that, Including: Obtain the target text to be detected, perform entity word segmentation processing on the content of the target text to be detected, and obtain a number of entity byte units after entity word segmentation processing; Perform entity prediction processing on the number of entity byte units to obtain the occurrence probability of the number of entity byte units in the target text to be detected, and determine the perplexity for measuring the semantic coherence of the target text to be detected according to the occurrence probability corresponding to the number of entity byte units; Perform semantic feature extraction processing on the target text to be detected to obtain the target global representation vector of the target text to be detected; Obtain the global representation vectors corresponding to at least two reference normal texts respectively, perform clustering processing on the global representation vectors corresponding to the at least two reference normal texts respectively to obtain at least two clustering clusters, and respectively use the global representation vectors located at the cluster centers in each clustering cluster as the reference global representation vectors of the normal text types; each clustering cluster respectively represents different normal text types; Determine the normal text vector distance for measuring the normal semanticity of the target text to be detected according to the target global representation vector and the reference global representation vectors of at least two normal text types; Determine whether the target text to be detected belongs to a normal text or an abnormal text according to the perplexity and the normal text vector distance.
2. The method according to claim 1, characterized in that, The performing entity prediction processing on the number of entity byte units to obtain the occurrence probability of the number of entity byte units in the target text to be detected, and determining the perplexity for measuring the semantic coherence of the target text to be detected according to the occurrence probability corresponding to the number of entity byte units includes: Obtain a start tag word and an end tag word, determine the start tag word, the end tag word, and the number of entity byte units as words to be input, and generate a set of words to be input according to the words to be input; Obtain the word index corresponding to each word to be input in a word dictionary; Identify the sorting position of each word to be input in the set of words to be input, call a local semantic model, and predict the word index corresponding to each word to be input according to the sorting position in the local semantic model to obtain the word prediction distribution corresponding to each entity byte unit; the word prediction distribution is used to represent the occurrence probability of the words in the word dictionary in the target text to be detected; Determine the occurrence probability of each entity byte unit in the target text to be detected according to the word prediction distribution corresponding to each entity byte unit; Determine the perplexity for measuring the semantic coherence of the target text to be detected according to the occurrence probability of each entity byte unit in the target text to be detected.
3. The method according to claim 2, characterized in that The word prediction distribution corresponding to each entity byte unit includes a forward prediction distribution and a backward prediction distribution; the local semantic model includes a first embedding layer, a bidirectional memory network layer, and a normalization layer; Identifying the sorting position of each word to be input in the set of words to be input, and calling the local semantic model to predict the word index corresponding to each word to be input according to the sorting position in the local semantic model, so as to obtain the word prediction distribution corresponding to each entity byte unit, including: Calling the first embedding layer to perform embedding feature processing on the word index corresponding to each word to be input, so as to obtain the word embedding vector corresponding to each word to be input; Identifying the sorting position of each word to be input in the set of words to be input, and calling the bidirectional memory network layer to perform hidden layer feature processing on the word embedding vector corresponding to each word to be input according to the sorting position, so as to obtain the forward hidden layer representation vector and the backward hidden layer representation vector corresponding to each entity byte unit; Calling the normalization layer to perform normalization processing on the forward hidden layer representation vector and the backward hidden layer representation vector corresponding to each entity byte unit, so as to obtain the forward prediction distribution and the backward prediction distribution corresponding to each entity byte unit.
4. The method according to claim 3, characterized in that, The sorting position includes a forward sorting position and a backward sorting position; The identifying the sorting position of each word to be input in the set of words to be input, and calling the bidirectional memory network layer to perform hidden layer feature processing on the word embedding vector corresponding to each word to be input according to the sorting position, so as to obtain the forward hidden layer representation vector and the backward hidden layer representation vector corresponding to each entity byte unit, includes: Identifying the forward sorting position of each word to be input in the set of words to be input, and calling the bidirectional memory network layer to perform hidden layer feature processing on the word embedding vector corresponding to each word to be input according to the forward sorting position, so as to obtain the forward hidden layer representation vector corresponding to each entity byte unit; Identifying the backward sorting position of each word to be input in the set of words to be input, and calling the bidirectional memory network layer to perform hidden layer feature processing on the word embedding vector corresponding to each word to be input according to the backward sorting position, so as to obtain the backward hidden layer representation vector corresponding to each entity byte unit.
5. The method according to claim 3, characterized in that, The determining the occurrence probability of each entity byte unit in the text to be detected target according to the word prediction distribution corresponding to each entity byte unit, includes: Generating the average prediction distribution corresponding to each entity byte unit according to the forward prediction distribution and the backward prediction distribution corresponding to each entity byte unit; the several entity byte units include entity byte unit M; Searching for the word that matches the word index corresponding to the entity byte unit M in the average prediction distribution corresponding to the entity byte unit M as the target word, and determining the occurrence probability corresponding to the target word in the average prediction distribution corresponding to the entity byte unit M as the occurrence probability of the entity byte unit M in the text to be detected target.
6. The method according to claim 1, characterized in that The performing semantic feature extraction processing on the text to be detected target to obtain the target global representation vector of the text to be detected target, includes: Obtain the head tag word and the tail tag word, determine the head tag word, the tail tag word, and the several entity byte units as words to be processed, and generate a text word set according to the words to be processed; Obtain the word index corresponding to each word to be processed in the word dictionary, and generate a set of word indexes to be processed; Obtain the word position information of each word to be processed in the text word set respectively, and generate a set of word position information; Obtain the sentence position information of each word to be processed in the text word set respectively, and generate a set of sentence position information; Call the global semantic model, and perform semantic feature extraction processing on the set of word indexes to be processed, the set of word position information, and the set of sentence position information in the global semantic model to obtain the target global representation vector of the text to be detected; 7. The method according to claim 6, characterized in that, The global semantic model includes a second embedding layer and a self-attention network layer; The step of calling the global semantic model and performing semantic feature extraction processing on the set of word indexes to be processed, the set of word position information, and the set of sentence position information in the global semantic model to obtain the target global representation vector of the text to be detected includes: Call the second embedding layer to perform embedding feature processing on the set of word indexes to be processed, the set of word position information, and the set of sentence position information, and obtain the word index embedding vector to be processed, the word position information embedding vector, and the sentence position information embedding vector corresponding to each word to be processed respectively; Generate the fusion embedding vector corresponding to each word to be processed according to the word index embedding vector to be processed, the word position information embedding vector, and the sentence position information embedding vector corresponding to each word to be processed respectively; the fusion embedding vector corresponding to a word to be processed is generated by the word index embedding vector to be processed, the word position information embedding vector, and the sentence position information embedding vector corresponding to this word to be processed; Call the self-attention network layer to perform semantic representation processing on the fusion embedding vector corresponding to each word to be processed respectively, and obtain the hidden layer vector corresponding to each word to be processed respectively; Among the hidden layer vectors corresponding to each word to be processed respectively, determine the hidden layer vector corresponding to the head tag word as the target global representation vector of the text to be detected; 8. The method according to claim 1, wherein The step of determining the normal text vector distance for measuring the normal semantics of the text to be detected according to the target global representation vector and the reference global representation vectors of at least two normal text types includes: Respectively obtain the vector distance between the reference global representation vector of each normal text type and the target global representation vector as the reference distance; Obtain Q reference global representation vectors in order from the at least two reference global representation vectors according to the reference distance; Q is a positive integer, and Q is less than or equal to the number of normal text types; Generate the average vector distance between the target global representation vector and the Q reference global representation vectors, and use the average vector distance as the normal text vector distance for measuring the normal semantics of the text to be detected.
9. The method according to claim 1, wherein The determining whether the text to be detected is normal text or abnormal text according to the perplexity and the normal text vector distance includes: Harmonize the perplexity and the normal text vector distance according to the reference weights corresponding to the perplexity and the normal text vector distance respectively, to obtain the text semantic relevance of the text to be detected; If the text semantic relevance is greater than the standard relevance threshold, determine that the text to be detected is abnormal text; If the text semantic relevance is less than or equal to the standard relevance threshold, determine that the text to be detected is normal text.
10. The method according to claim 2, wherein It further includes: Obtain a first normal sample text, perform entity word segmentation processing on the content of the first normal sample text to obtain a number of first sample entity byte units after entity word segmentation processing; Determine the start tag word, the end tag word, and the number of first sample entity byte units as the words to be input samples, and generate a set of words to be input samples according to the words to be input samples; Obtain the word index corresponding to each word to be input sample in the word dictionary; Identify the sample sorting position of each word to be input sample in the set of words to be input samples, call the local semantic initial model, and predict the word index corresponding to each word to be input sample according to the sample sorting position in the local semantic initial model, to obtain the sample word prediction distribution corresponding to each first sample entity byte unit; Train the local semantic initial model based on the word label distribution corresponding to each first sample entity byte unit and the sample word prediction distribution to obtain the local semantic model.
11. The method according to claim 6, characterized in that, It further includes: Obtain a second normal sample text, perform entity word segmentation processing on the content of the second normal sample text to obtain a number of second sample entity byte units after entity word segmentation processing; Determine the head tag word, the tail tag word, and the number of second sample entity byte units as the first sample words, and generate a set of sample text words according to the first sample words; Obtain N random tag words, and replace N first sample words in the set of sample text words with the N random tag words respectively to obtain a set of sample text words after replacement; N is a positive integer less than the total number of the second sample entity byte units in the set of sample text words; Obtain the tag position information of the N random tag words in the set of sample text words after replacement, and generate a set of tag position information according to the tag position information; Train the global semantic initial model according to the set of sample text words after replacement, the set of tag position information, and the random tag words to generate the global semantic model.
12. A text detection device, characterized in that, It includes: A target text processing module, configured to obtain a target text to be detected, perform entity word segmentation processing on the content of the target text to be detected, and obtain a number of entity byte units after entity word segmentation processing; A perplexity determination module, configured to perform entity prediction processing on the number of entity byte units, obtain the occurrence probability of the number of entity byte units in the target text to be detected, and determine the perplexity for measuring the semantic coherence of the target text to be detected according to the occurrence probability corresponding to the number of entity byte units; A vector acquisition module, configured to perform semantic feature extraction processing on the target text to be detected, and obtain a target global representation vector of the target text to be detected; The vector acquisition module is further configured to obtain global representation vectors corresponding to at least two reference normal texts respectively, perform clustering processing on the global representation vectors corresponding to the at least two reference normal texts respectively, obtain at least two clustering clusters, and respectively use the global representation vectors located at the cluster centers in each clustering cluster as reference global representation vectors of normal text types; each clustering cluster respectively represents a different normal text type; A vector distance determination module, configured to determine a normal text vector distance for measuring the normal semantic property of the target text to be detected according to the target global representation vector and the reference global representation vectors of at least two normal text types; A target text determination module, configured to determine whether the target text to be detected belongs to a normal text or an abnormal text according to the perplexity and the normal text vector distance.
13. A computer device, characterized in that, Comprising: A processor, a memory, and a network interface; The processor is connected to the memory and the network interface, wherein the network interface is used to provide network communication functions, the memory is used to store program codes, and the processor is used to call the program codes to execute the method according to any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is suitable for being loaded and executed by the processor to execute the method according to any one of claims 1-11.
15. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in the computer-readable storage medium, and are suitable for being read and executed by the processor, so that a computer device having the processor executes the method according to any one of claims 1-11.
Citation Information
Patent Citations
Text detection method, device, electronic device, and computer-readable storage medium
CN109271526A
Error correction processing method and device, storage medium and processor
CN110457688A