Text Detection Method and Device
Through the text detection method combined with Solr search engine and local text matching algorithm, the problems of low efficiency and high memory usage of traditional sensitive word matching algorithms are solved, and efficient sensitive word detection and personalized text filtering are achieved.
Patent Information
- Application Number
- CN202411864499.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2044-12-17
AI Technical Summary
The traditional sensitive word matching algorithm is inefficient and has high memory occupancy, so it cannot efficiently process large amounts of sensitive vocabulary and text data.
The Solr search engine is used to generate text detection requests through q parameters, combining it with a query parser and a preset query parser to achieve efficient detection of adjacent sensitive words and long sensitive words, and processing specific types of sensitive text in combination with local text matching algorithm.
It improves the efficiency of sensitive word detection, reduces the memory usage of the user side, supports users to customize sensitive text libraries, and meets personalized needs.
Smart Images

Figure CN119782538B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a text detection method and apparatus. Background Art
[0002] Sensitive word filtering is a technology for implementing content censorship in websites, applications or platforms, and is used to prevent websites and users from publishing content containing inappropriate, illegal or policy-violating information.
[0003] Traditional sensitive word filtering methods include the following steps:
[0004] Step 1: Construction of a sensitive word library. The sensitive word library is the core of sensitive word filtering technology and contains various words that may cause controversy or discomfort, such as swear words, insulting language, politically sensitive words, etc. The construction of the sensitive word library needs to comprehensively consider factors such as regional culture, laws and regulations, and social morality to ensure that it neither overly restricts freedom of speech nor effectively prevents the spread of bad information. The sensitive word library can be manually collected, downloaded from the Internet or obtained by system synchronization.
[0005] Step 2: At a preset filtering time, call a matching algorithm to match the text with the sensitive words in the sensitive word library to obtain a matching result. Among them, the matching algorithms generally include the following:
[0006] 1. Simple matching algorithm: Check each word or phrase in the text one by one to see if it is the same as the sensitive word. Although this matching algorithm is simple, its efficiency is low, and as the text length or the sensitive word library size increases, the matching efficiency further decreases.
[0007] 2. Prefix Tree (Trie tree) algorithm: Also known as a dictionary tree, it is a tree structure used to quickly retrieve strings. When performing sensitive word matching, construct a Trie tree to store the sensitive word library, and then traverse the text and use the Trie tree to quickly find the matching items to obtain the matching result. Since the Trie tree can check multiple prefixes while traversing the text, this method is more efficient than simple matching and can improve the query efficiency.
[0008] However, on the one hand, the Trie tree needs to store a large number of pointers for each node, and in the worst case, the storage space requirement may be equal to the number of characters, which results in the Trie tree requiring a large amount of memory when storing a large number of strings.
[0009] On the other hand, compared with the hash table, the Trie tree has a lower search efficiency. The search time complexity of an effectively constructed hash table (i.e., using a good hash function and a reasonable load factor) for searching is O(1), while the search time complexity of the Trie tree is O(l), where l is the length of the string.
[0010] On the other hand, when dealing with strings without a common prefix, since each string needs to be stored starting from the root node, the Trie tree will consume more memory space. In addition, the implementation complexity of the Trie tree is relatively high, and various boundary conditions and special cases need to be handled.
[0011] 3. Deterministic Finite Automaton (DFA) algorithm: It is a non-recursive automaton that determines the next state based on an event and the current state, i.e., event + state = next state. It is a pattern matching algorithm based on a finite state machine and is commonly used in application scenarios such as sensitive word filtering and string search. This method has an efficient and fixed time complexity.
[0012] However, when the sensitive word library is very large, the performance of the DFA algorithm will significantly decline, and the efficiency of simple matching is relatively low. At the same time, the DFA algorithm lacks the necessary flexibility and can only perform maximum matching. For example, in sensitive word filtering, if two sensitive words "Hello" and "Hello World" are set, the DFA algorithm can only determine that "Hello World" is a sensitive word and cannot determine that "Hello" is a sensitive word.
[0013] 4. Aho-Corasick (AC) automaton algorithm: The AC automaton is an extension of the Trie tree and is used to quickly find multiple pattern strings (i.e., strings) in the input text. It supports searching for multiple sensitive words simultaneously. The AC automaton optimizes the search process by constructing a failure pointer array, enabling it to jump to the search of other sensitive words when searching for a sensitive word.
[0014] However, on the one hand, constructing an AC automaton requires more memory, especially when there are a large number or very long pattern strings. At the same time, the Chinese character set is larger than the English character set, resulting in an increase in the number of nodes in the Trie tree and further increasing memory consumption. On the other hand, the preprocessing process of constructing a finite state automaton is relatively complex and requires additional time and space overhead. Moreover, the process of constructing the Trie tree and the failure pointer takes a certain amount of time, which may lead to a long initialization time in scenarios where the word list often changes.
[0015] 5. Regular expression algorithm: It can match complex patterns, including variations or variants of sensitive words. However, regular expressions may reduce the matching efficiency due to excessive complexity.
[0016] In summary, how to reduce the memory occupied by the algorithm and improve the text detection efficiency during the text detection process is an urgent problem to be solved. Summary of the Invention
[0017] In view of this, the present disclosure provides a text detection method and apparatus, which can solve the problems of low algorithm efficiency and high memory occupation of traditional sensitive word matching algorithms.
[0018] According to one aspect of the present disclosure, a text detection method is provided, and the method includes:
[0019] Obtain a target text to be detected;
[0020] Determine a target sensitive text for this text detection from a pre-configured sensitive text library; wherein, the sensitive text library includes at least one sensitive text;
[0021] When the text type of the target sensitive text is the first type, use the target sensitive text as the q parameter in the Solr search engine, and generate a text detection request based on the text information of the target text and the q parameter; wherein, the q parameter is a query parameter in the Solr search engine;
[0022] Send the text detection request to the Solr server, so that the Solr server detects the target sensitive text indicated by the q parameter in the target text to obtain a detection result;
[0023] Obtain and display the obtained detection result.
[0024] In a possible implementation, the first type includes the adjacent sensitive word type; the target sensitive text of the adjacent sensitive word type includes two sensitive words, as well as distance indication information and order indication information between the two sensitive words;
[0025] Accordingly,
[0026] The sending the text detection request to the Solr server includes:
[0027] Send the text detection request to the Solr server's around query parser, so that the around query parser detects two sensitive words that meet the distance indication information and order indication information in the target text to obtain the detection result.
[0028] In a possible implementation, the first type includes the long sensitive word type; the target sensitive text of the long sensitive word type is a sensitive word with a character number greater than or equal to a preset character number; accordingly,
[0029] The sending the text detection request to the Solr server includes:
[0030] Send the text detection request to a preset query parser of the Solr server, so that the preset query parser tokenizes the target text based on the text detection request, and detects whether each token obtained by tokenization matches the target sensitive text, to obtain the detection result;
[0031] Wherein, the preset query parser is the same as or different from the surrounding query parser.
[0032] In a possible implementation manner, the method further includes:
[0033] When the text type of the target sensitive text is the second type, call a local text matching algorithm to detect the target sensitive text in the target text, to obtain the detection result.
[0034] In a possible implementation manner, the second type includes a sensitive statement type; the target sensitive text of the sensitive statement type is a preset statement; correspondingly,
[0035] When the text type of the target sensitive text is the second type, calling a local text matching algorithm to detect the target sensitive text in the target text, to obtain the detection result, includes:
[0036] For each statement in the target text, use the statement and the target sensitive text as the row information and column information of a matrix respectively;
[0037] Compare the row information of the i-th row and the column information of the j-th column in the matrix to obtain the element value of the i-th row and j-th column; both i and j are positive integers;
[0038] Based on the element values of each element of the matrix, determine the detection result.
[0039] In a possible implementation manner, the determining the detection result based on the element values of each element of the matrix includes:
[0040] Determine the number of elements with an element value of 1 in each set of diagonal elements of the matrix; in the diagonal elements with the largest number of elements, determine the string formed by the row information or column information corresponding to the element value of 1, to obtain the detection result; wherein, when the row information of the i-th row is the same as the column information of the j-th column, the element value of the i-th row and j-th column is 1; when the row information of the i-th row is different from the column information of the j-th column, the element value of the i-th row and j-th column is 0;
[0041] Or,
[0042] Determine the maximum element value of the matrix; among a group of diagonal elements where the maximum element value is located, determine the string formed by the row information or column information of the elements with element values greater than 0 to obtain the detection result; where, when the row information of the i-th row is the same as the column information of the j-th column, the element value of the element in the i-th row and j-th column is the sum of 1 and the element value of the element in the (i - 1)-th row and (j - 1)-th column; when the row information of the i-th row is different from the column information of the j-th column, the element value of the element in the i-th row and j-th column is 0;
[0043] Wherein, a group of diagonal elements refers to the elements on the same diagonal slanting from the upper left to the lower right of the matrix.
[0044] In a possible implementation manner, the method further includes:
[0045] For each statement in the target text, determine whether the statement meets a preset statement filtering condition;
[0046] When the statement meets the statement filtering condition, filter out the statement, and execute the step of determining whether the next statement meets the preset statement filtering condition and subsequent steps;
[0047] When the statement does not meet the statement filtering condition, if the statement matches at least two target sensitive texts, after de-duplicating the at least two target sensitive texts, trigger the execution of the step of obtaining and displaying the detection result;
[0048] Wherein, the statement filtering condition includes at least one of the following:
[0049] The character length of the statement is less than the first character threshold;
[0050] Other characters in the statement except punctuation marks are exactly the same as the target sensitive text;
[0051] The output result of the local text matching algorithm indicates that the character length of the same characters in the statement and the target sensitive text is less than the second character threshold.
[0052] In a possible implementation manner, the second type includes short sensitive word types; the target sensitive text of the short sensitive word type is a sensitive word with a character count less than a preset character count; correspondingly,
[0053] When the text type of the target sensitive text is the second type, calling the local text matching algorithm to detect the target sensitive text in the target text to obtain the detection result includes:
[0054] Call the String.Contains function to detect whether the target text contains the target sensitive text to obtain the detection result.
[0055] In a possible implementation, before determining the target sensitive text used for this text detection from the pre-configured sensitive text library, it further includes:
[0056] Load the sensitive text library configured in the Excel vocabulary list;
[0057] Verify the sensitive text in the sensitive text library;
[0058] After the verification passes, store the loaded sensitive text library in the cache so as to call the target sensitive text stored in the cache during text detection.
[0059] According to another aspect of the present disclosure, there is provided a text detection device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to implement the above method when executing the instructions stored in the memory.
[0060] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0061] According to another aspect of the present disclosure, there is provided a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above method.
[0062] By obtaining the target text to be detected; determining the target sensitive text used for this text detection from the pre-configured sensitive text library; in the case where the text type of the target sensitive text is the first type, taking the target sensitive text as the q parameter in the Solr search engine, and generating a text detection request based on the text information of the target text and the q parameter; sending the text detection request to the Solr server so that the Solr server detects the target sensitive text indicated by the q parameter in the target text to obtain a detection result; obtaining and displaying the obtained detection result; it is possible to solve the problems of low algorithm efficiency and high memory occupancy of the traditional sensitive word matching algorithm. Since the Solr server has the characteristics of efficiently processing and searching a large amount of data, therefore, by querying whether the target text to be detected includes the target sensitive text through the Solr server, the efficiency of sensitive word detection can be improved. At the same time, less memory of the user side will be occupied during the detection process.
[0063] In addition, this embodiment supports a user-defined sensitive text library, thereby meeting the diverse and personalized text detection needs of users.
[0064] Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] The accompanying drawings, which are included in and constitute a part of this specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure together with the specification.
[0066] Figure 1 Schematic diagram showing a text detection system according to an embodiment of the present disclosure;
[0067] Figure 2 Flowchart showing a text detection method according to an embodiment of the present disclosure;
[0068] Figure 3 Schematic diagram showing the classification of sensitive text according to an embodiment of the present disclosure;
[0069] Figure 4 Schematic diagram showing a sensitive text configuration page according to an embodiment of the present disclosure;
[0070] Figure 5 Schematic diagram showing a sensitive text library according to an embodiment of the present disclosure;
[0071] Figure 6 Schematic diagram showing a sensitive text library according to another embodiment of the present disclosure;
[0072] Figure 7 Schematic diagram showing a response code of a verification result according to an embodiment of the present disclosure;
[0073] Figure 8 Schematic diagram showing a response code of a detection result according to an embodiment of the present disclosure;
[0074] Figure 9 Schematic diagram showing a response code of a detection result according to another embodiment of the present disclosure;
[0075] Figure 10 Schematic diagram showing a response code of a detection result according to yet another embodiment of the present disclosure;
[0076] Figure 11 Schematic diagram showing a response code of a detection result according to still another embodiment of the present disclosure;
[0077] Figure 12 Schematic diagram showing a local text matching algorithm according to an embodiment of the present disclosure;
[0078] Figure 13 Schematic diagram showing a local text matching algorithm according to another embodiment of the present disclosure;
[0079] Figure 14 A block diagram showing a text detection device according to an embodiment of the present disclosure;
[0080] Figure 15 A block diagram showing a text detection device according to another embodiment of the present disclosure. Detailed implementation manners
[0081] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Identical reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0082] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments.
[0083] In addition, for a better description of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some of these specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0084] First, several terms related to this application are introduced.
[0085] Solr search application server (The Apache Solr Project, Solr): It is an independent enterprise-level search application server that provides API interfaces similar to Web services externally. The client can submit an XML file in a certain format to the search engine server (i.e., the Solr server hereinafter) through an http request to generate an index; it can also submit a search request through an Http Get operation and obtain a return result in XML format. Solr is an open-source search platform built on the Lucene library, providing powerful full-text search capabilities, having advantages such as efficient handling of large traffic, friendly interfaces and easy to use, scalability and flexibility, high reliability and data recovery capabilities, real-time updates, and support for multiple data formats and rich search functions. It is widely used in enterprise-level search solutions, capable of handling complex queries and supporting multiple data formats and index types.
[0086] Among them, Lucene is a high-performance and scalable search engine library that provides the ability for full-text search. Lucene is mainly used to build search engine applications and can index and search text content. The core of Lucene is the information retrieval technology based on inverted index, which allows for the rapid retrieval of a collection of documents containing specific words.
[0087] With the development of the Internet, information dissemination and communication have become more convenient. However, the information may contain some text that needs to be filtered out. Based on this, it is necessary to detect the text to determine whether there is sensitive text (including sensitive words and / or sensitive sentences) in the text.
[0088] Generally, text detection scenarios include but are not limited to the following:
[0089] 1. Complexity of the network environment and regulatory scenarios:
[0090] With the rapid development of the Internet, websites have become important platforms for information dissemination and communication. To maintain network order and social stability, countries and regions have successively introduced relevant laws and regulations, requiring websites to review and filter content to prevent the spread of sensitive information.
[0091] 2. User security and privacy protection scenarios:
[0092] When users post and view information on websites, they may come into contact with inappropriate content, such as porn, violence, politically sensitive topics, etc. These contents may have a negative impact on the physical and mental health of users. Therefore, it is necessary to filter these inappropriate contents.
[0093] 3. Brand image and reputation management scenarios:
[0094] Improper content on websites may affect the brand image and reputation of enterprises. For example, the spread of sensitive information may cause dissatisfaction and complaints from users, thereby damaging the credibility and brand value of the website.
[0095] By filtering sensitive words, websites can maintain a positive brand image and enhance user trust and loyalty.
[0096] After the text is detected for sensitive words, there are at least the following benefits:
[0097] 1. Comply with laws and regulations: Complying with laws and regulations is a basic requirement for data media operations. By filtering sensitive words, data resources in the Internet scenario can ensure that the published content meets the requirements of laws and regulations and avoid legal liability.
[0098] 2. Protecting user rights and interests: Sensitive word filtering helps protect users' legitimate rights and interests and prevent them from being harmed by bad information. This includes protecting users' legitimate rights such as the right to privacy and reputation, and providing users with a safe and healthy online environment.
[0099] 3. Maintaining the quality of platform content: Sensitive word filtering can improve the quality of content on the data platform, enabling the resources in the Internet data platform to have higher-quality and low-risk content. This helps build a positive and healthy content platform and enhance users' usage experience and satisfaction.
[0100] 4. Preventing network risks: Sensitive word filtering is an effective means of preventing bad behaviors such as cyberbullying, malicious attacks, and rumor spreading. By promptly identifying and filtering sensitive words, websites can timely detect and respond to potential network risks, maintaining network order and social stability.
[0101] 5. Meeting compliance requirements: For the data resources involved in various industries of society, if there is sensitive information related to politically sensitive and regulated expressions, after being processed by sensitive word filtering, the data can meet the corresponding compliance requirements, protecting user privacy and ensuring data security.
[0102] In summary, sensitive word filtering has very important background and necessity for websites. It is not only a requirement for compliance with laws and regulations, but also an important means to protect user rights and interests, maintain the quality of website content, prevent network risks, and meet compliance requirements. Therefore, it is a very important task to perform text detection on the text to filter sensitive words.
[0103] Next, the text detection system involved in this application will be introduced.
[0104] Figure 1 The schematic diagram of the text detection system according to an embodiment of the present disclosure is shown. As Figure 1 shown, the system at least includes: a user terminal 110 and a Solr server 120 communicatively connected to the user terminal 110.
[0105] The user terminal 110 and the Solr server 120 run a text detection platform built based on the Solr search engine to provide text detection services through this text detection platform.
[0106] The user terminal 110 is used to run the front-end module of the text detection platform. The front-end module supports human-computer interaction with users, such as: obtaining user operations, displaying information sent by the Solr server 120, etc. The front-end module is also used for interaction with the Solr server 120.
[0107] Optionally, the front-end module can be a website or an independent application program. This embodiment does not limit the implementation manner of the front-end module.
[0108] Optionally, the client 110 may be an electronic device such as a computer, a tablet computer, a notebook, etc. that has processing capabilities and human-computer interaction capabilities. The implementation manner of the client 110 is not limited in this embodiment.
[0109] The Solr server 120 is used to run the backend module of the text detection platform. The backend module is used to interact with the frontend module (such as interact based on the HTTP protocol), and perform sensitive word matching on the target text to be detected according to the information obtained from the interaction to obtain the detection result.
[0110] Optionally, the Solr server 120 may be a single server, or may also be a server cluster composed of multiple distributed servers. The implementation manner of the Solr server 120 is not limited in this embodiment.
[0111] In this embodiment, the client 110 is used to: obtain the target text to be detected; determine the target sensitive text used for this text detection from the pre-configured sensitive text library; when the text type of the target sensitive text is the first type, use the target sensitive text as the q parameter in the Solr search engine, and generate a text detection request based on the text information of the target text and the q parameter; send the text detection request to the Solr server.
[0112] Optionally, the target text to be detected may be a text stored locally on the client 110; or, it may also be an online text provided by the content server. At this time, the system further includes a content server communicatively connected to the client 110. Optionally, the content server may belong to the same server platform as the Solr server, or may also belong to different server platforms. The implementation manner of the content server is not limited in this embodiment.
[0113] Correspondingly, the Solr server is used to: obtain the text detection request; detect the target sensitive text indicated by the q parameter in the target text to obtain the detection result;
[0114] The client 110 is further used to: obtain and display the obtained detection result.
[0115] Among them, the sensitive text library includes at least one sensitive text; the q parameter is a query parameter in the Solr search engine, and is used to indicate the query task for the Solr server 120.
[0116] Since the Solr server has the characteristics of efficiently processing and searching a large amount of data, therefore, in this embodiment, by sending the target sensitive text from the client to the Solr server, and having the Solr server query whether the target text to be detected includes the target sensitive text, the efficiency of sensitive word detection can be improved. At the same time, less memory of the client will be occupied during the detection process.
[0117] Next, a detailed introduction to the text detection method involved in this application will be given.
[0118] Figure 2 The flowchart of the text detection method according to an embodiment of the present disclosure is shown. In this embodiment, it is described by taking the method being used in the client in the system shown as an example, as Figure 1 shown, the method includes: Figure 2 shown, the method includes:
[0119] Step 201, obtain the target text to be detected.
[0120] The target text to be detected refers to a document, or a statement, or a word that needs to be detected for the existence of sensitive text. In this embodiment, the implementation manner of the target text is not limited. Additionally, the number of target texts can be one, such as: a document; or it can also be at least two, such as: 10 documents, etc. In this embodiment, the number of target texts is not limited.
[0121] Optionally, the ways to obtain the target text to be detected include but are not limited to at least one of the following:
[0122] The first way: display a text upload control on the text detection page; in the case of receiving an upload operation acting on the text upload control, read the target text stored locally indicated by the upload operation.
[0123] The second way: obtain the target text sent by the content server.
[0124] Step 202, determine the target sensitive text used for this text detection from the pre-configured sensitive text library.
[0125] Among them, the sensitive text library includes at least one sensitive text. Sensitive text refers to text that is not expected to be visible to users. Schematically, sensitive text is text that does not conform to laws and regulations, and / or text that does not conform to public order and good customs, etc. Sensitive text includes sensitive statements and / or sensitive words. In this embodiment, the implementation manner of sensitive text is not limited.
[0126] Optionally, the sensitive text library is configured in the client, or it can also be configured in other devices and then sent to the client. In this embodiment, the way for the client to obtain the sensitive text library is not limited.
[0127] Exemplarily, the ways to configure the sensitive text library include: displaying a sensitive text configuration page; receiving a sensitive text configuration operation on the sensitive text configuration page; storing the sensitive text information indicated by the sensitive text configuration operation into a preset Excel word list to obtain the sensitive text library. Among them, the sensitive text information includes at least the content of the sensitive text. In other embodiments, the sensitive text information may further include the classification of the sensitive text, and / or modification suggestions for the sensitive text, etc. This embodiment does not limit the content included in the sensitive text information.
[0128] Optionally, the classification of the sensitive text is determined based on the detection requirements of text detection. For example, the sensitive text is divided into six classifications: politically sensitive, technology sensitive, copyright rights and interests, standard expressions, bad guidance, and proofreading and correction. Different classifications correspond to different Excel word lists. Accordingly, a sensitive text library corresponding to each classification is obtained. Refer to Figure 3 the Excel word lists corresponding to the six classifications shown.
[0129] Optionally, the classification of each sensitive text may further include at least one level of sub-classification under this classification. For example, for the classification "copyright rights and interests", it includes the sub-classification "sensitive word classification", this sub-classification includes a first-level sub-classification, this first-level sub-classification includes a second-level sub-classification, and this second-level sub-classification includes a third-level sub-classification. In actual implementation, the classification of the sensitive text may also be other ways. This embodiment does not limit the classification method of the sensitive text.
[0130] The modification suggestion is used to indicate the modification method of changing the sensitive text to a non-sensitive text. For example, the modification suggestion is to delete the sensitive words that do not conform to laws and regulations, etc.
[0131] For example: Refer to Figure 4 the sensitive text configuration page 401 shown. This sensitive text configuration page 401 includes a text input box 402, a classification input box 403, a modification suggestion input box 404, and a note input box 405 for the sensitive text. After receiving the sensitive text configuration operation on the text input box 402, the sensitive text indicated by this sensitive text configuration operation is obtained; after receiving the sensitive text configuration operation on the classification input box 403, the classification of the sensitive text indicated by this sensitive text configuration operation is obtained; after receiving the sensitive text configuration operation on the modification suggestion input box 404, the modification suggestion indicated by this sensitive text configuration operation is obtained; after receiving the sensitive text configuration operation on the note input box 405, the note of the sensitive text indicated by this sensitive text configuration operation is obtained. Then, in the case of receiving the trigger operation on the "confirm" control, the current sensitive text configuration is successful, and the sensitive text information obtained from the current configuration is obtained.
[0132] The sensitive text information is stored in a preset Excel word list, and the obtained sensitive text library is for reference Figure 5 as shown. According to Figure 5 it can be known that the Excel word list to which the sensitive text information belongs includes serial number 501, retrieval formula 502, Chinese / foreign language 503, sensitive text classification 504, modification suggestions 505, and remarks 506. Among them, the serial number 501 is used to indicate the configuration order of the sensitive text information. The content in the column where the retrieval formula 502 is located is used to store the content of the sensitive text corresponding to each serial number. For example, the content of the sensitive text is "ID number", "within the classified level" NOT "classified level public", etc. The content in the column where the Chinese / foreign language 503 is located is used to store the language type of the sensitive text corresponding to each serial number. For example, the language type of the sensitive content is Chinese or foreign language (it can also be refined to a specific language, such as English, Spanish, etc.). The content in the column where the sensitive text classification 504 is located is used to store the classification of the sensitive text corresponding to each serial number. Figure 5 Taking the classification of sensitive text including word library classification, sensitive word classification, first-level classification, selectable classification, second-level classification, and third-level classification as an example for illustration. In actual implementation, the classification method of sensitive text can also be other methods, and this embodiment does not limit this. The content in the column where the modification suggestions 505 is located is used to store the modification suggestions for the sensitive text corresponding to each serial number. The content in the column where the remarks 506 is located is used to store the content of the remarks for the sensitive text corresponding to each serial number.
[0133] For another example: Refer to Figure 6 the sensitive text library in the case where the sensitive text is a sensitive sentence as shown. At this time, the retrieval formula is a sentence. For example, the sentence is: "We firmly believe that as long as... unite... around, firm our confidence and forge ahead, we will surely be able to win victory."
[0134] Since after step 202, the target sensitive text of the first type needs to be sent to the Solr server for text detection. Based on this, exemplarily, at least the sensitive text of the first type in the sensitive text library is edited according to the Solr syntax. In other implementation manners, each sensitive text in the sensitive text library can also be edited according to the Solr syntax. The retrieval operators specified by the Solr syntax rules at least include:
[0135] 1. ":" is used to specify the field and the search value. For example, querying *.* will match any value in any field. Querying FieldName:* will return all documents with any value in the FieldName field; querying -FieldName:* will return all documents without any value in the FieldName field.
[0136] 2. "?" represents a wildcard for a single arbitrary character. For example, when retrieving "te?t", the retrieval results may include: test, text, etc.
[0137] 3. "*" represents a wildcard for multiple arbitrary characters (but cannot be used at the beginning of the retrieval term). For example, when retrieving "tes*", the retrieval results may include: test, testing, tester, etc.; when retrieving "te*t", the retrieval results may include: test, text, etc.; when retrieving "*est" (*est is not the first word of the retrieval term), the retrieval results may include: pest, test, etc.
[0138] 4. "~" represents fuzzy retrieval. Specifically, the Damerau-Levenshtein Distance algorithm is used, and the edit distance can be specified. The default edit distance is 2 characters. For example, if you need to retrieve characters similar to "roam", you can query "roam~", and the retrieval results may include: roam, roams, foam, foams, etc.; when querying "roam~1", the retrieval results may include: roams, foam, but not foams (the edit distance is 2). "~" can also be used for proximity retrieval. For example, to retrieve "apache" and "jakarta" that are 10 words apart, you can query "apachejakarta"~10.
[0139] 5. "^" is used to control the relevance of the retrieval. For example, when retrieving "jakarta apache" and hoping to make the relevance of the retrieval results better with jakarta, then you can add the "^" symbol after jakarta and add an increment value after this symbol. This increment value represents the minimum number of characters in the retrieval results that are the same as jakarta. For example, "jakarta^4apache" will result in the first string in the retrieval results having at least 4 characters the same as jakarta.
[0140] 6. Boolean operators: "AND" and "||" indicate whether to take the intersection of two query units; "OR" and "&&" indicate whether to take the difference set of two query units; "NOT", "!", and "-" represent negation; "+" represents an existence operator, requiring the term after the "+" symbol to exist in the corresponding field of the document.
[0141] 7. "()" is used to form a subquery.
[0142] 8. "[]" represents a range retrieval including boundaries. For example, to retrieve records in a certain time period including the start and end, you can query date: [200707 TO 200710].
[0143] 9. "{}" represents a range search that does not include the boundaries. For example, to retrieve records within a certain time period without including the start and end, you can query date: {200707 TO 200710}.
[0144] 10. " / " represents an escape operator. For example, when special characters such as "+-&&||!(){}^"~*?:" are preceded by " / ", the special character can be escaped to a meaning readable by the Solr search engine.
[0145] In one example, since the sensitive text library includes sensitive text edited according to Solr syntax, if there are syntax errors in this part of the sensitive text, it will cause the problem that the Solr server fails in text detection. Based on this, before determining the target sensitive text used for this text detection from the pre-configured sensitive text library, it also includes: loading the sensitive text library configured in the Excel vocabulary; validating the sensitive text in the sensitive text library; after the validation passes, storing the loaded sensitive text library in the cache for calling the target sensitive text stored in the cache during text detection.
[0146] In this example, by validating the sensitive text in the sensitive text library to determine whether the sensitive text matches the Solr syntax, the effectiveness of text detection can be guaranteed. Optionally, the sensitive text to be validated can be the first type of sensitive text, or all the sensitive text in the sensitive text library.
[0147] Exemplarily, the method for validating the sensitive text in the sensitive text library includes: when first reading the content of the Excel vocabulary, taking each sensitive text as the q parameter in the Solr search engine to generate a validation request; sending the validation request to the Solr server. Correspondingly, the Solr server is used to determine whether the q parameter carried in the validation request conforms to the Solr syntax to obtain a validation result; returning the validation result to the client; and the client displays the validation result.
[0148] For example, a sensitive text to be validated is: (test; then the q parameter is (test, and the validation request can be represented by the following Uniform Resource Locator (URL):
[0149] https: / / solr.xxx.cn / solr / SingleDection / select?q=(test
[0150] According to "q=(test" in the above URL, it can be seen that the sensitive text "(test" is carried as the q parameter in the validation request.
[0151] Accordingly, after the Solr server verifies the sensitive text indicated by the q parameter, the verification result (or verification response) returned to the client is as follows for reference Figure 7 as shown Figure 7 The meanings of the response codes in each row are as follows:
[0152] "zkConnected": true, indicating that the Solr is normally connected to the ZooKeeper cluster;
[0153] "tatus": 400, indicating the HTTP status code, which usually means that there is an error in the verification request from the client;
[0154] "QTime": 1 indicates that the query time is 1 millisecond;
[0155] "params" is the abbreviation of query parameters, which contains the parameters used when initiating a query, such as q (i.e., the q parameter), paging parameters (start, rows), etc.;
[0156] "q": "(test", indicating that the q parameter is (test;
[0157] "forwardedCount": "1" indicates the number of times the query is forwarded;
[0158] error represents the error message;
[0159] metadata, representing the metadata of the error message, containing information about the error class and the root error class;
[0160] error-class: "org.apache.solr.common.SolrException" indicates the type of exception thrown.
[0161] root-error-class: "org.apache.solr.parser.ParseException" indicates that the root error type is a parsing exception.
[0162] msg: The error message is used to explain the specific reason for the error. org.apache.solr.search.SyntaxError indicates a syntax error and cannot parse the query string "(test", because when parsing, <eof>(End-of-file), indicating that the query string is not properly closed.
[0163] code: Indicates that the HTTP status code is 400, indicating that the client's verification request is incorrect.
[0164] Based on the above verification results, the verification results displayed on the client side can be: Serial number: 1, Retrieval formula: (test, Abnormal reason: There is a syntax error, please check!
[0165] In other embodiments, the display method of the verification results can also be other methods, and this embodiment does not limit the display method of the verification results.
[0166] In another example, the method for verifying sensitive text in a sensitive text library includes: when first reading the content of the Excel word list, verifying the format of the Excel word list to obtain a verification result; displaying the verification result. Verifying the format of the Excel word list includes, but is not limited to: verifying the absence of a header title, verifying that the content of the word list is not empty, verifying that the required items in each row of the word list are not empty, verifying the format of the serial number column, etc. This embodiment does not limit the content of the format verification of the Excel word list.
[0167] Since the sensitive text library is stored in the Excel word list, the client needs to read the sensitive text from the Excel file. In the file reading operation, to reduce input / output (Input / Output, IO) operations, a cache structure is adopted in this embodiment, that is, after the sensitive text in the Excel word list passes the verification, the content of the Excel word list is stored in the cache, so that the file reading efficiency can be improved, the number of disk IOs can be reduced, and thus the retrieval efficiency can be improved.
[0168] Exemplarily, after obtaining the sensitive text library, determining the target sensitive text used for the current text detection from the pre-configured sensitive text library includes: displaying a sensitive text selection control on the text detection page; in the case of receiving a selection operation acting on the sensitive text selection control, determining the sensitive text indicated by the selection operation as the target sensitive text.
[0169] Optionally, the target sensitive text can be at least one. In other embodiments, in the case of not receiving a selection operation acting on the sensitive text selection control, all the sensitive texts in the sensitive text library can also be used as the target sensitive texts.
[0170] Step 203, in the case where the text type of the target sensitive text is the first type, use the target sensitive text as the q parameter in the Solr search engine, and generate a text detection request based on the text information of the target text and the q parameter.
[0171] Among them, the q parameter is a query parameter in the Solr search engine.
[0172] In one example, when the client receives a detection instruction, step 203 is executed. Optionally, the ways to obtain the retrieval instruction include but are not limited to the following:
[0173] 1. Display a start detection control on the text detection page. When a detection operation on the start detection control is received, a detection instruction is generated.
[0174] At this time, the target text obtained through the text detection page can be text-detected to filter out the target sensitive text.
[0175] 2. Generate a detection instruction every preset scanning period. At this time, the client scans the target text in the text database regularly to batch-filter the target sensitive text in each piece of target text.
[0176] 3. When the target text sent online by the content server is received, a detection instruction is generated. At this time, when the client displays the target text, it will automatically detect and filter out the target sensitive text.
[0177] In other embodiments, the way to obtain the retrieval instruction can also be other ways, which are not listed one by one in this embodiment.
[0178] In this embodiment, the text type of the target sensitive text and the classification of the sensitive text are different concepts. The text type is used to indicate the detection method of the target sensitive text. For example: The target sensitive text of the first type is the sensitive text detected through the Solr server.
[0179] Exemplarily, the target sensitive text of the first type includes at least one of the following:
[0180] The first type: The first type includes the adjacent sensitive word type; the target sensitive text of the adjacent sensitive word type includes two sensitive words, as well as the distance indication information and order indication information between the two sensitive words.
[0181] For each sentence in the target text, some sensitive words do not necessarily need to be filtered or highlighted as soon as they appear. Only when multiple sensitive words appear in the same sentence at the same time, does it become necessary to filter or highlight that group of sensitive words. For example: Only when the two words "New World" and "Entertainment City" appear at the same time, "New World" must appear before "Entertainment City", and the number of characters in the middle does not exceed 19 characters, then "New World" and "Entertainment City" need to be highlighted. Currently, the traditional matching algorithms cannot meet the above retrieval requirements.
[0182] Based on this, the present application performs text detection through the Surround Query Parser in the Solr server to achieve sensitive text detection for adjacent sensitive word types. The Surround Query Parser supports the surround query syntax and provides a proximity search function. Specifically, the Surround Query Parser includes two operations: Operation 1, creating an ordered span (distance) through "w"; Operation 2, creating an unordered span through "n". Both of these operations declare the distance between two sensitive words through a numerical value, with the default being one character and the maximum being 99 characters.
[0183] According to the above Solr syntax of the Surround Query Parser, for the target sensitive text of the adjacent sensitive word type, the order of two sensitive words is indicated by "w", and the maximum distance between the two sensitive words is indicated by the numerical value before "w". Alternatively, the lack of order of two sensitive words is indicated by "n", and the maximum distance between the two sensitive words is indicated by the numerical value before "n".
[0184] For example, the target sensitive text of the adjacent sensitive word type is:
[0185] {!surround}3w(foo,bar)
[0186] Among them, 3w means that the sensitive word "foo" needs to appear before the sensitive word "bar", and the distance between the sensitive word "foo" and the sensitive word "bar" does not exceed 3 characters.
[0187] Among them, "w" and "n" can be in uppercase or lowercase, and this embodiment does not limit the editing method of "w" and "n".
[0188] Another example: The target sensitive text of the adjacent sensitive word type is:
[0189] {!surround}19w("New World","Entertainment City")
[0190] It means that "New World" must appear before "Entertainment City", and the interval between the two sensitive words is less than or equal to 19 characters. Only the text that meets the above conditions will be detected.
[0191] Another example: The target sensitive text of the adjacent sensitive word type is:
[0192] {!surround}19n("New World","Entertainment City")
[0193] It means that "New World" and "Entertainment City" appear in the sentence (the order between the two is not limited), and the interval between the two sensitive words is within 19 characters. Only the text that meets the above conditions will be detected.
[0194] Exemplarily, based on the text information of the target text and the q parameter, a text detection request is generated, including: obtaining the text identifier of the target text to get the text information; carrying the text information and the q parameter in the text detection request to obtain the text detection request.
[0195] Optionally, the text identifier includes the article identifier generated by the client for each target text respectively, or the article identifier and the batch number. The article identifiers and batch numbers of different target texts are different, and the article identifiers and batch numbers generated for the same target text during different text detections are different. In other words, for each target text, the text identifier of this target text during this text detection is unique.
[0196] Exemplarily, the text identifier is generated when a retrieval instruction is obtained, and the text identifier and the corresponding target text are uploaded to the Solr server for storage.
[0197] Optionally, the text detection request further includes the field name to be highlighted, the display mode of the retrieval result, etc. For example: the field name to be highlighted is: FullText. Set the highlighting label prefix (hl.simple.pre): <em>, the highlight label suffix (hl.simple.post) is: < / em> .
[0198] For example: the text detection request is represented by the following URL:
[0199] https: / / solr.xxx.cn / SingleDection / select?hl.fl=FullText&hl.simple.post=<%2Fem>&hl.simple.pre= <em>&hl=on&q=%7B!surround%7D%209n(ice%20and%20snow,%20economy)&rows=1;
[0200] Among them, "hl.fl = FullText" means specifying that the FullText field should be highlighted; "hl.simple.post = <%2Fem>" means the highlight label suffix is< / em> ;"hl.simple.pre= <em>” indicates that the highlight tag prefix is <em>;"hl=on" indicates that the highlighting function is enabled; "q=%7B!surround%7D%209n(ice%2C%20snow%2C%20economy)" indicates that the q parameter is {!surround}9n(economy,ice,snow); "rows=1" indicates that the number of search results to be returned is specified as 1.
[0201] In the case where the text identifier is not included in the text detection request, it means that text detection is performed on all target texts indicated by the client in the Solr server.
[0202] Second type: The first type includes the long sensitive word type; the target sensitive text of the long sensitive word type is a sensitive word whose character count is greater than or equal to the preset character count.
[0203] Since the efficiency of the traditional matching algorithm for matching the target sensitive text of the long sensitive word type is low, based on this, in this embodiment, the target sensitive text of the long sensitive word type is sent to the Solr server for text detection. Since the Solr server has the characteristic of high text matching efficiency, the efficiency of text detection can be improved.
[0204] Among them, the preset character count can be 4, or 5, 6, etc. The preset character count can be determined based on the text retrieval requirements, and the value of the preset character count is not limited in this embodiment.
[0205] The generation method of the text detection request corresponding to the target sensitive text of the long sensitive word type is the same as the generation method of the text detection request corresponding to the target sensitive text of the adjacent sensitive word type, and will not be elaborated here in this embodiment.
[0206] For example: Set the highlighting field to FullText, and the highlighting label prefix (hl.simple.pre): <em>, the highlight label suffix (hl.simple.pre hl.simple.post) is: < / em> , the target sensitive text of the long sensitive word type is: ABCDE, and the text identifier (ArticleId) is 684C8BC2-F022-4447-96C0-4C049D729D4B202409200954127955.
[0207] Correspondingly, the text detection request is represented by the following URL:
[0208] https: / / solr.xxx.cn / SingleDection / select?hl.fl=FullText&hl.simple.post=< / em> &hl.simple.pre= <em>&hl=on&q=ABCDE&ArticleId=684C8BC2-F022-4447-96C0-4C049D729D4B202409200954127955。
[0209] Step 204, send a text detection request to the Solr server so that the Solr server detects the target sensitive text indicated by the q parameter in the target text and obtains a detection result.
[0210] Since the function of providing query proximity sensitive word types is centered around the query parser, for the proximity sensitive word type, sending a text detection request to the Solr server includes: sending a text detection request to the query parser around the Solr server for the query parser around the Solr server to detect two sensitive words that meet the distance indication information and order indication information in the target text and obtain a detection result.
[0211] For example: The text detection request corresponding to the above proximity sensitive word type is: https: / / solr.xxx.cn / SingleDection / select?hl.fl=FullText&hl.simple.post=<%2Fem>&hl.simple.pre= <em>&hl=on&q=%7B!surround%7D%209n(ice%2C%20snow%2C%20economy)&rows=1; The detection results returned by the Solr server are for reference Figure 8 as shown Figure 8 The meanings of each line of response code in
[0212] zkConnected:true, indicating that Solr is normally connected to the ZooKeeper cluster;
[0213] status:0, indicating that the query was successful and no error occurred;
[0214] QTime:40, indicating that the query took 40 milliseconds;
[0215] params: contains the parameters used during the query.
[0216] q: represents the q parameter Figure 8 in which, according to the corresponding URL, the q parameter is {!surround}9n(economy,ice,snow);
[0217] hl=on, indicating that the highlighting function is enabled;
[0218] hl.simple.post=< / em> , indicating that the highlight label suffix is< / em> ;
[0219] hl.simple.pre= <em>, indicating that the highlight label prefix is <em>;
[0220] hl.fl = FullText indicates that the specified FullText field should be highlighted;
[0221] rows:1 indicates that the number of returned results is limited to 1;
[0222] forwardedCount:1 indicates that the number of times the query is forwarded is 1;
[0223] _: 1730707648679 is an identifier automatically generated by the Solr server, and this parameter does not need to be configured in this embodiment;
[0224] response represents the response content. Among them, numFound:6 indicates that a total of 6 matching target texts are found; start:0 indicates the first target text; maxScore:4.293788 indicates the score of the target text with the highest score; docs: contains the returned text list, and only one target text is returned here.
[0225] SourceSign:F represents the source flag, which is an identifier used to distinguish the source of the target text;
[0226] DBID:WF_QK represents the database identifier, which is used to identify the database or dataset to which the target text belongs;
[0227] BatchId:1718787613470 represents the batch identifier, which is used to identify the batch of data processing.
[0228] ArticleId:684c8bc2 - f022 - 4447 - 96c0 - 4c049d729d4b202406190900134709 represents the text identifier of the target text.
[0229] version:1799922775019749376 represents the version number of the target text;
[0230] FullText: contains the full text content of the matching target text, (III) Sports Industry Chengdu will transform "cold ice and snow" into "hot economy" "It's getting colder, and the desire to ski is burning. Just take advantage of the weekend to go skiing;
[0231] highlighting indicates the highlighted results;
[0232] 684c8bc2-f022-4447-96c0-4c049d729d4b202406190900134709, is the text identifier of the highlighted target text. FullText: contains the highlighted text fragment, where <em>and< / em> The tag marks the position of the text in the target text that is the same as the target sensitive text.
[0233] According to Figure 8 the red line part, it can be seen that there are 6 characters between the ice and snow and the economy, so it is <em>< / em> tagged, that is, it matches the target sensitive text.
[0234] And if the Figure 8 q parameter in the corresponding URL is modified to {!surround}9w(economy, ice and snow), the detection result returned by the Solr server is as shown in Figure 9 According to Figure 9 it can be known that at this time, no target sensitive text is detected in the target text.
[0235] For the long sensitive word type, sending the text detection request to the Solr server includes: sending the text detection request to the preset query parser of the Solr server, so that the preset query parser tokenizes the target text based on the text detection request and detects whether each token obtained by tokenization matches the target sensitive text to obtain the detection result.
[0236] Optionally, the preset query parser is the same as or different from the surround query parser. Since the surround query parser can also support boolean operations, such as: AND, OR, NOT, or uppercase, lowercase, etc., as well as wildcards, quotes for phrase search, and weighting, therefore, detecting the target sensitive text of the long sensitive word type in the target text can be achieved through the surround query parser.
[0237] In other embodiments, if the preset query parser is different from the surround query parser, for example: the preset query parser is the Standard Query Parser in the Solr server, or the Lucene Query Parser, etc., this embodiment does not limit the implementation manner of the preset query parser. Correspondingly, the Solr server determines the text type of the target sensitive text by parsing the format of the target sensitive text indicated by the q parameter; sends the target sensitive text of this text type to the corresponding query parser. Specifically, it sends the target sensitive text of the adjacent sensitive word type to the surround query parser and sends the target sensitive text of the long sensitive word type to the preset query parser.
[0238] As can be seen from the above, the format of the target sensitive text of the adjacent sensitive word type is different from that of the target sensitive text of the long sensitive word type. Specifically, the front end of the target sensitive text of the adjacent sensitive word type includes the string "{!surround}". Therefore, the Solr server can determine the text type of the target sensitive word indicated by the q parameter according to whether the "{!surround}" is included in the q parameter.
[0239] In this embodiment, the preset query parser includes a tokenizer, which is used to tokenize the target text, so that each token obtained by tokenization can be matched with the target sensitive text respectively.
[0240] For example: If the text detection request is: https: / / solr.xxx.cn / SingleDection / select?hl.fl=FullText&hl.simple.post=< / em> &h l.simple.pre= <em>&hl=on&q=ABCDE
[0241] &ArticleId=684C8BC2-F022-4447-96C0-4C049D729D4B202409200954127955, where "ABCDE" is the target sensitive word, then the detection result returned by the olr server is for reference Figure 10 as shown Figure 10 in the same Figure 8 as the same code fields have the same meaning, which will not be elaborated in this embodiment. According to Figure 10 the underlined part in it, it can be seen that the target sensitive text "ABCDE" is marked by <em>< / em> the
[0242] In this embodiment, the target text and the text information of the target text are pre-transmitted to the Solr server. Among them, the text content of the target text is stored in the FullText field of the Solr server, the article identifier in the text information is stored in the ArticleId field of the Solr server, and the batch number in the text information is stored in the BatchId field of the Solr server. Based on this, the pre-designed index storage structure of the Solr server includes the following index fields: FullText field, ArticleId field, and BatchId field. In other embodiments, the Solr server may also include other index fields, such as: SourceSign field, DBID field, etc. This embodiment does not limit the field types included in the index storage structure in the Solr server.
[0243] Exemplarily, the index storage structure in the Solr server is configured through the core configuration file schema.xml in Solr. schema.xml is used to define important information such as field types, field attributes, default search fields, and copy fields in the index. schema.xml can also define the domains in the index data, specifically including domain names, domain types, whether the domain is indexed, whether it is tokenized, and whether it is stored.
[0244] Optionally, the target text and the text information are obtained by the client and then uploaded to the Solr server after generating the text information of the target text; or, the target text and the text information are uploaded to the Solr server when sending a text detection request. This embodiment does not limit the upload timing of the target text and the text information.
[0245] Reference Figure 11 The response content of the search results returned by the Solr server shown. According to this response content, the index storage structure in the Solr server includes a SourceSign field, a DBID field, an ArticleId field, a BatchId field, and a FullText field.
[0246] Step 205, obtain and display the obtained detection result.
[0247] Exemplarily, in the case of obtaining the search results returned by the Solr server, the target sensitive text marked by the highlighting tags in the detection result is highlighted. For example: each highlighting tag includes a highlighting tag prefix and a highlighting tag suffix, and the characters between the highlighting tag prefix and the highlighting tag suffix are used as the target sensitive text to be highlighted.
[0248] In other embodiments, the way to display the detection result can also be other ways. For example: the field where the target sensitive text is located is highlighted in a first color, and the target sensitive text is highlighted in a second color in this field, and the first color is different from the second color; the display method of the detection result is not limited in this embodiment.
[0249] In summary, the text detection method provided in this embodiment, by obtaining the target text to be detected; determining the target sensitive text used for this text detection from the pre-configured sensitive text library; in the case where the text type of the target sensitive text is the first type, using the target sensitive text as the q parameter in the Solr search engine, and generating a text detection request based on the text information of the target text and the q parameter; sending the text detection request to the Solr server so that the Solr server detects the target sensitive text indicated by the q parameter in the target text to obtain a detection result; obtaining and displaying the obtained detection result; can solve the problems of low algorithm efficiency and high memory occupancy of the traditional sensitive word matching algorithm. Since the Solr server has the characteristics of efficiently processing and searching a large amount of data, therefore, by querying whether the target text to be detected includes the target sensitive text through the Solr server, the efficiency of sensitive word detection can be improved. At the same time, less memory of the user side will be occupied during the detection process.
[0250] In addition, this embodiment supports a user-defined sensitive text library, so as to meet the diverse and personalized text detection needs of users.
[0251] In a possible implementation manner, the pre-configured sensitive text library further includes sensitive text of the second type, and the sensitive text of the second type refers to sensitive text that does not need to be detected through the Solr server. At this time, after step 202, it further includes:
[0252] When the text type of the target sensitive text is the second type, call the local text matching algorithm to detect the target sensitive text in the target text to obtain a detection result.
[0253] Exemplarily, the second type includes a sensitive statement type; the target sensitive text of this sensitive statement type is a preset statement. For example: Figure 6 At least one sensitive statement in the shown sensitive text library.
[0254] Correspondingly, when the text type of the target sensitive text is the second type, call the local text matching algorithm to detect the target sensitive text in the target text to obtain a detection result, including: for each statement in the target text, use the statement and the target sensitive text as the row information and column information of the matrix respectively; compare the row information of the i-th row and the column information of the j-th column in the matrix to obtain the element value of the i-th row and j-th column; determine the detection result based on the respective element values of the matrix. Wherein, both i and j are positive integers.
[0255] In one example, when the row information of the i-th row is the same as the column information of the j-th column, the element value of the i-th row and j-th column is 1; when the row information of the i-th row is different from the column information of the j-th column, the element value of the i-th row and j-th column is 0. At this time, determining the detection result based on the respective element values of the matrix includes: determining the number of elements with a value of 1 in each set of diagonal elements of the matrix; in the diagonal elements with the largest number of elements, determining the string formed by the row information or column information corresponding to the element value of 1 to obtain the detection result.
[0256] Wherein, a set of diagonal elements refers to the elements on the same diagonal slanting from the upper left to the lower right of the matrix.
[0257] For example: Refer to Figure 12 The schematic diagram of the shown local text matching algorithm, it can be seen Figure 12 that the diagonal elements shown by the dotted line include 4 element values of "1", indicating that the statement of the target text and the target sensitive text include the same string "CDEF", and the detection result is obtained.
[0258] In another example, when the row information of the i-th row is the same as the column information of the j-th column, the element value of the i-th row and j-th column is the sum of the element value of 1 and the element value of the (i - 1)-th row and (j - 1)-th column; when the row information of the i-th row is different from the column information of the j-th column, the element value of the i-th row and j-th column is 0. At this time, determining the detection result based on the respective element values of the matrix includes: determining the maximum element value of the matrix; in a set of diagonal elements where the maximum element value is located, determining the string formed by the row information or column information with an element value greater than 0 to obtain the detection result.
[0259] For example: Refer to Figure 13 Another schematic diagram of the shown local text matching algorithm, it can be seen Figure 13 It can be seen that the diagonal element shown by the dotted line includes the maximum element value "4" in the matrix. At this time, the string formed by the row information greater than 0 in this diagonal element is "CDEF", and the detection result is obtained.
[0260] Since the target text includes some statements that do not need to be matched with the target sensitive text of the sensitive statement type, such as: statements that are exactly the same as the target sensitive text of the sensitive statement type, or statements with too short length, etc. At this time, if these statements that do not need to be matched are also detected and output, on the one hand, it will waste computing resources, and on the other hand, it will affect the accuracy of the detection result.
[0261] Based on the above technical problems, optionally, this embodiment further includes: for each statement in the target text, determining whether the statement meets a preset statement filtering condition; in the case where the statement meets the statement filtering condition, filtering out the statement, and performing the steps of determining whether the next statement meets the preset statement filtering condition and subsequent steps on the next statement.
[0262] Among them, the statement filtering condition includes at least one of the following:
[0263] 1. The character length of the statement is less than the first character threshold.
[0264] Optionally, the first character threshold can be a fixed value, or a value determined based on the length of the target sensitive text of the sensitive statement type. For example: the first character threshold is 60% or 70% of the length of the target sensitive text of the sensitive statement type, etc. This embodiment does not limit the setting method of the first character threshold.
[0265] 2. Other characters in the statement except punctuation marks are exactly the same as the target sensitive text.
[0266] At this time, all punctuation marks in the statement and punctuation marks in the target sensitive text are removed for matching to determine that other characters in the statement except punctuation marks are exactly the same as the target sensitive text.
[0267] 3. The output result of the local text matching algorithm indicates that the length of the characters in the statement that are the same as the target sensitive text is less than the second character threshold.
[0268] Optionally, the second character threshold can be a fixed value, or a value determined based on the length of the target sensitive text of the sensitive statement type. For example: the second character threshold is 60% or 70% of the length of the target sensitive text of the sensitive statement type, etc. This embodiment does not limit the setting method of the second character threshold.
[0269] Optionally, when the statement does not meet the statement filtering condition, if the statement matches at least two target sensitive texts, after de-duplicating the at least two target sensitive texts, the step of obtaining and displaying the detection result is triggered to be executed.
[0270] Exemplarily, de-duplicating the at least two target sensitive texts includes: determining the target sensitive text with the longest character length identical to the statement among the at least two target sensitive texts, and determining the characters identical to the statement in the determined one target sensitive text as the retrieval result.
[0271] The following gives an example to illustrate the detection process of the target sensitive text of the sensitive statement type. After step 202, the text detection method of the present application further includes the following steps:
[0272] Step 1, when the text type of the target sensitive text is the sensitive statement type, read the target sensitive text of the sensitive statement type from the cache and start at least one thread;
[0273] Wherein, each thread is used to perform text detection based on at least one target sensitive text of the sensitive statement type to improve the text detection efficiency.
[0274] Step 2, split the target text into statements;
[0275] Step 3, for each statement obtained by splitting, determine whether the length of the statement is less than 60% of the target sensitive text of the sensitive statement type; if so, filter out the statement; if not, execute Step 4;
[0276] Step 4, compare the statement after removing punctuation marks with the target sensitive text of the sensitive statement type after removing punctuation marks; if the two are the same, filter out the statement; if the two are different, execute Step 5;
[0277] Optionally, the filtering order between Step 3 and Step 4 can be swapped, and the present embodiment does not limit the execution order between Step 3 and Step 4.
[0278] Step 5, use the local text matching algorithm corresponding to the sensitive statement type to determine the identical string between the statement and the target sensitive text of the sensitive statement type to obtain the output result; determine whether the length of the identical characters indicated by the output result is less than the second character threshold; if so, filter out the statement; if not, execute Step 6;
[0279] Step 6, mark the output result obtained by the local text matching algorithm with the highlight label prefix and the highlight label suffix;
[0280] Step 7, when the characters in the statement are the same as the characters in at least two target sensitive texts (i.e., multiple target sensitive texts are hit), perform deduplication on the at least two target sensitive texts to obtain a detection result, and execute step 205.
[0281] In this embodiment, when detecting sensitive statements in the target text, the statements in the target text are filtered. On the one hand, it can save the computing resources consumed by running the local text matching algorithm corresponding to the sensitive statement type. On the other hand, it can improve the accuracy of the detection result.
[0282] In a possible implementation manner, the second type includes short sensitive word types; the target sensitive text of the short sensitive word type is a sensitive word with a character count less than a preset character count.
[0283] Correspondingly, when the text type of the target sensitive text is the second type, call the local text matching algorithm to detect the target sensitive text in the target text, and obtain a detection result, including: call the String.Contains function to detect whether the target text contains the target sensitive text, and obtain a detection result.
[0284] Exemplarily, when the target text contains the target sensitive text, obtain the characters in the target text that are the same as the target sensitive text of the short sensitive word type, and obtain a detection result.
[0285] At this time, the characters in the target text that are the same as the target sensitive text can be marked with a highlight label prefix and a highlight label suffix to display the detection result.
[0286] In this embodiment, for the target sensitive text of the short sensitive word type with a small number of characters, text detection is performed locally. On the one hand, it can ensure the detection efficiency. On the other hand, it can also save transmission resources.
[0287] Figure 14 The block diagram of a text detection device according to an embodiment of the present disclosure is shown. This embodiment takes the method being used in Figure 1 the user side in the system shown as an example for illustration, as Figure 14 shown, the method includes: a target text acquisition module 1410, a sensitive text configuration module 1420, a detection request generation module 1430, a detection request sending module 1440, and a detection result acquisition module 1450.
[0288] The target text acquisition module 1410 is used to acquire the target text to be detected;
[0289] The sensitive text configuration module 1420 is used to determine the target sensitive text used for this text detection from a pre-configured sensitive text library; wherein, the sensitive text library includes at least one sensitive text;
[0290] A detection request generation module 1430, configured to, when the text type of the target sensitive text is the first type, use the target sensitive text as the q parameter in the Solr search engine, and generate a text detection request based on the text information of the target text and the q parameter; wherein, the q parameter is a query parameter in the Solr search engine;
[0291] A detection request sending module 1440, which sends the text detection request to a Solr server, so that the Solr server detects the target sensitive text indicated by the q parameter in the target text to obtain a detection result;
[0292] A detection result acquisition module 1450, configured to acquire and display the obtained detection result.
[0293] For relevant details, refer to the above method embodiments.
[0294] In some embodiments, the functions or modules included in the device provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be elaborated here.
[0295] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0296] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to implement the above methods when executing the instructions stored in the memory.
[0297] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above methods.
[0298] Figure 15 It is a block diagram of a text detection device 800 shown according to an exemplary embodiment. For example, the device 800 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0299] Refer to Figure 15 , the device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output interface 812 (I / O interface), a sensor component 814, and a communication component 816.
[0300] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0301] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0302] The power component 806 provides power to the various components of the device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.
[0303] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0304] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0305] The input / output interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0306] The sensor component 814 includes one or more sensors for providing status assessments of various aspects of the device 800. For example, the sensor component 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and keypad of the device 800. The sensor component 814 can also detect a change in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor component 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0307] The communication component 816 is configured to facilitate communication, either wired or wirelessly, between the device 800 and other devices. The device 800 may access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0308] In an exemplary embodiment, the device 800 may be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.
[0309] In an exemplary embodiment, a non-transitory computer-readable storage medium is also provided, such as a memory 804 including computer program instructions, which can be executed by the processor 820 of the device 800 to complete the above-described method.
[0310] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the disclosed embodiments.< / em> < / em> < / em> < / eof>
Claims
1. A text detection method, characterized in that, The method includes: Obtaining a target text to be detected; Determining a target sensitive text for this text detection from a pre-configured sensitive text library; wherein, the sensitive text library includes at least one sensitive text; When the text type of the target sensitive text is the first type, using the target sensitive text as the q parameter in the Solr search engine, and generating a text detection request based on the text information of the target text and the q parameter; wherein, the q parameter is a query parameter in the Solr search engine; Sending the text detection request to the Solr server, so that the Solr server detects the target sensitive text indicated by the q parameter in the target text to obtain a detection result; Obtaining and displaying the obtained detection result; The first type includes the adjacent sensitive word type; the target sensitive text of the adjacent sensitive word type includes two sensitive words, as well as distance indication information and order indication information between the two sensitive words; Accordingly, The sending the text detection request to the Solr server includes: Sending the text detection request to the around query parser of the Solr server, so that the around query parser detects two sensitive words that meet the distance indication information and order indication information in the target text to obtain the detection result.
2. The method according to claim 1, wherein The first type includes the long sensitive word type; the target sensitive text of the long sensitive word type is a sensitive word with a character count greater than or equal to a preset character count; accordingly, The sending the text detection request to the Solr server includes: Sending the text detection request to the preset query parser of the Solr server, so that the preset query parser tokenizes the target text based on the text detection request and detects whether each tokenized word matches the target sensitive text to obtain the detection result; Wherein, the preset query parser is the same as or different from the around query parser.
3. The method according to claim 1, wherein The method further includes: When the text type of the target sensitive text is the second type, calling a local text matching algorithm to detect the target sensitive text in the target text to obtain the detection result.
4. The method according to claim 3, wherein The second type includes the sensitive statement type; the target sensitive text of the sensitive statement type is a preset statement; accordingly, When the text type of the target sensitive text is the second type, calling a local text matching algorithm to detect the target sensitive text in the target text to obtain the detection result, includes: For each statement in the target text, using the statement and the target sensitive text as the row information and column information of a matrix respectively; Comparing the row information of the i-th row and the column information of the j-th column in the matrix to obtain the element value of the i-th row and j-th column; both i and j are positive integers; Determining the detection result based on the respective element values of the matrix.
5. The method according to claim 4, wherein The determining the detection result based on the respective element values of the matrix includes: Determine the number of elements with a value of 1 in each set of diagonal elements of the matrix; in the diagonal elements with the largest number of elements, determine the string formed by the row information or column information corresponding to the element value of 1 to obtain the detection result; wherein, when the row information of the i-th row is the same as the column information of the j-th column, the element value of the element in the i-th row and j-th column is 1; when the row information of the i-th row is different from the column information of the j-th column, the element value of the element in the i-th row and j-th column is 0; Or, Determine the maximum element value of the matrix; in a set of diagonal elements where the maximum element value is located, determine the string formed by the row information or column information of the elements with an element value greater than 0 to obtain the detection result; wherein, when the row information of the i-th row is the same as the column information of the j-th column, the element value of the element in the i-th row and j-th column is the sum of the element value of 1 and the element value of the element in the (i - 1)-th row and (j - 1)-th column; when the row information of the i-th row is different from the column information of the j-th column, the element value of the element in the i-th row and j-th column is 0; Wherein, a set of diagonal elements refers to the elements on the same diagonal slanting from the upper left to the lower right of the matrix.
6. The method according to claim 4, characterized in that The method further includes: For each statement in the target text, determine whether the statement meets a preset statement filtering condition; When the statement meets the statement filtering condition, filter out the statement, and execute the step of determining whether the statement meets the preset statement filtering condition and subsequent steps for the next statement; When the statement does not meet the statement filtering condition, if the statement matches at least two target sensitive texts, after de-duplicating the at least two target sensitive texts, trigger the execution of the step of obtaining and displaying the detection result; Wherein, the statement filtering condition includes at least one of the following: The character length of the statement is less than the first character threshold; Other characters in the statement except punctuation marks are exactly the same as the target sensitive text; The output result of the local text matching algorithm indicates that the length of the characters in the statement that are the same as the target sensitive text is less than the second character threshold.
7. The method according to claim 3, wherein The second type includes short sensitive word types; the target sensitive text of the short sensitive word type is a sensitive word with a character count less than a preset character count; correspondingly, When the text type of the target sensitive text is of the second type, calling a local text matching algorithm to detect the target sensitive text in the target text to obtain the detection result includes: Calling the String.Contains function to detect whether the target text contains the target sensitive text to obtain the detection result.
8. The method according to any one of claims 1 to 7, characterized in that Before determining the target sensitive text used in this text detection from a pre-configured sensitive text library, it further includes: Loading the sensitive text library configured in the Excel word list; Verifying the sensitive texts in the sensitive text library; After passing the verification, storing the loaded sensitive text library in the cache to call the target sensitive text stored in the cache during text detection.
9. A text detection device, characterized in that, Includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to implement the method according to any one of claims 1 to 8 when executing the instructions stored in the memory.
Citation Information
Patent Citations
Network crawler test method and device, server and storage medium
CN107766237A
Case recognition method based on question and answer template
CN110196897A