Sensitive word detection method, storage medium and program product
By using the method of segmenting words on the detected text and building a chain graph of sensitive words, the problem of inefficient detection of text sensitive words in the prior art is solved, and a fast and efficient detection effect is achieved.
Patent Information
- Application Number
- CN202510184735.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is inefficient when detecting sensitive words on large-scale text data and cannot meet application scenarios with high real-time requirements.
By segmenting words on the text to be detected, a chain diagram of sensitive words is constructed, and the overlapping degree of the word segmentation to be detected and the chain diagram of sensitive words is detected to generate sensitive word detection results.
It realizes rapid and efficient detection of sensitive words in text, improves detection efficiency, and can meet application scenarios with high real-time requirements.
Smart Images

Figure CN120144745A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of text processing, and particularly relates to a method for detecting sensitive words, a storage medium, and a program product. Background Art
[0002] With the development of technology, it has become increasingly easier for users to interact through the network, and now it can support more and more users to achieve group interaction. In order to maintain civilized speech in network interaction, therefore, when users speak through the network, sensitive word detection is usually performed. If the detection fails, the speech cannot be posted on the network, and only after passing the check can it be posted on the network, so as to ensure civilized speech on the network.
[0003] However, in the current sensitive word filtering technology, simple string matching algorithms (such as brute-force matching algorithms) are usually used to filter sensitive words. When dealing with large-scale text data and complex sensitive word libraries, its efficiency is significantly reduced. For example, when detecting sensitive words in a text containing a large number of characters, the brute-force matching algorithm needs to compare each sensitive word with every possible position in the text one by one, resulting in an exponential increase in time complexity, seriously affecting the filtering efficiency and unable to meet application scenarios with high real-time requirements.
[0004] Therefore, how to quickly and efficiently detect sensitive words in text is a technical problem that urgently needs to be solved currently. Summary of the Invention
[0005] The purpose of this application is to quickly and efficiently filter sensitive words in text.
[0006] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.
[0007] According to one aspect of the embodiments of this application, a method for detecting sensitive words is provided, including:
[0008] Performing word segmentation on the obtained text to be detected to obtain at least one word segment to be detected;
[0009] Gradually determining the corresponding target sensitive word chain graph according to the characters in the word segment to be detected;
[0010] Determining the sensitive words in the word segment to be detected according to the coincidence degree between the word segment to be detected and the target sensitive word chain graph;
[0011] Generating a sensitive word detection result of the text to be detected according to the sensitive words in the word segment to be detected.
[0012] According to one aspect of the embodiments of the present application, before tokenizing the obtained text to be detected to obtain at least one token to be detected, the method further includes:
[0013] Construct sample nodes corresponding to the characters in the sensitive word sample, and link the sample nodes to obtain a sensitive word chain graph corresponding to each sensitive word sample. The sequential connection order of the sample nodes is consistent with the sequential order of the characters in the sensitive word sample.
[0014] Establish index data corresponding to the sample nodes according to the sample nodes. The index data is used to index the sensitive word chain graph having the sample nodes.
[0015] According to one aspect of the embodiments of the present application, the characters in the sample nodes can exist in the form of pinyin.
[0016] According to one aspect of the embodiments of the present application, the method further includes:
[0017] Obtain similar words for the sensitive word sample;
[0018] Construct a sensitive word chain graph of the similar words.
[0019] According to one aspect of the embodiments of the present application, after constructing the sensitive word chain graph, the method further includes:
[0020] Merge and store the sensitive word chain graphs having the same sample nodes according to each sample node of the sensitive word chain graphs.
[0021] According to one aspect of the embodiments of the present application, tokenizing the obtained text to be detected to obtain at least one token to be detected includes:
[0022] Tokenize the obtained text to be detected to obtain a plurality of initial tokens to be detected;
[0023] Use the initial tokens to be detected that meet the set conditions as tokens that do not need to be detected, and use the initial tokens to be detected other than the tokens that do not need to be detected as tokens to be detected.
[0024] According to one aspect of the embodiments of the present application, gradually determining the corresponding target sensitive word chain graph according to the characters in the token to be detected includes:
[0025] Obtain at least one detection node from the characters in the token to be detected, and determine a target sample node corresponding to the detection node in the sample nodes;
[0026] Based on the index data corresponding to the target sample nodes, determine the target sensitive word chain graph corresponding to the token to be detected in any order according to the index data corresponding to each of the target sample nodes.
[0027] According to one aspect of the embodiments of the present application, gradually determining the corresponding target sensitive word chain graph according to the characters in the token to be detected further includes:
[0028] Obtain at least one detection node from the characters in the token to be detected, and determine the arrangement order of each of the detection nodes according to the arrangement order of the characters;
[0029] Determine the target sample node corresponding to the detection node;
[0030] According to the arrangement order of the detection nodes, sequentially determine the index data corresponding to the target sample nodes until the index data corresponding to the last target sample node is obtained, so as to determine the target sensitive word chain graph corresponding to the token to be detected according to the index data corresponding to the last detection node.
[0031] According to one aspect of the embodiments of the present application, there is provided a readable storage medium, on which a readable program / instruction is stored, and when the readable program / instruction is executed by a processor, the method described in any one of the above is implemented.
[0032] According to one aspect of the embodiments of the present application, there is provided a program product, including a readable program / instruction, and when the readable program / instruction is executed by a processor, the method described in any one of the above is implemented.
[0033] In the present application, by tokenizing the obtained text to be detected, at least one token to be detected is obtained. Then, the corresponding sensitive word chain graph is gradually determined according to the characters in the token to be detected. Then, according to the coincidence degree between the token to be detected and the sensitive word chain graph, the sensitive word in the token to be detected is determined. Finally, the detection result of the text to be detected is generated according to the sensitive word in the token to be detected. In the embodiments of the present application, by tokenizing the text to be detected to obtain the token to be detected, and gradually matching the pre-constructed sensitive word chain graph with the token to be detected to determine whether the token to be detected is a sensitive word. By determining multiple tokens to be detected, multiple tokens to be detected can be detected simultaneously, increasing the efficiency of the text to be detected. In addition, gradually determining the sensitive word chain graph can gradually reduce the eligible sensitive word chain graphs, and thus it is not necessary to screen all the sensitive word chain graphs each time, so as to speed up the detection of the token to be detected. Therefore, the embodiments of the present application can quickly and efficiently detect sensitive words in the text.
[0034] Other features and advantages of the present application will become apparent from the following detailed description, or will be learned in part through the practice of the present application.
[0035] It should be understood that the foregoing general description and the following detailed description are merely exemplary and explanatory, and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0037] Figure 1 A flowchart showing a method for detecting sensitive words according to an embodiment of the present application is shown.
[0038] Figure 2 A flowchart showing tokenizing the obtained text to be detected to obtain at least one token to be detected according to an embodiment of the present application is shown.
[0039] Figure 3 A flowchart showing constructing a sensitive word chain graph according to an embodiment of the present application is shown.
[0040] Figure 4 A flowchart showing constructing a sensitive word chain graph of similar words of a sensitive word sample according to an embodiment of the present application is shown.
[0041] Figure 5 A flowchart showing gradually determining a corresponding target sensitive word chain graph according to characters in the token to be detected according to an embodiment of the present application is shown.
[0042] Figure 6 A flowchart showing gradually determining a corresponding target sensitive word chain graph according to characters in the token to be detected according to another embodiment of the present application is shown.
[0043] Figure 7 A block diagram of a computer system for implementing a method for detecting sensitive words according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0044] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0045] In addition, the described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be employed. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0046] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0047] The flowcharts shown in the drawings are only illustrative and not necessarily include all the content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined, so the actual execution order may change according to the actual situation.
[0048] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0049] Please refer to Figure 1 , Figure 1 , which shows a flowchart of a method for detecting sensitive words according to an embodiment of the present application. The embodiments of the present application provide the steps of a method for detecting sensitive words, including:
[0050] Step S110, performing word segmentation on the obtained text to be detected to obtain at least one word to be detected;
[0051] Step S120, gradually determining the corresponding target sensitive word chain graph according to the characters in the word to be detected;
[0052] Step S130, determining the sensitive words in the word to be detected according to the coincidence degree between the word to be detected and the target sensitive word chain graph.
[0053] Step S140: Generate a sensitive word detection result for the text to be detected based on the sensitive words in the segmented words to be detected.
[0054] The above four steps will be described in detail below.
[0055] In step S110, the text to be detected is segmented, and the text to be detected is divided into at least one segmented word to be detected. In this way, the smallest unit for detecting the text to be detected is determined. After obtaining the segmented words to be detected, only the segmented words to be detected need to be detected.
[0056] In some embodiments, for the same segmented word to be detected that appears multiple times, a segmented word to be detected is randomly selected for detection, and the detection result of this segmented word to be detected is used as the detection result of all the same segmented words to be detected. In this way, repeated detection of the same segmented word to be detected can be avoided to save computing resources.
[0057] In some embodiments, after obtaining the text to be detected, the format of the detection text is converted. In some embodiments, the detection text is converted to the UTF-8 encoding format to improve the segmentation efficiency and detection efficiency of the text to be detected.
[0058] Please refer to Figure 2 , Figure 2 FIG. shows a flowchart of segmenting the obtained text to be detected into at least one segmented word to be detected according to an embodiment of the present application. The embodiment of the present application provides step S110 for segmenting the obtained text to be detected into at least one segmented word to be detected, including:
[0059] Step S111: Segment the obtained text to be detected to obtain a plurality of initial segmented words to be detected;
[0060] Step S112: Use the initial segmented words to be detected that meet the set conditions as the segmented words that do not need to be detected, and use the initial segmented words to be detected other than the segmented words that do not need to be detected as the segmented words to be detected.
[0061] The above two steps will be described in detail below.
[0062] In step S111, the obtained text to be detected is segmented to obtain a plurality of initial segmented words to be detected. Among them, the segmentation method includes, but is not limited to, a segmentation algorithm based on a statistical model or a segmentation algorithm combining dictionary matching and rules.
[0063] In step S112, it should be made clear in advance that there are some characters or words in the language system used in daily life that cannot form sensitive words. Exemplarily, modal particles such as "de", "ne", "na" cannot form sensitive words. Collect and store some characters or words that cannot form sensitive words to obtain a non-inspection library. If the initial token to be detected is the same as the characters or words in the non-inspection library, it is considered that the initial token to be detected meets the set conditions, and then the initial token to be detected is used as a token that does not need to be detected. Finally, among all the initial tokens to be detected, except for the tokens that do not need to be detected, the remaining initial tokens to be detected are used as tokens to be detected.
[0064] In this embodiment, by screening the initial tokens to be detected, the number of tokens to be detected is reduced as much as possible, avoiding waste of computing resources and improving the detection efficiency of the text to be detected.
[0065] In step S120, it should be made clear in advance that one sample node or multiple sample nodes are linked in sequence to form a sensitive word chain graph. That is, the sensitive word chain graph includes at least one sample node. If there are multiple sample nodes in the sensitive word chain graph, these sample nodes are connected in sequence to represent the order between the sample nodes in the sensitive word chain graph. It should be further clarified that each sample node corresponds to at least one character.
[0066] According to the order between the characters in the token to be detected, sequentially find the target sample nodes corresponding to the characters in the token to be detected, and then determine the sensitive word chain graph matching the token to be detected according to the target sample nodes. For example, obtain the sensitive word chain graph including the target sample nodes as the sensitive word chain graph matching the token to be detected.
[0067] Exemplarily, if there is one character in the token to be detected, directly find the target sample node corresponding to the character, and use the sensitive word chain graph with the target sample node as the corresponding sensitive word chain graph. Exemplarily, find the sensitive word chain graph with the target sample node in all the sensitive word chain graphs as the sensitive word chain graph corresponding to the token to be detected.
[0068] Exemplarily, if there are multiple characters in the segment to be detected, the characters in the segment to be detected are divided into at least one character group, and each character group includes at least one character. According to the order of the characters in the segment to be detected, the order between the character groups is determined. Exemplarily, in the order of character arrangement, the character groups determined successively are the first character group, the second character group, and so on. According to the order between the character groups, the target sample nodes corresponding to the character groups are determined successively. Exemplarily, first, the target sample node corresponding to the first character group is determined among multiple sample nodes. According to the target sample node corresponding to the first character group, the sensitive word chain graph in which the first sample node is the sensitive word chain graph of the first target sample node is determined among all the sensitive word chain graphs. Secondly, the second target sample node corresponding to the second character group is determined. According to the target sample node corresponding to the second character group, in the sensitive word chain graph determined according to the first target sample node, the sensitive word chain graph in which the second sample node is the sensitive word chain graph of the second target sample node is determined. The above process is executed for subsequent character groups. In this process, the number of sensitive word chain graphs will become smaller and smaller, and finally, the sensitive word chain graph that meets the target sample nodes corresponding to all the character groups is determined as the sensitive word chain graph corresponding to the segment to be detected.
[0069] If the sensitive word chain graph of the segment to be detected cannot be determined, it is directly considered that the segment to be detected is not a sensitive word. For example, when determining the target sample node corresponding to the character group or character in the segment to be detected, if the target sample node cannot be determined among all the sample nodes, it means that this character group cannot form a sensitive word. Therefore, it can be directly considered that the segment to be detected is not a sensitive word. Another example is that after determining the target sample node corresponding to the character group or character in the segment to be detected, if the sensitive word chain graph corresponding to the current target sample node cannot be determined in the sensitive word chain graph determined according to the previous target sample node, that is, the current target sample node cannot form a sensitive word with the previous target sample node at the current position. Therefore, it can be directly considered that the segment to be detected is not a sensitive word.
[0070] In some embodiments, it is not necessary to perform the detection in the order of the character groups. Exemplarily, the sensitive word chain graphs corresponding to each character group can be determined separately without considering the position of the target node in the sensitive word chain graph. As long as there is a target sample node corresponding to a character group in the sensitive word chain graph, the sensitive word chain graph is added to the set of sensitive word chain graphs corresponding to this character group. Thus, multiple sets of sensitive word chain graphs can be obtained. That is, each character group should correspond to a set of sensitive word chain graphs (it is possible to have an empty set. As long as there is an empty set, it means that the segment to be detected is not a sensitive word), and then the intersection of all the sets of sensitive word chain graphs is obtained as the sensitive word chain graph corresponding to the segment to be detected.
[0071] In some embodiments, multiple participles to be detected can be detected simultaneously to increase the efficiency and speed of detecting sensitive words in the text to be detected.
[0072] Please refer to Figure 3 , Figure 3 which shows a flowchart of constructing a sensitive word chain graph according to an embodiment of the present application. The embodiments of the present application provide the steps of the flowchart of constructing a sensitive word chain graph, including:
[0073] Step S201: Construct sample nodes corresponding to the characters in the sensitive word sample according to the characters, and link the sample nodes to obtain the sensitive word chain graph corresponding to each sensitive word sample. The sequential connection order of the sample nodes is consistent with the sequential order of the characters in the sensitive word sample.
[0074] Step S202: Establish index data corresponding to the sample nodes according to the sample nodes. The index data is used to index the sensitive word chain graph having the sample nodes.
[0075] The above two steps will be described in detail below.
[0076] In step S201, it should be noted that the sensitive word sample refers to the sensitive words that need to be blocked or not allowed to exist in the text. The sensitive word chain graph is constructed through the sensitive word sample to detect the sensitive words in the text to be detected.
[0077] The construction process of the sensitive word chain graph can be carried out in the following manner. First, divide the characters in the sensitive word sample into at least one character group, and each character group includes at least one character. In some embodiments, the character group can be one character. In other embodiments, the character group can be multiple characters, and different character groups can include different numbers of characters. Based on the sequential order of the characters in the sensitive word sample, it is used as the arrangement order between the character groups. For example, if the sensitive word sample is "1234567", the character groups are: "12", "3", "456", "7". Then the sorting result of the character groups is the first: "12"; the second: "3"; the third: "456"; the fourth: "7"; construct the sample nodes corresponding to each character group respectively. Each sample node corresponds to a character group. Finally, link the sample nodes corresponding to each character group in the arrangement order between the character groups to obtain the sensitive word chain graph. It should be further clarified that in some embodiments, each character group can include different numbers of characters.
[0078] In some embodiments, when constructing the corresponding sample nodes according to the character groups, the pinyin of the characters in the character group is used as the sample node. It should be clear that the pinyin used as the sample node can have tones or no tones, and no restrictions are imposed here.
[0079] In some embodiments, the content in the character group can be directly used as a sample node, that is, the content of the sample node is a character.
[0080] In some embodiments, a corresponding sample node is established for each character in the sensitive word sample, that is, each sample node corresponds to a character. Then the order before and after the characters is the order before and after each sample node.
[0081] In step S202, index data can be established for the characters corresponding to each sample node in each sensitive word chain graph. The index data is used to index the sensitive word chain graph having a specific sample node at a specific position. For example, through the index data, the sensitive word chain graph representing the character "456" at the third sample node can be indexed. That is, the index data includes two attributes: "content" and "arrangement position". Through the index data, the sensitive word chain graph having a specific sample node at a specific position can be quickly determined. By establishing the index data, "brute-force matching" for the to-be-detected segmented words can be avoided, the sensitive word chain graph corresponding to the to-be-detected segmented words can be increased, and thus the detection efficiency of the to-be-detected text can be increased.
[0082] In some embodiments, the sensitive word samples are classified. Exemplarily, they are classified according to the length of the sensitive word samples. Or they are classified according to the attributes of the sensitive word samples (such as verbs, nouns). Similarly, the sensitive word chain graphs corresponding to each sensitive word are classified according to the classification results of the sensitive word samples. In this way, when detecting the to-be-detected segmented words, according to the length and / or attributes of the to-be-detected segmented words, the sensitive word chain graphs can be screened first, and thus quick matching between the to-be-detected segmented words and the sensitive word chain graphs can be realized.
[0083] In some embodiments, when the sensitive word sample is a non-Chinese language, the abbreviated forms of some characters in the sensitive word sample are separately used as sample nodes to be used as the sensitive word chain graph or as a part of the sensitive word chain graph. In this way, the size of the sensitive word chain graph is reduced. At the same time, it avoids users avoiding sensitive word detection through abbreviations.
[0084] In some embodiments, when the sensitive word sample is a non-Chinese language, the abbreviated letters of some characters in the sensitive word sample are split and used as sample nodes respectively to be used as the sensitive word chain graph or as a part of the sensitive word chain graph. To improve the sensitive word detection rules for the to-be-detected segmented words.
[0085] Please refer to Figure 4 , Figure 4 shows a flowchart of constructing a sensitive word chain graph of similar words of a sensitive word sample according to an embodiment of the present application. The embodiments of the present application provide steps for constructing a sensitive word chain graph of similar words of a sensitive word sample, including:
[0086] Step S301, obtaining similar words for sensitive word samples;
[0087] Step S302, constructing a sensitive word chain graph of similar words.
[0088] The above two steps are described in detail below.
[0089] In step S301, in order to prevent users from constructing sensitive words through similar words with different pronunciations, the embodiment of the present application obtains similar characters for each character in the sensitive word sample, and then arranges and combines them to obtain similar words of the sensitive word sample.
[0090] In step S302, the technical means described in steps S201 and S202 are used to obtain a sensitive word chain graph of similar words to the sensitive word sample.
[0091] In some embodiments, since there are many sensitive word chain graphs, which will occupy a large storage space, the same sample nodes in the sensitive word chain graph are stored together.
[0092] In some embodiments, multiple sensitive word chain graphs share a "prefix" sample node. That is, if there are multiple sensitive word chain graphs with the same number of sample nodes in the front order, then the multiple sensitive word chain graphs share the same sample node. Figure 1 and sensitive word chain Figure 2 If the first N (N is a positive integer greater than or equal to the set number) sample nodes are the same, then the sensitive word chain Figure 1 and sensitive word chain Figure 2 Share the first N sample nodes.
[0093] The embodiments of the present application can greatly save the storage space of the sensitive word chain graph by sharing the same sample nodes.
[0094] See also Figure 5 , Figure 5 A flowchart of gradually determining a corresponding target sensitive word chain graph according to characters in a to-be-detected word segment according to an embodiment of the present application is shown. The embodiment of the present application provides a step of gradually determining a corresponding target sensitive word chain graph according to characters in a to-be-detected word segment, including:
[0095] Step S121a, obtaining at least one detection node from the characters in the word to be detected, and determining a target sample node corresponding to the detection node in the sample nodes according to the detection node;
[0096] Step S122a, based on the index data corresponding to the target sample node, determine the target sensitive word chain graph corresponding to the to-be-detected segmented word in accordance with the index data corresponding to each target sample node in an arbitrary order.
[0097] The above two steps will be described in detail below.
[0098] In step S121a, at least one detection node is obtained from the segmentation to be detected. The detection node corresponds to at least one character. For example, the detection node can directly be the corresponding character.
[0099] It should be clear that the detection node may not exist in the form of a character. The detection node can exist in the form of the pinyin of the corresponding character, and the pinyin can have tones or no tones.
[0100] In some embodiments, at least one character group is determined from the characters in the segmentation to be detected. Each character group includes at least one character. It should be clear that different numbers of characters can be included between each character group. A corresponding detection node is constructed for each character group, such as directly using the character group as the detection node.
[0101] In addition, the detection node may not exist in the form of a character. The detection node exists in the form of the pinyin of the corresponding character group, and the pinyin can have tones or no tones.
[0102] In some embodiments, for each character in the segmentation to be detected, a corresponding detection node is established in sequence. That is, each character in the segmentation to be detected corresponds to a detection node.
[0103] It should be clear that in some embodiments, the detection node and the sample node represent the corresponding characters in the same way. Exemplarily, both the detection node and the sample node represent the corresponding characters in pinyin. Both the detection node and the sample node represent the corresponding characters in numbers. It should be clear that in some embodiments, when using numbers to represent the characters corresponding to the detection node and the sample node, the numbers corresponding to the characters with the same pinyin (which can have tones or no tones) are the same.
[0104] It should be further clear that when the detection node and the sample node use the same representation method, the detection node and the sample node use the same representation for representing the same character. For example, for a character or a character group, if the detection node and the sample node use the same representation method for representation, then the detection node and the sample node have the same representation.
[0105] In other embodiments, the detection node and the sample node can use different representation methods to represent the corresponding characters. However, when the detection node and the sample node are matched, the representation methods of the detection node and the sample node are converted so that the representation methods of the detection node and the sample node are the same. So as to facilitate the matching between the detection node and the sample node.
[0106] Finally, according to each detection node, the target sample nodes corresponding to each detection node are determined respectively. Exemplarily, for any detection node, a sample node with the same content as it in the sample nodes is selected as the target sample node corresponding to the detection node.
[0107] In step S122a, after determining the target sample nodes corresponding to each detection node, the index data of the target sample nodes corresponding to the check nodes can be determined in any order, and then the sensitive word chain graph corresponding to the to-be-detected segmented word can be determined in turn.
[0108] Exemplarily, first, any one detection node is used as the first detection node for matching to determine the index data of the target sample node corresponding to the first detection node. This index data is used to represent the sensitive word chain graph including the first target sample node. It should be clear that the position of the first target sample node in this sensitive word chain graph is not limited, as long as the first target sample node exists. Then, the sensitive word chain graph corresponding to the first detection node is determined. The sensitive word chain graph corresponding to the first detection node refers to the sensitive word chain graph with the first target sample node existing at any position. Then, any other unmatched detection node is used as the second detection node. The second detection node corresponds to the second target sample node. The index data corresponding to the second target sample node is determined in the index data of the first target sample node. That is, at this time, the sensitive word chain graph represented by the index data corresponding to the second target sample node includes both the first target sample node and the second target sample node. Repeat the above steps, and match any detection node that has not been matched in turn. Finally, a sensitive word chain graph including all target sample nodes is obtained as the target sensitive chain graph corresponding to the to-be-detected segmented word.
[0109] It should be clear that if, during the matching process, the index data of the target sample node corresponding to any detection node is a null value, it directly indicates that the to-be-detected segmented word is not a sensitive word.
[0110] Through the method of this embodiment of the present application, it is possible to prevent users from avoiding the detection of sensitive words in a disordered order, making the detection of sensitive words in the to-be-detected text by the present application more accurate and more conducive to shaping a civilized language environment.
[0111] Please refer to Figure 6 , Figure 6 which shows a flowchart of gradually determining the corresponding target sensitive word chain graph according to the characters in the to-be-detected segmented word according to another embodiment of the present application. The embodiment of the present application provides steps for gradually determining the corresponding target sensitive word chain graph according to the characters in the to-be-detected segmented word, including:
[0112] Step S121b: Obtain at least one detection node from the characters in the to-be-detected segmented word, and determine the arrangement order of each detection node according to the arrangement order of the characters.
[0113] Step S122b: Determine the target sample node corresponding to the detection node.
[0114] Step S123b: According to the arrangement order of the detection nodes, sequentially determine the index data corresponding to the target sample nodes until the index data corresponding to the last target sample node is obtained, so as to determine the target sensitive word chain graph corresponding to the to-be-detected segmented word according to the index data corresponding to the last detection node.
[0115] The above three steps will be described in detail below.
[0116] In step S121b, obtain at least one detection node from the characters in the to-be-detected segmented word, and determine the arrangement order of each detection node according to the arrangement order of the characters. The order between the detection nodes is the same as the order of the characters corresponding to the detection nodes in the to-be-detected segmented word.
[0117] In step S122b, obtain the sample node representing the same character as the detection node as the target sample node of the detection node. In this way, the target sample nodes corresponding to each detection node can be determined.
[0118] In step S123b, according to the order between the detection nodes, determine the order between the detection nodes, that is, it can be clearly known which is the first detection node, the second detection node...
[0119] Exemplarily, first match the first detection node to determine the index data of the target sample node corresponding to the first detection node. This index data is used to represent the sensitive word chain graph including the first target sample node. It should be clear that this sensitive word chain graph defines the position where the first target sample node is located, and the first target sample node needs to be the sample node ranked first in the sensitive word chain graph. Then determine the sensitive word chain graph corresponding to the first detection node. The sensitive word chain graph corresponding to the first detection node refers to the sensitive word chain graph where the sample node ranked first is the first target sample node. Then match the second detection node. The second detection node corresponds to the second target sample node. Determine the index data corresponding to the second target sample node in the index data of the first target sample node. That is, at this time, the sensitive word chain graph represented by the index data corresponding to the second target sample node includes both the first target sample node and the second target sample node, and the sample node ranked second in the sensitive word chain graph is the second target sample node. Repeat the above steps to sequentially match the remaining detection nodes that have not been matched before, and finally obtain a sensitive word chain graph that includes all target sample nodes and the arrangement order of the target sample nodes is the same as the arrangement order of the detection nodes. Use this sensitive word chain graph as the target sensitive chain graph corresponding to the to-be-detected segmented word.
[0120] It should be clear that if, during the matching process, the index data of the target sample node corresponding to any detection node is a null value, it directly indicates that the to-be-detected segmented word is not a sensitive word.
[0121] In step S130, if the characters corresponding to each sample node in the target sensitive word chain graph overlap with the characters in the to-be-detected segmented word to reach the set threshold, then regard the to-be-detected segmented word as a sensitive word.
[0122] In some embodiments, if the sample nodes in the target sensitive word chain graph include each detection node in the to-be-detected segmented word, but there are other sample nodes different from the detection nodes, then according to the detection accuracy, the to-be-detected segmented word can be regarded as a sensitive word or not regarded as a sensitive word.
[0123] In some embodiments, if the characters corresponding to each sample node in the target sensitive word chain graph, when combined according to the sample nodes, result in a segmented word that is the same as the to-be-detected segmented word. Then regard the to-be-detected segmented word as a sensitive word.
[0124] In some embodiments, if in the target sensitive word chain graph corresponding to the finally determined to-be-detected segmented word, the existing sample nodes are exactly the same (in terms of quantity and content) as the detection nodes, it indicates that the to-be-detected segmented word is a sensitive word. Otherwise, the to-be-detected segmented word is not a sensitive word.
[0125] In some embodiments, if in the target sensitive word chain graph corresponding to the to-be-detected segmented word finally determined, the sample nodes existing are exactly the same (in terms of quantity and content) as the detection nodes, and the arrangement order among the detection nodes is the same as the arrangement order of the corresponding sample nodes in the target sensitive word chain graph, that is, the word formed by assembling the characters corresponding to the sample nodes according to the arrangement order of the sample nodes in the target sensitive word chain graph is the same as the to-be-detected segmented word, then the to-be-detected segmented word is considered a sensitive word.
[0126] Using the above technical means, the to-be-detected segmented words are respectively detected, and thus the sensitive words in the to-be-detected segmented words.
[0127] In step S140, a detection result of the to-be-detected text is generated according to the sensitive words in the to-be-detected segmented word, so as to indicate to the user which words are sensitive words.
[0128] In some embodiments, the sensitive words in the to-be-detected text are converted into set symbols to mask the sensitive words. In other embodiments, the sensitive words in the to-be-detected text are marked to prompt the user which characters form the sensitive words.
[0129] Figure 7 The block diagram of the computer system structure for implementing the method for detecting sensitive words according to an embodiment of the present application is shown.
[0130] It should be noted that Figure 7 The shown computer system 800 is only an example and should not bring any restrictions to the functions and usage scopes of the embodiments of the present application.
[0131] As Figure 7 shown, the computer system 800 includes a central processing unit 801 (Central Processing Unit, CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 802 (Read-Only Memory, ROM) or the program loaded from the storage section 808 into the random access memory 803 (Random Access Memory, RAM). In the random access memory 803, various programs and data required for system operation are also stored. The central processing unit 801, the read-only memory 802, and the random access memory 803 are connected to each other through a bus 804. The input / output interface 805 (Input / Output interface, that is, I / O interface) is also connected to the bus 804.
[0132] The following components are connected to the input / output interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a local area network card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output interface 805 as needed. A removable medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 810 as needed so that a computer program read therefrom is installed into the storage section 808 as needed.
[0133] Specifically, according to an embodiment of the present application, the processes described in each of the method flowcharts can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product including a computer program carried on a computer-readable medium, the computer program including program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the central processing unit 801, various functions defined in the system of the present application are executed.
[0134] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable medium may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0136] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0137] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by the way of software combined with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0138] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include the well-known common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0139] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only defined by the appended claims.
Claims
1. A method for detecting sensitive words, characterized in that: include: Performing word segmentation on the acquired text to be detected to obtain at least one word to be detected; Step by step, according to the characters in the to-be-detected segmented words, determine the corresponding target sensitive word chain graph; Determine the sensitive words in the to-be-detected segmented words according to the overlap degree between the to-be-detected segmented words and the target sensitive word chain graph; A sensitive word detection result of the text to be detected is generated according to the sensitive words in the to-be-detected segmented words.
2. The method according to claim 1, characterized in that Before performing word segmentation on the acquired text to be detected to obtain at least one word to be detected, the method further includes: Constructing sample nodes corresponding to the characters according to the characters in the sensitive word samples, and linking the sample nodes to obtain a sensitive word chain graph corresponding to each of the sensitive word samples, wherein the order of connection of the sample nodes is consistent with the order of characters in the sensitive word samples; Index data corresponding to the sample node is established according to the sample node, and the index data is used to index a sensitive word chain graph having the sample node.
3. The method according to claim 2, characterized in that The characters in the sample node may exist in the form of pinyin.
4. The method according to claim 2, characterized in that: The method further comprises: Acquire similar words for the sensitive word sample; Construct a sensitive word chain graph of the similar words.
5. The method according to claim 2, characterized in that: After constructing the sensitive word chain graph, the method further includes: According to each sample node of each of the sensitive word chain graphs, the sensitive word chain graphs with the same sample nodes are merged and stored.
6. The method according to claim 1, characterized in that Segmenting the acquired text to be detected to obtain at least one segmented word to be detected includes: Segment the acquired text to be detected to obtain multiple initial segmented words to be detected; The initial to-be-detected segmented words that meet the set conditions are regarded as the segmented words that do not need to be detected, and the initial to-be-detected segmented words other than the segmented words that do not need to be detected are regarded as the segmented words to be detected.
7. The method according to claim 2, characterized in that The corresponding target sensitive word chain graph is gradually determined according to the characters in the to-be-detected segmented word, including: At least one detection node is obtained from the characters in the word to be detected, and a target sample node corresponding to the detection node is determined in the sample nodes according to the detection node; Based on the index data corresponding to the target sample node, the target sensitive word chain graph corresponding to the to-be-detected segmented word is determined in any order according to the index data corresponding to each target sample node.
8. The method according to claim 2, characterized in that: The method further comprises: determining the corresponding target sensitive word chain graph step by step according to the characters in the to-be-detected segmented words; Obtaining at least one detection node from the characters in the word to be detected, and determining the arrangement order of each of the detection nodes according to the arrangement order of the characters; Determine a target sample node corresponding to the detection node; According to the arrangement order of the detection nodes, the index data corresponding to the target sample nodes are determined in sequence until the index data corresponding to the last target sample node is obtained, so as to determine the target sensitive word chain graph corresponding to the to-be-detected word segmentation according to the index data corresponding to the last detection node.
9. A readable storage medium, characterized in that: A readable program / instruction is stored thereon, and when the readable program / instruction is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
10. A program product, comprising a readable program / instruction, characterized in that: When the readable program / instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.