Sensitive word detection method and sensitive word detection device
By constructing a sensitive word tree and combining context similarity calculation, the existing sensitive word detection methods have solved the problems of high misjudgment rate and low efficiency, and achieved more efficient and accurate sensitive word detection.
Patent Information
- Application Number
- CN202510646547.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The existing sensitive word detection methods lack understanding of context semantics, resulting in high misjudgment rates and low word-by-word comparison efficiency, making it difficult to deal with large-scale texts.
By grouping preset sensitive words, a sensitive word tree is constructed, word segmentation units are obtained from the text to be detected, and the context similarity calculation is combined to improve detection accuracy and efficiency.
It greatly improves the accuracy and efficiency of sensitive word detection, can quickly identify conventional sensitive words, and deal with deformation, inversion, insertion and other situations, taking into account the balance between false alarm rate and missed rate.
Smart Images

Figure CN120163154A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a sensitive word detection method and a sensitive word detection device. Background Art
[0002] With the rapid development of the Internet, the network has become an important platform for people to obtain information and exchange ideas. However, the openness and anonymity of the network environment have also brought many challenges. One of them is the phenomenon that users deliberately use sensitive words to damage the network environment. In order to maintain the harmony and health of the network space and ensure the legality and appropriateness of network content, detecting sensitive words in text content has become an important technical requirement.
[0003] Currently, content review technology mainly relies on building a sensitive word library. By splitting the text content to be reviewed into word groups and then comparing them one by one with the sensitive words in the library, potential sensitive content can be identified and filtered out. Although this method is simple to implement, it has obvious limitations in practical applications. First of all, it lacks the understanding of the context semantics and cannot accurately judge the true meaning of words in a specific context, which is prone to misjudgment. Secondly, the method of comparing word by word is inefficient, especially when dealing with large-scale text, it takes a long time, resulting in unsatisfactory detection effects. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a sensitive word detection method and a sensitive word detection device. By grouping preset sensitive words, constructing a sensitive word tree, using the sensitive word tree to obtain word segmentation units from the text to be detected, and combining the calculation of context similarity, the accuracy and efficiency of sensitive word detection are greatly improved. It can not only quickly identify conventional sensitive words, but also handle various situations such as deformation, inversion, and insertion, while taking into account the balance between false alarm rate and missed alarm rate.
[0005] In the first aspect, an embodiment of this application provides a sensitive word detection method, and the sensitive word detection method includes: Classify multiple preset sensitive words and construct at least one sensitive word tree using the multiple preset sensitive words; Split the obtained text to be detected to determine multiple split units; For each target sensitive word tree whose root node is the same as the split unit, determine the length of each branch link in the target sensitive word tree, and based on the length of the branch link, determine at least one first word segmentation unit from the multiple split units; For each first word segmentation unit, compare the node in the target sensitive word tree with the first word segmentation unit to determine the first similarity corresponding to the first word segmentation unit, and based on the first similarity, determine whether there is a sensitive word in the text to be detected.
[0006] Further, classifying the multiple preset sensitive words and constructing at least one sensitive word tree by using the multiple preset sensitive words includes: Determining at least one sensitive word group from the multiple preset sensitive words; wherein, the first characters of each preset sensitive word in the sensitive word group have the same element; For each sensitive word group, taking the element corresponding to the first character of the preset sensitive word in the sensitive word group as the root node, and constructing corresponding child nodes for each root node in the order of the characters included in the preset sensitive word in the sensitive word group to generate the sensitive word tree corresponding to the sensitive word group.
[0007] Further, determining at least one first word segmentation unit from the multiple splitting units based on the length of the branch link includes: Determining at least one target splitting unit having the same element as the root node of the target sensitive word tree from the multiple splitting units; For each target splitting unit, combining multiple splitting units corresponding to the length of at least one branch link of the target sensitive word tree forward and / or backward starting from the target splitting unit according to the order of the multiple splitting units in the text to be detected to obtain the first word segmentation unit.
[0008] Further, comparing the nodes in the target sensitive word tree with the first word segmentation unit to determine the first similarity corresponding to the first word segmentation unit includes: Comparing the nodes of each branch link in the target sensitive word tree with the characters in the first word segmentation unit to determine the number of characters in the first word segmentation unit that have the same element as the nodes of each branch link in the target sensitive word tree; Taking the ratio between the number and the length of the corresponding branch link of the target sensitive word tree as the first similarity; Or, For each branch link in the target sensitive word tree, taking the ratio between the number and the length of the branch link as the branch similarity; Determining the first similarity corresponding to the first word segmentation unit by using the multiple branch similarities.
[0009] Further, after determining the first similarity corresponding to the first word segmentation unit, the sensitive word detection method further includes: If the first similarity is greater than or equal to the first similarity threshold, determining a first sensitive word corresponding to the first word segmentation unit from the target sensitive word tree and determining whether there is an exemption phrase for the first sensitive word; If so, based on the phrase length of the exemption phrase, and in the order of multiple splitting units in the text to be detected, starting from each target splitting unit, obtain multiple splitting units corresponding to the phrase length forward and / or backward for combination, so as to obtain a second segmented unit; Compare the exemption phrase with the second segmented unit to determine the second similarity corresponding to the second segmented unit, so as to judge whether to regard the first segmented unit as a sensitive word appearing in the text to be detected based on the second similarity and the second similarity threshold.
[0010] Further, after determining the first similarity corresponding to the first segmented unit, the sensitive word detection method further includes: Determine from the target sensitive word tree a second sensitive word whose first similarity corresponding to the first segmented unit is greater than or equal to the first similarity threshold and whose sensitive word category belongs to non-essential filtering; Based on a first preset segmentation length, determine a third segmented unit corresponding to the second sensitive word from the text to be detected, and input the third segmented unit into a pre-trained sensitive value detection model to determine the third similarity corresponding to the third segmented unit, so as to judge whether to regard the first segmented unit as a sensitive word appearing in the text to be detected based on the third similarity, or judge whether to regard the first segmented unit as a sensitive word appearing in the text to be detected based on the first similarity and the third similarity.
[0011] Further, after determining the first similarity corresponding to the first segmented unit, the sensitive word detection method further includes: Determine from the target sensitive word tree a second sensitive word whose first similarity corresponding to the first segmented unit is greater than or equal to the first similarity threshold and whose sensitive word category belongs to non-essential filtering; Determine the exclusion weight corresponding to the second sensitive word, and use the product of the exclusion weight and the first similarity as the fourth similarity, so as to judge whether to regard the first segmented unit as a sensitive word appearing in the text to be detected based on the fourth similarity.
[0012] Further, after determining the first similarity corresponding to the first segmented unit, the sensitive word detection method further includes: Determine from the target sensitive word tree a third sensitive word whose first similarity corresponding to the first segmented unit is greater than or equal to the second similarity threshold and whose sensitive word category belongs to a combined phrase; Based on the second preset word segmentation length, determine the fourth word segmentation unit corresponding to the third sensitive word from the text to be detected, and determine whether there are sensitive words in the text to be detected based on the judgment result of whether the fourth word segmentation unit contains the fourth sensitive word.
[0013] Further, classifying the multiple preset sensitive words and constructing at least one sensitive word tree by using the multiple preset sensitive words includes: Classify the multiple preset sensitive words according to any one or more of pronunciation, character shape, and letters, and construct at least one sensitive word tree by using the multiple preset sensitive words. The splitting of the obtained text to be detected to determine multiple splitting units includes: Split the obtained text to be detected according to any one or more of the pronunciation, character shape, and letters of each character in the text to be detected, and determine multiple splitting units.
[0014] In a second aspect, an embodiment of the present application further provides a sensitive word detection device, and the sensitive word detection device includes: A sensitive word tree construction module, configured to classify multiple preset sensitive words and construct at least one sensitive word tree by using the multiple preset sensitive words; A splitting unit determination module, configured to split the obtained text to be detected to determine multiple splitting units; A word segmentation unit determination module, configured to, for each target sensitive word tree with the same root node as the splitting unit, determine the length of each branch link in the target sensitive word tree, and determine at least one first word segmentation unit from the multiple splitting units based on the length of the branch link; A first detection module, configured to, for each first word segmentation unit, compare the node in the target sensitive word tree with the first word segmentation unit to determine the first similarity corresponding to the first word segmentation unit, and determine whether there are sensitive words in the text to be detected based on the first similarity.
[0015] A sensitive word detection method and a sensitive word detection device provided by an embodiment of the present application. First, classify multiple preset sensitive words and construct at least one sensitive word tree using the multiple preset sensitive words; then, split the obtained text to be detected to determine multiple split units; for each target sensitive word tree whose root node is the same as the split unit, determine the length of each branch link in the target sensitive word tree, and determine at least one first word segmentation unit from the multiple split units based on the length of the branch link; finally, for each first word segmentation unit, compare the node in the target sensitive word tree with the first word segmentation unit to determine the first similarity corresponding to the first word segmentation unit, and determine whether there is a sensitive word in the text to be detected based on the first similarity.
[0016] In the present application, by grouping preset sensitive words, constructing a sensitive word tree, obtaining corresponding word segmentation units from the text to be detected using the sensitive word tree, and combining the calculation of context similarity, the accuracy and efficiency of sensitive word detection are greatly improved. It can not only quickly identify conventional sensitive words, but also handle various situations such as deformation, inversion, insertion, etc., while taking into account the balance of false alarm rate and missed alarm rate.
[0017] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a flowchart of a sensitive word detection method provided by an embodiment of the present application; Figure 2 It is a structural schematic diagram of a sensitive word tree provided by an embodiment of the present application; Figure 3 It is a structural schematic diagram of a sensitive word tree with exemption phrases provided by an embodiment of the present application; Figure 4 It is a structural schematic diagram of a sensitive word detection device provided by an embodiment of the present application; Figure 5 It is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part rather than all of the embodiments of this application. Components of the embodiments of this application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without creative efforts falls within the scope of protection of this application.
[0021] First, the applicable application scenarios of this application will be introduced. This application can be applied to the field of computer technology.
[0022] With the rapid development of the Internet, the network has become an important platform for people to obtain information and exchange ideas. However, the openness and anonymity of the network environment have also brought many challenges, one of which is the phenomenon that users deliberately use sensitive words to disrupt the network environment. In order to maintain the harmony and health of the network space and ensure the legality and appropriateness of network content, sensitive word detection of text content has become an important technical requirement.
[0023] It has been found through research that currently, content review technology mainly relies on building a sensitive word library. By splitting the text content to be reviewed into word groups and then comparing them one by one with the sensitive words in the library, potential sensitive content can be identified and filtered out. Although this method is simple to implement, it has obvious limitations in practical applications. First, it lacks the understanding of context semantics and cannot accurately judge the true meaning of words in a specific context, which is prone to misjudgment. Second, the method of comparing word by word is inefficient, especially when dealing with large-scale text, it takes a long time and the detection effect is not ideal.
[0024] Based on this, the embodiments of this application provide a sensitive word detection method, which greatly improves the accuracy and efficiency of sensitive word detection. It can not only quickly identify conventional sensitive words, but also handle various situations such as deformation, inversion, and insertion, while taking into account the balance between false alarm rate and missed alarm rate.
[0025] Please refer to Figure 1 , Figure 1 which is a flowchart of a sensitive word detection method provided by the embodiments of this application. As shown in Figure 1 the sensitive word detection method provided by the embodiments of this application includes: S101, classify multiple preset sensitive words and construct at least one sensitive word tree by using the multiple preset sensitive words.
[0026] Here, the preset sensitive words are pre-set, manually or algorithmically defined vocabulary sets that need to be blocked. The sensitive word tree is a multi-branch tree structure constructed by decomposing the preset sensitive words into tree levels by characters. Here, a sensitive word tree contains a root node and at least one child node. One character represents one node. One character can be represented by a word, radical, phrase, or pinyin, or a letter in a node. The root node and the child nodes of the root node at all levels form a preset sensitive word.
[0027] Regarding the above step S101, in a specific implementation, a plurality of preset sensitive words are obtained, the plurality of preset sensitive words are classified, and at least one sensitive word tree is constructed using the plurality of preset sensitive words.
[0028] As an optional embodiment, with respect to the above step S101, the step of classifying a plurality of preset sensitive words and constructing at least one sensitive word tree using the plurality of preset sensitive words includes: Step 1011, determining at least one sensitive word group from a plurality of preset sensitive words.
[0029] Here, the first character of each preset sensitive word in the sensitive word group has the same element. Specifically, the same element can be that the first character of each preset sensitive word in the same sensitive word group is exactly the same.
[0030] For the above step 1011, in the specific implementation, preset sensitive words with the same elements in the first characters are determined from multiple preset sensitive words, and the preset sensitive words with the same elements in the first characters are used as a sensitive phrase group. Here, the same element is used as an example that the first characters of each preset sensitive word are exactly the same. For example, the preset sensitive words are ABB, ABC, BBA, BDA, CDA, DDA, AB, BD, ADCB, and CBAD. Then, according to the first character in the preset sensitive words, the preset sensitive words with the same first character are classified into a sensitive phrase group, and multiple sensitive phrase groups can be obtained. The first sensitive phrase group includes ABB, ABC, AB, and ADCB, the second sensitive phrase group includes BBA, BDA, and BD, the third sensitive phrase group includes CDA and CBAD, and the fourth sensitive phrase group includes DDA. Here, as an optional embodiment, the preset sensitive words with the same first character can also be classified into a large category first, and then the characters with the same characters in the subsequent corresponding positions can be classified into the same sensitive phrase group, or the characters with the same characters in the subsequent corresponding positions and the same character length can be classified into the same sensitive phrase group.
[0031] Step 1012: For each sensitive phrase, use the element corresponding to the first character of the preset sensitive word in the sensitive phrase as the root node, and construct corresponding child nodes for each root node in the order of the characters included in the preset sensitive word in the sensitive phrase, so as to generate a sensitive word tree corresponding to the sensitive phrase.
[0032] Regarding the above Step 1012, in specific implementation, group the preset sensitive words to obtain multiple sensitive words. Then, for each sensitive phrase, split each preset sensitive phrase in the sensitive phrase into characters, and then construct a tree structure layer by layer, using the element corresponding to the first character of the preset sensitive word in the sensitive phrase as the root node. Here, when the sensitive phrases in the above Step 1011 are grouped according to the same characters, the root node is the first character of the preset sensitive word. Then, construct corresponding child nodes for each root node in the order of the characters included in the preset sensitive word in the sensitive phrase, so as to generate a sensitive word tree corresponding to the sensitive phrase. Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a sensitive word tree provided by an embodiment of the present application. Continuing the embodiment in the above Step 1011, disassemble the preset sensitive words in each sensitive phrase into a tree-like hierarchy, and construct a multi-way tree structure. Each node in the sensitive word tree represents a character, and the path from the root node to the leaf node forms a complete sensitive word.
[0033] S102: Split the obtained text to be detected to determine multiple split units.
[0034] Here, the text to be detected can refer to any text content input by the user or any text content output by the large language model, and it is necessary to detect sensitive words through the sensitive word tree. For example, the text to be detected can be a comment or chat message input by the user, or other types of text data, and the present application does not make specific limitations on this.
[0035] Regarding the above Step S102, in specific implementation, obtain the text to be detected input by the user, and split each character in the obtained text to be detected to determine multiple split units. Here, as an example, when the obtained text to be detected is "AC,DBDBA.CCA#BBCD@D*AA!B! ", split each character in the text to be detected, and the multiple split units obtained include: A, C, D, B, D, B, A, C, C, A, B, B, C, D, D, A, A, B. Among them, interference elements such as spaces, special characters, and punctuation marks can be filtered out, and only meaningful characters are retained for content review.
[0036] According to the sensitive word detection method provided by this application, the preset sensitive words can also be classified according to the pronunciation, shape, and / or letters of the characters. As another alternative embodiment, for the above step S101, the classification of multiple preset sensitive words and the construction of at least one sensitive word tree using the multiple preset sensitive words include: Classify multiple preset sensitive words according to any one or more of pronunciation, shape, and letters, and construct at least one sensitive word tree using the multiple preset sensitive words.
[0037] For the above steps, in specific implementation, classify the preset sensitive words according to any one or more of pronunciation, shape, and letters, and based on the classification results, construct at least one sensitive word tree using the multiple preset sensitive words. Here, for example, classify the preset sensitive words based on the pronunciation of the characters in the obtained preset sensitive words, group the preset sensitive words with the same pronunciation of the first character into one category, and the pronunciation of the first character of the preset sensitive words in each sensitive word group is exactly the same. In this way, the root node of the generated sensitive word tree is the pronunciation of the first character of the preset sensitive word, and subsequent sensitive word judgment is based on pronunciation to solve the problem of identifying sensitive words with the same pronunciation but different characters.
[0038] Further, for the above step S102, the splitting of the obtained text to be detected to determine multiple splitting units includes.
[0039] Split the obtained text to be detected according to any one or more of the pronunciation, shape, and letters of each character in the text to be detected to determine multiple splitting units.
[0040] Here, after classifying the preset sensitive words according to any one or more of pronunciation, shape, and letters, the text to be detected also needs to be split according to any one or more of pronunciation, shape, and letters. For the above steps, in specific implementation, split the obtained text to be detected according to any one or more of the pronunciation, shape, and letters of each character in the text to be detected to determine multiple splitting units. Continuing the example in the above steps, after classifying the preset sensitive words based on their pronunciation, the text to be detected also needs to be split into multiple splitting units according to pronunciation.
[0041] S103. For each target sensitive word tree with the same root node as the splitting unit, determine the length of each branch link in the target sensitive word tree, and based on the length of the branch link, determine at least one first word segmentation unit from the multiple splitting units.
[0042] For the above step S103, in specific implementation, first, based on the multiple splitting units obtained in the above step S102, a target sensitive word tree with the same root node as the splitting unit is determined from multiple sensitive word trees. Here, continuing the above embodiment, the selected target sensitive word trees are the sensitive word tree with root node A, the sensitive word tree with root node B, the sensitive word tree with root node C, and the sensitive word tree with root node D. This can greatly reduce the matching workload of the review content in the thesaurus, save matching time, and improve the content review efficiency. Then, for each target sensitive word tree, the length of each branch link in the target sensitive word tree is determined, and at least one first word segmentation unit is determined from the multiple splitting units based on the length of each branch link. As an alternative embodiment, the maximum length can also be determined from the lengths of the multiple branch links, and the first word segmentation unit is determined from the multiple splitting units based on the maximum length.
[0043] As an alternative embodiment, for the above step S103, the determining of at least one first word segmentation unit from the multiple splitting units based on the length of the branch link includes: Step 1031, determine at least one target splitting unit having the same element as the root node of the target sensitive word tree from the multiple splitting units.
[0044] For the above step 1031, in specific implementation, determine at least one target splitting unit having the same element as the root node of the target sensitive word tree from the multiple splitting units. Here, as an example, when the root node of the target sensitive word tree is A, the splitting unit that is A among the multiple splitting units is used as the target splitting unit.
[0045] Step 1032, for each target splitting unit, in accordance with the order of the multiple splitting units in the text to be detected, starting from the target splitting unit, obtain and combine multiple splitting units corresponding to the length of at least one branch link of the target sensitive word tree forward and / or backward to obtain the first word segmentation unit.
[0046] Regarding the above step 1032, in specific implementation, for each target splitting unit, according to the order of multiple splitting units in the text to be detected, starting from the target splitting unit, multiple splitting units corresponding to the lengths of at least one branch link of the target sensitive word tree are obtained forward and / or backward for combination to obtain a first word segmentation unit. Here, for the obtained multiple splitting units, they can be combined in order or randomly without order to overcome the problem of deliberately disrupting the order to avoid sensitive word detection. Continuing the above embodiment, when the splitting unit with A among multiple splitting units is used as the target splitting unit and the lengths of each branch link are 2, 3, and 4 respectively, multiple first word segmentation units, namely ACDB, ACCA, ABBC, AAB, and AB, can be extracted from the text to be detected by combining them in order.
[0047] S104. For each first word segmentation unit, use the nodes in the target sensitive word tree to compare with the first word segmentation unit to determine the first similarity corresponding to the first word segmentation unit, and based on the first similarity, determine whether there are sensitive words in the text to be detected.
[0048] Regarding the above step S104, in specific implementation, for each first word segmentation unit, use the nodes in the target sensitive word tree to compare with the first word segmentation unit to determine the first similarity corresponding to the first word segmentation unit. Here, according to the embodiments provided in the present application, the first similarity can be the similarity of the first word segmentation unit for each specific branch in the target sensitive word tree or the similarity for the target sensitive word tree. Then, based on the first similarity corresponding to the first word segmentation unit, determine whether there are sensitive words in the text to be detected. Specifically, compare the determined first similarity corresponding to the first word segmentation unit with a preset first similarity threshold. If the first similarity is greater than or equal to the first similarity threshold, it is determined that there are sensitive words in the text to be detected.
[0049] Here, when determining the first similarity corresponding to the word segmentation unit, the embodiments of the present application provide two methods for determining the first similarity. Specifically, in Method 1, regarding the above step S104, the step of using the nodes in the target sensitive word tree to compare with the first word segmentation unit to determine the first similarity corresponding to the first word segmentation unit includes: A: Compare the nodes of each branch link in the target sensitive word tree with the characters in the first word segmentation unit to determine the number of characters in the first word segmentation unit that have the same elements as the nodes of each branch link in the target sensitive word tree.
[0050] For the above-mentioned step A, in specific implementation, compare the nodes of each branch link in the target sensitive word tree with the characters in the first word segmentation unit to determine the number of characters in the first word segmentation unit that have the same elements as the nodes of each branch link in the target sensitive word tree. Here, the nodes of each branch link include the root node of the target sensitive word tree. Specifically, when comparing the nodes with the characters, the comparison can be carried out in order. Since in the above step of determining the first word segmentation unit, the first character of the first word segmentation unit is the same as the root node of the target sensitive word tree, when comparing in order, first use each branch link node in the second layer of the target sensitive word tree to compare with the second character in the first word segmentation unit, and so on. It can also be not compared in order. For example, use each branch link node in the second layer of the target sensitive word tree to compare with the 2nd / 3rd / 4th characters in the first word segmentation unit, which can avoid the problem of missing sensitive words caused by word order. It can also be that the nodes of each branch link are compared with multiple characters in the first word segmentation unit. For example, each branch link node in the second layer of the target sensitive word tree is compared with the 2nd and 3rd characters in the first word segmentation unit. It can also be to only compare the characters in the corresponding positions. For example, each branch link node in the second layer of the target sensitive word tree is compared with the second character in the first word segmentation unit.
[0051] B: Take the ratio between the quantity and the length of the corresponding branch link of the target sensitive word tree as the first similarity.
[0052] For the above-mentioned step B, in specific implementation, divide the number of the same characters by the length of the corresponding branch link in the target sensitive word tree to determine the first similarity. In this way, the order of character appearance can be ignored, as long as they can match one by one, without the need for sequential matching, which can overcome the problems of word order inversion and missing sensitive words.
[0053] Method 2. For the above-mentioned step S104, the use of the nodes in the target sensitive word tree to compare with the first word segmentation unit to determine the first similarity corresponding to the first word segmentation unit includes: a: Compare the nodes of each branch link in the target sensitive word tree with the characters in the first word segmentation unit to determine the number of characters in the first word segmentation unit that have the same elements as the nodes of each branch link in the target sensitive word tree.
[0054] Here, the description of the above step a can refer to the description of the above step A and can achieve the same technical effect, which will not be elaborated here.
[0055] b: For each branch link in the target sensitive word tree, take the ratio between the quantity and the length of the branch link as the branch similarity.
[0056] c: Determine the first similarity corresponding to the first word segmentation unit by using multiple branch similarities.
[0057] For the above steps b - c, in specific implementation, for each branch link in the target sensitive word tree, the ratio between the number of characters in the first word segmentation unit that have the same elements as the nodes in the target sensitive word tree and the length of the branch link is used as the branch similarity. Then, the first similarity corresponding to the first word segmentation unit is determined by using multiple branch similarities. Here, as an example, the average value of multiple branch similarities can be taken as the first similarity corresponding to the first word segmentation unit, or the maximum similarity, minimum similarity, or median of multiple branch similarities can be taken as the first similarity. This application does not make specific limitations on this.
[0058] According to the sensitive word detection method provided by this application, after determining the similarity corresponding to the first word segmentation unit, the first word segmentation unit can be further detected based on the sensitive word tree to determine whether the first word segmentation unit is a sensitive word that needs to be blocked in the text to be detected. As an optional embodiment, after determining the first similarity corresponding to the first word segmentation unit, the sensitive word detection method further includes: I: If the first similarity is greater than or equal to the first similarity threshold, determine the first sensitive word corresponding to the first word segmentation unit from the target sensitive word tree, and determine whether there is an exemption phrase for the first sensitive word.
[0059] Here, the exemption phrase refers to a sensitive phrase with more characters extended from a preset sensitive word.
[0060] For the above step I, in specific implementation, when the first similarity of the determined first word segmentation unit is greater than the first similarity threshold, the first sensitive word corresponding to the first word segmentation unit is determined from the target sensitive word tree, and it is determined whether there is an exemption phrase for the first sensitive word. Here, as an example, the first sensitive word is obtained through the similarity between each sensitive word in the target sensitive word tree and the first word segmentation unit. When the similarity between a certain sensitive word in the target sensitive word tree and the first word segmentation unit is greater than or equal to a certain value, then this sensitive word is used as the first sensitive word. For example, when the first word segmentation unit is ABBC, the first sensitive word ABB with a relatively high similarity to the first word segmentation unit can be determined from the target sensitive word tree with the root node A. Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a sensitive word tree with exemption phrases provided by an embodiment of this application. As Figure 3 shown, the exemption phrase corresponds to 3 groups of characters bound to the ABB branch, and the exemption phrase is bound to the last leaf node.
[0061] II: If so, based on the phrase length of the exemption phrase, in the order of multiple splitting units in the text to be detected, starting from each target splitting unit, obtain multiple splitting units corresponding to the phrase length forward and / or backward for combination to obtain a second segmented unit.
[0062] Regarding the above step II, in specific implementation, when there is an exemption phrase for the first sensitive word, based on the phrase length of the exemption phrase, in the order of multiple splitting units in the text to be detected, starting from each target splitting unit, obtain multiple splitting units corresponding to the phrase length forward and / or backward, and perform combination to obtain a second segmented unit. Here, for the multiple splitting units corresponding to the obtained phrase length, they can be combined in order or randomly without following the order to overcome the problem of deliberately disrupting the order to avoid sensitive word detection. Here, continuing with the example in the above step I, for the 3 exemption phrases bound to the above ABB, the longest phrase is 8 characters. Then, taking A as the target splitting unit, obtain the second segmented unit formed by the 8 characters before and after A in order from the text to be detected, that is, CDBDBACCABBCDDAAB.
[0063] III: Compare the exemption phrase with the second segmented unit to determine the second similarity corresponding to the second segmented unit, so as to judge whether to regard this first segmented unit as the sensitive word appearing in the text to be detected based on the second similarity and the second similarity threshold.
[0064] Regarding the above step III, in specific implementation, compare the exemption phrase with the second segmented unit to determine the second similarity corresponding to the second segmented unit. Here, the method for calculating the second similarity is the same as the method for calculating the first similarity in the above embodiment, and the same technical effect can be achieved, which will not be elaborated here. Then, based on the second similarity threshold and the second similarity, judge whether to regard this first segmented unit as the sensitive word appearing in the text to be detected. Here, as an example, when there are 3 exemption phrases for the sensitive word ABB, calculate the similarity between the second segmented unit and each exemption phrase respectively, and then take the average of the three similarities as the second similarity, or take the maximum similarity, minimum similarity or median of the three similarities as the second similarity. At this time, compare the second similarity with the second similarity threshold. If the second similarity is greater than or equal to the second similarity threshold, regard this first segmented unit as the sensitive word appearing in the text to be detected.
[0065] As an optional embodiment, after determining the first similarity corresponding to this first segmented unit, the sensitive word detection method further includes: i: Determine, from the target sensitive word tree, a second sensitive word whose first similarity corresponding to the first word segmentation unit is greater than or equal to the first similarity threshold and whose sensitive word category belongs to non-essential filtering.
[0066] Regarding step i above, in specific implementation, determine, from the target sensitive word tree, a second sensitive word whose first similarity corresponding to the first word segmentation unit is greater than or equal to the first similarity threshold and whose sensitive word category belongs to non-essential filtering. Here, as an example, when the first word segmentation unit is ABBC, a second sensitive word ABB whose first similarity corresponding to the first word segmentation unit in the target sensitive word tree with root node A is greater than or equal to the first similarity threshold and whose sensitive word category belongs to non-essential filtering can be determined.
[0067] ii: Based on the first preset word segmentation length, determine, from the text to be detected, a third word segmentation unit corresponding to the second sensitive word, and input the third word segmentation unit into a pre-trained sensitive value detection model to determine the third similarity corresponding to the third word segmentation unit, so as to judge whether to regard the first word segmentation unit as a sensitive word appearing in the text to be detected based on the third similarity, or judge whether to regard the first word segmentation unit as a sensitive word appearing in the text to be detected based on the first similarity and the third similarity.
[0068] Here, the preset word segmentation length can be set to 8, and the present application does not make specific limitations thereon.
[0069] For step ii above, in specific implementation, based on a preset word segmentation length, a third word segmentation unit corresponding to the second sensitive word is determined from the text to be detected. Here, a phrase of the preset word segmentation length before and / or after the second sensitive word in the text to be detected is used as the third word segmentation unit. Then, the third word segmentation unit is input into a pre-trained sensitive value detection model to determine the third similarity corresponding to the third word segmentation unit. Here, the sensitive value detection module is trained through sample sensitive words and the sample sensitive values corresponding to the sample sensitive words, and its training method is the same as the training method for neural network models in the prior art, which will not be elaborated here. Then, based on the third similarity of the third word segmentation unit, it is determined whether to regard the first word segmentation unit as a sensitive word appearing in the text to be detected, or based on the first similarity of the first word segmentation unit and the third similarity of the third word segmentation unit, it is determined whether to regard the first word segmentation unit as a sensitive word appearing in the text to be detected. Here, as an optional embodiment, the third similarity of the third word segmentation unit is compared with a threshold. If it is greater than or equal to the threshold, it is considered that the first word segmentation unit is a sensitive word appearing in the text to be detected. Alternatively, the judgment result of the sensitive word detection model, that is, the third similarity of the third word segmentation unit, can also be multiplied by the first similarity of the first word segmentation unit, and it is judged whether the multiplied similarity is greater than the threshold. If so, it is considered that the first word segmentation unit is a sensitive word appearing in the text to be detected.
[0070] As an optional embodiment, after determining the first similarity corresponding to the first word segmentation unit, the sensitive word detection method further includes: Determining from the target sensitive word tree a second sensitive word whose first similarity corresponding to the first word segmentation unit is greater than or equal to a first similarity threshold and whose sensitive word category belongs to non-essential filtering; determining the exclusion weight corresponding to the second sensitive word, and using the product of the exclusion weight and the first similarity as the fourth similarity, so as to determine whether to regard the first word segmentation unit as a sensitive word appearing in the text to be detected based on the fourth similarity.
[0071] For the above two steps, in specific implementation, different sizes of exclusion weights can also be set for sensitive words whose sensitive word category belongs to non-essential filtering, and it is determined whether to exclude the sensitive word based on the exclusion weight, the first similarity, and an exclusion threshold. Specifically, a second sensitive word whose first similarity corresponding to the first word segmentation unit is greater than or equal to a first similarity threshold and whose sensitive word category belongs to non-essential filtering is determined from the target sensitive word tree. The exclusion weight corresponding to the second sensitive word is determined, and the product of the exclusion weight and the first similarity of the first word segmentation unit is used as the fourth similarity. If the fourth similarity is greater than or equal to the threshold, it is considered that the first word segmentation unit is a sensitive word appearing in the text to be detected.
[0072] As an alternative embodiment, after determining the first similarity corresponding to the first word segmentation unit, the sensitive word detection method further includes: Determine a third sensitive word from the target sensitive word tree, where the first similarity corresponding to the first word segmentation unit is greater than or equal to a second similarity threshold and the sensitive word category belongs to a combined phrase; based on a second preset word segmentation length, determine a fourth word segmentation unit corresponding to the third sensitive word from the text to be detected, and determine whether there is a sensitive word in the text to be detected based on the judgment result of whether the fourth word segmentation unit contains a fourth sensitive word.
[0073] Here, when performing sensitive word detection in this application, combined phrase sensitive words are also included. For the above two steps, in specific implementation, first determine a third sensitive word from the target sensitive word tree, where the first similarity corresponding to the first word segmentation unit is greater than or equal to the second similarity threshold and the sensitive word category belongs to a combined phrase. Here, for example, set AB......CD as the third sensitive word whose sensitive word category belongs to a combined phrase. Then, based on the second preset word segmentation length, determine the fourth word segmentation unit corresponding to the third sensitive word from the text to be detected. For example, if it is detected that AB appears in the text to be detected, then forward and / or backward from AB in the text to be detected to find the second preset word segmentation length, and determine whether there is a character C. If there is a character C, then continue to forward and / or backward from C in the text to be detected to determine whether there is a character D. If there is a character D, then take the found content as the fourth word segmentation unit. Or, if it is detected that AB appears in the text to be detected, then forward and / or backward from AB in the text to be detected to find the second preset word segmentation length, and determine whether characters C and D are found, and take the found content as the fourth word segmentation unit. Then, determine whether there is a sensitive word in the text to be detected based on the judgment result of whether the fourth word segmentation unit contains a fourth sensitive word.
[0074] The sensitive word detection method provided by the embodiment of this application first classifies multiple preset sensitive words and constructs at least one sensitive word tree using the multiple preset sensitive words; then splits the obtained text to be detected to determine multiple splitting units; for each target sensitive word tree whose root node is the same as the splitting unit, determine the length of each branch link in the target sensitive word tree, and determine at least one first word segmentation unit from the multiple splitting units based on the length of the branch link; finally, for each first word segmentation unit, compare the node in the target sensitive word tree with the first word segmentation unit to determine the first similarity corresponding to the first word segmentation unit, and determine whether there is a sensitive word in the text to be detected based on the first similarity.
[0075] This application groups preset sensitive words to construct a sensitive word tree, uses the sensitive word tree to obtain corresponding word segmentation units from the text to be detected, and combines the calculation of context similarity, which greatly improves the accuracy and efficiency of sensitive word detection. It can not only quickly identify conventional sensitive words, but also handle various situations such as deformation, inversion, and insertion, while taking into account the balance between false positive rate and false negative rate.
[0076] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a sensitive word detection device provided by an embodiment of this application. As Figure 4 shown in the sensitive word tree construction module 401 is used to classify multiple preset sensitive words and construct at least one sensitive word tree by using the multiple preset sensitive words; the splitting unit determination module 402 is used to split the obtained text to be detected and determine multiple splitting units; the word segmentation unit determination module 403 is used to, for each target sensitive word tree with the same root node as the splitting unit, determine the length of each branch link in the target sensitive word tree, and determine at least one first word segmentation unit from the multiple splitting units based on the length of the branch link; the first detection module 404 is used to, for each first word segmentation unit, compare the node in the target sensitive word tree with the first word segmentation unit to determine the first similarity corresponding to the first word segmentation unit, and determine whether there are sensitive words in the text to be detected based on the first similarity.
[0077] Further, when the sensitive word tree construction module 401 is used to classify multiple preset sensitive words and construct at least one sensitive word tree by using the multiple preset sensitive words, the sensitive word tree construction module 401 is further used to: determine at least one sensitive word group from the multiple preset sensitive words; wherein, the first characters of each preset sensitive word in the sensitive word group have the same element; for each sensitive word group, use the element corresponding to the first character of the preset sensitive word in the sensitive word group as the root node, and construct corresponding child nodes for each root node in the order of the characters included in the preset sensitive word in the sensitive word group to generate the sensitive word tree corresponding to the sensitive word group.
[0078] Further, when the word segmentation unit determination module 403 is used to determine at least one first word segmentation unit from the multiple splitting units based on the length of the branch link, the word segmentation unit determination module 403 is further used to: determine at least one target splitting unit having the same element as the root node of the target sensitive word tree from the multiple splitting units; For each target splitting unit, in accordance with the order of the multiple splitting units in the text to be detected, starting from the target splitting unit, combine multiple splitting units corresponding to the lengths of at least one branch link of the target sensitive word tree forward and / or backward to obtain the first word segmentation unit.
[0079] Further, when the first detection module 404 is used to compare the nodes in the target sensitive word tree with the first word segmentation unit to determine the first similarity corresponding to the first word segmentation unit, the first detection module 404 is further used for: Compare the nodes of each branch link in the target sensitive word tree with the characters in the first word segmentation unit to determine the number of characters in the first word segmentation unit that have the same elements as the nodes of each branch link in the target sensitive word tree; Take the ratio between the quantity and the length of the corresponding branch link of the target sensitive word tree as the first similarity; Or, For each branch link in the target sensitive word tree, take the ratio between the quantity and the length of the branch link as the branch similarity; Use multiple branch similarities to determine the first similarity corresponding to the first word segmentation unit.
[0080] Further, the sensitive word detection device 400 further includes a second detection module. After determining the first similarity corresponding to the first word segmentation unit, the second detection module is used for: If the first similarity is greater than or equal to the first similarity threshold, determine the first sensitive word corresponding to the first word segmentation unit from the target sensitive word tree, and determine whether there is an exemption phrase for the first sensitive word; If so, based on the phrase length of the exemption phrase, in accordance with the order of the multiple splitting units in the text to be detected, starting from each target splitting unit, combine multiple splitting units corresponding to the phrase length forward and / or backward to obtain a second word segmentation unit; Compare the exemption phrase with the second word segmentation unit to determine the second similarity corresponding to the second word segmentation unit, and based on the second similarity and the second similarity threshold, determine whether to regard the first word segmentation unit as the sensitive word appearing in the text to be detected.
[0081] Further, the sensitive word detection device 400 further includes a third detection module. After determining the first similarity corresponding to the first word segmentation unit, the third detection module is used for: Determine a second sensitive word from the target sensitive word tree, where the first similarity corresponding to the first word segmentation unit is greater than or equal to the first similarity threshold, and the sensitive word category belongs to non-essential filtering; Based on the first preset word segmentation length, determine a third word segmentation unit corresponding to the second sensitive word from the text to be detected, and input the third word segmentation unit into a pre-trained sensitive value detection model to determine the third similarity corresponding to the third word segmentation unit, so as to judge whether to regard the first word segmentation unit as a sensitive word appearing in the text to be detected based on the third similarity, or judge whether to regard the first word segmentation unit as a sensitive word appearing in the text to be detected based on the first similarity and the third similarity.
[0082] Further, the sensitive word detection device 400 further includes a fourth detection module. After determining the first similarity corresponding to the first word segmentation unit, the fourth detection module is configured to: Determine a second sensitive word from the target sensitive word tree, where the first similarity corresponding to the first word segmentation unit is greater than or equal to the first similarity threshold, and the sensitive word category belongs to non-essential filtering; Determine the exclusion weight corresponding to the second sensitive word, and use the product of the exclusion weight and the first similarity as the fourth similarity, so as to judge whether to regard the first word segmentation unit as a sensitive word appearing in the text to be detected based on the fourth similarity.
[0083] Further, the sensitive word detection device 400 further includes a fifth detection module. After determining the first similarity corresponding to the first word segmentation unit, the fifth detection module is configured to: Determine a third sensitive word from the target sensitive word tree, where the first similarity corresponding to the first word segmentation unit is greater than or equal to the second similarity threshold, and the sensitive word category belongs to a combined phrase; Based on the second preset word segmentation length, determine a fourth word segmentation unit corresponding to the third sensitive word from the text to be detected, and determine whether there is a sensitive word in the text to be detected based on the judgment result of whether the fourth word segmentation unit contains a fourth sensitive word.
[0084] When the sensitive word tree construction module 401 is used to classify multiple preset sensitive words and construct at least one sensitive word tree using the multiple preset sensitive words, the sensitive word tree construction module 401 is further configured to: Classify multiple preset sensitive words according to any one or more of pronunciation, glyph, and letters, and construct at least one sensitive word tree using the multiple preset sensitive words; When the splitting unit determination module 402 is used to split the obtained text to be detected and determine multiple splitting units, the splitting unit determination module 402 is further configured to: Split the obtained text to be detected according to any one or more of the pronunciation, glyph, and letters of each character in the text to be detected, and determine multiple splitting units.
[0085] Please refer to Figure 5 , Figure 5 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 5 shown in the figure, the electronic device 500 includes a processor 510, a memory 520, and a bus 530.
[0086] The memory 520 stores machine-readable instructions executable by the processor 510. When the electronic device 500 runs, the processor 510 communicates with the memory 520 through the bus 530. When the machine-readable instructions are executed by the processor 510, the steps of the sensitive word detection method in the method embodiment as described above can be executed. For the specific implementation manner, reference can be made to the method embodiment, which will not be elaborated here. Figure 1 shown in the figure, the steps of the sensitive word detection method in the method embodiment as described above can be executed. For the specific implementation manner, reference can be made to the method embodiment, which will not be elaborated here.
[0087] An embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the sensitive word detection method in the method embodiment as described above can be executed. For the specific implementation manner, reference can be made to the method embodiment, which will not be elaborated here. Figure 1 shown in the figure, the steps of the sensitive word detection method in the method embodiment as described above can be executed. For the specific implementation manner, reference can be made to the method embodiment, which will not be elaborated here.
[0088] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated here.
[0089] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0090] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0091] In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, may exist physically alone for each unit, or two or more units may be integrated in one unit.
[0092] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0093] Finally, it should be noted that: the above-described embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed in the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A sensitive word detection method, characterized in that: The sensitive word detection method comprises: Classifying a plurality of preset sensitive words, and constructing at least one sensitive word tree using the plurality of preset sensitive words; Split the acquired text to be detected into multiple split units; For each target sensitive word tree whose root node is the same as the splitting unit, determine the length of each branch link in the target sensitive word tree, and determine at least one first word segmentation unit from the multiple splitting units based on the length of the branch link; For each first word segmentation unit, a node in the target sensitive word tree is used to compare with the first word segmentation unit to determine a first similarity corresponding to the first word segmentation unit, and based on the first similarity, it is determined whether there is a sensitive word in the text to be detected.
2. The sensitive word detection method according to claim 1, characterized in that: The method of classifying a plurality of preset sensitive words and constructing at least one sensitive word tree using the plurality of preset sensitive words includes: Determine at least one sensitive word group from a plurality of preset sensitive words; wherein the first characters of each preset sensitive word in the sensitive word group have the same element; For each sensitive phrase, the element corresponding to the first character of the preset sensitive word in the sensitive phrase is taken as the root node, and corresponding child nodes are constructed for each root node according to the order of the characters contained in the preset sensitive word in the sensitive phrase to generate a sensitive word tree corresponding to the sensitive phrase.
3. The sensitive word detection method according to claim 1, characterized in that: The determining at least one first word segmentation unit from a plurality of splitting units based on the length of the branch link comprises: Determine at least one target splitting unit having the same elements as the root node of the target sensitive word tree from the multiple splitting units; For each target splitting unit, in the order of multiple splitting units in the text to be detected, multiple splitting units corresponding to the length of at least one branch link of the target sensitive word tree are obtained forward and / or backward with the target splitting unit as the starting point, and then combined to obtain the first word segmentation unit.
4. The sensitive word detection method according to claim 1, characterized in that: The comparing the nodes in the target sensitive word tree with the first word segmentation unit to determine the first similarity corresponding to the first word segmentation unit includes: Compare the nodes of each branch link in the target sensitive word tree with the characters in the first word segmentation unit to determine the number of characters in the first word segmentation unit that have the same elements as each branch link node in the target sensitive word tree; The ratio between the number and the length of the corresponding branch link of the target sensitive word tree is used as the first similarity; or, For each branch link in the target sensitive word tree, the ratio between the number and the length of the branch link is used as the branch similarity; A first similarity corresponding to the first word segmentation unit is determined by using the multiple branch similarities.
5. The sensitive word detection method according to claim 3, characterized in that: After determining the first similarity corresponding to the first word segmentation unit, the sensitive word detection method further includes: If the first similarity is greater than or equal to a first similarity threshold, determining a first sensitive word corresponding to the first word segmentation unit from the target sensitive word tree, and determining whether there is an exemption phrase for the first sensitive word; If yes, based on the phrase length of the exempted phrase, in accordance with the order of multiple split units in the text to be detected, taking each target split unit as a starting point, forward and / or backward obtain multiple split units corresponding to the phrase length and combine them to obtain a second word segmentation unit; The exemption phrase is compared with the second word segmentation unit to determine a second similarity corresponding to the second word segmentation unit, so as to determine whether to use the first word segmentation unit as a sensitive word appearing in the text to be detected based on the second similarity and a second similarity threshold.
6. The sensitive word detection method according to claim 1, characterized in that: After determining the first similarity corresponding to the first word segmentation unit, the sensitive word detection method further includes: Determine from the target sensitive word tree a second sensitive word whose first similarity corresponding to the first word segmentation unit is greater than or equal to a first similarity threshold and whose sensitive word category belongs to non-necessary filtering; Based on a first preset word segmentation length, a third word segmentation unit corresponding to the second sensitive word is determined from the text to be detected, and the third word segmentation unit is input into a pre-trained sensitivity value detection model to determine a third similarity corresponding to the third word segmentation unit, so as to judge whether to use the first word segmentation unit as a sensitive word appearing in the text to be detected based on the third similarity, or to judge whether to use the first word segmentation unit as a sensitive word appearing in the text to be detected based on the first similarity and the third similarity.
7. The sensitive word detection method according to claim 1, characterized in that: After determining the first similarity corresponding to the first word segmentation unit, the sensitive word detection method further includes: Determine from the target sensitive word tree a second sensitive word whose first similarity corresponding to the first word segmentation unit is greater than or equal to a first similarity threshold and whose sensitive word category belongs to non-necessary filtering; An exclusion weight corresponding to the second sensitive word is determined, and the product of the exclusion weight and the first similarity is used as a fourth similarity, so as to determine whether to use the first word segmentation unit as a sensitive word appearing in the text to be detected based on the fourth similarity.
8. The sensitive word detection method according to claim 1, characterized in that: After determining the first similarity corresponding to the first word segmentation unit, the sensitive word detection method further includes: Determine from the target sensitive word tree a third sensitive word whose first similarity corresponding to the first word segmentation unit is greater than or equal to a second similarity threshold and whose sensitive word category belongs to a combined word group; Based on the second preset word segmentation length, a fourth word segmentation unit corresponding to the third sensitive word is determined from the text to be detected, and based on the judgment result of whether the fourth word segmentation unit contains the fourth sensitive word, it is determined whether there is a sensitive word in the text to be detected.
9. The sensitive word detection method according to claim 1, characterized in that: The method of classifying a plurality of preset sensitive words and constructing at least one sensitive word tree using the plurality of preset sensitive words comprises: Classify multiple preset sensitive words according to any one or more of pronunciation, character shape, and letter, and construct at least one sensitive word tree using the multiple preset sensitive words; The step of splitting the acquired text to be detected to determine a plurality of split units includes: According to any one or more of the pronunciation, shape, and letter of each character in the text to be detected, the acquired text to be detected is split to determine a plurality of split units.
10. A sensitive word detection device, characterized in that: The sensitive word detection device comprises: A sensitive word tree construction module is used to classify multiple preset sensitive words and construct at least one sensitive word tree using the multiple preset sensitive words; A split unit determination module is used to split the acquired text to be detected and determine multiple split units; A word segmentation unit determination module, configured to determine, for each target sensitive word tree whose root node is the same as the splitting unit, the length of each branch link in the target sensitive word tree, and determine at least one first word segmentation unit from the plurality of splitting units based on the length of the branch link; The first detection module is used to compare each first word segmentation unit with the node in the target sensitive word tree, determine the first similarity corresponding to the first word segmentation unit, and determine whether there is a sensitive word in the text to be detected based on the first similarity.
Citation Information
Patent Citations
Sensitive word detection method and device and sensitive word tree construction method and device
CN112328732A
Sensitive word detection method and device, computer equipment, storage medium and product
CN115391524A
Sensitive word detection method and device, equipment and storage medium
CN118069809A
Text classification method, electronic device and computer-readable storage medium
US20230015054A1
Cited By
Word library construction method and word query method based on double-array tree
CN121166940A