Keyword extraction method, apparatus, device and computer-readable medium
By creating an index for the target text and recording the character mapping relationship, the problem of inaccurate keywords being able to be accurately extracted and blocked in the existing technology is solved, and the effect of accurately extracting and blocking keywords in the original text is achieved.
Patent Information
- Application Number
- CN202111671895.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-31
AI Technical Summary
The prior art cannot accurately extract keywords in the original text for blocking, resulting in only the entire sentence being blocked.
By creating an index for each character in the target text, recording the character mapping relationship before and after each transformation, and matching with the preset keyword database after each transformation, determine the corresponding characters of the hit word in the original text.
It realizes precisely extracting keywords in the original text and blocking them, retaining effective information and improving user experience.
Smart Images

Figure CN116431754B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing technology, and in particular to a keyword extraction method, apparatus, device, and computer-readable medium. Background Art
[0002] With the continuous development of online games, social features have become indispensable. Some games incorporate in-game social systems, allowing players to communicate directly within the game; others foster communication through forums and videos. These communication channels provide a low-cost platform for studios, influencers, and other advertisers to reach audiences. However, some individuals with ulterior motives exploit this vulnerability, using shocking language to incite uninformed players, inciting them to boycott or denigrate the game. Some competitors also maliciously incite players to post provocative and abusive comments within the game or on external forums, ultimately disrupting the gaming environment and attempting to stifle successful games in their early stages. Therefore, reviewing game-related content and issuing warnings or penalties to those who post inappropriate content are essential components of game operations.
[0003] At present, in the related technologies, a combination of keyword blocking and machine learning text classification is usually used to conduct game content review. Keyword blocking is to use the original text or a variant of the original text to match an established keyword library to directly determine whether the content has violated the rules. Machine learning text classification uses machine learning methods, such as text classification methods, to classify text into normal text and various abnormal texts through a trained model to determine whether the text may violate the rules. However, when matching keywords, there are cases where the position of the hit part is significantly different from the position of the original text. In this case, the entire sentence can often only be blocked. Similarly, the machine learning method only has classification results and corresponding confidence levels for text recognition. In this way, sentences classified as spam text can only be completely blocked. Therefore, the related technologies are unable to accurately extract keywords from the original text for blocking processing. They can only determine whether there are keywords in the sentence and block the entire sentence.
[0004] Currently, there is no effective solution to the problem of being unable to accurately extract keywords from the original text for blocking. Summary of the Invention
[0005] The present application provides a keyword extraction method, apparatus, device and computer-readable medium to solve the technical problem of being unable to accurately extract keywords from the original text for shielding processing.
[0006] According to one aspect of the embodiments of the present application, the present application provides a keyword extraction method, comprising:
[0007] Create an index for each character in the target text, where the target text is the text corresponding to the game player;
[0008] Transform the characters in the target text at least once, record the character mapping relationship before and after each transformation through an index mapping table, and match the transformation result with a preset keyword library after each transformation;
[0009] When the matching result indicates that the transformed result hits a keyword, the corresponding characters of the hit word in the transformed result in the original target text are determined based on the index mapping table.
[0010] Optionally, after determining based on the index mapping table that the hit word in the conversion result is after the corresponding character in the original target text, the method further includes:
[0011] Replace the corresponding characters with the target characters, and display the replaced target text on the game platform.
[0012] Optionally, the transformation operation includes at least one of the following: character deletion, one-to-one character replacement, one-to-many character replacement, many-to-one character replacement, many-to-many character replacement, pinyin replacement, and regular expression replacement.
[0013] Optionally, recording the character mapping relationship before and after each transformation through an index mapping table includes:
[0014] Determine the character / substring corresponding to each character in the target text after transformation, and arrange all the corresponding characters / substrings after transformation in the order of the corresponding characters in the target text before transformation to obtain a transformation map, wherein the transformation map records the mapping relationship between the index of each character before transformation and the index of the corresponding character / substring after transformation;
[0015] Delete the empty characters in the transformation map and expand the remaining characters to obtain an expanded map, wherein the expanded map records the mapping relationship between each character / substring index before expansion and the index of the corresponding character after expansion;
[0016] Merge the transformation mapping and the expansion mapping to obtain a forward mapping table and a reverse mapping table, wherein the forward mapping table is used to represent the mapping relationship from the index of the character before the transformation to the index of the character after the transformation, and the reverse mapping table is used to represent the mapping relationship from the index of the character after the transformation to the index of the character before the transformation. The index mapping table includes the forward mapping table and the reverse mapping table.
[0017] Optionally, after recording the character mapping relationship before and after each transformation by using the index mapping table, the method further includes:
[0018] Merge the forward mapping tables obtained from each transformation in the transformation order to obtain a forward chain mapping table;
[0019] The reverse mapping tables obtained from each transformation are merged in the reverse order of the transformation to obtain a reverse chain mapping table.
[0020] Optionally, when the matching result indicates that the transformed result hits a keyword, determining the corresponding characters of the hit word in the transformed result in the original target text based on the index mapping table includes:
[0021] Determine the starting point of each character in the hit word in the target text based on the reverse chain mapping table;
[0022] Determine whether other characters between two adjacent starting characters are mapped to the hit word based on the forward chain mapping table, wherein the determination range of the last starting character is the remaining characters starting from the last starting character in the target text;
[0023] The last character mapped to the hit word between two adjacent starting characters is determined as the ending character corresponding to the current starting character;
[0024] Each character in the hit word is mapped to all characters from the starting character to the ending character in the target text original text to determine as the corresponding character of the hit word in the target text original text.
[0025] Optionally, when the number of transformations reaches a threshold and all transformation results do not hit the keyword, the method further includes:
[0026] Input the target text marked with the index into the target neural network model, and use the network structure of each layer of the target neural network model to extract the feature vector of the target text step by step, and record the feature mapping relationship before and after feature extraction of each layer structure through the index mapping table;
[0027] When the target neural network model determines that the target text classification hits the keyword based on the feature vector, the corresponding characters in the target text original text whose contribution to the classification result is greater than a preset threshold are determined based on the index mapping table.
[0028] Optionally, determining, based on the index mapping table, corresponding characters in the target text whose contribution to the classification result is greater than a preset threshold includes:
[0029] Get the output of the fully connected layer of the target neural network model;
[0030] Determine from the output results the target feature dimension whose contribution to the classification result is greater than a preset threshold;
[0031] Based on the reverse chain mapping table in the index mapping table, the feature elements in the feature vector of the target feature dimension before being processed by each layer of the network are reversed step by step until the first target feature element corresponding to the target feature dimension output by the embedding layer is obtained;
[0032] According to the feature mapping relationship between each character in the target text and the output feature vector of the embedding layer recorded in the index mapping table, the corresponding character of the first target feature element in the original target text is determined.
[0033] Optionally, the method further includes:
[0034] Determine the gradient of the output of the fully connected layer on each feature element output by the embedding layer;
[0035] Determine a feature element whose gradient is greater than or equal to a preset gradient threshold as a second target feature element;
[0036] Determine the corresponding character of the second target feature element in the original target text according to the feature mapping relationship between each character in the target text and the feature vector output by the embedding layer recorded in the index mapping table;
[0037] The corresponding character determined by the first target feature element and the corresponding character determined by the second target feature element are combined to obtain a finally determined corresponding character.
[0038] Optionally, before replacing the corresponding character with the target character, the method further includes verifying the corresponding character in at least one of the following ways:
[0039] Convert the target text into pinyin, and if the pinyin constituting the hit word is actually the pinyin of a non-corresponding text, determine that the hit fails, otherwise the hit succeeds;
[0040] The target text is segmented, and if the characters constituting the hit word are actually characters in multiple segmented words, it is determined that the hit fails, otherwise the hit succeeds.
[0041] Optionally, if multiple keywords are hit, the corresponding characters of the multiple keywords in the original target text are checked for overlap.
[0042] According to another aspect of the embodiments of the present application, the present application provides a keyword extraction device, comprising:
[0043] An index creation module, configured to create an index for each character in a target text, wherein the target text is a text corresponding to a game player;
[0044] A text transformation module is used to transform the characters in the target text at least once, record the character mapping relationship before and after each transformation through an index mapping table, and match the transformation result with a preset keyword library after each transformation;
[0045] The index reverse lookup module is used to determine the corresponding characters of the hit words in the transformation result in the original target text based on the index mapping table when the matching result indicates that the transformation result hits the keyword.
[0046] According to another aspect of an embodiment of the present application, the present application provides an electronic device, including a memory, a processor, a communication interface and a communication bus, wherein the memory stores a computer program that can be run on the processor, the memory and the processor communicate through the communication bus and the communication interface, and the steps of the above method are implemented when the processor executes the computer program.
[0047] According to another aspect of an embodiment of the present application, the present application further provides a computer-readable medium having a non-volatile program code executable by a processor, where the program code enables the processor to execute the above method.
[0048] The above technical solution provided by the embodiment of the present application has the following advantages compared with the related art:
[0049] The technical solution of this application is to create an index for each character in the target text, wherein the target text is the text corresponding to the game player; transform the characters in the target text at least once, record the character mapping relationship before and after each transformation through an index mapping table, and match the transformation result with a preset keyword library after each transformation; when the matching result indicates that the transformation result hits the keyword, determine the corresponding character of the hit word in the transformation result in the original target text based on the index mapping table. This application uses an index to record the mapping relationship between text characters before and after text transformation, accurately recording the changes in characters during each transformation, so that when a keyword is hit, the corresponding character of the hit word in the original text can be reversely checked based on the index mapping table, and then only the corresponding character of the hit word in the original text needs to be replaced. The final displayed text can retain valid information, solving the technical problem of being unable to accurately extract keywords from the original text for shielding processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0052] Figure 1 A schematic diagram of a hardware environment for an optional keyword extraction method provided according to an embodiment of the present application;
[0053] Figure 2 A flowchart of an optional keyword extraction method provided according to an embodiment of the present application;
[0054] Figure 3 This is a schematic diagram of an optional character replacement provided according to an embodiment of the present application;
[0055] Figure 4 This is a block diagram of an optional keyword extraction device provided according to an embodiment of the present application;
[0056] Figure 5 A schematic diagram of an optional electronic device structure provided in an embodiment of the present application. DETAILED DESCRIPTION
[0057] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0058] In the subsequent description, the suffixes such as "module", "component" or "unit" used to represent elements are only used to facilitate the description of this application and have no specific meaning. Therefore, "module" and "component" can be used interchangeably.
[0059] In the related art, a combination of keyword blocking and machine learning text classification is usually used to conduct game content review. Keyword blocking is to use the original text or a variant of the original text to match the established keyword library to directly determine whether the content has violated the rules. Machine learning text classification uses machine learning methods, such as text classification methods, to classify text into normal text and various abnormal texts through a trained model to determine whether the text may violate the rules. However, when matching keywords, there are cases where the position of the hit part is significantly different from the position of the original text. In this case, the entire sentence can often only be blocked. Similarly, the machine learning method only has classification results and corresponding confidence levels for text recognition. In this way, sentences classified as spam text can only be completely blocked. Therefore, the related art cannot accurately extract keywords from the original text for blocking processing. It can only determine the presence of keywords in the sentence and block the entire sentence.
[0060] In order to solve the problems mentioned in the background technology, according to one aspect of the embodiments of the present application, an embodiment of a keyword extraction method is provided.
[0061] Optionally, in the embodiment of the present application, the above keyword extraction method can be applied to Figure 1 In the hardware environment composed of the terminal 101 and the server 103 shown in FIG. Figure 1 As shown, the server 103 is connected to the terminal 101 via a network, and can be used to provide services (such as character conversion, keyword matching, machine learning, index mapping construction services, etc.) for the terminal or the client installed on the terminal. A database 105 can be set on the server or independently of the server to provide data storage services for the server 103. The above-mentioned network includes but is not limited to: wide area network, metropolitan area network or local area network, and the terminal 101 includes but is not limited to PC, mobile phone, tablet computer, etc.
[0062] A keyword extraction method in the embodiment of the present application can be executed by the server 103, or can be executed by the server 103 and the terminal 101 together, such as Figure 2 As shown, the method may include the following steps:
[0063] Step S202: Create an index for each character in the target text, where the target text is the text corresponding to the game player.
[0064] Social interaction has become indispensable in online games today. Some games have built-in social systems that allow players to communicate directly within the game, while others use forums, videos, and other methods to establish communication channels between players.
[0065] The keyword extraction method provided in this embodiment can be applied to gaming platforms such as online games and game-related forums. The target text is speech text from players participating in online games, or speech text posted on game forums or game videos. For example, if the target text is abcdefgh, the index created for each character in the target text can be: a-0, b-1, c-2, d-3, e-4, f-5, g-6, h-7.
[0066] In games and forums, some studios and product promotion teams engage in traffic-driving speeches in an attempt to gain audiences at a low cost. Some individuals with ulterior motives attempt to incite uninformed players through shocking language, posting provocative and abusive comments in-game or on forums in an attempt to disrupt the gaming environment. To curb this behavior, this application requires reviewing player comments within gaming platforms, such as games and forums, and blocking any offending keywords. In addition, offending keywords can include uncivilized language and politically charged terms.
[0067] Step S204 : transform the characters in the target text at least once, record the character mapping relationship before and after each transformation through an index mapping table, and match the transformation result with a preset keyword library after each transformation.
[0068] Studios, sales teams, and some people with ulterior motives may use variations of their statements to avoid keywords.
[0069] In the embodiment of the present application, keyword matching can adopt a text matching algorithm. The difficulty of keyword matching lies in the identification of variants. For example, the possible variants of adding WeChat may be +WeiXin, JiaWangXin, etc. Therefore, the target text needs to be transformed multiple times to find the corresponding keywords as much as possible. Among them, since each transformation design changes the character form, such as one character is transformed into multiple characters, or multiple characters are transformed into one character, or even from Chinese characters to English characters, pinyin characters, etc., after multiple transformations, it is difficult to find the part corresponding to the hit word in the original text when hitting the keyword. Therefore, the present application records the character mapping relationship before and after each transformation through an index mapping table.
[0070] Step S206 : when the matching result indicates that the transformed result hits the keyword, the corresponding characters of the hit word in the transformed result in the original target text are determined based on the index mapping table.
[0071] In an embodiment of the present application, by querying the mapping relationship of the current character in the index mapping table, the character form of the current character before the last transformation can be found. By continuing to search for the mapping relationship of the character before the last transformation, the character form before the last two transformations can be found. Based on this, the corresponding character of the hit word can finally be found in the original text of the target text.
[0072] Optionally, after determining the corresponding characters of the hit words in the transformation result in the original target text based on the index mapping table, the method further includes replacing the corresponding characters with target characters and displaying the replaced target text on the game platform.
[0073] In the embodiment of the present application, since the corresponding characters associated with the transformed hit word in the original target text are located, it is only necessary to replace the corresponding characters to modify the target text into a compliant text, and then the modified target text can be displayed on the game platform, without having to block the entire target text containing the variant of the illegal keyword. The output of valid information is retained for other players and users, and only the illegal part is blocked. For example, in the sentence "AA, don't come over! There are enemies ahead!!", after character transformation processing, the variant "aa" of "AA" is determined to be the hit word of the keyword, and then only "AA" needs to be blocked in the original text, that is, replaced with the target character, such as "**". The final output text is "**, don't come over! There are enemies ahead!!" instead of "**********************!". In this way, the output part blocks the illegal part and retains the valid information, greatly improving the user experience.
[0074] In addition, after blocking the keywords in the corresponding text of the game player, the game player can also be punished, such as restricted from speaking. Specifically, the game player can be prohibited from speaking for a period of time according to the category of the blocked keywords. The severity of the situation can be judged according to the category, and the restriction time can be determined according to the severity of the situation. If the situation is particularly serious, the game player can be banned.
[0075] Through steps S202 to S206, the present application records the mapping relationship between text characters before and after text transformation through an index, accurately records the changes in characters during each transformation, so that when a keyword is hit, the corresponding characters of the hit word in the original text can be reversely checked based on the index mapping table, and then only the corresponding characters of the hit word need to be replaced in the original text. The final displayed text can retain valid information, solving the technical problem of being unable to accurately extract keywords from the original text for shielding processing.
[0076] Optionally, the transformation operation includes at least one of the following: character deletion, one-to-one character replacement, one-to-many character replacement, many-to-one character replacement, many-to-many character replacement, pinyin replacement, and regular expression replacement.
[0077] In the embodiment of this application, Figure 3 As shown, the transformation operation includes the following cases:
[0078] (1) The character does not involve transformation, such as a in the figure;
[0079] (2) One-to-one character replacement, where c is replaced by k in the figure, representing a character replacement operation;
[0080] (3) Character deletion, such as the character b in the figure, which represents the character deletion operation;
[0081] (4) One-to-many character replacement, such as d is replaced by dd, and many-to-many character replacement, such as ef is replaced by III, representing operations such as pinyin conversion and regular expression replacement;
[0082] (5) Many-to-one character replacement, that is, character strings are merged into one, such as g and h are merged into g, which represents a word segmentation operation.
[0083] Optionally, recording the character mapping relationship before and after each transformation through an index mapping table includes:
[0084] Step 1: Determine the character / substring corresponding to each character in the target text after transformation, and arrange all the corresponding characters / substrings after transformation in the order of the corresponding characters in the target text before transformation to obtain a transformation map, wherein the transformation map records the mapping relationship between the index of each character before transformation and the index of the corresponding character / substring after transformation;
[0085] Step 2: Delete the empty characters in the transformation map and expand the remaining characters to obtain an expanded map, wherein the expanded map records the mapping relationship between the index of each character / substring before expansion and the index of the corresponding character after expansion;
[0086] Step 3, merge the transformation mapping and the expansion mapping to obtain a forward mapping table and a reverse mapping table, wherein the forward mapping table is used to represent the mapping relationship from the index of the character before the transformation to the index of the character after the transformation, and the reverse mapping table is used to represent the mapping relationship from the index of the character after the transformation to the index of the character before the transformation, and the index mapping table includes the forward mapping table and the reverse mapping table.
[0087] In the embodiment of this application, Figure 3 As shown, the first line of character strings is the character form before transformation, the second line of character strings is the character form before transformation and before expansion, and the third line is the character form after expansion output.
[0088] At the beginning of the transformation, each part of the string is transformed, and substring replacement is performed as a whole. With this process, characters a and c still correspond to 1 character after the transformation, with indices 0<->0 and 2<->2 respectively; character d corresponds to dd after the transformation, with indices 3<->3; e and f correspond to III after the transformation, with indices 4<->4 and 5<->4 respectively; g and h correspond to g after the transformation, with indices 6<->6 and 7<->6 respectively; and b corresponds to nothing, with the index mapping set to a one-way 1->-1. This is the transformation mapping.
[0089] After establishing the transformation mapping, expand the transformed portion with a length greater than 1 and delete the portion with a length of 0. The resulting mapping relationship between the transformed portion and the expanded output is 0 <-> 0, 2 <-> 1, 3 <-> (2, 3), 4 <-> (4, 5, 6), and 6 <-> 7. This is the expanded mapping.
[0090] Merging the transformed mapping with the expanded mapping yields the forward mapping tables 0->0, 1->-1, 2->1, 3->2, 4->4, 5->4, 6->7, and 7->7, and the reverse mapping tables 0->0, 1->2, 2->3, 3->3, 4->4, 5->4, 6->4, and 7->6. The mapping positions are all the starting points of the corresponding character positions.
[0091] In an embodiment of the present application, if the length of the original string is n and the length of the transformed string is m, during the transformation process, the index operation is only performed on the transformed part, so the time complexity does not exceed O(n); when expanding, the expanded characters need to be traversed, so the time complexity is O(m); considering that the time complexity of reading in characters is O(n), the time complexity when establishing the mapping relationship is O(max(n,m)).
[0092] Optionally, after recording the character mapping relationship before and after each transformation by using the index mapping table, the method further includes:
[0093] Merge the forward mapping tables obtained from each transformation in the transformation order to obtain a forward chain mapping table;
[0094] The reverse mapping tables obtained from each transformation are merged in the reverse order of the transformation to obtain a reverse chain mapping table.
[0095] In the above process, after each step, the operations are directly merged according to the chain rule, with a complexity of O(n). When the keyword is finally reverse-looked, the merged forward chain mapping table and reverse chain mapping table are directly used, so the time complexity is O(n).
[0096] In the embodiment of the present application, the mapping relationship of each transformation needs to be merged with the existing mapping relationship. The forward mapping tables are merged in the order of transformation to obtain a forward chain mapping table, and the reverse mapping tables are merged in the reverse order of transformation to obtain a reverse chain mapping table. Taking two adjacent transformations as an example, if the previous transformation obtains the forward mapping table f1 and the reverse mapping table b1, and the next transformation obtains the forward mapping table f2 and the reverse mapping table b2, then based on the principle that the mapping position is the starting point of the corresponding character position, the forward chain mapping table f(x) = f2(f1(x)) and the reverse chain mapping table b(x) = b1(b2(x)) are obtained.
[0097] Optionally, when the matching result indicates that the transformed result hits a keyword, determining the corresponding characters of the hit word in the transformed result in the original target text based on the index mapping table includes:
[0098] Step 1: determining the starting point of each character in the hit word in the target text based on the reverse chain mapping table;
[0099] Step 2: Determine whether other characters between two adjacent starting characters are mapped to the hit word based on the forward chain mapping table, wherein the determination range of the last starting character is the remaining characters starting from the last starting character in the target text;
[0100] Step 3: Determine the last character between two adjacent starting characters that is mapped to the hit word as the ending character corresponding to the current starting character;
[0101] Step 4: Map each character in the hit word to all characters from the starting character to the ending character in the target text to determine them as corresponding characters of the hit word in the target text.
[0102] In the embodiment of the present application, after performing multiple transformations, a sentence hits certain keywords, and the hit results need to be reverse-checked in the original text to determine the characters related to the keywords in the original text. First, the starting point of each character in the hit portion in the original text can be found through the reverse mapping table. Secondly, based on the starting point, the following part is queried in the forward mapping table to see if it maps to the hit portion, and the end point is found. Based on the starting point and end point, the form of the hit portion in the original text can be obtained. In the most extreme case, the algorithm complexity of this part is O(m+n).
[0103] In the embodiments of the present application, both the forward and reverse mapping tables are linear structures that can be directly stored using arrays. Combined with chain processing logic, this ensures that in actual implementation, there is no interference with the highly modularized transformations. The combination of the forward and reverse mapping tables not only enables reverse lookups in most scenarios but also plays an important role in various verification processes. The combination of the two also reduces average complexity.
[0104] The key point of keyword matching is the number of words in the preset keyword library. The more words there are, the larger the search range. However, in the face of endless variants, it is difficult to cover the increasing number of variants by relying solely on the keyword library for matching. Therefore, in addition to the direct keyword matching hit method described above, a model inference hit method can also be added, that is, a neural network model is trained through machine learning and the neural network model is used to perform keyword recognition. However, the neural network model can only output text classification, that is, to determine whether the target text carries keywords, and cannot accurately determine the characters that cause violations. Therefore, based on index mapping, this application provides a method for keyword extraction using a neural network model, so that the neural network model can output text classification while also outputting one or more characters in the original text that have the greatest impact on the violation judgment. This method can be used as a supplement to the above-mentioned keyword matching. When the number of text transformations reaches a threshold and all transformation results do not hit the keyword, the neural network model is used to extract keywords. The method is described in detail below.
[0105] Optionally, when the number of transformations reaches a threshold and all transformation results do not hit the keyword, the method further includes:
[0106] Step 1: Input the target text marked with the index into the target neural network model, and use the network structure of each layer of the target neural network model to extract the feature vector of the target text step by step, and record the feature mapping relationship before and after feature extraction of each layer structure through the index mapping table;
[0107] Step 2: When the target neural network model determines that the target text classification hits the keyword based on the feature vector, the corresponding characters in the target text original text whose contribution to the classification result is greater than a preset threshold are determined based on the index mapping table.
[0108] Optionally, determining, based on the index mapping table, corresponding characters in the target text whose contribution to the classification result is greater than a preset threshold includes:
[0109] Step 1: Obtain the output result of the fully connected layer of the target neural network model;
[0110] Step 2: Determine the target feature dimension whose contribution to the classification result is greater than a preset threshold from the output results;
[0111] Step 3: Inversely deduce the feature elements in the feature vector of the target feature dimension before being processed by each layer of the network based on the reverse chain mapping table in the index mapping table, until the first target feature element corresponding to the target feature dimension output by the embedding layer is obtained;
[0112] Step 4: Determine the corresponding character of the first target feature element in the original target text according to the feature mapping relationship between each character in the target text and the output feature vector of the embedding layer recorded in the index mapping table.
[0113] In an embodiment of the present application, the target text with an index is input into the target neural network model, the target text is converted into a feature vector by the embedding layer, and then the feature vector is subjected to feature extraction layer by layer through each layer of the network structure in the target neural network model, and finally a feature vector with deep features of the target text is obtained. The fully connected layer then performs feature recognition on the feature vector to finally determine the classification of the target text.
[0114] Before and after the embedding layer performs feature vector conversion, the mapping relationship between each character of the target text and the feature elements (i.e., vector elements) in the feature vector is recorded in an index mapping table. This establishes a mapping relationship between the index of the converted feature and the index of the character, allowing the feature elements in the feature vector to be traced back to the characters in the original target text. Similarly, the feature elements of each layer of the target neural network model before and after feature extraction are also mapped through indices, allowing the feature elements after deep extraction to be traced back to the feature elements before extraction.
[0115] Accordingly, when the target neural network model classifies the target text as masked text, we can first determine the feature dimension in the output of the fully connected layer whose contribution to the classification result is greater than a preset threshold, and then, based on the index mapping table, find the corresponding characters of the features under this feature dimension in the original target text through feature reverse search.
[0116] In the embodiment of the present application, index back-checking and gradient calculation can also be combined to improve the accuracy of keyword extraction.
[0117] Optionally, the method further includes:
[0118] Step 1: Determine the gradient of the output of the fully connected layer on each feature element output by the embedding layer;
[0119] Step 2: determining a feature element whose gradient is greater than or equal to a preset gradient threshold as a second target feature element;
[0120] Step 3: Determine the corresponding character of the second target feature element in the original target text according to the feature mapping relationship between each character in the target text and the feature vector output by the embedding layer recorded in the index mapping table;
[0121] Step 4 combines the corresponding character determined by the first target feature element and the corresponding character determined by the second target feature element to obtain a final corresponding character.
[0122] In this embodiment of the application, the connection between the fully connected layer and the word embedding can be used to find the index position where the gradient calculation has the greatest impact on the result. Suppose the embedding layer output is
[0123]
[0124] where X i is the feature of each word after being converted into a word vector, and n is the sentence length. Let the output of the fully connected layer be W, and a certain input be X0. Then, linearize it according to Taylor's formula around the input, and we get
[0125]
[0126] Here b is a constant. Find the term that makes W(X0) as large as possible, that is, find Items exceeding the threshold. Output the i that exceeds the threshold, which is the hit part.
[0127] The results of index reverse lookup and gradient calculation are combined based on the normalized confidence sum to obtain the final index result. For example, "Are you there? I'm the gang leader. You can add me to the gang on WeChat." After index reverse lookup, the result can be "**, * ...
[0128] In an embodiment of the present application, the target neural network model can be deployed on tf-serving, and the result after word segmentation of the example sentence "Are you there? I am the gang leader. You WeChat me + you join the gang skirt." is "Are you there? I am the gang leader. You WeChat me + you join the gang skirt." The output classification result is "advertisement" and a score of 0.99999893, and the index information [0,0,0,0,1,0,1,0,1,0,1] after merging the two methods is attached. Based on this index information, the word segmentation marked as "1" can be shielded by the target character such as "*" in the original text, and the final output result is "Are you there? I am **you**I*you***".
[0129] This application also provides a method for verifying the hit word to check whether the hit is successful.
[0130] Optionally, before replacing the corresponding character with the target character, the method further includes verifying the corresponding character in at least one of the following ways:
[0131] Convert the target text into pinyin, and if the pinyin constituting the hit word is actually the pinyin of a non-corresponding text, determine that the hit fails, otherwise the hit succeeds;
[0132] The target text is segmented, and if the characters constituting the hit word are actually characters in multiple segmented words, it is determined that the hit fails, otherwise the hit succeeds.
[0133] In the embodiments of the present application, verification can be performed through pinyin matching in combination with an index. For example, the phrase "add WeChat" is set as a keyword in some games. Generally, this phrase appears in some traffic diversion scenarios, and some malicious players deceive players into joining promotion groups by mass - sending "add WeChat" and accounts. However, there are many homophonic variants of "add WeChat", such as "add WeiChat", "add JiaWeiChat", "jia WeiChat", etc. Pinyin matching is a reasonable solution to solve such variants. But after adding "jiaweixin" to the keyword library, the sentence "deejia is the first in the southwest region" will also match the first part of "jia is in the southwest", that is, "jiaweixinan", where "deejia" is the player name, "southwest region" is the southwest part of a certain map in the game, and "the first" is the first place in the PVP gameplay in the game scene.
[0134] After adding index information, verification of pinyin matching can be performed according to the index of the matched part. For example, in the above example, according to the information in the positive and reverse mapping tables, it is found that the last letter "n" of the matched part "xin" is part of the pinyin "nan" of the next character "south", so it is considered that the match fails, and "deejia is in the southwest" should be a compliant text.
[0135] In the embodiments of the present application, verification can also be performed through word segmentation. For example, in an auto - chess game, there may be a sentence "Please give me another B迦", where "B迦" is a character name in the game. However, "BB" is homophonic with the inappropriate term "bb" and will be matched. But the word segmentation result of this sentence is "Please, again, give me, one B, B迦". Therefore, the matched "BB" is actually part of two words, so this text should be judged as compliant.
[0136] Optionally, if multiple keywords are matched by the above method, coincidence verification needs to be performed on the corresponding characters of the multiple keywords in the original target text.
[0137] In the embodiments of the present application, in the case of multiple keyword matches, coincidence verification needs to be performed to avoid the problem of repeated output where one matched keyword is included in another matched keyword. For such a situation where one matched keyword is included in another matched keyword, the matched results need to be merged before output. For example, "Click on my avatar to add WeChat 123456" may match "Click on my avatar to add WeChat" and "add WeChat". At this time, check the matching situation of the two keywords through the index system, and after interval merging, the output result is: "Click on my avatar to add WeChat".
[0138] In the embodiments of the present application, for advertisements that commonly send contact information in games, we added an interval matching function based on the indexing system, modified the minimum window subsequence algorithm, and under the complexity of O(nk), calculated whether a string of length n with a window size of k contains contact information with added interval characters. For example, the case where Chinese character "一" is added between "123456" and it is modified to "1一2一3一4一5一6" can be detected and restored to "123456", and then "1一2一3一4一5一6" is hit.
[0139] Optionally, after determining the corresponding characters of the keyword in the original text of the target text, the characters can also be masked or replaced.
[0140] The present application records the mapping relationship between text characters before and after text transformation through indexing, accurately records the character changes during each transformation, so that when a keyword is hit, the corresponding characters of the hit word in the original text can be retrieved based on the index mapping table. Furthermore, only the corresponding characters of the hit word need to be replaced in the original text, and the finally displayed text retains valid information, solving the technical problem of being unable to accurately extract keywords in the original text for masking processing.
[0141] According to another aspect of the embodiments of the present application, as Figure 4 shown, a keyword extraction device is provided, including:
[0142] An index creation module 401, configured to create an index for each character in the target text, where the target text is the text corresponding to the game player;
[0143] A text transformation module 403, configured to perform at least one transformation on the characters in the target text, record the character mapping relationship before and after each transformation through an index mapping table, and match the transformation result with a preset keyword library after each transformation;
[0144] An index reverse lookup module 405, configured to, when the matching result indicates that the transformation result hits a keyword, determine the corresponding characters of the hit word in the transformation result in the original text of the target text based on the index mapping table.
[0145] It should be noted that the index creation module 401 in this embodiment can be used to execute step S202 in the embodiments of the present application, the text transformation module 403 in this embodiment can be used to execute step S204 in the embodiments of the present application, and the index reverse lookup module 405 in this embodiment can be used to execute step S206 in the embodiments of the present application.
[0146] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiments. Figure 1 In the hardware environment shown, it can be implemented by software or by hardware.
[0147] Optionally, the text transformation module is used to:
[0148] Determine the character / substring corresponding to each character in the target text after transformation, and arrange all the corresponding characters / substrings after transformation in the order of the corresponding characters in the target text before transformation to obtain a transformation map, wherein the transformation map records the mapping relationship between the index of each character before transformation and the index of the corresponding character / substring after transformation;
[0149] Delete the empty characters in the transformation map and expand the remaining characters to obtain an expanded map, wherein the expanded map records the mapping relationship between each character / substring index before expansion and the index of the corresponding character after expansion;
[0150] Merge the transformation mapping and the expansion mapping to obtain a forward mapping table and a reverse mapping table, wherein the forward mapping table is used to represent the mapping relationship from the index of the character before the transformation to the index of the character after the transformation, and the reverse mapping table is used to represent the mapping relationship from the index of the character after the transformation to the index of the character before the transformation. The index mapping table includes the forward mapping table and the reverse mapping table.
[0151] Optionally, the text transformation module is further configured to:
[0152] Merge the forward mapping tables obtained from each transformation in the transformation order to obtain a forward chain mapping table;
[0153] The reverse mapping tables obtained from each transformation are merged in the reverse order of the transformation to obtain a reverse chain mapping table.
[0154] Optionally, the index reverse lookup module is used to:
[0155] Determine the starting point of each character in the hit word in the target text based on the reverse chain mapping table;
[0156] Determine whether other characters between two adjacent starting characters are mapped to the hit word based on the forward chain mapping table, wherein the determination range of the last starting character is the remaining characters starting from the last starting character in the target text;
[0157] The last character mapped to the hit word between two adjacent starting characters is determined as the ending character corresponding to the current starting character;
[0158] Each character in the hit word is mapped to all characters from the starting character to the ending character in the target text original text to determine as the corresponding character of the hit word in the target text original text.
[0159] Optionally, the keyword extraction device further includes a machine learning model processing module, which is used to:
[0160] When the number of transformations reaches the threshold and none of the transformation results hits the keyword, the target text marked with the index is input into the target neural network model, so as to extract the feature vector of the target text step by step by using the network structure of each layer of the target neural network model, and record the feature mapping relationship before and after feature extraction of each layer structure through the index mapping table;
[0161] When the target neural network model determines that the target text classification hits the keyword based on the feature vector, the corresponding character in the target text original text whose contribution to the classification result is greater than a preset threshold is determined based on the index mapping table;
[0162] Replace the corresponding characters with the target characters, and display the replaced target text on the game platform.
[0163] Optionally, the machine learning model processing module is further used to:
[0164] Get the output of the fully connected layer of the target neural network model;
[0165] Determine from the output results the target feature dimension whose contribution to the classification result is greater than a preset threshold;
[0166] Based on the reverse chain mapping table in the index mapping table, the feature elements in the feature vector of the target feature dimension before being processed by each layer of the network are reversed step by step until the first target feature element corresponding to the target feature dimension output by the embedding layer is obtained;
[0167] According to the feature mapping relationship between each character in the target text and the output feature vector of the embedding layer recorded in the index mapping table, the corresponding character of the first target feature element in the original target text is determined.
[0168] Optionally, the machine learning model processing module is further used to:
[0169] Determine the gradient of the output of the fully connected layer on each feature element output by the embedding layer;
[0170] Determine a feature element whose gradient is greater than or equal to a preset gradient threshold as a second target feature element;
[0171] Determine the corresponding character of the second target feature element in the original target text according to the feature mapping relationship between each character in the target text and the feature vector output by the embedding layer recorded in the index mapping table;
[0172] The corresponding character determined by the first target feature element and the corresponding character determined by the second target feature element are combined to obtain a finally determined corresponding character.
[0173] Optionally, the keyword extraction device further includes a verification module, which is used to:
[0174] Convert the target text into pinyin, and if the pinyin constituting the hit word is actually the pinyin of a non-corresponding text, determine that the hit fails, otherwise the hit succeeds;
[0175] The target text is segmented, and if the characters constituting the hit word are actually characters in multiple segmented words, it is determined that the hit fails, otherwise the hit succeeds.
[0176] Optionally, the verification module is further configured to: if multiple keywords are hit, perform a coincidence check on corresponding characters of the multiple keywords in the original target text.
[0177] According to another aspect of the embodiment of the present application, the present application provides an electronic device, such as Figure 5 As shown, it includes a memory 501, a processor 503, a communication interface 505 and a communication bus 507. The memory 501 stores a computer program that can be run on the processor 503. The memory 501 and the processor 503 communicate through the communication interface 505 and the communication bus 507. When the processor 503 executes the computer program, the steps of the above method are implemented.
[0178] The memory and processor in the electronic device communicate via a communication bus and a communication interface. The communication bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The communication bus may be divided into an address bus, a data bus, a control bus, and the like.
[0179] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.
[0180] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0181] According to another aspect of the embodiments of the present application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of any of the above embodiments.
[0182] Optionally, in an embodiment of the present application, the computer-readable medium is configured to store program codes for the processor to execute the following steps:
[0183] Create an index for each character in the target text, where the target text is the text corresponding to the game player;
[0184] Transform the characters in the target text at least once, record the character mapping relationship before and after each transformation through an index mapping table, and match the transformation result with a preset keyword library after each transformation;
[0185] When the matching result indicates that the transformed result hits a keyword, the corresponding characters of the hit word in the transformed result in the original target text are determined based on the index mapping table.
[0186] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and this embodiment will not be described in detail here.
[0187] When implementing the embodiments of the present application, reference may be made to the above embodiments, which have corresponding technical effects.
[0188] It is understood that the embodiments described herein may be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or a combination thereof.
[0189] For software implementation, the technology described herein can be implemented by a unit that performs the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0190] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0191] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0192] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0193] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0194] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0195] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application are essentially or partly contributed to the prior art or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard drive, a ROM, a RAM, a magnetic disk, or an optical disk. It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device. Without further constraints, an element defined by the phrase "comprises a..." does not preclude the existence of additional identical elements in the process, method, article or apparatus that includes the element.
[0196] The foregoing is merely a list of specific embodiments of the present application, intended to enable those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the broadest scope consistent with the principles and novel features of the present application.
Claims
1. A keyword extraction method, characterized in that, it includes: Create an index for each character in the target text, where the target text is the text corresponding to a game player; Perform at least one transformation on the characters in the target text, record the character mapping relationship before and after each transformation through an index mapping table, and match the transformation result with a preset keyword library after each transformation; When the matching result indicates that the transformation result hits a keyword, determine the corresponding character in the original target text of the hit word in the transformation result based on the index mapping table; The transformation operation includes at least one of the following: character deletion, one-to-one character replacement, one-to-many character replacement, many-to-one character replacement, many-to-many character replacement, pinyin replacement, and regular replacement; recording the character mapping relationship before and after each transformation through an index mapping table includes: Determine the corresponding character / substring after each character transformation in the target text, and arrange all the corresponding characters / substrings after transformation in the order of the corresponding characters in the target text before transformation to obtain a transformation mapping, where the transformation mapping records the mapping relationship from the index of each character before transformation to the index of the corresponding character / substring after transformation; Delete the empty characters in the transformation mapping, and expand the remaining characters / substrings to obtain an expansion mapping, where the expansion mapping records the mapping relationship from the index of each character / substring before expansion to the index of the corresponding character after expansion, and the mapping positions are all the starting points of the positions of the corresponding characters / substrings; Merge the transformation mapping and the expansion mapping to obtain a forward mapping table and a reverse mapping table, where the forward mapping table is used to represent the mapping relationship from the index of the character before transformation to the index of the character after transformation, the reverse mapping table is used to represent the mapping relationship from the index of the character after transformation to the index of the character before transformation, and the index mapping table includes the forward mapping table and the reverse mapping table.
2. The method according to claim 1, characterized in that, after recording the character mapping relationship before and after each transformation through the index mapping table, the method further includes: Merge the forward mapping tables obtained from each transformation in the order of transformation to obtain a forward chain mapping table; Merge the reverse mapping tables obtained from each transformation in the reverse order of transformation to obtain a reverse chain mapping table.
3. The method according to claim 2, characterized in that, when the matching result indicates that the transformation result hits a keyword, determining the corresponding character in the original target text of the hit word in the transformation result based on the index mapping table includes: Determine the starting character of each character in the hit word in the original target text based on the reverse chain mapping table; Based on the forward chain mapping table, determine whether the other characters between two adjacent starting characters are mapped to the hit word, where the judgment range of the last starting character is the remaining characters starting from the last starting character in the target text; Determine the ending character corresponding to the current starting character as the last character that is mapped to the hit word between two adjacent starting characters; Map each character in the hit word to all characters from the starting character to the ending character in the original target text, and determine them as the corresponding characters of the hit word in the original target text.
4. The method according to claim 1, wherein, when the number of transformations reaches the number threshold and all transformation results do not hit the keyword, the method further includes: Input the target text marked with indexes into the target neural network model to sequentially extract the feature vectors of the target text by each layer network structure of the target neural network model, and record the feature mapping relationship before and after feature extraction by each layer structure through the index mapping table; When the target neural network model determines that the target text classification hits the keyword according to the feature vector, determine the corresponding characters in the original target text whose contribution degree to the classification result is greater than the preset threshold based on the index mapping table.
5. The method according to claim 4, wherein, Determining the corresponding characters in the original target text whose contribution degree to the classification result is greater than the preset threshold based on the index mapping table includes: Obtain the output result of the fully connected layer of the target neural network model; Determine the target feature dimensions in the output result whose contribution degree to the classification result is greater than the preset threshold; Based on the reverse chain mapping table in the index mapping table, sequentially backtrack the feature elements in the feature vector before being processed by each layer network for the target feature dimension until obtaining the first target feature element corresponding to the target feature dimension output by the embedding layer; According to the feature mapping relationship between each character in the target text recorded in the index mapping table and the feature vector output by the embedding layer, determine the corresponding characters of the first target feature element in the original target text.
6. The method according to claim 5, wherein, The method further includes: Determine the gradient of the output result of the fully connected layer on each feature element output by the embedding layer; Determine the feature elements whose gradient is greater than or equal to the preset gradient threshold as the second target feature elements; According to the feature mapping relationship between each character in the target text recorded in the index mapping table and the feature vector output by the embedding layer, determine the corresponding characters of the second target feature elements in the original target text; Merge the corresponding characters determined by the first target feature element and the corresponding characters determined by the second target feature element to obtain the finally determined corresponding characters.
7. The method according to any one of claims 1 to 6, wherein, Before replacing the corresponding characters with the target characters, the method further includes verifying the corresponding characters in at least one of the following ways: Convert the target text into pinyin, and determine that the hit fails when the pinyin forming the hit word is actually not the pinyin of the corresponding text, otherwise the hit is successful; Segment the target text, and determine that the hit fails when the characters forming the hit word are actually characters in multiple segments, otherwise the hit is successful.
8. The method according to claim 7, wherein, The method further includes: If multiple keywords are hit, perform coincidence verification on the corresponding characters of the multiple keywords in the original text of the target text.
9. A keyword extraction device, characterized in that, it includes: An index creation module, configured to create an index for each character in the target text, where the target text is the text corresponding to a game player; A text transformation module, configured to perform at least one transformation on the characters in the target text, record the character mapping relationship before and after each transformation through an index mapping table, and match the transformation result with a preset keyword library after each transformation; An index reverse lookup module, configured to, when the matching result indicates that the transformation result hits a keyword, determine the corresponding characters of the hit word in the transformation result in the original text of the target text based on the index mapping table; The transformation operation includes at least one of the following: character deletion, one-to-one character replacement, one-to-many character replacement, many-to-one character replacement, many-to-many character replacement, pinyin replacement, and regular replacement; specifically, the text transformation module is configured to: Determine the corresponding character / substring after each character transformation in the target text, and arrange all the corresponding characters / substrings after transformation in the order of the corresponding characters in the target text before transformation to obtain a transformation mapping, where the transformation mapping records the mapping relationship from the index of each character before transformation to the index of the corresponding character / substring after transformation; Delete the empty characters in the transformation mapping, and expand the remaining characters / substrings to obtain an expansion mapping, where the expansion mapping records the mapping relationship from the index of each character / substring before expansion to the index of the corresponding character after expansion, and the positions of the mappings are all the starting points of the positions of the corresponding characters / substrings; Merge the transformation mapping and the expansion mapping to obtain a forward mapping table and a reverse mapping table, where the forward mapping table is used to represent the mapping relationship from the index of the character before transformation to the index of the character after transformation, the reverse mapping table is used to represent the mapping relationship from the index of the character after transformation to the index of the character before transformation, and the index mapping table includes the forward mapping table and the reverse mapping table.
10. An electronic device, including a memory, a processor, a communication interface, and a communication bus, where a computer program that can run on the processor is stored in the memory, and the memory and the processor communicate through the communication bus and the communication interface, characterized in that, when the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8 above.
11. A computer-readable medium having non-volatile program code executable by a processor, characterized in that, the program code causes the processor to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Hybrid text sensitive word variant recognition method and device
CN111259151A
Keyword extraction method and device based on neural network and electronic equipment
CN111611807A