Text processing method and apparatus, and electronic device
Patent Information
- Application Number
- CN202210741603.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-28
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-06-28
AI Technical Summary
[0003]相关技术中,通常使用深度模型对文本进行内容审核,以识别出垃圾文本,但是该模型的响应时间较长,不能满足对高查询量的文本进行内容审核的需求
[0075]根据本申请的第三方面,提供了一种电子设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,所述处理器执行所述程序时,实现上述第一方面所述的方法。
Smart Images

Figure CN117349411B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information processing technology, and in particular to a text processing method, apparatus and electronic device. Background Technology
[0002] With the development of the internet and the increase in internet users, people are becoming increasingly reliant on the internet for the dissemination of various information. To ensure the legality of information, all business modules need to conduct text content review.
[0003] In related technologies, deep learning models are often used to perform content moderation on text in order to identify spam text. However, the response time of this model is relatively long, which cannot meet the needs of content moderation for text with high query volume. Summary of the Invention
[0004] To address the aforementioned problems, this application provides a text processing method, apparatus, and electronic device.
[0005] According to a first aspect of this application, a text processing method is provided, comprising:
[0006] Get the text to be processed;
[0007] Determine the text length of the text to be processed;
[0008] If the length of the text to be processed is less than or equal to the short text length threshold, determine the N words that the text to be processed can form; where N is a positive integer;
[0009] Based on a pre-defined trie, keyword matching is performed on each of the words to determine i target positive keywords and j target verification keywords that match each of the words; wherein, the trie is constructed based on multiple positive keywords and at least one verification keyword corresponding to each positive keyword; i is a natural number and j is a natural number;
[0010] The category of the text to be processed is identified based on i target positive keywords and j target verification keywords that match each of the aforementioned words.
[0011] The step of identifying the category of the text to be processed based on the i target positive keywords and j target verification keywords that match each of the words includes:
[0012] Based on a preset keyword matching information table, each target positive keyword is matched with the j target verification keywords; wherein, the keyword matching information table includes the correspondence between all positive keywords and verification keywords in the trie;
[0013] If each of the target positive keywords matches the corresponding target verification keyword, determine the first position information of each of the target positive keywords in the text to be processed, and determine the second position information of the target verification keyword corresponding to each of the target positive keywords in the text to be processed.
[0014] Based on the first location information and the second location information, determine whether there is a location inclusion relationship between each target positive keyword and its corresponding target verification keyword;
[0015] If each target positive keyword and its corresponding target verification keyword have a positional inclusion relationship, the text to be processed is identified as legitimate text.
[0016] In some embodiments of this application, it also includes:
[0017] If there are target positive keywords that do not match the corresponding target verification keyword, or if each of the target positive keywords matches the corresponding target verification keyword and at least one of the target positive keywords does not have a positional inclusion relationship with the matched target verification keyword, then the target positive keywords to be matched with the template are obtained from the i target positive keywords; wherein, the target positive keywords to be matched with the template include target positive keywords that do not match the corresponding target verification keyword, and target positive keywords that do not have a positional inclusion relationship with the matched target verification keyword;
[0018] Obtain the third position information of the positive keyword of the target template in the text to be processed;
[0019] Based on a preset template information table, the target positive keywords for template matching, and the third position information, template matching is performed on the target positive keywords for template matching; the template information table contains the content information of the template corresponding to each template identifier;
[0020] Based on the template matching results, the category of the text to be processed is identified.
[0021] In some embodiments of this application, the keyword matching information table further includes a list of template identifiers corresponding to each positive keyword; the step of performing template matching on the positive keyword to be matched based on the preset template information table, the target positive keyword to be matched, and the third position information includes:
[0022] Based on the target positive keywords to be matched by the template, determine the template identifier list corresponding to the target positive keywords to be matched by the template in the template identifier list corresponding to each positive keyword;
[0023] Based on the list of template identifiers corresponding to the target positive keywords to be matched, at least one target template identifier is determined;
[0024] Based on the template information table, obtain the content information of the target template corresponding to each target template identifier;
[0025] Based on the target positive keywords to be matched, the third position information, and the content information of each target template, template matching is performed on the target positive keywords to be matched.
[0026] As one possible implementation, the content information of the template includes template keywords and the relative position information between the template keywords; the step of performing template matching on the target positive keywords based on the target positive keywords to be matched, the third position information, and the content information of each target template includes:
[0027] Based on the target positive keywords for matching the template to be matched and the template keywords in the content information of each target template, at least one first target template is determined from at least one target template; wherein, the template keywords in the content information of the first target template are exactly the same as the target positive keywords for matching the template to be matched;
[0028] Based on the third position information and the relative position information between template keywords in the content information of each first target template, it is determined whether there is a second target template in the at least one first target template that matches the second target positive keyword.
[0029] The step of identifying the category of the text to be processed based on the template matching result includes:
[0030] If the second target template exists in at least one of the first target templates, the text to be processed is identified as illegal text.
[0031] If the second target template is not present in at least one of the first target templates, the category of the text to be processed is identified by a preset statistical model.
[0032] In some embodiments of this application, after performing keyword matching on each word based on a preset trie, the method further includes:
[0033] If none of the stated words match the target positive keyword, the text to be processed is identified as legitimate text.
[0034] In some other embodiments of this application, before performing keyword matching on each word based on a preset trie to determine i target positive keywords and j target verification keywords that match each word, the method further includes:
[0035] Each of the aforementioned words is matched with keywords in a preset word list;
[0036] If the vocabulary contains keywords that match each of the words, the text to be processed is identified as illegal text.
[0037] If no keyword matches each of the words in the vocabulary, the step of performing keyword matching on each word based on a preset trie to determine i target positive keywords and j target verification keywords that match each word is executed.
[0038] In some other embodiments of this application, the method further includes:
[0039] If the length of the text to be processed is greater than the short text length threshold, the category of the text to be processed is identified by a preset statistical model.
[0040] According to a second aspect of this application, a text processing apparatus is provided, comprising:
[0041] The first acquisition module is used to acquire the text to be processed.
[0042] The first determining module is used to determine the text length of the text to be processed;
[0043] The second determining module is used to determine N words that the text to be processed can form if the text length of the text to be processed is less than or equal to a short text length threshold; wherein, N is a positive integer;
[0044] The third determining module is used to perform keyword matching on each of the words based on a preset trie, and determine i target positive keywords and j target verification keywords that match each of the words; wherein, the trie is constructed based on multiple positive keywords and at least one verification keyword corresponding to each positive keyword; i is a natural number and j is a natural number;
[0045] The first identification module is used to identify the category of the text to be processed based on i target positive keywords and j target verification keywords that match each of the words.
[0046] In some embodiments of this application, the first identification module includes:
[0047] A keyword matching unit is used to match each of the target positive keywords with the j target verification keywords based on a preset keyword matching information table; wherein, the keyword matching information table includes the correspondence between all positive keywords and verification keywords in the trie;
[0048] The determining unit is configured to, if each of the target positive keywords matches a corresponding target verification keyword, determine the first position information of each of the target positive keywords in the text to be processed, and determine the second position information of the target verification keyword corresponding to each of the target positive keywords in the text to be processed.
[0049] The judgment unit is used to determine, based on the first location information and the second location information, whether there is a location inclusion relationship between each target positive keyword and its corresponding target verification keyword;
[0050] The first identification unit is used to identify the text to be processed as legal text if each of the target positive keywords and its corresponding target verification keywords have a positional inclusion relationship.
[0051] In some embodiments of this application, the first identification module further includes:
[0052] The first acquisition unit is configured to acquire target positive keywords to be matched with the template from the i target positive keywords if there are target positive keywords that have not been matched with the corresponding target verification keyword, or if each of the target positive keywords is matched with the corresponding target verification keyword and at least one of the target positive keywords has no positional inclusion relationship with the matched target verification keyword; wherein the target positive keywords to be matched with the template include target positive keywords that have not been matched with the corresponding target verification keyword, and target positive keywords that have no positional inclusion relationship with the matched target verification keyword;
[0053] The second acquisition unit is used to acquire the third position information of the positive keyword of the template matching target in the text to be processed;
[0054] The template matching unit is used to perform template matching on the positive keywords of the target template based on a preset template information table, the positive keywords of the target template to be matched, and the third position information; the template information table contains the content information of the template corresponding to each template identifier;
[0055] The second recognition unit is used to identify the category of the text to be processed based on the template matching result.
[0056] In some embodiments of this application, the keyword matching information table further includes a template identifier list corresponding to each positive keyword; the template matching unit is specifically used for:
[0057] Based on the target positive keywords to be matched by the template, determine the template identifier list corresponding to the target positive keywords to be matched by the template in the template identifier list corresponding to each positive keyword;
[0058] Based on the list of template identifiers corresponding to the target positive keywords to be matched, at least one target template identifier is determined;
[0059] Based on the template information table, obtain the content information of the target template corresponding to each target template identifier;
[0060] Based on the target positive keywords to be matched, the third position information, and the content information of each target template, template matching is performed on the target positive keywords to be matched.
[0061] As one possible implementation, the template content information includes template keywords and the relative position information between the template keywords; the template matching unit is specifically used for:
[0062] Based on the target positive keywords for matching the template to be matched and the template keywords in the content information of each target template, at least one first target template is determined from at least one target template; wherein, the template keywords in the content information of the first target template are exactly the same as the target positive keywords for matching the template to be matched;
[0063] Based on the third position information and the relative position information between template keywords in the content information of each first target template, it is determined whether there is a second target template in the at least one first target template that matches the positive keyword of the target to be matched.
[0064] Specifically, the second identification unit is used for:
[0065] If the second target template exists in at least one of the first target templates, the text to be processed is identified as illegal text.
[0066] If the second target template is not present in at least one of the first target templates, the category of the text to be processed is identified by a preset statistical model.
[0067] In some embodiments of this application, the device further includes:
[0068] The second identification module is used to identify the text to be processed as legal text if, after keyword matching is performed on each word based on a preset trie, no target positive keyword is matched for each word.
[0069] In other embodiments of this application, the device further includes:
[0070] The keyword matching module is used to match each word with keywords in a preset word list before performing keyword matching on each word based on a preset trie and determining the i target positive keywords and j target verification keywords that match each word.
[0071] The third identification module is used to identify the text to be processed as illegal text if there are keywords in the vocabulary that match each of the words.
[0072] The second determining module is used to perform keyword matching on each word based on a preset trie if there is no keyword matching for each word in the word list, and to determine i target positive keywords and j target verification keywords that match each word.
[0073] In some other embodiments of this application, the device further includes:
[0074] The fourth identification module is used to identify the category of the text to be processed by using a preset statistical model if the text length of the text to be processed is greater than the short text length threshold.
[0075] According to a third aspect of this application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method described in the first aspect above.
[0076] According to a fourth aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the method described in the first aspect above.
[0077] According to the technical solution of this application, when the text to be processed is short, N words that can be formed from the text are determined, and keyword matching is performed on each word based on a preset trie. i target positive keywords and j target verification keywords that match each word are determined. Based on the i target positive keywords and j target verification keywords that match each word, the category of the text to be processed is identified. Since most texts in typical text processing are short, this solution directly identifies the text category through keyword matching. This not only reduces the workload of using models to review text but also significantly shortens the average response time for text content review, thereby effectively improving the efficiency of text content review.
[0078] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0079] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0080] Figure 1 A flowchart illustrating a text processing method provided in an embodiment of this application;
[0081] Figure 2 This is a flowchart illustrating how to identify the category of text to be processed based on target positive keywords and target verification keywords in an embodiment of this application.
[0082] Figure 3 This is a flowchart illustrating another method for identifying the category of the text to be processed based on target positive keywords and target verification keywords in an embodiment of this application.
[0083] Figure 4 This is a flowchart illustrating a template matching method in an embodiment of this application.
[0084] Figure 5 A flowchart illustrating another text processing method provided in this application embodiment;
[0085] Figure 6 This application provides a structural block diagram of a text processing device according to an embodiment of the present application;
[0086] Figure 7 This is a structural block diagram of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0087] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0088] With the development of the internet and the increase in internet users, people are becoming increasingly reliant on the internet for the dissemination of various information. To ensure the legitimacy of information, all business modules need to review the text content. In related technologies, deep learning models are commonly used for text content review to identify spam; however, these models have long response times and cannot meet the needs of reviewing text with high query volumes.
[0089] To address the aforementioned problems, this application provides a text processing method, apparatus, and electronic device.
[0090] Figure 1 This is a flowchart illustrating a text processing method provided in an embodiment of this application. It should be noted that the text processing method in this embodiment can be used in the text processing device described in this application, and the text processing device can be configured in an electronic device. For example... Figure 1 As shown, the method may include the following steps:
[0091] Step 101: Obtain the text to be processed.
[0092] In some embodiments of this application, the text to be processed can be text information published by users in various business modules, or it can include searched text submitted by users in various business modules, or it can include other text to be processed. As an example, the text to be processed can be obtained by calling the relevant interfaces for user-published information and retrieval in various business modules to obtain the text information submitted by users in real time. As another example, each business module can uniformly place the text to be processed in a certain temporary storage space, so that the text to be processed can be read directly from that temporary storage space.
[0093] Step 102: Determine the text length of the text to be processed.
[0094] In some embodiments of this application, the text to be processed can be preprocessed first, including the removal of punctuation marks, conversion between simplified and traditional Chinese characters, conversion between uppercase and lowercase English characters, and conversion of consecutive spaces; then the preprocessed text content is processed by character segmentation, that is, the Chinese characters in the text are split into single characters and the English characters are split into words to obtain the text length of the text to be processed.
[0095] Furthermore, short text can refer to text whose length is less than or equal to a preset short text length threshold. The inventors of this application have statistically found that short texts involving search terms account for as much as 90% of the texts to be processed in each business module. Therefore, in order to shorten the average response time for text review, short texts in the text to be processed can be quickly filtered. That is, by comparing the text length of the text to be processed with the preset short text length threshold, it is determined whether the text to be processed is short text. If the text length of the text to be processed is less than or equal to the short text length threshold, it is considered short text; if the text length of the text to be processed is greater than the short text length threshold, it is considered long text. The short text length threshold can be determined based on the statistical results of the lengths of frequently occurring short texts in actual scenarios, and this application does not limit it.
[0096] Step 103: If the length of the text to be processed is less than or equal to the short text length threshold, determine the N words that the text to be processed can form; where N is a positive integer.
[0097] Here, the N words that the text to be processed can form refer to all possible words that can be formed in the text. As an example, the process of determining the N words that the text to be processed can form includes: traversing the text according to a preset window size range and the order of the preprocessed text to obtain all possible words. For example, if the preset window size range is 2-4, and the preprocessed text to be processed is "our country", then traversing according to a window size of 2 will yield four words: "we", "our", "country", and "country"; traversing according to a window size of 3 will yield three words: "our", "our country", and "country"; and traversing according to a window size of 4 will yield two words: "our country" and "our country".
[0098] Step 104: Based on the preset trie, perform keyword matching on each word to determine i target positive keywords and j target verification keywords that match each word; wherein, the trie is constructed based on multiple positive keywords and at least one verification keyword corresponding to each positive keyword; i is a natural number and j is a natural number.
[0099] In some embodiments of this application, positive keywords and verification keywords correspond to each other. Positive keywords refer to keywords related to spam text, while the corresponding verification keywords are information within the normal range, and the verification keywords and their corresponding positive keywords have an inclusion relationship. For example, the positive keyword is "marijuana," while its corresponding verification keywords are "marijuana town" and "big trouble."
[0100] The pre-defined trie is constructed based on multiple positive keywords and at least one verification keyword corresponding to each positive keyword, and the constructed trie is at the character-level granularity. Keyword matching is performed on each word in the text to be processed based on the trie. This matching process is exact matching; that is, when a word in the text to be processed completely matches a positive keyword and / or a verification keyword in the trie, the target positive keyword and / or target verification keyword that matches that word are returned. By constructing a character-level trie for keyword matching on each word in the text to be processed, the efficiency of the matching process can be improved, the time consumption of the matching process can be reduced, and thus the response time of the text processing process can be shortened.
[0101] Step 105: Identify the category of the text to be processed based on the i target positive keywords and j target verification keywords that match each word.
[0102] It is understandable that if the keywords matched with each word only contain the target positive keyword, it means that the text to be processed may be illegal text. If the target positive keyword matched with each word corresponds to the target verification keyword, it means that the keyword matched with each word is actually the target verification keyword, that is, the text to be processed is legal text.
[0103] In some embodiments of this application, the correspondence between each positive keyword and verification keyword in the trie can be used to determine whether i target positive keywords and j target verification keywords can be completely matched. If the target positive keywords and target verification keywords are completely matched, the text to be processed can be identified as legal text. If the target positive keywords and target verification keywords are not completely matched, the category of the text to be processed can be identified by a preset statistical model. The preset statistical model can be a text review model in related technologies, such as a Long Short-Term Memory model or a text classification model like TextCNN. After the text to be processed is input into the preset statistical model, the model outputs the predicted category result of the text to be processed.
[0104] In some other embodiments of this application, after performing keyword matching on each word based on a preset trie, the method may further include: if no target positive keyword is matched for each word, the text to be processed is identified as legal text.
[0105] It should be noted that in some embodiments of the text processing method in this application, keyword matching is used to identify the category of the text to be processed for short texts, while statistical models in related technologies can be used to identify the category of the text to be processed for long texts. That is to say, the method may also include: if the length of the text to be processed is greater than the short text length threshold, identifying the category of the text to be processed through a preset statistical model.
[0106] According to the text processing method of this application embodiment, when the text to be processed is short text, N words that can be formed from the text to be processed are determined, and keyword matching is performed on each word based on a preset trie. i target positive keywords and j target verification keywords that match each word are determined. Based on the i target positive keywords and j target verification keywords that match each word, the category of the text to be processed is identified. Since most texts in typical text processing are short texts, this solution directly identifies the text category through keyword matching, which not only reduces the workload of using models to review text, but also greatly shortens the average response time for text content review, thereby effectively improving the efficiency of text content review.
[0107] The following section will provide a detailed description of the steps in the above embodiments for identifying the category of the text to be processed based on the i target positive keywords and j target verification keywords that match each word.
[0108] Figure 2 This is a flowchart illustrating how to identify the category of text to be processed based on target positive keywords and target verification keywords, as described in an embodiment of this application. Figure 2 As shown, based on the above embodiments, Figure 1 The implementation of step 105 may include the following steps:
[0109] Step 201: Based on the preset keyword matching information table, match each target positive keyword with j target verification keywords respectively; wherein, the keyword matching information table includes the correspondence between all positive keywords and verification keywords in the trie.
[0110] In some embodiments of this application, the keyword matching information table is a table used to store the correspondence between all positive keywords and verification keywords in the trie. Based on the keyword matching information table, the process of matching each target positive keyword with j target verification keywords may include: for each target positive keyword, querying the keyword matching information table for the corresponding verification keyword; iterating through the j target verification keywords to determine whether a corresponding verification keyword exists for the target positive keyword, thus determining whether the target positive keyword can be matched with a corresponding target verification keyword.
[0111] Step 202: If each target positive keyword matches the corresponding target verification keyword, determine the first position information of each target positive keyword in the text to be processed, and determine the second position information of the target verification keyword corresponding to each target positive keyword in the text to be processed.
[0112] It's understandable that each target positive keyword matches its corresponding target verification keyword. Furthermore, the position of each target positive keyword in the processed text and the position of its corresponding target verification keyword in the processed text are related by an inclusion relationship, indicating that the keywords actually matching the words are all target verification keywords. However, if the position of a target positive keyword in the processed text and the position of its corresponding target verification keyword in the processed text are not related by an inclusion relationship, it means that words matching the target positive keyword in the processed text cannot be canceled out by the corresponding verification keywords.
[0113] In some embodiments of this application, if each target positive keyword matches a corresponding target verification keyword, then i target positive keywords and j target verification keywords are mutually matched. Here, the target verification keyword corresponding to each target positive keyword refers to the target verification keyword matched by each target verification keyword. For example, if target positive keyword 1 matches target verification keyword 1, and target positive keyword 2 matches target verification keywords 2 and 3, then the target verification keyword corresponding to target positive keyword 1 is target verification keyword 1, and the target verification keywords corresponding to target positive keyword 2 are target verification keywords 2 and 3.
[0114] Furthermore, the first position information of each target positive keyword in the text to be processed may include the starting position information of each target positive keyword in the text to be processed, and may also include the length information of each target positive keyword. If the target positive keyword is "marijuana" and the text to be processed is "XXX marijuana XXXX", then the first position information of the target positive keyword in the text to be processed is (4,2), where 4 is the starting position information of the target positive keyword in the text to be processed, and 2 is the length information of the target positive keyword. The second position information of the target verification keyword corresponding to each target positive keyword in the text to be processed has the same format as the first position information.
[0115] Step 203: Based on the first position information and the second position information, determine whether there is a positional inclusion relationship between each target positive keyword and its corresponding target verification keyword.
[0116] In some embodiments of this application, determining whether there is a positional inclusion relationship between each target positive keyword and its corresponding target verification keyword is equivalent to determining whether there is a partial positional overlap between each target positive keyword and its corresponding target verification keyword. As an example, assuming that the format of both the first and second positional information is (starting position information, length information), the first positional information of target positive keyword 1 is (3, 2), and the second positional information of the target verification keyword 1 corresponding to target positive keyword 1 is (3, 3), then it indicates that there is a partial positional overlap between target positive keyword 1 and its corresponding target verification keyword 1, i.e., there is a positional inclusion relationship between target positive keyword 1 and its corresponding target verification keyword 1.
[0117] Step 204: If each target positive keyword and its corresponding target verification keyword have a positional inclusion relationship, the text to be processed is identified as legal text.
[0118] It is understandable that if each target positive keyword and its corresponding target verification keyword have a positional inclusion relationship, it means that the keywords matching each word are actually j target verification keywords. Since the verification keywords are all within the normal range of information, if each target positive keyword and its corresponding target verification keyword have a positional inclusion relationship, it means that the text to be processed is legal text.
[0119] In some embodiments of this application, if there is a target positive keyword among the i target positive keywords that does not match the corresponding target verification keyword, and / or, there is a target positive keyword among the i target positive keywords that does not have a positional inclusion relationship with the matched target verification keyword, it indicates that there are words in the text to be processed that match the positive keywords, that is, the text to be processed may be illegal text. At this time, the text to be processed can be input into a preset statistical model to identify the category of the text to be processed.
[0120] According to the text processing method of this application embodiment, when identifying the category of the text to be processed based on the target positive keywords and target verification keywords, each target positive keyword is matched with j target verification keywords based on the keyword matching information table. When each target positive keyword matches a corresponding target verification keyword, based on the positional information of each target positive keyword and its corresponding verification keyword, it is determined whether there is a positional inclusion relationship between each target positive keyword and its corresponding target verification keyword. If all positional inclusion relationships exist, the text to be processed is identified as legal text. This is equivalent to using the correspondence between positive keywords and verification keywords, as well as the positional relationship between target positive keywords and target verification keywords, to identify the category of the text to be processed, thereby improving the efficiency of text processing.
[0121] To further reduce the response time of text processing, this application proposes yet another embodiment.
[0122] Figure 3 This is a flowchart illustrating another method for identifying the category of text to be processed based on target positive keywords and target verification keywords, as described in this application embodiment. Figure 3 As shown, based on the above embodiments, Figure 1 The implementation of step 105 may include the following steps:
[0123] Step 301: Based on the preset keyword matching information table, match each target positive keyword with j target verification keywords respectively; wherein, the keyword matching information table includes the correspondence between all positive keywords and verification keywords in the trie.
[0124] Step 302: Determine whether each target positive keyword can match the corresponding target verification keyword.
[0125] If each target positive keyword can be matched with the corresponding target verification keyword, proceed to step 303; otherwise, proceed to step 306.
[0126] Step 303: Determine the first position information of each target positive keyword in the text to be processed, and determine the second position information of the target verification keyword corresponding to each target positive keyword in the text to be processed.
[0127] Step 304: Based on the first position information and the second position information, determine whether each target positive keyword and its corresponding target verification keyword have a positional inclusion relationship.
[0128] If each target positive keyword and its corresponding target verification keyword have a positional inclusion relationship, then proceed to step 305; otherwise, proceed to step 306.
[0129] Step 305: Identify the text to be processed as valid text.
[0130] Step 306: Obtain the target positive keywords to be matched by the template from the i target positive keywords; wherein, the target positive keywords to be matched by the template include target positive keywords that have not been matched with the corresponding target verification keywords, and target positive keywords that do not have a positional inclusion relationship with the matched target verification keywords.
[0131] In some embodiments of this application, if there are target positive keywords that do not match the corresponding target verification keyword, obtaining target positive keywords to be matched by the template from the i target positive keywords may include: if there are target positive keywords that do not match the corresponding target verification keyword, then the target positive keywords that do not match the corresponding target verification keyword can be used as target positive keywords to be matched by the template. Furthermore, obtaining target positive keywords to be matched by the template from the i target positive keywords may also include: determining the matched target positive keywords that have matched the corresponding target verification keyword from the i target positive keywords, and determining the fourth position information of each matched target positive keyword in the text to be processed, and the fifth position information of the target verification keyword corresponding to each matched target positive keyword in the text to be processed; determining, based on the fourth position information and the fifth position information, whether there is a positional inclusion relationship between each matched target positive keyword and its corresponding target verification keyword; and also using the matched target positive keywords with which there is no positional inclusion relationship as target positive keywords to be matched by the template. The number of target positive keywords to be matched by the template can be one or more.
[0132] In some other embodiments of this application, if each target positive keyword matches the corresponding target verification keyword and at least one target positive keyword does not have a positional inclusion relationship with the matched target verification keyword, the implementation of obtaining the target positive keyword to be matched from the i target positive keywords may include: during the judgment process in step 303, obtaining the target positive keywords that do not have a positional inclusion relationship with the matched target verification keyword, and using these target positive keywords as the target positive keywords to be matched by the template.
[0133] Step 307: Obtain the third position information of the positive keyword of the target to be matched in the text to be processed.
[0134] In the embodiments of this application, the third position information of the positive keyword of the template to be matched in the text to be processed is the same as the first position information and the second position information in the above embodiments, and will not be repeated here.
[0135] Step 308: Based on the preset template information table, the target positive keywords of the template to be matched, and the third position information, perform template matching on the target positive keywords of the template to be matched; the template information table contains the content information of the template corresponding to each template identifier.
[0136] In some embodiments of this application, the template information table is obtained based on statistics of historical illegal text. The template information table contains the content information of the template corresponding to each template identifier. The content information of each template may include one or more keywords, as well as the position information of each keyword, etc. As an example, the process of template matching for the target positive keyword of the template to be matched may include: comparing the target positive keyword of the template to be matched with the content information of each template in the template information table, and obtaining at least one first template whose keywords in the content information of the template are consistent with the target positive keyword of the template to be matched; comparing the position information of the keywords in the content information of the first template with third position information, and determining a second template whose position information matches each other from the first templates. This second template is the template that matches the target positive keyword of the template to be matched.
[0137] Step 309: Identify the category of the text to be processed based on the template matching results.
[0138] It is understandable that if a template exists in the template information table that matches the positive keyword of the target template, the text to be processed is identified as illegal text; if no template exists in the template information table that matches the positive keyword of the target template, the text to be processed is identified as legal text.
[0139] According to the text processing method in the application embodiment, for target positive keywords that do not match corresponding target verification keywords, or where each positive keyword matches a corresponding target verification keyword but there is no positional inclusion relationship between the positive keyword and the matched target verification keyword, the method obtains the target positive keyword to be matched from i target positive keywords and obtains the third position information of the target positive keyword to be matched in the text to be processed. Based on a preset template information table, the target positive keyword to be matched, and the third position information, template matching is performed on the target positive keyword to be matched, and the category of the text to be processed is identified according to the template matching result. In other words, by combining the matching of positive keywords, verification keywords, and template matching, the category of short text is identified, thereby further reducing the workload of using models to review text and further shortening the average response time of text content review.
[0140] Next, we will introduce in detail the template matching process for positive keywords of the target template.
[0141] Figure 4This is a flowchart illustrating template matching in one embodiment of this application. In some embodiments of this application, the keyword matching information table further includes a list of template identifiers corresponding to each positive keyword, wherein each positive keyword's list of template identifiers contains at least one template identifier. For example... Figure 4 As shown, based on the above embodiments, Figure 3 Step 308 may include the following steps:
[0142] Step 401: Based on the target positive keywords to be matched by the template, determine the template identifier list corresponding to the target positive keywords to be matched by the template in the template identifier list corresponding to each positive keyword.
[0143] In some embodiments of this application, if there is only one target positive keyword for the template to be matched, the target positive keyword can be searched in the keyword matching information table, and the template identifier list corresponding to the target positive keyword can be obtained. If there are multiple target positive keywords for the template to be matched, a template identifier list corresponding to each target positive keyword can be determined from the template identifier list corresponding to each positive keyword.
[0144] Step 402: Based on the list of template identifiers corresponding to the target positive keywords to be matched, determine at least one target template identifier.
[0145] In some embodiments of this application, if there is only one target positive keyword to be matched by the template, then at least one target template identifier is a template identifier from the template identifier list corresponding to that target positive keyword. If there are multiple target positive keywords to be matched by the template, then the intersection of the template identifier lists corresponding to each target positive keyword is taken to obtain at least one target template identifier. In other words, by using the template identifier list corresponding to each positive keyword in the keyword matching table, the speed of obtaining the target template can be improved, and the amount of computation in the template matching process can be reduced.
[0146] Step 403: Based on the template information table, obtain the content information of the target template corresponding to each target template identifier.
[0147] Since the template information table contains the content information of the template corresponding to each template identifier, the content information of the target template corresponding to each target template identifier can be obtained based on the template information table. Because there is at least one target template identifier, the number of target templates is also at least one.
[0148] Step 404: Based on the target positive keywords to be matched, the third position information, and the content information of each target template, perform template matching on the target positive keywords to be matched.
[0149] In other words, the target positive keyword and third position information of the template to be matched are compared with the content information of each target template in turn. If the content information of a target template matches the target positive keyword and third position information of the template to be matched, then a template matching the target positive keyword of the template to be matched is obtained. If the content information of the target template fails to match the target positive keyword and third position information of the template to be matched, then there is no template matching the target positive keyword of the template to be matched in the template information table.
[0150] In some embodiments of this application, the content information of the template in the template information table may include template keywords and their relative positions. There may be one or more template keywords, and the relative positions may include the order of the keywords, the interval between them, and other information. Thus, the implementation of step 404 may include the following steps:
[0151] Step 404-1: Based on the target positive keywords to be matched and the template keywords in the content information of each target template, determine at least one first target template from at least one target template; wherein, the template keywords in the content information of the first target template are exactly the same as the target positive keywords to be matched.
[0152] In other words, the target positive keyword of the template to be matched is compared sequentially with the template keywords in the content information of each target template. If the template keywords in the content information of a target template are exactly the same as the target positive keyword of the template to be matched, then that target template is the first target template. For example, if there is only one target positive keyword to be matched, and the content information of a target template also contains only one template keyword, and that template keyword matches the target positive keyword of the template to be matched, then that target template is the first target template. If there are multiple second target positive keywords, and the number of template keywords in the content information of a target template matches the number of target positive keywords of the template to be matched, and multiple template keywords are exactly the same as multiple target positive keywords of the template to be matched, then that target template is the first target template.
[0153] Step 404-2: Based on the third position information and the relative position information between template keywords in the content information of each first target template, determine whether there is a second target template in at least one first target template that matches the positive keyword of the target to be matched.
[0154] In other words, the third position information of the target positive keyword in the text to be processed is compared with the relative position information between the template keywords in the content information of each first target template. If the relative position information between the template keywords of a first target template satisfies the third position information, it means that at least one first target template contains a second target template that matches the target positive keyword of the target.
[0155] As an example, the content information of a first target template is formatted as [uid,Len1,KeyWord1,Len2,KeyWord2,Len3,KeyWord3,Len4], where uid is the template identifier, KeyWord1, KeyWord2, and KeyWord3 are the three template keywords in the content information, Len1 is the maximum interval length between KeyWord1 and the beginning of the sentence, Len2 is the maximum interval length between KeyWord1 and KeyWord2, Len3 is the maximum interval length between KeyWord2 and KeyWord3, and Len4 is the maximum interval length between KeyWord3 and the end of the sentence. If the target positive keywords for template matching include target positive keywords A, B, and C, and target positive keywords A matches Keyword1, B matches Keyword2, and C matches Keyword3, then based on the third position information of target positive keywords A, determine whether the interval length between target positive keywords A and the beginning of the sentence is less than Len1; based on the third position information of target positive keywords A and B... The system determines whether the interval between the target positive keyword A and the target positive keyword B is less than Len2; based on the third position information of the target positive keyword B and the target positive keyword C, it determines whether the interval between the target positive keyword B and the target positive keyword C is less than Len3; based on the third position information of the target positive keyword C, it determines whether the interval between the target positive keyword C and the end of the sentence is less than Len4; if all of the above are satisfied, then the first target template is the second target template that matches the target positive keyword.
[0156] It should be noted that, in the embodiments of this application, Figure 3The implementation of step 309 may include: if a second target template exists in at least one first target template, identifying the text to be processed as illegal text; if a second target template does not exist in at least one first target template, identifying the category of the text to be processed through a preset statistical model.
[0157] In some other embodiments of this application, the content information of each template in the template information table contains a corresponding category. If the target positive keyword of the template to be matched matches the second target template, the category in the content information of the second target template is used as the recognition result of the text to be processed.
[0158] According to the text processing method of this application embodiment, since the keyword matching information table also includes a list of template identifiers corresponding to each positive keyword, the list of template identifiers corresponding to the target positive keyword to be matched can be determined in the keyword matching information table based on the target positive keyword to be matched, and at least one target template identifier can be determined. The content information of each target template is obtained, and template matching is performed on the target positive keyword to be matched based on the target positive keyword to be matched, the third position information, and the content information of each target template. This is equivalent to only matching the target positive keyword to be matched with the target template, which can greatly improve the efficiency of template matching, reduce the amount of computation in the template matching process, and further shorten the average response time of the text processing process.
[0159] To further improve the efficiency of text processing, this application proposes yet another embodiment.
[0160] Figure 5 A flowchart illustrating yet another text processing method provided in an embodiment of this application. For example... Figure 5 As shown, the method may include the following steps:
[0161] Step 501: Obtain the text to be processed.
[0162] Step 502: Determine the text length of the text to be processed.
[0163] Step 503: If the text length of the text to be processed is less than or equal to the short text length threshold, determine the N words that the text to be processed can form; where N is a positive integer.
[0164] Step 504: Match each word with the keywords in the preset word list.
[0165] In some embodiments of this application, the keywords in the preset thesaurus are hard keywords, meaning that if the text to be processed contains a word that matches a keyword in the preset thesaurus, then the text to be processed is considered illegal text. The keywords in the preset thesaurus can be derived from historical statistics of illegal words or can be manually maintained. Furthermore, the process of matching each word with the keywords in the preset thesaurus is a precise matching process.
[0166] Step 505: If there are keywords in the vocabulary that match each word, the text to be processed is identified as illegal text.
[0167] In other words, if the vocabulary contains keywords that match one or more of the N words, the text to be processed is identified as illegal text. This allows for rapid identification of the text's category based on keywords in a pre-defined vocabulary. If the vocabulary contains keywords that match each of the N words, the text is directly identified as illegal text, and subsequent steps are not performed.
[0168] In some other embodiments of this application, illegal text can be further categorized. The preset vocabulary includes not only keywords but also the category corresponding to each keyword. Therefore, if there is a keyword in the vocabulary that matches a certain word, the category corresponding to that keyword is taken as the category of the text to be processed.
[0169] Step 506: If there is no keyword matching each word in the vocabulary, perform keyword matching on each word based on the preset trie to determine i target positive keywords and j target verification keywords that match each word; wherein, the trie is constructed based on multiple positive keywords and at least one verification keyword corresponding to each positive keyword; i is a natural number and j is a natural number.
[0170] Step 507: Identify the category of the text to be processed based on the i target positive keywords and j target verification keywords that match each word.
[0171] According to the text processing method of this application embodiment, before performing keyword matching on each word based on a preset trie, each word is matched with keywords in a preset vocabulary. If a keyword matching the word exists in the vocabulary, the text to be processed is identified as illegal text; otherwise, keyword matching is performed on each word based on the preset trie. This solution directly identifies illegal text based on keywords in a preset vocabulary, which not only reduces the computational load in the text processing process but also further improves the efficiency of text processing.
[0172] To implement the above embodiments, this application proposes a text processing apparatus.
[0173] Figure 6 This is a structural block diagram of a text processing device provided in an embodiment of this application. Figure 6 As shown, the device may include:
[0174] The first acquisition module 610 is used to acquire the text to be processed;
[0175] The first determining module 620 is used to determine the text length of the text to be processed;
[0176] The second determining module 630 is used to determine the N words that the text to be processed can form if the text length of the text to be processed is less than or equal to the short text length threshold; where N is a positive integer.
[0177] The third determining module 640 is used to perform keyword matching on each word based on a preset trie, and determine i target positive keywords and j target verification keywords that match each word; wherein, the trie is constructed based on multiple positive keywords and at least one verification keyword corresponding to each positive keyword; i is a natural number and j is a natural number;
[0178] The first recognition module 650 is used to identify the category of the text to be processed based on the i target positive keywords and j target verification keywords that match each word.
[0179] In some embodiments of this application, the first identification module 650 includes:
[0180] The keyword matching unit 651 is used to match each target positive keyword with j target verification keywords based on a preset keyword matching information table; wherein, the keyword matching information table includes the correspondence between all positive keywords and verification keywords in the trie;
[0181] The determining unit 652 is used to determine the first position information of each target positive keyword in the text to be processed and the second position information of the target verification keyword corresponding to each target positive keyword in the text to be processed if each target positive keyword matches the corresponding target verification keyword.
[0182] The judgment unit 653 is used to determine whether there is a positional inclusion relationship between each target positive keyword and its corresponding target verification keyword based on the first position information and the second position information.
[0183] The first identification unit 654 is used to identify the text to be processed as legal text if each target positive keyword and its corresponding target verification keyword have a positional inclusion relationship.
[0184] In some embodiments of this application, the first identification module 650 further includes:
[0185] The first acquisition unit 655 is used to acquire target positive keywords to be matched with the template from i target positive keywords if there are target positive keywords that have not been matched with the corresponding target verification keyword, or if each target positive keyword is matched with the corresponding target verification keyword and at least one target positive keyword has no positional inclusion relationship with the matched target verification keyword; wherein, the target positive keywords to be matched with the template include target positive keywords that have not been matched with the corresponding target verification keyword, and target positive keywords that have no positional inclusion relationship with the matched target verification keyword;
[0186] The second acquisition unit 656 is used to acquire the third position information of the positive keyword of the template matching target in the text to be processed;
[0187] The template matching unit 657 is used to perform template matching on the target positive keywords based on a preset template information table, the target positive keywords to be matched, and third position information; the template information table contains the content information of the template corresponding to each template identifier;
[0188] The second recognition unit 658 is used to identify the category of the text to be processed based on the template matching result.
[0189] In some embodiments of this application, the keyword matching information table further includes a list of template identifiers corresponding to each positive keyword; the template matching unit 657 is specifically used for:
[0190] Based on the target positive keywords to be matched by the template, determine the template identifier list corresponding to the target positive keywords in the template identifier list corresponding to each positive keyword;
[0191] Based on the list of template identifiers corresponding to the target positive keywords to be matched, at least one target template identifier is determined;
[0192] Based on the template information table, obtain the content information of the target template corresponding to each target template identifier;
[0193] Based on the target positive keywords to be matched, the third position information, and the content information of each target template, template matching is performed on the target positive keywords to be matched.
[0194] As one possible implementation, the template content information includes template keywords and the relative position information between template keywords; the template matching unit 657 is specifically used for:
[0195] Based on the target positive keywords for matching the template to be matched and the template keywords in the content information of each target template, at least one first target template is determined from at least one target template; wherein, the template keywords in the content information of the first target template are exactly the same as the target positive keywords for matching the template to be matched;
[0196] Based on the relative position information between template keywords in the third position information and the content information of each first target template, determine whether there is a second target template in at least one first target template that matches the positive keyword of the target to be matched.
[0197] Specifically, the second identification unit 658 is used for:
[0198] If a second target template exists in at least one first target template, the text to be processed is identified as illegal text.
[0199] If at least one first target template does not contain a second target template, the category of the text to be processed is identified by a preset statistical model.
[0200] In some embodiments of this application, the device further includes:
[0201] The second recognition module 660 is used to identify the text to be processed as legal text if no target positive keyword is matched for each word after keyword matching based on a preset dictionary tree.
[0202] In other embodiments of this application, the device further includes:
[0203] The keyword matching module 670 is used to match each word with keywords in a preset vocabulary before performing keyword matching on each word based on a preset trie and determining the i target positive keywords and j target verification keywords that match each word.
[0204] The third identification module 680 is used to identify the text to be processed as illegal text if there are keywords in the vocabulary that match each word.
[0205] The third determining module 640 is used to perform keyword matching on each word based on a preset trie if there is no keyword matching for each word in the vocabulary, and to determine i target positive keywords and j target verification keywords that match each word.
[0206] In some other embodiments of this application, the device further includes:
[0207] The fourth identification module 690 is used to identify the category of the text to be processed by using a preset statistical model if the length of the text to be processed is greater than the short text length threshold.
[0208] According to the text processing apparatus of this application embodiment, when the text to be processed is short text, it determines N words that the text to be processed can form, and performs keyword matching on each word based on a preset trie, determining i target positive keywords and j target verification keywords that match each word. Based on the i target positive keywords and j target verification keywords that match each word, it identifies the category of the text to be processed. Since most texts in typical text processing are short texts, this solution directly identifies the text category of short texts through keyword matching, which not only reduces the workload of using models to review texts, but also greatly shortens the average response time of text content review, thereby effectively improving the efficiency of text content review.
[0209] According to embodiments of this application, this application also provides an electronic device and a computer-readable storage medium.
[0210] like Figure 7 The diagram shown is a structural block diagram of an electronic device for a text processing method according to an embodiment of this application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described and / or claimed herein.
[0211] like Figure 7 As shown, the electronic device includes one or more processors 701, a memory 702, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise as required. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 7 Take the 701 processor as an example.
[0212] The memory 702 is the non-transitory computer-readable storage medium provided in this application. The memory stores instructions executable by at least one processor to cause the at least one processor to perform the text processing method provided in this application. The non-transitory computer-readable storage medium of this application stores computer instructions for causing a computer to perform the text processing method provided in this application. The computer program product of this application includes a computer program that, when executed by processor 701, implements the text processing method provided in this application.
[0213] The memory 702, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the text processing method in the embodiments of this application. The processor 701 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 702, thereby implementing the text processing method in the above method embodiments.
[0214] Memory 702 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the use of the electronic device according to the text processing method. Furthermore, memory 702 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory 702 may optionally include memory remotely located relative to processor 701, and these remote memories can be connected to the text processing electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0215] The electronic device for implementing the text processing method may further include an input device 703 and an output device 704. The processor 701, memory 702, input device 703, and output device 704 can be connected via a bus or other means. Figure 7 Taking the example of a connection between China and Israel via a bus.
[0216] Input device 703 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of an electronic device used to implement text processing methods, such as touch screens, keypads, mice, trackpads, touchpads, joysticks, one or more mouse buttons, trackballs, joysticks, etc. Output device 704 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The display device may include, but is not limited to, liquid crystal displays (LCDs), light-emitting diode (LED) displays, and plasma displays. In some embodiments, the display device may be a touch screen.
[0217] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, application-specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.
[0218] These computational programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0219] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0220] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0221] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0222] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, and this is not limited herein.
[0223] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0224] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A text processing method, characterized in that, include: Get the text to be processed; Determine the text length of the text to be processed; If the length of the text to be processed is less than or equal to the short text length threshold, determine the N words that the text to be processed can form; where N is a positive integer; Based on a pre-defined trie, keyword matching is performed on each of the words to determine i target positive keywords and j target verification keywords that match each of the words; wherein, the trie is constructed based on multiple positive keywords and at least one verification keyword corresponding to each positive keyword; i is a natural number and j is a natural number; The category of the text to be processed is identified based on i target positive keywords and j target verification keywords that match each of the aforementioned words. This includes: matching each target positive keyword with the j target verification keywords based on a preset keyword matching information table; wherein the keyword matching information table includes the correspondence between all positive keywords and verification keywords in the trie; if each target positive keyword matches a corresponding target verification keyword, determining the first position information of each target positive keyword in the text to be processed, and determining the second position information of the target verification keyword corresponding to each target positive keyword in the text to be processed; determining, based on the first position information and the second position information, whether there is a positional inclusion relationship between each target positive keyword and its corresponding target verification keyword; if there is a positional inclusion relationship between each target positive keyword and its corresponding target verification keyword, identifying the text to be processed as legitimate text. The positive keywords refer to keywords related to spam text, the verification keywords corresponding to the positive keywords are information within the normal range, and the verification keywords and their corresponding positive keywords have an inclusion relationship.
2. The method according to claim 1, characterized in that, Also includes: If there are target positive keywords that do not match the corresponding target verification keyword, or if each of the target positive keywords matches the corresponding target verification keyword and at least one of the target positive keywords does not have a positional inclusion relationship with the matched target verification keyword, then the target positive keywords to be matched with the template are obtained from the i target positive keywords; wherein, the target positive keywords to be matched with the template include target positive keywords that do not match the corresponding target verification keyword, and target positive keywords that do not have a positional inclusion relationship with the matched target verification keyword; Obtain the third position information of the positive keyword of the target template in the text to be processed; Based on a preset template information table, the target positive keywords for template matching, and the third position information, template matching is performed on the target positive keywords for template matching; the template information table contains the content information of the template corresponding to each template identifier; Based on the template matching results, the category of the text to be processed is identified.
3. The method according to claim 2, characterized in that, The keyword matching information table also includes a list of template identifiers corresponding to each positive keyword; the template matching process based on the preset template information table, the target positive keyword to be matched, and the third position information includes: Based on the target positive keywords to be matched by the template, determine the template identifier list corresponding to the target positive keywords to be matched by the template in the template identifier list corresponding to each positive keyword; Based on the list of template identifiers corresponding to the target positive keywords to be matched, at least one target template identifier is determined; Based on the template information table, obtain the content information of the target template corresponding to each target template identifier; Based on the target positive keywords to be matched, the third position information, and the content information of each target template, template matching is performed on the target positive keywords to be matched.
4. The method according to claim 3, characterized in that, The template content information includes template keywords and the relative position information between the template keywords; the step of performing template matching on the target positive keywords to be matched based on the target positive keywords to be matched, the third position information, and the content information of each target template includes: Based on the target positive keywords for matching the template to be matched and the template keywords in the content information of each target template, at least one first target template is determined from at least one target template; wherein, the template keywords in the content information of the first target template are exactly the same as the target positive keywords for matching the template to be matched; Based on the third position information and the relative position information between template keywords in the content information of each first target template, it is determined whether there is a second target template in the at least one first target template that matches the positive keyword of the target to be matched.
5. The method according to claim 4, characterized in that, The step of identifying the category of the text to be processed based on the template matching result includes: If the second target template exists in at least one of the first target templates, the text to be processed is identified as illegal text. If the second target template is not present in at least one of the first target templates, the category of the text to be processed is identified by a preset statistical model.
6. The method according to claim 1, characterized in that, After performing keyword matching on each word based on a preset trie, the method further includes: If none of the stated words match the target positive keyword, the text to be processed is identified as legitimate text.
7. The method according to any one of claims 1 to 6, characterized in that, Before performing keyword matching on each word based on a preset trie to determine the i target positive keywords and j target verification keywords that match each word, the method further includes: Each of the aforementioned words is matched with keywords in a preset word list; If the vocabulary contains keywords that match each of the words, the text to be processed is identified as illegal text. If no keyword matches each of the words in the vocabulary, the step of performing keyword matching on each word based on a preset trie to determine i target positive keywords and j target verification keywords that match each word is executed.
8. The method according to claim 1, characterized in that, Also includes: If the length of the text to be processed is greater than the short text length threshold, the category of the text to be processed is identified by a preset statistical model.
9. A text processing device, characterized in that, include: The first acquisition module is used to acquire the text to be processed. The first determining module is used to determine the text length of the text to be processed; The second determining module is used to determine N words that the text to be processed can form if the text length of the text to be processed is less than or equal to a short text length threshold; wherein, N is a positive integer; The third determining module is used to perform keyword matching on each of the words based on a preset trie, and determine i target positive keywords and j target verification keywords that match each of the words; wherein, the trie is constructed based on multiple positive keywords and at least one verification keyword corresponding to each positive keyword; i is a natural number and j is a natural number; The first identification module is used to identify the category of the text to be processed based on i target positive keywords and j target verification keywords that match each of the words. The first identification module includes: A keyword matching unit is used to match each of the target positive keywords with the j target verification keywords based on a preset keyword matching information table; wherein, the keyword matching information table includes the correspondence between all positive keywords and verification keywords in the trie; The determining unit is configured to, if each of the target positive keywords matches a corresponding target verification keyword, determine the first position information of each of the target positive keywords in the text to be processed, and determine the second position information of the target verification keyword corresponding to each of the target positive keywords in the text to be processed. The judgment unit is used to determine, based on the first location information and the second location information, whether there is a location inclusion relationship between each target positive keyword and its corresponding target verification keyword; The first identification unit is configured to identify the text to be processed as legal text if each of the target positive keywords and its corresponding target verification keywords have a positional inclusion relationship. The positive keywords refer to keywords related to spam text, the verification keywords corresponding to the positive keywords are information within the normal range, and the verification keywords and their corresponding positive keywords have an inclusion relationship.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Text classification method and device
CN111767403A