Text message management strategy generating method and apparatus, electronic device, and storage medium

By constructing keyword knowledge graphs and identifying keyword variants in SMS, the problem of low spam recognition accuracy in the prior art is solved, and more efficient SMS interception is achieved.

WO2025108155A1PCT designated stage expired Publication Date: 2025-05-30CHINA MOBILE GROUP DESIGN INST +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/131871
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-11-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has low accuracy in identifying and intercepting spam messages, making it difficult to effectively identify variants, extended or substituted keywords.

Method used

By constructing a keyword knowledge graph based on preset keywords and their conjunctions, combining character substring extraction in the pending text messages, keyword matching is performed, target keywords are identified, and SMS interception strategies are determined based on these keywords.

Benefits of technology

Improves the accuracy of spam blocking, enabling quick identification and interception of text messages containing variants, extensions or alternative keywords.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024131871_30052025_PF_FP_ABST
    Figure CN2024131871_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of data processing, and provides a text message management strategy generating method and apparatus, an electronic device, and a storage medium. The method comprises: obtaining a text message to be processed; performing character sub-string extraction on the basis of the text message to be processed, to obtain a sub-string set; matching the sub-string set with a keyword knowledge graph to obtain a target keyword, wherein the keyword knowledge graph is constructed on the basis of preset keywords and associated words thereof obtained by variants, extensions and substitutions; and determining a text message blocking strategy on the basis of the target keyword and the keyword knowledge graph, so as to perform text message blocking on the basis of the text message blocking strategy. According to the present disclosure, a new keyword formed by a variant, extension or substitution of a keyword in a text message to be processed can be quickly and accurately identified, and then a text message blocking strategy can be quickly and accurately determined, thereby facilitating relevant personnel to refer to the text message blocking strategy to perform spam text message blocking and thus improving the accuracy of spam text message blocking.
Need to check novelty before this filing date? Find Prior Art

Description

SMS management strategy generation method, device, electronic device and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This disclosure is based on and claims the priority of Chinese patent application with application number 202311546858.3 and application date November 20, 2023. The entire content of the Chinese patent application is hereby incorporated into this disclosure as a reference. Technical Field

[0003] The present disclosure relates to the field of data processing technology, and in particular to a method, device, electronic device, and storage medium for generating a short message management strategy. Background Art

[0004] SMS messages are used for daily communication, business promotion, and notifications. However, they can also be exploited by criminals to disseminate illegal information, causing distress to users and even leading to financial losses. Therefore, analyzing SMS content and identifying spam messages is essential.

[0005] Existing SMS content analysis technologies mainly include SMS text classification technology, SMS text clustering technology, and keyword combination strategy analysis technology. However, these technologies have low accuracy in identifying and blocking spam SMS.

[0006] Summary of the Invention

[0007] The embodiments of the present disclosure provide a method, device, electronic device, and storage medium for generating a short message management strategy, so as to solve the problem of low accuracy in the current interception of spam short messages.

[0008] In a first aspect, an embodiment of the present disclosure provides a method for generating a text message management strategy, comprising: obtaining a text message to be processed; extracting character substrings based on the text message to be processed to obtain a subset set; matching the subset set with a keyword knowledge graph to obtain a target keyword, wherein the associated words are variants, extensions or substitutes of the preset keywords; the keyword knowledge graph is constructed based on preset keywords and their associated words; determining a text message interception strategy based on the target keywords and the keyword knowledge graph, so as to perform text message interception based on the text message interception strategy.

[0009] In one embodiment, matching the subset set with the keyword knowledge graph to obtain the target keyword includes: obtaining the keyword knowledge graph; performing variant character matching on the character substrings in the subset set in the keyword knowledge graph to obtain a first candidate keyword; if there is at least one target character substring that fails to match in the variant character matching, performing character similarity matching based on each target character substring to obtain a second candidate keyword; determining at least one initial keyword based on the first candidate keyword, or the first candidate keyword and the second candidate keyword; and performing keyword matching based on each initial keyword to obtain a target keyword.

[0010] In one embodiment, the method of performing character similarity matching based on each of the target character substrings to obtain a second candidate keyword includes: performing character rareness detection on each of the target character substrings to obtain character rareness information; determining rare substrings based on the character rareness information; and performing character similarity matching based on the rare substrings to obtain a second candidate keyword.

[0011] In one embodiment, the character similarity matching based on the uncommon substring to obtain the second candidate keyword includes: constructing a pinyin and stroke order sequence dictionary of the keyword based on the keyword knowledge graph; and performing homophonic and similar matching on the uncommon substring based on the pinyin and stroke order sequence dictionary to obtain the second candidate keyword.

[0012] In one embodiment, the character substring extraction based on the SMS to be processed to obtain a substring set includes: cleaning invalid characters from the SMS to be processed to obtain a target character string; the invalid characters include at least one of punctuation marks, English symbols, and emoticons; and character segmentation of the target character string, forming a substring set from each character substring formed by the segmentation.

[0013] In one embodiment, the keyword matching is performed based on each of the initial keywords to obtain the target keyword, including: determining the character position information of each character in the target string corresponding to the SMS to be processed in the target string; determining the start and end position information of the initial keyword in the target string; based on each of the initial keywords and their start and end position information, the target string and its character position information, performing global optimal word combination detection to obtain the target keyword.

[0014] In one embodiment, after performing keyword matching based on the subset set and the keyword knowledge graph to obtain the target keyword, the method further includes: expanding the target keyword to the keyword knowledge graph.

[0015] In the second aspect, an embodiment of the present disclosure provides a device for generating a text message management strategy, including: an acquisition module for acquiring text messages to be processed; an extraction module for extracting character substrings based on the text messages to be processed to obtain a subset of substrings; a matching module for matching the subset of substrings with a keyword knowledge graph to obtain target keywords; the keyword knowledge graph is constructed based on preset keywords and their associated words, and the associated words are variants, extensions or substitutes of the preset keywords; a determination module for determining a text message interception strategy based on the target keywords and the keyword knowledge graph, so as to perform text message interception based on the text message interception strategy.

[0016] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising a processor and a memory storing a computer program, wherein when the processor executes the program, the method for generating a short message management policy described in the first aspect is implemented.

[0017] In a fourth aspect, an embodiment of the present disclosure provides a storage medium, which is a computer-readable storage medium and includes a computer program. When the computer program is executed by a processor, the method for generating a short message management strategy described in the first aspect is implemented.

[0018] The SMS management strategy generation method, device, electronic device and storage medium provided by the embodiments of the present disclosure can quickly and accurately identify new keywords in the SMS to be processed that are formed by variants, extensions or substitutions of keywords by combining a keyword knowledge graph constructed by preset keywords and their variants, extensions and substitutions with a substring set obtained by extracting character substrings from the SMS to be processed for keyword matching. Then, the SMS interception strategy can be quickly and accurately determined based on the target keywords obtained by matching and combined with the keyword knowledge graph, so that relevant personnel can refer to the SMS interception strategy to intercept junk SMS, thereby improving the accuracy of junk SMS interception. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present disclosure or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] FIG1 is a flow chart of a method for generating a short message management strategy according to an embodiment of the present disclosure;

[0021] FIG2 is a schematic diagram of a scenario of a method for generating a short message management strategy according to an embodiment of the present disclosure;

[0022] FIG3 is a schematic diagram of functional modules of an embodiment of a device for generating a short message management strategy according to the present disclosure;

[0023] FIG4 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] To make the objectives, technical solutions, and advantages of this disclosure more clear, the technical solutions of this disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of this disclosure. The described embodiments are part of the embodiments of this disclosure, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of this disclosure without making any creative efforts shall fall within the scope of protection of this disclosure.

[0025] First, the existing SMS content analysis technology related to the present disclosure and its existing defects are described as follows:

[0026] 1. SMS Text Classification Technology: SMS can be classified using artificial intelligence models. These supervised AI models extract features from the current message and classify it into its corresponding category. Text classification technology is suitable for text with distinct classification features. However, SMS texts are short and have sparse features, making it difficult to achieve good classification results. Furthermore, spam text features often evolve, and new spam content can appear. These changes in SMS text features can lead to poor classification results.

[0027] 2. SMS Text Clustering Technology: Text clustering compares the similarity of large numbers of SMS messages in an unsupervised manner and groups highly similar messages into the same category. Text clustering algorithms are suitable for processing big data and analyzing text with multiple categories and high uncertainty. In other words, text clustering technology is suitable for offline analysis of large amounts of SMS messages. However, variant spam messages contain a wide variety of keyword variations, and clustering may result in spam messages belonging to the same category being grouped into different categories.

[0028] 3. Keyword Combination Strategy Analysis: Keyword combination strategies are developed by experienced strategy experts and incorporate practical knowledge for identifying spam messages. Keyword combination strategy analysis uses "and" and "or" logic to match keywords appearing in messages to identify spam messages. However, in the era of artificial intelligence, criminals use automated methods to generate numerous and rapidly changing combinations of undesirable keyword variants. Manual detection of new variants or replacements is difficult, making it difficult to promptly identify and implement strategies for these new variants. This leads to missed messages containing these new variants, further reducing the accuracy of current spam message blocking.

[0029] The method, device, electronic device and storage medium for generating a short message management strategy provided by the present invention are described in detail below with reference to the embodiments and FIG. 1 to FIG. 4 .

[0030] 1 is a flow chart of a method for generating a short message management policy according to an embodiment of the present disclosure; and FIG. 2 is a scenario chart of a method for generating a short message management policy according to an embodiment of the present disclosure.

[0031] Specifically, referring to FIG. 1 , an embodiment of the present disclosure provides a method for generating a short message management policy. The method may include the following steps 100 to 400 .

[0032] In step 100, a short message to be processed is obtained.

[0033] It should be noted that the execution subject of the SMS management policy generation method provided in the embodiments of the present disclosure may be a server, a computer device, etc., such as a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA). The server or computer device of the present disclosure may be provided with or connected to a SMS management policy generation device, so that the SMS management policy generation method of the present disclosure can be completed by controlling the SMS management policy generation device.

[0034] The SMS to be processed in the present disclosure may be a SMS containing illegal words or a SMS without illegal words. The SMS to be processed may also be a SMS identified as spam.

[0035] In step 200, character substrings are extracted based on the short message to be processed to obtain a substring set.

[0036] The present disclosure can pre-process the SMS to be processed, specifically, can perform data cleaning on the SMS to be processed, and form a target character string from the characters remaining after the data cleaning.

[0037] Furthermore, character segmentation is performed on the target character string according to different number of character combination requirements, and the target character string is segmented into multiple character substrings, and a substring set is formed by each character substring.

[0038] In step 300, the subset set is matched with the keyword knowledge graph to obtain the target keyword.

[0039] It should be noted that, in the present disclosure, a keyword knowledge graph can be constructed based on pre-set keywords for illegal words. In the present disclosure, the pre-set keywords for illegal words can be words summarized based on manual experience.

[0040] Specifically, a keyword knowledge graph can be formed by presetting keywords and the relationships between keywords. The relationships between keywords include variant relationships, substitution relationships, and extended relationships.

[0041] The present disclosure uses statistical analysis and keyword combination strategies to determine the relationship between keywords, thereby constructing a keyword knowledge graph. The edge between two keywords in the keyword knowledge graph of the present disclosure is a unidirectional edge.

[0042] Suppose there is a set of keyword combination strategies S. Taking the strategies (A|B|C)&(D) as an example, the concepts of variant relationship, substitution relationship, and extended relationship between keywords are introduced.

[0043] Variant relationship between keywords: If there are two keywords A and B, A is a common word, some or all of the characters in B are different from A, but they have the same pinyin or similar glyph features. After reading B, people can associate B with A. Then there is a variant relationship between the two keywords, where the common word A is the entity (with the highest word frequency, and the word frequency is obtained by counting the keyword combination strategy set S), and B is the variant, and the direction is from the entity to the variant.

[0044] Substitution relationship between keywords: If two keywords A and C have "OR" logic in the keyword combination strategy, belong to the same spam message topic, and the cosine similarity of the vectors of A and C is high, then A and C are in a substitution relationship with each other, and the direction is from the word with higher frequency to the word with lower frequency.

[0045] Extended relationship between keywords: If two keywords A and D have "and" logic in the keyword combination strategy, and the number of times they appear in the strategy keyword set is less than a certain threshold, then there is an extended relationship between A and D, and the direction is from the lower frequency word to the higher frequency word.

[0046] Therefore, the present disclosure can first accurately match each character substring in the substring set with the keyword entry in the keyword knowledge graph, and use the successfully matched character substring as the first candidate keyword.

[0047] Furthermore, because variant spam text messages contain a wide variety of variant keywords, these variant keywords often include uncommon characters, and each Chinese character in the variant keyword has the same pronunciation or similar glyphs as the corresponding Chinese character in the original word, exact matching cannot discover new variant keywords. Therefore, for character substrings in the substring set that fail to match, a fuzzy matching algorithm can be used to further match keyword entries in the keyword knowledge graph to discover new variant keywords, and the matched keywords can be determined as the second candidate keywords.

[0048] Furthermore, a dynamic programming algorithm may be used to search for an optimal word combination based on the first candidate keyword, or both the first candidate keyword and the second candidate keyword, and the obtained keyword may be determined as the target keyword.

[0049] It should be noted that after performing keyword matching based on the subset set and the keyword knowledge graph to obtain the target keyword, it also includes: expanding the target keyword to the keyword knowledge graph.

[0050] After obtaining the target keywords, the present disclosure can automatically expand the target keywords to the keyword knowledge graph to enrich the keyword knowledge in the keyword knowledge graph, which helps to enhance the ability to identify new variant spam text messages.

[0051] In step 400, a text message interception strategy is determined based on the target keyword and the keyword knowledge graph, so as to perform text message interception based on the text message interception strategy.

[0052] After obtaining the target keywords, the present disclosure can query the knowledge of related words such as variants, extensions and substitutions of keywords in the target keywords based on the keyword knowledge graph, and generate SMS interception strategies that can intercept new variant SMS messages, so that strategy specialists can refer to them when formulating strategies, thereby improving the efficiency and quality of strategy formulation.

[0053] In the SMS management strategy generation method provided in the embodiment of the present disclosure, by constructing a keyword knowledge graph composed of preset keywords and their variants, extensions, and replacement related words, combined with keyword matching of a substring set obtained by extracting character substrings from the SMS to be processed, new keywords formed by variants, extensions, or replacements of keywords in the SMS to be processed can be quickly and accurately identified, and then the SMS interception strategy can be quickly and accurately determined based on the target keywords obtained by matching and combined with the keyword knowledge graph, so that relevant personnel can refer to the SMS interception strategy to intercept spam SMS, thereby improving the accuracy of spam SMS interception.

[0054] In one embodiment, character substrings are extracted based on the short message to be processed to obtain a substring set, including the following steps 201-202 (not shown in the drawings).

[0055] In step 201, invalid characters are cleared from the short message to be processed to obtain a target character string.

[0056] In step 202, character segmentation is performed on the target character string, and each character substring formed by the segmentation forms a substring set.

[0057] After obtaining the SMS to be processed, the present disclosure can perform data cleaning on invalid characters in the SMS to be processed, wherein the invalid characters include at least one of punctuation marks, English symbols, and emoticons.

[0058] Specifically, in the present disclosure, punctuation marks, English symbols, and emoticons in the short message to be processed can be removed, and only Chinese characters can be retained, and the target character string can be formed from the remaining Chinese characters.

[0059] Furthermore, the target string can be segmented into characters using K-shingle technology, and the character substrings formed by the segmentation form a substring set. K-shingle is a text processing technology used to split text data into continuous short segments.

[0060] In the present disclosure, for a target string, a shingle may be taken for every K consecutive characters, and a K-shingle is a set of all substrings consisting of K consecutive characters. In one embodiment, a 2-shingle, a 3-shingle, or a 4-shingle of the target string may be extracted.

[0061] This embodiment can remove noise information in the text messages to be processed through data cleaning, and can extract key features in the text messages to be processed through K-shingle extraction, so that the character substrings in the obtained substring set are more accurate, and thus the text message interception strategy determined based on the substring set is more accurate, which helps to improve the accuracy of spam text message interception.

[0062] Furthermore, the subset set is matched with the keyword knowledge graph to obtain the target keyword, including the following steps 301-305 (not shown in the accompanying drawings).

[0063] In step 301, a keyword knowledge graph is obtained.

[0064] In step 302, in the keyword knowledge graph, variant character matching is performed on the character substrings in the substring set to obtain the first candidate keyword.

[0065] In step 303, if there is at least one target character substring that fails to match in the variant character matching, character similarity matching is performed based on each target character substring to obtain a second candidate keyword.

[0066] Step 304: Determine at least one initial keyword based on the first candidate keyword, or the first candidate keyword and the second candidate keyword.

[0067] Step 305: Perform keyword matching based on the initial keywords to obtain target keywords.

[0068] A pre-built keyword knowledge graph can be obtained in the present disclosure.

[0069] Furthermore, character substrings such as 2-shingle, 3-shingle, and 4-shingle in the substring set can be precisely matched against pre-set keyword entries in the keyword knowledge graph. If the precisely matched entry is a variant of a keyword in the keyword knowledge graph, the variant relationship is used to automatically restore the entry to the keyword itself as the first candidate keyword. Successfully matched entries serve as important keyword features for subsequent SMS analysis.

[0070] Furthermore, for at least one target character substring that fails to match in the variant character matching process, similarity matching based on homophones or similar shapes can be performed based on each target character substring, thereby obtaining a second candidate keyword.

[0071] Furthermore, if only the first candidate keyword is matched based on the subset set, the first candidate keyword is used as the initial keyword.

[0072] If the first candidate keyword and the second candidate keyword are obtained based on the subset set matching, the first candidate keyword and the second candidate keyword are used together as initial keywords.

[0073] Furthermore, a dynamic programming algorithm can be used to find the best word combination from each initial keyword. Specifically, from each exact match and fuzzy match initial keyword, the word combination with the largest sum of characters in the matching text message can be found as the target keyword.

[0074] This embodiment performs keyword matching based on the substring set obtained by extracting character substrings from the SMS to be processed, and can quickly and accurately identify new keywords formed by variants, extensions or substitutions of keywords in the SMS to be processed. Then, the SMS interception strategy can be quickly and accurately determined based on the matched target keywords combined with the keyword knowledge graph, making it convenient for relevant personnel to refer to the SMS interception strategy to intercept spam SMS, thereby improving the accuracy of spam SMS interception.

[0075] Furthermore, character similarity matching is performed based on each target character substring to obtain a second candidate keyword, including the following steps 3031-3033 (not shown in the drawings).

[0076] In step 3031, character rareness detection is performed on each target character substring to obtain character rareness information.

[0077] In step 3032, uncommon substrings are determined based on the character uncommonness information.

[0078] In step 3033, character similarity matching is performed based on the uncommon substring to obtain a second candidate keyword.

[0079] In this disclosure, a fuzzy matching algorithm can be used to further match keyword entries in the knowledge graph to discover new variant keywords. This can be divided into two steps: filtering for uncommon character substrings (step 1 below) and fuzzy matching for homophones and similar shapes (step 2 below).

[0080] Step 1: Filter substrings containing uncommon characters: For each target character substring in the unmatched 2-shingle, 3-shingle, and 4-shingle, calculate the character uncommonness information to determine whether the target character substring contains uncommon characters. Target character substrings containing uncommon characters, hereinafter referred to as uncommon substrings, are retained. Perform keyword fuzzy matching on all uncommon substrings. The formula for calculating character uncommonness is as follows:

[0081] Where ch represents any Chinese character, and n is the number of times the Chinese character ch appears in the dictionary of common characters. The dictionary of common characters can be constructed using an open source news dataset. m is the critical frequency for distinguishing rare characters from common characters. When n is less than m, it is more likely to be an uncommon character. The uncommonness range of Chinese characters is between (0, 1], and the closer to 1, the higher the uncommonness of the character. In some embodiments of the present disclosure, m is set to 20, and the uncommon character threshold is set to 0.4, that is, if the uncommonness is greater than 0.4, the Chinese character is an uncommon character.

[0082] After obtaining the uncommon substrings, the second candidate keywords can be obtained by performing similarity matching on the uncommon substrings based on homophones or similar shapes.

[0083] Furthermore, character similarity matching is performed based on the uncommon substring to obtain a second candidate keyword, including the following steps 30331-30332 (not shown in the accompanying drawings).

[0084] In step 30331, based on the keyword knowledge graph, a dictionary of pinyin and stroke order sequences of keywords is constructed.

[0085] In step 30332, based on the pinyin and stroke order column dictionary, the uncommon substring is matched with homophones and similar shapes to obtain the second candidate keyword.

[0086] Step 2: Fuzzy matching of variant words with similar pronunciation and form: For the uncommon substrings selected in step 1, the fuzzy matching method of “similar pronunciation and form” is used to automatically discover variant keywords.

[0087] Specifically, first, all central node words (i.e., keywords) in the keyword knowledge graph are extracted to construct a dictionary of pinyin and stroke order sequences. Then, homophone or shape similarity matching is performed on each rare substring, and the matching ones are new variant words and are determined as the second candidate keywords. The shape similarity degree of two Chinese characters can be obtained by calculating the edit distance of the stroke order sequences of the two Chinese characters. If the edit distance is less than a set threshold (for example, 4), it is considered that the two Chinese characters are similar. For example, the pinyin of the keyword entry "baccarat" in the keyword knowledge graph is "baijiale", and the pinyin of the rare substring "栢迦泺" is "baijialuo". Among the two words, "baijia" and "栢迦" are homophone matches, and "le" and "泺" are shape similarity matches.

[0088] Examples of exact and fuzzy matching of keyword entries: Based on the keyword knowledge graph, exact and fuzzy matching of keyword entries are performed on 2-shingle, 3-shingle, and 4-shingle. In one example, the keywords hit in the exact matching step are: "preferential", "please keep", "address". Then, rare substrings are screened from the non-exactly matched substrings to obtain 2-shingle, 3-shingle, and 4-shingle containing rare Chinese characters; finally, "homophone and shape similarity" variant word fuzzy matching is performed to obtain keyword variants.

[0089] This embodiment can automatically identify keywords in new variant spam messages based on the "homophone and shape similarity" characteristics of variant keywords, so that the target keywords obtained by matching based on the identified keywords can be combined with the keyword knowledge graph to quickly and accurately determine the short message interception strategy, facilitating relevant personnel to refer to the short message interception strategy for spam message interception, and thus improving the accuracy of spam message interception.

[0090] In one embodiment, keyword matching is performed based on each initial keyword to obtain a target keyword, including the following steps 3051-3053 (not shown in the drawings).

[0091] In step 3051, the character position information of each character in the target string corresponding to the to-be-processed short message is determined.

[0092] In step 3052, the start and end position information of the initial keyword in the target string is determined.

[0093] In step 3053, based on each initial keyword and its start and end position information, the target string and its character position information, global best word combination detection is performed to obtain the target keyword.

[0094] The present disclosure can find the best word combination from each initial keyword as the target keyword through a dynamic programming algorithm.

[0095] In one embodiment, for the short message "Good news, Pakalo Limited-time Offer, deposit 999 and enjoy up to 9999, address: ****.space", the set of keyword entries with exact matches is {offer, save, address}, and the set of variant keyword entries with fuzzy matches is {news, limited-time, enjoy, good news, Pakalo, limited-time offer}. Among them, there are overlapping words. For example, "news" and "good news", where "good news" is the best match, and "limited-time offer" and "save", where "limited-time offer" is the correct match.

[0096] The matching scheme for the best word combination should find a word combination in the set of keywords with exact and fuzzy matches (i.e., the initial keywords) that has the maximum sum of characters matched in the short message. Specifically, there are:

[0097] First, use the AC (Aho-Corasick) automaton to find the start and end positions of all candidate words (i.e., the initial keywords) in the preprocessed short message (i.e., the target string), and determine the character position information of each character in the target string corresponding to the short message to be processed in the target string.

[0098] Furthermore, based on each initial keyword and its start and end position information, the target string and its character position information, a table containing a digital matrix can be generated.

[0099] The first column of the table is all the initial keywords and their start and end positions. The initial keywords are arranged in ascending order of the start position. When the start positions of two initial keywords are the same, they are arranged in ascending order of the end position. The second row of the table is the position label of the short message, and the first row is the content of the preprocessed short message (i.e., the target string). For example: for the entry "good news", the start position index in the preprocessed short message sequence is 0, and the end position index of "good news" is 2, then the start and end positions of "good news" are recorded as (0, 3), which is a list notation. Similarly, the start and end positions of "news" are (1, 3).

[0100] In one embodiment, a 9*17 digital matrix can be formed in the table. The element value at the i-th row and j-th column of the matrix is denoted as a i,j , representing that when the short message content is from the 0th position to the jth position, and the keywords to be matched are the first 0th word to the i-th word, the maximum sum of characters matched in the short message. For example, a 0,0 = 0, representing that when the short message content is "good", and the keyword to be matched is only "good news", the maximum sum of characters matched in the short message is 0. And for example, a 1,2 = 3, representing that when the short message is "good news", and the keywords to be matched are "good news" and "news", the maximum sum of characters matched in the short message is 3. It can be obtained that after filling this digital matrix, the value in the lower right corner a8,16 It is the maximum value of the sum of the characters in the text message that can be matched by all the words to be matched.

[0101] The digital matrix can be gradually filled through the following recurrence formula:

[0102] For the 0th row (i = 0) of the digital matrix:

[0103] When j < end(w0), a 0,j = 0; It means that there is no matching word at this time, and the sum of the characters matched in the text message is equal to 0;

[0104] When j = end(w0), a 0,j = len(w0); It means that the first word is successfully matched, and at this time the sum of the characters matched in the text message is equal to the total number of characters of this word;

[0105] When j > end(w0), a 0,j = a 0,j-1 ; It means that after the word being matched ends in the text message, there is no more word to be matched in the 0th row, and at this time the sum of the characters matched in the text message no longer changes.

[0106] Where a 0,j represents the element in the 0th row and the jth column of the digital matrix; w0 represents the 0th word to be matched, which is "好消螅" in this example; end(w0) is the ending position of the word w0, for example, the ending position of "好消螅(0, 3)" is 3 - 1 = 2; len(w0) is the number of characters of the word w0, which is 3.

[0107] For the ith row (i > 0) of the digital matrix, there is the following recurrence formula:

[0108] When j < end(w j ), a i,j = a i-1,j ; It means that when the word w j has not been successfully matched, the sum of the characters matched in the text message at this time is equal to the matching status of the jth column in the previous row.

[0109] When j = end(w j ), if k = j - len(w j ) < 0, a i,j = len(w j ), otherwise a i,j = max{a i,k + len(w j ), a i-1,j}; It means that the word w j is successfully matched. At this time, it is necessary to compare and select the sum of the characters in the text message when the word w j is selected with the sum of the characters in the text message when the word w j is not selected, ai,j Choose a larger value;

[0110] When j>end(w j ), a i,j =a i,j-1 ; means that when the matched word ends in the text message, the sum of the matched characters in the text message is equal to the value on its left.

[0111] After the digital matrix is ​​filled in, we can start from the lower right corner of the digital matrix and backtrack upwards and then leftwards (that is, first determine whether the word in the current row is selected. If it is selected, backtrack forward to remove the length of the current word and continue to determine), and output the best word combination. Assume that the current element value is a ij :

[0112] 1. If a i-1,j =a i,j Then i=i-1. Repeat this cycle until a i-1, j≠a i,j ;

[0113] 2. If a i,j-1 =a i,j Then j=j-1. Repeat this cycle until a i,j-1 ≠a i,j , at this time the word in row i is selected as a candidate word, and output is w i . At this time, let j=j-len(w i ), and then repeat the above steps 1 and 2 until the upper left corner of the digital matrix is ​​traversed.

[0114] Through the above cycle, several optimal word combinations can be output as target keywords.

[0115] This embodiment searches for the best word combination as the target keyword from each initial keyword based on a dynamic programming algorithm, and can quickly and accurately identify new keywords formed by variants, extensions or substitutions of keywords in the text messages to be processed. Then, based on the matched target keywords and combined with the keyword knowledge graph, the text message interception strategy can be quickly and accurately determined, making it convenient for relevant personnel to refer to the text message interception strategy to intercept spam text messages, thereby improving the accuracy of spam text message interception.

[0116] FIG2 is a schematic diagram of a scenario of a method for generating a short message management policy according to an embodiment of the present disclosure. Referring to FIG2 , in one embodiment, the method for generating a short message management policy according to the present disclosure may include the following steps:

[0117] Get an input text message, remove punctuation marks, English, emoticons, etc. from the text message, and only retain Chinese characters, thereby completing the text message preprocessing.

[0118] Furthermore, K-shingle is extracted from the information of the preprocessed SMS, and the character substrings in the substring set extracted by K-shingle are precisely matched with the keyword entries in the keyword knowledge graph, and the character substrings that are not precisely matched are fuzzy matched with the keyword entries.

[0119] For exact matching and fuzzy matching keyword sets, a dynamic programming algorithm is used to find the best word combination as the target keyword.

[0120] Furthermore, the target keyword is automatically expanded into the keyword knowledge graph as variant knowledge of the knowledge graph.

[0121] In addition, based on the keyword knowledge graph, the variants, extensions, and alternative relationship word knowledge of the target keyword are queried, and the SMS interception strategy that can intercept new variant SMS messages is automatically generated and the process is ended. This is used as a reference for strategy specialists when formulating strategies to improve the efficiency and quality of strategy formulation.

[0122] The SMS management strategy generation method disclosed in the present invention can make up for the shortcomings of SMS classification and clustering technologies in the existing technology, which have poor classification and clustering effects due to sparse and scattered short text features. In addition, it can use the variant, substitution, and extended relationship knowledge of keywords in the keyword knowledge graph to automatically analyze SMS texts, discover new variant words in SMS, and automatically generate strategies that can intercept new variant SMS messages for reference by strategy specialists, thereby improving strategy formulation efficiency and enhancing the strategy's recall and precision rates.

[0123] Furthermore, the present disclosure also provides a device for generating a short message management strategy.

[0124] Refer to FIG3 , which is a schematic diagram of functional modules of an embodiment of a device for generating a short message management strategy according to the present disclosure.

[0125] The short message management strategy generating device includes:

[0126] The acquisition module 310 is used to acquire the short message to be processed;

[0127] An extraction module 320 is configured to extract character substrings based on the short message to be processed to obtain a substring set;

[0128] A matching module 330 is configured to match the subset set with a keyword knowledge graph to obtain a target keyword; the keyword knowledge graph is constructed based on preset keywords and their associated words, where the associated words are variants, extensions, or replacements of the preset keywords;

[0129] The determination module 340 is used to determine a text message interception strategy based on the target keyword and the keyword knowledge graph, so as to perform text message interception based on the text message interception strategy.

[0130] The SMS management strategy generation device provided by the embodiment of the present disclosure can quickly and accurately identify new keywords in the SMS to be processed that are formed by variants, extensions or substitutions of keywords by combining a keyword knowledge graph constructed from preset keywords and their variants, extensions, and substitutions with a substring set obtained by extracting character substrings from the SMS to be processed for keyword matching. Then, the SMS interception strategy can be quickly and accurately determined based on the target keywords obtained by matching and combined with the keyword knowledge graph, making it convenient for relevant personnel to refer to the SMS interception strategy to intercept junk SMS, thereby improving the accuracy of junk SMS interception.

[0131] In one embodiment, the extraction module 320 is specifically used to: clean up invalid characters in the SMS to be processed to obtain a target character string; the invalid characters include at least one of punctuation marks, English symbols, and emoticons; perform character segmentation on the target character string, and form a substring set from each character substring formed by the segmentation.

[0132] In one embodiment, the matching module 330 is specifically used to: obtain a keyword knowledge graph; perform variant character matching on the character substrings in the substring set in the keyword knowledge graph to obtain a first candidate keyword; if there is at least one target character substring that fails to match in the variant character matching, perform character similarity matching based on each target character substring to obtain a second candidate keyword; determine at least one initial keyword based on the first candidate keyword, or the first candidate keyword and the second candidate keyword; perform keyword matching based on each initial keyword to obtain a target keyword.

[0133] In one embodiment, the matching module 330 includes a first matching unit, which is used to: perform character rareness detection on each target character substring to obtain character rareness information; determine rare substrings based on the character rareness information; and perform character similarity matching based on the rare substrings to obtain a second candidate keyword.

[0134] In one embodiment, the matching module 330 includes a second matching unit, which is used to: construct a pinyin and stroke order sequence dictionary of the keyword based on the keyword knowledge graph; based on the pinyin and stroke order sequence dictionary, perform homophonic and similar matching on the uncommon substring to obtain a second candidate keyword.

[0135] In one embodiment, the matching module 330 includes a third matching unit, which is used to: determine the character position information of each character in the target string corresponding to the short message to be processed in the target string; determine the start and end position information of the initial keyword in the target string; and perform global optimal word combination detection based on each of the initial keywords and their start and end position information, the target string and its character position information to obtain the target keyword.

[0136] In one embodiment, the matching module 330 is further configured to expand the target keyword into the keyword knowledge graph.

[0137] Figure 4 illustrates a schematic diagram of the physical structure of an electronic device. As shown in Figure 4, the electronic device may include: a processor (processor) 410, a communication interface (Communication Interface) 420, a memory (memory) 430 and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440.

[0138] The processor 410 may call the computer program in the memory 430 to execute the steps of the method for generating a short message management policy, for example, including:

[0139] Get pending SMS messages;

[0140] Extracting character substrings based on the short message to be processed to obtain a substring set;

[0141] Matching the subset set with a keyword knowledge graph to obtain target keywords, wherein the keyword knowledge graph is constructed based on preset keywords and associated words, and the associated words are variants, extensions, or replacements of the preset keywords;

[0142] A text message interception strategy is determined based on the target keyword and the keyword knowledge graph, so as to intercept text messages based on the text message interception strategy.

[0143] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0144] On the other hand, an embodiment of the present disclosure further provides a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores a computer program, which is used to cause a processor to execute the steps of the methods provided in the above embodiments, for example, including:

[0145] Get pending SMS messages;

[0146] Extracting character substrings based on the short message to be processed to obtain a substring set;

[0147] Matching the subset set with a keyword knowledge graph to obtain target keywords, wherein the keyword knowledge graph is constructed based on preset keywords and associated words, and the associated words are variants, extensions, or replacements of the preset keywords;

[0148] A text message interception strategy is determined based on the target keyword and the keyword knowledge graph, so as to intercept text messages based on the text message interception strategy.

[0149] The computer-readable storage medium may be any available medium or data storage device that can be accessed by a processor, including but not limited to magnetic storage (e.g., floppy disks, hard disks, magnetic tapes, magneto-optical disks (MOs), etc.), optical storage (e.g., CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (e.g., ROMs, EPROMs, EEPROMs, non-volatile memories (NAND FLASHs), solid-state drives (SSDs), etc.). In one embodiment of the present disclosure, the computer-readable storage medium may be a non-transitory computer-readable storage medium.

[0150] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0151] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present disclosure.

Claims

1. A method for generating a short message management strategy, comprising: Get pending SMS messages; Extracting character substrings based on the short message to be processed to obtain a substring set; Matching the subset set with a keyword knowledge graph to obtain a target keyword, wherein the keyword knowledge graph is constructed based on preset keywords and associated words, and the associated words are variants, extensions, or substitutes of the preset keywords; A text message interception strategy is determined based on the target keyword and the keyword knowledge graph, so as to perform text message interception based on the text message interception strategy.

2. The method for generating a short message management strategy according to claim 1, wherein: The matching of the subset set with the keyword knowledge graph to obtain the target keyword includes: Get keyword knowledge graph; In the keyword knowledge graph, performing variant character matching on the character substrings in the substring set to obtain a first candidate keyword; If there is at least one target character substring that fails to match in the variant character matching, then character similarity matching is performed based on each of the target character substrings to obtain a second candidate keyword; Determine at least one initial keyword based on the first candidate keyword, or the first candidate keyword and the second candidate keyword; Keyword matching is performed based on each of the initial keywords to obtain a target keyword.

3. The method for generating a short message management strategy according to claim 2, wherein: The performing character similarity matching based on each of the target character substrings to obtain a second candidate keyword includes: Performing character rareness detection on each of the target character substrings to obtain character rareness information; Determining an uncommon substring based on the character uncommonness information; Based on the uncommon substring, character similarity matching is performed to obtain a second candidate keyword.

4. The method for generating a short message management strategy according to claim 3, wherein: The method of performing character similarity matching based on the uncommon substring to obtain a second candidate keyword includes: Based on the keyword knowledge graph, construct a dictionary of the pinyin and stroke order of keywords; Based on the pinyin and stroke order sequence dictionary, the uncommon substring is matched with homophones and similar shapes to obtain a second candidate keyword.

5. The method for generating a short message management strategy according to claim 2, wherein: The character substring extraction is performed based on the short message to be processed to obtain a substring set, including: Cleaning invalid characters from the short message to be processed to obtain a target character string; the invalid characters include at least one of punctuation marks, English symbols, and emoticons; The target character string is segmented into characters, and each character substring formed by the segmentation forms a substring set.

6. The method for generating a short message management strategy according to claim 5, wherein: The keyword matching based on each of the initial keywords to obtain the target keyword includes: Determine character position information of each character in a target character string corresponding to the short message to be processed in the target character string; Determine the starting and ending position information of the initial keyword in the target character string; Based on the initial keywords and their start and end position information, the target character string and its character position information, a global optimal word combination detection is performed to obtain the target keyword.

7. The method for generating a short message management strategy according to any one of claims 1 to 6, wherein: After matching the subset set with the keyword knowledge graph to obtain the target keyword, the method further includes: Expand the target keyword to the keyword knowledge graph.

8. A device for generating a short message management strategy, comprising: The acquisition module is used to obtain the short messages to be processed; An extraction module, used for extracting character substrings based on the short message to be processed to obtain a substring set; A matching module, used for matching the subset set with a keyword knowledge graph to obtain a target keyword, wherein the keyword knowledge graph is constructed based on preset keywords and associated words, and the associated words are variants, extensions or substitutes of the preset keywords; A determination module is used to determine a text message interception strategy based on the target keyword and the keyword knowledge graph, so as to perform text message interception based on the text message interception strategy.

9. An electronic device, comprising a processor and a memory storing a computer program, wherein the processor implements the method for generating a short message management strategy according to any one of claims 1 to 7 when executing the computer program.

10. A storage medium, which is a computer-readable storage medium, comprising a computer program, wherein when the computer program is executed by a processor, the method for generating a short message management strategy according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method for generating a short message management strategy according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Character string processing method, device and apparatus

    CN113849706A

  • Method and device for generating fraud-related short message interception template

    CN114786184A

  • Short message management strategy generation method and device, electronic equipment and storage medium

    CN118803604A

  • Method and device for automatically proofreading chinese document

    JP1998269204A

  • Method and apparatus for identifying new words in spam message, and electronic device

    WO2020052547A1