Junk short message identification method and device, electronic equipment and storage medium

By using the dictionary tree structure for spam recognition, the recognition delay problem caused by the increase in the number of keyword rules is solved, and efficient and accurate spam recognition and timely alarms are achieved.

CN120087363APending Publication Date: 2025-06-03ULTRAPOWER SOFTWARE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510056619.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

As the number of keyword rules increases, the global search efficiency of the keyword matching algorithm and the multi-rule matching efficiency of the same keyword will be linearly reduced, resulting in delayed spam recognition and inability to trigger alarms in time.

Method used

The dictionary tree structure is used for spam recognition, and the first letter of the first word in the text message segmentation is obtained and the pre-constructed dictionary tree is matched layer by layer to determine whether the target text message is spam.

Benefits of technology

It improves the efficiency and accuracy of spam message recognition, reduces unnecessary full retrieval, and can trigger alarms in a timely manner, avoid identification delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087363A_ABST
    Figure CN120087363A_ABST
Patent Text Reader

Abstract

The invention relates to the field of information identification, in particular to a junk short message identification method and device, electronic equipment and a storage medium. The method comprises the following steps: firstly, obtaining a to-be-processed target short message, extracting the text content of the short message, and then carrying out word segmentation processing on the text content of the short message to obtain a set containing a plurality of short message segmented words; and then, extracting an initial letter of a first word in the short message segmented words, and performing layer-by-layer matching with a pre-constructed dictionary tree based on the initial letter to determine whether the target short message is a junk short message. By means of the mode, the situation that in a traditional multi-keyword multi-mode matching scheme, the global retrieval efficiency and the same keyword multi-rule matching efficiency are linearly reduced along with increase of the number of keyword rules is avoided. The structure of the dictionary tree enables the matching of the keywords to be more efficient, possible branches are quickly positioned through initials, and then step-by-step deep matching is performed, so that unnecessary full-amount retrieval is reduced, and the global retrieval efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information recognition, and particularly to a method, device, electronic device and storage medium for identifying junk messages. Background Art

[0002] With the development of the Internet industry, industry short messages are widely used in many fields, providing an effective means for product promotion and service maintenance. However, some merchants send mass advertising messages, resulting in a sharp increase in junk message complaints. At the same time, problems such as illegal debt collection are developing wildly, bringing a bad impact to borrowers. To cooperate with the Ministry of Industry and Information Technology's special governance work on junk messages, China Unicom's industry short message capability platform uses the method of configuring junk keyword rules to manage and warn about junk messages. In the existing junk keyword recognition scenario for junk messages, a multi-keyword and multi-pattern matching scheme is used. However, as the number of keyword rules increases, the global retrieval efficiency of the keyword matching algorithm and the multi-rule matching efficiency of the same keyword will linearly decrease, resulting in a delay in junk message recognition and an inability to trigger an alarm in a timely manner. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a method, device, electronic device and storage medium for identifying junk messages, so as to solve the problem that as the number of keyword rules increases, the global retrieval efficiency of the keyword matching algorithm and the multi-rule matching efficiency of the same keyword will linearly decrease, resulting in a delay in junk message recognition and an inability to trigger an alarm in a timely manner.

[0004] In a first aspect, an embodiment of the present invention provides a method for identifying junk messages, the method comprising:

[0005] Obtain a target message to be processed, and extract the message text content of the target message;

[0006] Perform word segmentation processing on the message text content to obtain a word segmentation set, where the word segmentation set includes a plurality of message word segments;

[0007] Successively extract the first letters of the first words in the message word segments, and perform layer-by-layer matching based on the first letters and a pre-constructed trie tree to determine whether the target message is a junk message; wherein the first layer of the pre-constructed trie tree is an empty node, the second layer of the trie tree includes a plurality of letter nodes, the third layer of the trie tree includes at least one pinyin node associated with each letter node, and the fourth layer of the trie tree includes at least one first keyword node associated with each pinyin node, and the pinyin represented by the pinyin node is the pinyin of the first character of the keyword represented by the first keyword node.

[0008] Further, the step of determining whether the target short message is a spam short message by performing layer-by-layer matching based on the first letter and a pre-constructed trie tree includes:

[0009] Matching the first letter of the short message word segmentation with the letters in the second level of the trie tree to obtain the target letter node hit by the first letter;

[0010] Obtaining the pinyin of the first character in the short message word segmentation, and traversing the third level of the trie tree to determine whether there is a target pinyin node that matches the pinyin of the first word;

[0011] If there is the target pinyin node, traverse the fourth level corresponding to the target pinyin node to determine whether there is a first keyword node that matches the short message word segmentation;

[0012] If there is a first keyword node that matches the short message word segmentation, and the first keyword node matched by the short message word segmentation carries a preset label, then determine that the target short message is a spam short message;

[0013] If there is no pinyin node in the third level of the trie tree that matches the pinyin of the first word, or if there is no matching first keyword node in the fourth level corresponding to the target pinyin node, then determine that the target short message is not a spam short message.

[0014] Further, the first keyword node is associated with multiple keyword combinations, and the multiple keyword combinations are stored in the form of a subtree structure of the trie tree. Each level of the subtree structure includes at least one keyword; the first level of the subtree structure corresponds to the first keyword node in the trie tree;

[0015] The method further includes:

[0016] If the first keyword node that matches the short message word segmentation does not carry a preset label, then obtain the second keyword located in the second level of the subtree structure;

[0017] Determine whether there is a short message word segmentation in the word segmentation set that matches the second keyword;

[0018] If there is a short message word segmentation that matches the second keyword, then detect whether the matching second keyword carries a preset label;

[0019] If the matching second keyword carries a preset label, then determine that the target short message is a spam short message;

[0020] Or, if there is no short message word segmentation that matches the second keyword, then obtain the next short message word segmentation in the word segmentation set, and repeat the step of matching the next short message word segmentation with the first keyword located in the first level of the subtree structure.

[0021] Further, the method further includes:

[0022] If the matched second keyword does not carry a preset label, it is detected whether there is a corresponding third layer for the matched second keyword in the subtree structure;

[0023] If there is, the third keyword at the third level corresponding to the matched second keyword is obtained;

[0024] It is determined whether there is a text message word segmentation in the word segmentation set that matches the third keyword;

[0025] If there is a text message word segmentation that matches the third keyword, it is detected whether the matched third keyword carries a preset label. If the matched third keyword carries a preset label, it is determined that the target text message is a spam text message.

[0026] Further, the method for constructing the trie tree includes:

[0027] The at least one spam keyword rule is split to obtain corresponding keyword combinations; the first spam keyword in the keyword combination is determined, and a subtree structure is constructed by using the keyword combination and the corresponding first spam keyword;

[0028] The pinyin and the first letter of the first character of the first spam keyword in the keyword combination are determined;

[0029] An initial tree structure is obtained, and an empty node is created in the initial tree structure;

[0030] An alphabet node corresponding to the first letter of the first character of the first spam keyword is created with the empty node as the root node;

[0031] Corresponding pinyin nodes are created in each alphabet node based on the pinyin of the first character in each first spam keyword;

[0032] The subtree structure corresponding to the first spam keyword is added under the pinyin node corresponding to the initial tree to obtain the trie tree.

[0033] Further, after obtaining the trie tree, it further includes:

[0034] The proportion of keyword combinations under each alphabet node in the trie tree is calculated, and the alphabet nodes with a proportion higher than the preset threshold are used as the first alphabet nodes and the alphabet nodes with a proportion lower than the preset threshold are used as the second alphabet nodes;

[0035] The first spam keyword in the keyword combination under the first alphabet node is re-determined to adjust the keyword combination in the first alphabet node to the second alphabet node to obtain an updated trie tree.

[0036] Further, after obtaining the trie tree, the method further includes:

[0037] Obtaining a new junk keyword rule corresponding to each letter node;

[0038] Splitting the junk keyword rule to obtain a new keyword combination;

[0039] Updating the trie tree by using the new keyword combination.

[0040] In a third aspect, an embodiment of the present invention provides a spam message recognition device, the device includes:

[0041] An obtaining module, configured to obtain a target message to be processed and extract the message text content of the target message;

[0042] A processing module, configured to perform word segmentation processing on the message text content to obtain a word segmentation set, where the word segmentation set includes a plurality of message word segments;

[0043] An analysis module, configured to sequentially extract the first letter of the first word in the message word segments, and perform layer-by-layer matching based on the first letter and a pre-constructed trie tree to determine whether the target message is a spam message; wherein, the first layer of the pre-constructed trie tree is an empty node, the second layer of the trie tree includes a plurality of letter nodes, the third layer of the trie tree includes at least one pinyin node associated with each letter node, and the fourth layer of the trie tree includes at least one first keyword node associated with each pinyin node, and the pinyin represented by the pinyin node is the pinyin of the first character of the keyword represented by the first keyword node.

[0044] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the method according to the first aspect or any corresponding embodiment thereof.

[0045] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the method according to the first aspect or any corresponding embodiment thereof.

[0046] In the embodiments of the present application, first, the target short message to be processed is obtained and its short message text content is extracted. Then, the short message text content is segmented to obtain a set containing multiple short message segments. Next, the first letter of the first word in the short message segments is extracted, and based on this first letter, a hierarchical matching is performed with a pre-constructed trie tree to determine whether the target short message is a spam message. This method avoids the situation where the global retrieval efficiency and the matching efficiency of the same keyword with multiple rules linearly decrease as the number of keyword rules increases in the traditional multi-keyword and multi-pattern matching scheme. The structure of the trie tree makes the matching of keywords more efficient. By quickly locating the possible branches through the first letter and then gradually delving deeper into the matching, unnecessary full-scale retrievals are reduced, thereby improving the global retrieval efficiency. For the multi-rule matching of the same keyword, more targeted searches can also be performed on specific paths of the trie tree, avoiding inefficient repeated searches, and thus reducing the latency of spam message identification, enabling timely triggering of alarms. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 is a schematic flowchart of a method for identifying spam messages according to some embodiments of the present invention;

[0049] Figure 2 is a schematic diagram of spam message identification according to some embodiments of the present invention;

[0050] Figure 3 is a schematic flowchart of a method for identifying spam messages according to some embodiments of the present invention;

[0051] Figure 4 is a schematic structural diagram of a trie tree according to some embodiments of the present invention;

[0052] Figure 5 is a structural block diagram of a spam message identification device according to an embodiment of the present invention;

[0053] Figure 6 is a schematic hardware structure diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] According to an embodiment of the present invention, there is provided a method, apparatus, electronic device, and storage medium for identifying spam messages. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0056] In this embodiment, there is provided a method for identifying spam messages. Figure 1 It is a flowchart of a method for identifying spam messages according to an embodiment of the present invention, as Figure 1 shown, and the process includes the following steps:

[0057] Step S101, obtain a target message to be processed, and extract the message text content of the target message.

[0058] In the embodiments of the present application, first, a target message to be processed is obtained. This target message can come from various channels, such as a message receiving system, a storage record in a database, etc. After the target message is determined, the message text content of the target message is further extracted. The pure text content part can be accurately separated from the data structure of the target message through a specific programming interface or data parsing method. This message text content will serve as the basic data for subsequent word segmentation and spam rule matching, providing the original material for determining whether the message is a spam message.

[0059] Step S102, perform word segmentation processing on the message text content to obtain a word segmentation set, where the word segmentation set includes multiple message word segments.

[0060] In the embodiment of the present application, when processing the extracted SMS text content, the constructed HanLP custom word segmenter is used to perform word segmentation on the SMS text content. This word segmenter is custom-configured according to the full set of spam SMS rule keywords and can more accurately divide the SMS text into meaningful word units. After word segmentation, a word segmentation set is obtained, which contains multiple SMS word segments. Each SMS word segment is an independent word segmented from the original SMS text, and these word segments will be used as the basic elements for subsequent retrieval and matching in the spam rule trie. For example, for an SMS text "Come and participate in the preferential activity", after word segmentation, multiple SMS word segments such as "Come", "participate", "preferential", and "activity" may be obtained, forming a word segmentation set, preparing for further determination of whether the SMS is a spam SMS.

[0061] Step S103: Sequentially extract the first letters of the first words in the SMS word segments, and perform layer-by-layer matching based on the first letters with the pre-constructed trie to determine whether the target SMS is a spam SMS; wherein, the first layer of the pre-constructed trie is an empty node, the second layer of the trie includes multiple letter nodes, the third layer of the trie includes at least one pinyin node corresponding to each letter node, and the fourth layer of the trie includes at least one first keyword node corresponding to each pinyin node. The pinyin represented by the pinyin node is the pinyin of the first character of the keyword represented by the first keyword node.

[0062] In the embodiment of the present application, during the process of determining whether the target SMS is a spam SMS, for each SMS word segment in the word segmentation set, the first letter of its first word is extracted. Starting from this first letter, layer-by-layer matching is performed in the pre-constructed trie. The trie has an empty node as the root node, and the second-layer tree nodes are letter nodes. Each letter node stores the child node information in a hash table data structure. First, the extracted first letter is searched in the second-layer letter nodes. If the letter node exists, proceed to the next step; if not, process the next SMS word segment. If the letter node exists, then obtain the pinyin of the first word of the word segment, and starting from the third-layer tree nodes, determine whether there is a corresponding pinyin node. If the pinyin node exists, search for the current word segment in the fourth-layer first keyword nodes. If the current word segment is found and the current node is marked as the end point of the spam rule or the non-leaf node is configured with the spam keyword identifier "spam", then it can be determined that the content of the target SMS is spam content.

[0063] First, construct the structure of the dictionary tree. Use an empty node as the first level of the dictionary tree, that is, the root node. This root node does not store specific data, but only serves as the starting point of the entire dictionary tree. For the second level of the dictionary tree, the 26 letters are used as independent nodes. The data storage structure of each letter node uses a hash table so that child node information can be quickly retrieved and whether a specific letter node exists can be quickly determined. The third level is determined based on the spam keywords of the spam text messages. For each letter node, based on the words that begin with the letter in the spam keyword, extract the pinyin of the first character as the pinyin node of the third level. Each pinyin node is associated with multiple keywords, which are obtained by analyzing and summarizing a large number of spam text messages.

[0064] For example, if there are junk keywords "free" and "receive", and the first letter is "M", extract the pinyin "mian" of the first letter "免", create a "mian" pinyin node under the "M" letter node, and associate the keyword "free" with the pinyin node, and then associate the keyword "receive" with the keyword "free".

[0065] Specifically, the first letter is matched layer by layer with the pre-built dictionary tree to determine whether the target SMS is a spam SMS, including the following steps A1-A5:

[0066] Step A1, using the pinyin of the first word in the short message segmentation, and extracting the first letter of the pinyin of the first word.

[0067] Specifically, for the SMS word segmentation, obtain its first character, and then find the pinyin corresponding to the first character through the pre-built Chinese character and pinyin mapping library. Extract the first letter from the obtained pinyin to complete the entire process, providing basic data for subsequent operations such as spam SMS judgment based on the dictionary tree. For example, for the SMS "Promotional activities are waiting for you", after word segmentation, we get ["discount", "activity", "wait", "you come"], the first character of "discount" "优", find the pinyin "you" through the mapping library, and then extract the first letter "y", and so on to process the entire SMS content.

[0068] Step A2, matching the first letter with the letter nodes in the second level of the dictionary tree to obtain the target letter node hit by the first letter.

[0069] Specifically, the second level of the trie contains multiple letter nodes. Then, obtain the first letter of the pinyin of the first character in the text segmentation of the short message. Next, starting from the second level of the trie, traverse each letter node in sequence. Compare the obtained first letter with the value of the currently traversed letter node. If the two match, then this letter node is the target letter node where the first letter hits. For example, if the first letter is 'y', when traversing the second level and encountering the letter node 'y', it is determined as the target letter node.

[0070] Step A3: Traverse the third level of the trie to determine whether there is a target pinyin node that matches the pinyin of the first character.

[0071] Specifically, after obtaining the target letter node where the first letter hits, enter the third level of the trie through the pointer or index of this node. The third level contains multiple pinyin nodes associated with the target letter node. Then, traverse these pinyin nodes in sequence, and compare the pinyin represented by each pinyin node with the pinyin of the first character in the text segmentation of the short message one by one. If during the traversal, it is found that the pinyin of a certain pinyin node is exactly the same as the pinyin of the first character, then this pinyin node is the target pinyin node to be found; if no match is found after traversing all the pinyin nodes in the third level, it means that there is no target pinyin node that matches the pinyin of the first character.

[0072] Step A4: If there is a target pinyin node, traverse the fourth level corresponding to the target pinyin node to determine whether there is a first keyword node that matches the text segmentation of the short message.

[0073] Specifically, following the association relationship of the target pinyin node, enter its corresponding fourth level, which contains several first keyword nodes associated with the target pinyin node. Subsequently, traverse each first keyword node in the fourth level in sequence, and compare the keywords represented by these keyword nodes with each word obtained from the text segmentation of the short message one by one. During the comparison process, as long as it is found that the keyword represented by a certain first keyword node is exactly the same as a certain word in the text segmentation of the short message, then this first keyword node is the node that matches the text segmentation of the short message; if after traversing all the first keyword nodes in this level, no exactly matching word is found, it means that there is no first keyword node that matches the text segmentation of the short message.

[0074] For example, the text segmentation of the short message has words such as 'preferential' and 'activity'. The letter represented by the target letter node is 'y', the target pinyin node associated with the target letter node 'y' in the third level is 'you', and the first keyword associated with the target pinyin node in the fourth level is 'preferential'. When traversing to the keyword node 'preferential', it is found that it matches the 'preferential' in the text segmentation of the short message, so it is determined that there is a matching first keyword node.

[0075] Step A5: If there is a first keyword node that matches the SMS segmentation, and the first keyword node that matches the SMS segmentation carries a preset tag, then the target SMS is determined to be a spam SMS.

[0076] Specifically, when it is determined that there is a first keyword node that matches the SMS segmentation, the preset tag carried by the matching first keyword node is immediately checked. The preset tag is used to mark the spam keyword rule. This preset tag is set in advance based on the relevant characteristics of spam SMS and is assigned to each keyword node. It should be understood that if the keyword node is a leaf node, it is considered to carry the preset tag. If the check finds that the preset tag carried by the matching first keyword node indicates that it is an identifier related to spam SMS (for example, the tag value is a set value that specifically represents spam SMS), then the target SMS can be determined to be a spam SMS.

[0077] Step A6, if there is no pinyin node matching the pinyin of the first character in the third level of the dictionary tree, or if there is no first keyword node matching the SMS segmentation in the fourth level corresponding to the target pinyin node, it is determined that the target SMS is not a spam SMS.

[0078] Specifically, if after a complete traversal, it is found that there is no pinyin node that matches the pinyin of the first character in the SMS segmentation, then it can be directly determined that the target SMS is not a spam message; and if the target pinyin node is found in the third level, then the fourth level corresponding to the target pinyin node is traversed, but after traversing this level, no first keyword node matching the SMS segmentation is found. In this case, it can also be determined that the target SMS is not a spam message.

[0079] As an example, suppose you receive a text message "Come and get the free benefits", first perform word segmentation processing on it, and get the text message segmentations such as "free", "benefits", "come", and "get". For the segmentation word "free", the pinyin of its first character "免" is "mian", and the first letter of the first pinyin is "m". Then search for letter nodes in the second level of the dictionary tree, and find the letter node "m", which is the target letter node hit by the first letter. Then traverse the third level corresponding to the letter node "m", and find that there is a pinyin node "mian", that is, the target pinyin node that matches the pinyin of the first character "免". Then enter the fourth level corresponding to the pinyin node "mian", and find the first keyword node "免费", which matches the text message segmentation word "免费". Assuming that the keyword node "免费" carries a preset label indicating spam text messages, it can be determined that the text message containing "Come and get the free benefits" is a spam text message.

[0080] In the embodiments of the present application, by using the pinyin and initials of the first character in the SMS word segmentation and matching them with the hierarchical structure of the trie tree, the SMS content can be screened efficiently and accurately. In this way, the keywords that may have problems can be quickly located, the efficiency of spam SMS detection is improved, and time and computing resources are saved. Secondly, according to the preset tags, it is determined whether the SMS is a spam SMS, so that the judgment criteria are clear and operable, the possibility of misjudgment is reduced, and the detection accuracy is improved. Furthermore, when there is no matching pinyin node in the third level of the trie tree or no matching keyword node in the fourth level corresponding to the target pinyin node, it can be accurately determined that the SMS is not a spam SMS, ensuring the normal circulation of normal SMS, avoiding interference with the normal communication of users, and improving the user experience.

[0081] In the embodiments of the present application, a first keyword node is associated with a plurality of keyword combinations, and the plurality of keyword combinations are stored in the form of a subtree structure of the trie tree. Each level in the subtree structure includes at least one keyword; the first level of the subtree structure corresponds to the first keyword node in the trie tree.

[0082] In the embodiments of the present application, the method further includes: if the first keyword node matching the SMS word segmentation does not carry a preset tag, it is determined that the spam keyword rule is not matched; obtain the second keyword located in the second level of the subtree structure, and determine whether there is an SMS word segmentation in the word segmentation set that matches the second keyword; if there is an SMS word segmentation that matches the second keyword, detect whether the matching second keyword carries a preset tag; if the matching second keyword carries a preset tag, determine that the target SMS is a spam SMS; or, if there is no SMS word segmentation that matches the second keyword, obtain the next SMS word segmentation in the word segmentation set, and repeat the step of matching the next SMS word segmentation with the first keyword located in the first level of the subtree structure.

[0083] Specifically, when a first keyword node matching the SMS word segmentation is found in the trie tree (for example, the SMS word segmentation is "free", and the first keyword node "free" is found in the trie tree), and this node does not carry a preset tag (indicating that it cannot be directly determined that the SMS is a spam SMS based on this node alone), according to the definition of the subtree structure, find the subtree structure corresponding to this first keyword node.

[0084] Then, obtain all the keywords in the second level of the subtree structure, and these keywords are the second keywords. For example, if the second level of the subtree structure corresponding to the first keyword node "free" has the keyword "receive", "receive" is the second keyword.

[0085] Traverse each word segment in the word segment set, and compare it with the second keyword obtained one by one to check if there is an exactly identical word segment. For example, in the word segment set, there are word segments such as "receive" and "gift". Match these word segments with the keyword "receive" to see if there is a match. If a text message word segment that matches the second keyword is found in the word segment set (for example, if the second keyword is "receive" and the word segment set also includes the word segment "receive"), then check whether the matching second keyword carries a preset label in the trie tree. The preset label is used to identify whether the keyword is related to spam. If it carries a preset label (for example, the label value is a specific identifier representing spam), then proceed to the next step.

[0086] When it is detected that the matching second keyword carries a preset label, it indicates that the target text message contains a keyword combination related to spam (such as the combination of "free" and "receive", that is, the spam keyword rule). According to the rule, it can be determined that the target text message is spam.

[0087] If after traversing all the second keywords, no matching text message word segment is found in the word segment set, then take out the next text message word segment from the word segment set. Then, repeat the matching process in the above embodiment for this new text message word segment until it is determined whether the target text message is spam or all the word segments in the word segment set have been traversed.

[0088] Through such a step-by-step and repetitive matching process, based on rules such as the trie tree structure and the preset labels of keywords, the text message content is comprehensively checked to accurately determine whether the target text message belongs to spam.

[0089] This application first obtains the target text message to be processed and extracts its text message content, and then performs word segmentation on the content to obtain a set containing multiple text message word segments, preparing for subsequent judgment. On this basis, the first letter of the first word in the text message word segments is extracted in sequence and matched layer by layer with a pre-constructed trie tree with a specific hierarchical structure (the first layer is an empty node, the second layer is multiple letter nodes, the third layer covers the pinyin nodes associated with each letter node, and the fourth layer contains the first keyword nodes associated with each pinyin node and the pinyin corresponds to the first character pinyin of the keyword) to determine whether the target text message is spam. This method effectively avoids the drawback in the traditional multi-keyword multi-pattern matching scheme that the global retrieval efficiency and the matching efficiency of the same keyword with multiple rules linearly decline due to the increase in the number of keyword rules. Thanks to the unique structure of the trie tree, it can quickly locate the possible branches based on the first letter and then gradually deepen the matching, greatly reducing unnecessary full-scale retrieval operations, and finally significantly improving the global retrieval efficiency, making the screening of spam more efficient and accurate.

[0090] In an embodiment of the present application, the method further includes: if the matched second keyword does not carry a preset tag, detecting whether there is an associated third level in the subtree structure for the matched second keyword; if there is a third level associated with the matched second keyword in the subtree structure, extracting a third keyword from the third level associated with the matched second keyword; determining whether there is a text message token in the token set that matches the third keyword; if there is a text message token that matches the third keyword, detecting whether the matched third keyword carries a preset tag, and if the matched third keyword carries a preset tag, determining that the target text message is a spam text message.

[0091] Specifically, when a second keyword that matches the text message token is found in the previous step and the second keyword does not carry a preset tag, according to the definition and storage method of the subtree structure, check whether there is a next level, i.e., a third level, in the subtree structure where the second keyword is located. This can be determined through the node association relationship of the subtree structure. For example, check whether the second keyword node has a pointer or link pointing to a child node.

[0092] If it is detected that there is a third level, then traverse all the nodes of the third level pointed to by the second keyword node, and extract the keywords represented by these nodes. These are the third keywords. For example, the third level associated with the second keyword "receive" has the keyword "gift".

[0093] After obtaining the extracted third keywords, traverse the token set obtained by tokenizing the target text message again. Compare each token in the token set with these third keywords one by one to see if there is an exactly identical token. For example, the token set contains tokens such as "gift" and "activity". Match the third keyword "gift" with these tokens to see if an identical one can be found.

[0094] If a text message token that matches a certain third keyword is found in the token set (for example, the third keyword is "gift" and the token set also includes the text message token "gift"), then check whether the matched third keyword carries a preset tag in the trie tree. If the third keyword carries a preset tag (for example, the tag value is a specific identifier representing a spam text message), then the target text message can be determined to be a spam text message according to the rule.

[0095] When the matching second keyword does not carry a preset label in this application, it further checks whether there is a third level, which is the next level. If it exists, it continues to obtain the third keyword at the third level and matches it with the word segmentation set. If there is no next level, it is regarded as carrying a preset label. By checking the SMS content in this way, layer by layer and progressively, without missing any clues of spam keywords that may be hidden in deeper-level associations, it can analyze the SMS more comprehensively and meticulously, effectively avoiding the situation of missing some spam word combination rules due to being limited to shallow matching only, thereby improving the accuracy of spam SMS determination, enabling real spam SMS to be accurately identified as much as possible, and reducing misjudgment and missed judgment phenomena.

[0096] As an example, Figure 2 As shown, for the word segmentation "free", the pinyin of the first character "mian" is "mian", and the first letter "m" is extracted. In the trie tree, the letter node "m" is found at the second level, which is the target letter node. Enter the third level corresponding to the letter node "m", and the pinyin node "mian" is found, and it is determined as the target pinyin node. Enter the fourth level corresponding to the "mian" pinyin node, and the first keyword node "free" is found, which matches the SMS word segmentation "free".

[0097] Assume that the first keyword node "free" does not carry a preset label, and proceed to the next step. There are second keywords such as "receive" and "welfare" at the second level of the subtree structure corresponding to "free", and these keywords are obtained. It is found that there is an SMS word segmentation in the word segmentation set that matches the second keyword "receive".

[0098] If there is an SMS word segmentation that matches the second keyword, it checks whether the matching second keyword carries a preset label. Assume that the second keyword "receive" also does not carry a preset label, and continue to the next step.

[0099] Check the third level associated with "receive" in the subtree structure. Third keywords such as "coupon" are extracted from the third level associated with "receive". It is found that there is an SMS word segmentation in the word segmentation set that matches the third keyword "coupon". Assume that the third keyword "coupon" carries a preset label, then it can be determined that this SMS is a spam SMS.

[0100] Through this complete example, it demonstrates the whole process of spam SMS judgment for SMS according to the given trie tree structure and rules, from the first letter matching to the detection of multi-level keywords, and gradually determines whether the SMS is a spam SMS.

[0101] In this example, after sequentially matching and detecting all the word segments according to the rules, it is finally found that there is no matching situation that can determine the text message as a spam message (that is, no keyword carrying a preset tag is successfully matched at the corresponding level). Therefore, it can be determined that this text message "The bank card under your name has been deactivated" is not a spam message.

[0102] In the embodiment of the present application, first, the target text message to be processed is obtained and its text content is extracted. Then, the text content of the text message is segmented to obtain a set containing multiple text message segments. Next, the first letter of the first word in the text message segment is extracted, and based on this first letter, a hierarchical match is performed with a pre-constructed trie to determine whether the target text message is a spam message. This method avoids the situation where the global retrieval efficiency and the matching efficiency of the same keyword with multiple rules linearly decrease as the number of keyword rules increases in the traditional multi-keyword multi-pattern matching scheme. The structure of the trie makes the matching of keywords more efficient. By quickly locating the possible branches through the first letter and then gradually delving deeper into the match, unnecessary full-scale retrievals are reduced, thereby improving the global retrieval efficiency. For the multi-rule matching of the same keyword, more targeted searches can also be performed on specific paths of the trie, avoiding inefficient repeated searches, and thus reducing the latency of spam message identification, enabling timely triggering of alarms.

[0103] In the embodiment of the present application, as Figure 3 shown, the method for constructing the trie includes:

[0104] Step S201, splitting at least one spam keyword rule to obtain a corresponding keyword combination; determining the first spam keyword in the keyword combination, and constructing a subtree structure using the keyword combination and the corresponding first spam keyword.

[0105] In the embodiment of the present application, spam keyword rules are collected from various channels, and these channels may include but are not limited to: spam message databases, from which common spam message content patterns and keywords are extracted. The blacklist of spam message keywords released by network security agencies. Analyzing and summarizing the spam messages feedback by users.

[0106] Each of the collected spam keyword rules is split to obtain a corresponding keyword combination. For example: for the rule "Free to receive a high bonus", it is split into "Free", "Receive", "High", "Bonus". For "Low-interest loan with quick disbursement", it is split into "Low-interest", "Loan", "Quick", "Disbursement".

[0107] In each split keyword combination, determine the first junk keyword according to the following principle: Select words that appear frequently in spam messages, have strong representativeness, or can clearly reflect the characteristics of spam messages. For example, in "Get a high bonus for free", "free" may be a more representative first junk keyword because many spam messages often use "free" as a bait to attract users.

[0108] Create the root node: Use the determined first junk keyword as the root node of the subtree structure. For example, for the keyword combination "Get a high bonus for free", create a subtree with "free" as the root node.

[0109] Add child nodes (second level): Add the other keywords in the keyword combination under the root node to form the second level of the subtree. For "Get a high bonus for free", add "get", "high", and "bonus" as child nodes under the "free" root node.

[0110] Expand the subtree (if there are more related combinations): If there are other keyword combinations related to the first junk keyword, continue to add nodes to the subtree. For example, if there is also "Free welfare broadcast", then add "welfare", "broadcast", etc. as child nodes under the "free" root node.

[0111] Through the above steps, the processing of the junk keyword rules can be completed, and a subtree structure for subsequent applications such as spam message detection can be constructed. In actual applications, when receiving a text message, after segmenting the text message, it can be matched and judged according to the constructed subtree structure to determine whether the text message is a spam message. For example, if the segmented text message contains "free", it will enter the subtree with "free" as the root node for further keyword matching and detection.

[0112] Step S202, determine the pinyin and initial letter of the first character of the first junk keyword in the keyword combination.

[0113] In the embodiment of the present application, based on the previously determined keyword combination, select the first junk keyword from it. For example, for the keyword combination "Get a high bonus for free", the first junk keyword is "free". Pre-create a mapping table containing common Chinese characters and their corresponding pinyin, which can use an existing Chinese character pinyin library or be sorted out according to needs. Extract the first character from the first junk keyword. For "free", the first character is "mian". Look up the pinyin corresponding to the character "mian" in the pinyin mapping table to get "mian". Use the string processing functions or methods provided by the programming language to extract the initial letter from the obtained pinyin string "mian". For example, in many programming languages, the initial letter "m" can be obtained using index operations or specific string truncation functions.

[0114] Through the above steps, the pinyin and first letter of the first word of the first junk keyword in the keyword combination can be accurately determined, providing basic data for subsequent operations such as building a dictionary tree. For example, when building a dictionary tree, this information will be used to determine the location and hierarchical relationship of the nodes, so as to more efficiently detect and filter junk text messages.

[0115] Step S203, obtaining an initial tree structure, and creating an empty node in the initial tree structure.

[0116] In the embodiment of the present application, the existing tree structure data structure class or library in the programming language is used, for example, in Java, data structures such as TreeSet and TreeMap can be used, or a custom tree node class can be used to build a tree structure. If a custom tree structure is selected, a node class can be defined, including properties such as node value and child node list, as well as methods for operations such as adding child nodes and traversing.

[0117] Depending on the implementation method you choose, create an empty tree structure object as the initial tree structure. For example, if it is a custom tree structure, create a tree object with an empty root node. If it is a custom tree structure, use the node class to create an empty node object. This empty node will serve as the starting point of the entire tree structure, and subsequent nodes will be built based on this node.

[0118] Step S204: Create a letter node corresponding to the first letter of the first character of the first junk keyword with the empty node as the root node.

[0119] In the embodiment of the present application, the pinyin and the first letter of the first word of the first junk keyword are determined. For example, for the first junk keyword "免" (free), the pinyin of the first word "免" (free) is "mian", and the first letter is "m". Use the node class to create a node object representing the first letter, for example, NodeletterNode = newNode('m');, where 'm' is the first letter. Add the created letter node to the child nodes of the empty node (root node).

[0120] Step S205: Create a corresponding pinyin node at each letter node based on the pinyin of the first character in each first junk keyword.

[0121] In the embodiment of the present application, it starts from a tree structure that has been constructed with an empty node as the root node and a letter node corresponding to the first letter of the first character of the first junk keyword under the root node.

[0122] Use a tree traversal algorithm (such as breadth-first traversal or depth-first traversal) to traverse all letter nodes under the root node. For example, in breadth-first traversal, a queue can be used to assist the implementation. First, the root node is queued, and then the nodes in the queue are taken out one by one, and its unvisited child nodes are queued until the queue is empty.

[0123] For each traversed letter node, determine all first junk keywords associated with it. This can be obtained through a pre-established mapping relationship or related information recorded when constructing the letter node. For example, the first junk keywords that may be associated with the letter node "m" include "free", "profit-making", etc.

[0124] For each related first junk keyword, extract the pinyin of its first character. For example, for "free", extract the pinyin "mian" of the first character "免"; for "赚利", extract the pinyin "mou" of the first character "牟". Use the node class (similar to the previous creation of letter nodes) to create a node object representing the pinyin. For example, for "mian", create NodepinyinNode = newNode ("mian");. Add the created pinyin node to the child nodes of the currently traversed letter node.

[0125] Step S206, adding the subtree structure corresponding to the first junk keyword to the pinyin node corresponding to the initial tree to obtain a dictionary tree.

[0126] In the embodiment of the present application, child nodes are added in sequence according to the keyword combination related to the first junk keyword. For example, for "free high bonus", "receive", "high amount", "bonus" and so on are added as child nodes under the "free" root node to form the second level of the subtree. If there are other related combinations, such as "free welfare big giveaway", then "welfare", "big giveaway" and so on are added under the "free" root node to enrich the subtree structure.

[0127] Use a tree traversal algorithm (such as breadth-first traversal or depth-first traversal) to traverse the previously constructed initial tree and find the pinyin node corresponding to the first pinyin of the first junk keyword. For example, for the first junk keyword "免免", the pinyin of its first character "免" is "mian", and the pinyin node "mian" is found in the initial tree.

[0128] The constructed subtree structure corresponding to the first junk keyword is added as a whole to the found pinyin node as the subtree of the pinyin node. The specific implementation method depends on the implementation method of the tree structure. If it is a custom tree structure, it may be necessary to add the root node (the first junk keyword node) and its child nodes of the subtree through the addChild method of the pinyin node or similar operations.

[0129] Through the above steps, the subtree structures corresponding to each first junk keyword can be added under the pinyin nodes corresponding to the initial tree, thereby obtaining a complete trie tree. This trie tree can be used for subsequent applications such as spam message detection.

[0130] In the embodiment of the present application, after obtaining the trie tree, the following steps B1 - B2 are further included:

[0131] Step B1, calculate the proportion of keyword combinations under each letter node in the trie tree, and regard the letter nodes with a proportion higher than the preset threshold as the first letter nodes and the letter nodes with a proportion lower than the preset threshold as the second letter nodes.

[0132] Specifically, first, traverse each letter node in the trie tree. For each letter node, count the number of all keyword combinations included in its subordinates, which is recorded as the total number of keyword combinations of this letter node. Then, obtain the total number of all keyword combinations in the entire trie tree (which can be obtained by accumulating the number of keyword combinations under all letter nodes). Next, divide the total number of keyword combinations of each letter node by the total number of keyword combinations in the entire trie tree, and the result obtained is the proportion of keyword combinations under this letter node. For example, in the trie tree, there are letter nodes "A", "B", "C", etc. There are 10 keyword combinations under the letter node "A", and the total number of keyword combinations under all letter nodes in the entire trie tree is 100. Then the proportion of keyword combinations under the letter node "A" is 10÷100 = 0.1 (i.e., 10%). In this way, calculate the proportion of keyword combinations under each letter node in the trie tree in turn.

[0133] Pre - set a threshold (this threshold can be determined according to actual needs, past experience or multiple tests, such as set to 30%). After obtaining the proportion of keyword combinations under each letter node, classify the letter nodes with a proportion higher than this preset threshold as the first letter nodes, which means that the keyword combinations under these nodes are relatively dense; and determine the letter nodes with a proportion lower than the preset threshold as the second letter nodes, that is, the keyword combinations under these nodes are relatively few and sparse. For example, if the preset threshold is 30%, the proportion of the letter node "A" is 40%, and the proportion of the letter node "B" is 20%, then the letter node "A" will be classified as the first letter node, and the letter node "B" will be classified as the second letter node.

[0134] Step B2, re - determine the first junk keywords in the keyword combinations under the first letter nodes, so as to adjust the keyword combinations in the first letter nodes to the second letter nodes, and obtain an updated trie tree.

[0135] Specifically, for each of the first-letter nodes that have been divided, re-examine the keyword combinations subordinate to it. The process of re-determining the first junk keyword may comprehensively consider multiple factors. For example, the change in the occurrence frequency of the keyword combination in recent actual application scenarios (such as spam message judgment, etc.). If the occurrence frequency of the keyword combination corresponding to a previously determined first junk keyword has decreased significantly recently, it may be necessary to re-select a more representative and higher-occurrence-frequency keyword as the new first junk keyword; or adjust according to the degree of association between the keyword combination and other relevant business rules and data characteristics, etc.

[0136] For example, for the letter node "Z", the keyword combinations under it include "online" and "daily settlement", etc. Calculate the total number of its keyword combinations, and then divide it by the total number of keyword combinations in the entire trie to obtain the proportion of the keyword combinations under the letter node "Z". At the same time, calculate the proportions of the keyword combinations under other letter nodes such as the letter node "R". Assume that the preset threshold is a certain value (such as 30%), and it is found through calculation that the proportion of the keyword combinations under the letter node "Z" is much lower than that of the letter node "R".

[0137] Then, for the first-letter node "Z", originally the keyword combinations "online" and "daily settlement" corresponding to the first junk keyword "online" were under it. However, due to the low proportion of the keyword combinations under the "Z" node, the first junk keyword is re-determined, and "daily settlement" is re-determined as the first junk keyword, and the keyword combinations "online" and "daily settlement" originally under the "Z" node are adjusted to the second-letter node "R", thus obtaining an updated trie. Such an adjustment can optimize the structure of the trie and make it more in line with the requirements in actual applications. For example, in scenarios such as spam message judgment, it can perform keyword matching and judgment more efficiently and accurately.

[0138] After re-determining the first junk keywords under the first-letter nodes, based on these newly determined first junk keywords, migrate the entire corresponding keyword combinations from the first-letter nodes to the second-letter nodes. After the operation of adjusting some keyword combinations in the first-letter nodes to the second-letter nodes, the internal structure of the trie has changed, and the distribution of the keyword combinations under each letter node has been optimized and adjusted, forming a new, updated trie. This updated trie may have higher efficiency and better accuracy in subsequent related applications such as spam message matching because of its more reasonable structure. For example, when searching for keywords, it can locate relevant nodes faster and reduce unnecessary search paths, etc.

[0139] In the embodiment of the present application, after obtaining the trie tree, the method further includes: obtaining the newly added junk keyword rules corresponding to each letter node; splitting the junk keyword rules to obtain newly added keyword combinations; and updating the trie tree by using the newly added keyword combinations.

[0140] First, it is necessary to clarify the source channels of the newly added junk keyword rules, which may come from analyzing and summarizing newly emerged types of junk information, based on new junk information cases feedback by users, referring to newly monitored trends of bad information in the industry, etc. For example, by collecting the content of junk text messages that users have complained about recently, extracting the common descriptions as the newly added junk keyword rules; or obtaining the relevant rule content from the newly released notices on the characteristics of junk information by the network security supervision department. For each newly added junk keyword rule obtained, determine the corresponding letter node according to the first letter of the keyword in the rule.

[0141] For example, assume that new junk information cases are collected from user feedback, and it is found that frequently appearing expressions such as "live streaming e-commerce link" and "virtual currency investment" are used as the newly added junk keyword rules. For the rule "live streaming e-commerce link", its first letter is "Z", so it corresponds to the letter node "Z". After splitting it, the newly added keyword combinations such as "live streaming", "e-commerce", and "link" can be obtained. And the first letter of "virtual currency investment" is "X", corresponding to the letter node "X", and after splitting, newly added keyword combinations such as "virtual", "currency", and "investment" can be obtained.

[0142] Then, for the letter node "Z", there are already a certain number of keyword combinations under it. Now, integrate the newly split keyword combinations of "live streaming", "e-commerce", and "link" into it. According to the construction logic of the trie tree, create nodes for them at the corresponding levels, improve the corresponding subtree structure, and make it organically combined with the original structure. Similarly, for the letter node "X", also add the newly added keyword combinations such as "virtual", "currency", and "investment", and update the corresponding subtree structure. In this way, the trie tree is updated by using the newly added keyword combinations.

[0143] In this way, the updated trie tree can cover more newly emerged characteristics of junk information. In subsequent actual application scenarios such as junk text message judgment, keywords can be matched more comprehensively and accurately based on the updated structure, and then it is more efficient and accurate to identify whether a text message is a junk text message, better meeting the changing needs in actual applications.

[0144] There are various ways to split the newly added spam keyword rules. Commonly, such as using lexical analysis methods in natural language processing to split according to natural separation marks such as spaces and punctuation marks between words. Using the selected splitting method, perform splitting operations on each newly added spam keyword rule one by one, and convert them all into corresponding newly added keyword combination forms. For example, the newly added spam keyword rules "live streaming with goods, goods link" and "purchase virtual currency, click the entrance" are obtained. To split these rules, a lexical analysis method based on spaces and punctuation marks as separation marks in natural language processing is adopted. For "live streaming with goods, goods link", according to the comma and space, it is split into two newly added keyword combinations: "live streaming with goods" and "goods link". For "purchase virtual currency, click the entrance", it is also split into two newly added keyword combinations: "purchase virtual currency" and "click the entrance" based on punctuation and space.

[0145] For the newly added keyword combinations obtained by splitting, first determine the first spam keyword in each combination (the determination method can refer to the one mentioned above, such as based on factors like the strength of the negative attribute reflected by the keyword and the key semantic position). Then, for the trie tree that needs to be updated, check the corresponding node positions (locate the corresponding letter nodes, pinyin nodes, etc. according to the first letter and pinyin of the first spam keyword in the newly added keyword combination). If there is no suitable subtree structure at the corresponding position to accommodate the newly added keyword combination, construct a new subtree structure according to the conventional method of constructing a subtree structure (using the first spam keyword as the root node and constructing a hierarchical relationship according to the association between keywords, etc.); if there is already a subtree structure but it is not perfect, supplement and improve it, such as adding new hierarchical nodes or adjusting the association relationship between existing nodes, etc., to ensure that the newly added keyword combination can be integrated into the trie tree in the form of a reasonable subtree structure.

[0146] According to the first letter of the first spam keyword in the newly added keyword combination, find the corresponding letter node, and then find the corresponding pinyin node according to the pinyin of its first character (if there is such a hierarchical structure requirement), and then add the newly added keyword combination under the corresponding pinyin node (as the relevant content under this pinyin node, jointly constituting the information set under this node with other existing keyword combinations).

[0147] After adding new keyword combinations, you may need to make appropriate adjustments to the overall structure of the dictionary tree, such as checking whether some nodes have become too large due to the addition of new content (too many keyword combinations, etc.). If so, you can follow the previously mentioned strategies such as balancing node loads (such as calculating the proportion of keyword combinations under each node, making adjustments between nodes, etc.) to re-layout some nodes of the dictionary tree and their subordinate keyword combinations to ensure that the updated dictionary tree structure is still reasonable and efficient, and facilitate subsequent related application operations such as spam matching.

[0148] As an example, Figure 4 As shown, create an empty node as the starting point of the entire dictionary tree, that is, the root node. This root node is the basis of the entire structure, and all subsequent nodes will be built based on it. Next, process each keyword. Take the keyword "long-term" as an example, first extract its first letter. Through simple character operations, the first letter is "c". At this time, create a new node under the root node, which represents the letter "c". Then, get the pinyin "chang" of the first word "long" in "long-term". Under the created "c" node, create another new node, which is used to represent the pinyin "chang". Finally, add the keyword "long-term" to the "chang" pinyin node to complete the preliminary construction of "long-term" in the dictionary tree. Repeat the above steps for the keyword "malicious". First extract the first letter "e" and create a node representing "e" under the root node. Next, get the pinyin "e" of the first word "evil", and create a node corresponding to the pinyin "e" under the "e" node. Then, add "malicious" to this pinyin node.

[0149] For other keywords, they are processed in this way. Continuously create the first letter node under the root node, create the corresponding pinyin node under the first letter node, and then add the keyword to the corresponding pinyin node. As the keywords are processed one by one, the dictionary tree gradually becomes richer, from only the root node at the beginning to a complete structure with different levels and branches. When all keywords are processed, a complete dictionary tree is constructed.

[0150] In the embodiment of the present application, an empty node is first created as the root node in the initial tree structure, providing a clear starting point for constructing the trie tree subsequently, making the entire structure have good hierarchy and organization. Based on the empty node, multiple letter nodes are created, and the corresponding junk keywords are obtained for each letter node, so that the junk keywords can be quickly classified according to the first letter, improving the initial efficiency of retrieval. The pinyin of the first word in the junk keyword is determined and the corresponding pinyin node is created, further refining the classification method and making the keyword matching more accurate and efficient. The junk keyword is split to obtain a keyword combination and a subtree structure is constructed, which can better handle complex keyword relationships and improve the recognition ability of multi-word combination junk keywords. Finally, the subtree structure is added to the initial tree structure and the letter nodes are adjusted, and the obtained trie tree structure is more optimized and balanced, avoiding the problem of low retrieval efficiency caused by over-concentration of certain nodes, and generally improving the accuracy, efficiency and flexibility of keyword matching in spam detection.

[0151] In this embodiment, an apparatus for identifying spam is also provided. This apparatus is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, the term "module" may be a combination of software and / or hardware that can achieve a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0152] This embodiment provides an apparatus for identifying spam, as Figure 5 shown, including:

[0153] An acquisition module 501, configured to acquire a target short message to be processed and extract the short message text content of the target short message;

[0154] A processing module 502, configured to perform word segmentation processing on the short message text content to obtain a word segmentation set, where the word segmentation set includes multiple short message word segments;

[0155] An analysis module 503, configured to sequentially extract the first letters of the first words in the short message word segments, and perform layer-by-layer matching based on the first letters and a pre-constructed trie tree to determine whether the target short message is spam; wherein, the first layer of the pre-constructed trie tree is an empty node, the second layer of the trie tree includes multiple letter nodes, the third layer of the trie tree includes at least one pinyin node associated with each letter node, and the fourth layer of the trie tree includes at least one first keyword node associated with each pinyin node, and the pinyin represented by the pinyin node is the pinyin of the first character of the keyword represented by the first keyword node.

[0156] In the embodiment of the present application, the analysis module 503 includes:

[0157] The first extraction sub-module is used to utilize the pinyin of the first character in the SMS word segmentation and extract the initial letter in the pinyin of the first character;

[0158] The first matching sub-module is used to match the initial letter with the letter nodes in the second level of the trie tree to obtain the target letter node that the initial letter hits;

[0159] The second matching sub-module is used to traverse the third level of the trie tree to determine whether there is a target pinyin node that matches the pinyin of the first character;

[0160] The first judgment sub-module is used to, if there is a target pinyin node, traverse the fourth level corresponding to the target pinyin node to determine whether there is a first keyword node that matches the SMS word segmentation;

[0161] The second judgment sub-module is used to, if there is a first keyword node that matches the SMS word segmentation and the first keyword node matched by the SMS word segmentation carries a preset label, determine that the target SMS is a spam SMS.

[0162] The second judgment sub-module is further used to, if there is no pinyin node in the third level of the trie tree that matches the pinyin of the first character, or if there is no first keyword node in the fourth level corresponding to the target pinyin node that matches the SMS word segmentation, determine that the target SMS is not a spam SMS.

[0163] In the embodiment of the present application, a first keyword node is associated with multiple keyword combinations, and the multiple keyword combinations are stored in the form of a subtree structure of the trie tree. Each level of the subtree structure includes at least one keyword; the first level of the subtree structure corresponds to the first keyword node in the trie tree;

[0164] The device further includes: a second matching sub-module, which is used to, if the first keyword node that matches the SMS word segmentation does not carry a preset label, obtain the second keyword located in the second level of the subtree structure; determine whether there is an SMS word segmentation in the segmentation set that matches the second keyword; if there is an SMS word segmentation that matches the second keyword, detect whether the matched second keyword carries a preset label; if the matched second keyword carries a preset label, determine that the target SMS is a spam SMS; or, if there is no SMS word segmentation that matches the second keyword, obtain the next SMS word segmentation of the SMS word segmentation from the segmentation set, and repeat the step of matching the next SMS word segmentation with the first keyword located in the first level of the subtree structure.

[0165] In an embodiment of the present application, the apparatus further includes: a third matching sub-module, configured to detect whether there is an associated third level in the sub-tree structure for a second keyword that matches if the second keyword that matches does not carry a preset tag; if there is a third level associated with the second keyword that matches in the sub-tree structure, extract a third keyword from the third level associated with the second keyword that matches; determine whether there is a text segment of a short message in the word segmentation set that matches the third keyword; if there is a text segment of a short message that matches the third keyword, detect whether the third keyword that matches carries a preset tag, and if the third keyword that matches carries a preset tag, determine that the target short message is a spam short message.

[0166] In an embodiment of the present application, the apparatus further includes: a construction module, configured to split at least one spam keyword rule to obtain a corresponding keyword combination; determine a first spam keyword in the keyword combination, and construct a sub-tree structure by using the keyword combination and the corresponding first spam keyword; determine the pinyin and the first letter of the first character of the first spam keyword in the keyword combination; obtain an initial tree structure, and create an empty node in the initial tree structure; create a letter node corresponding to the first letter of the first character of the first spam keyword with the empty node as the root node; create corresponding pinyin nodes in each letter node based on the pinyin of the first character in each first spam keyword; add the sub-tree structure corresponding to the first spam keyword to the pinyin node corresponding to the initial tree to obtain a dictionary tree.

[0167] In an embodiment of the present application, the apparatus further includes: an adjustment sub-module, configured to calculate the proportion of keyword combinations under each letter node in the dictionary tree, and use the letter node with a proportion higher than a preset threshold as the first letter node and the letter node with a proportion lower than the preset threshold as the second letter node; re-determine the first spam keyword in the keyword combination under the first letter node, so as to adjust the keyword combination in the first letter node to the second letter node to obtain an updated dictionary tree.

[0168] This application first obtains the target short message to be processed and extracts its short message text content. Subsequently, word segmentation is performed on the content to obtain a set containing multiple short message word segments, preparing for subsequent judgment. On this basis, the first letters of the first words in the short message word segments are sequentially extracted and matched layer by layer with a pre-constructed trie tree having a specific hierarchical structure (the first layer is an empty node, the second layer is multiple letter nodes, the third layer covers the pinyin nodes associated with each letter node, and the fourth layer contains the first keyword nodes associated with each pinyin node and the pinyin corresponds to the first character pinyin of the keyword) to determine whether the target short message is a spam message. This method effectively avoids the disadvantages in the traditional multi-keyword and multi-pattern matching scheme that the global retrieval efficiency and the matching efficiency of the same keyword with multiple rules linearly decrease due to the increase in the number of keyword rules. Thanks to the unique structure of the trie tree, it is possible to quickly locate the possible branches by the first letter and then gradually deepen the matching, greatly reducing unnecessary full-scale retrieval operations, and finally significantly improving the global retrieval efficiency, making the identification of spam messages more efficient and accurate.

[0169] Please refer to Figure 6 , Figure 6 FIG. is a schematic structural diagram of an electronic device provided by an optional embodiment of the present invention, as Figure 6 shown. The electronic device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common main board or installed in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system).

[0170] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.

[0171] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.

[0172] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device presented by a kind of landing page of a small program, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided relative to the processor 10, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0173] The memory 20 may include a volatile memory, for example, a random access memory; the memory may also include a non-volatile memory, for example, a flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memories.

[0174] The electronic device further includes a communication interface 30 for the electronic device to communicate with other devices or a communication network.

[0175] The embodiments of the present invention further provide a computer-readable storage medium. The methods according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented by downloading through a network the original computer code stored in a remote storage medium or a non-transitory machine-readable storage medium and to be stored in a local storage medium, so that the methods described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium may be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium may also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.

[0176] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for identifying spam text messages, characterized in that: The method comprises: Obtaining a target SMS to be processed, and extracting SMS text content of the target SMS; Performing word segmentation processing on the text content of the SMS to obtain a word segmentation set, wherein the word segmentation set includes a plurality of SMS word segments; The first letter of the first word in the SMS word segmentation is extracted in sequence, and based on the first letter, a layer-by-layer match is performed with a pre-constructed dictionary tree to determine whether the target SMS is a spam SMS; wherein the first level of the pre-constructed dictionary tree is an empty node, the second level of the dictionary tree includes multiple letter nodes, the third level of the dictionary tree includes at least one pinyin node associated with each letter node, and the fourth level of the dictionary tree includes at least one first keyword node associated with each pinyin node, and the pinyin represented by the pinyin node is the pinyin of the first character of the keyword represented by the first keyword node.

2. The method according to claim 1, characterized in that The step of matching the first letter with a pre-built dictionary tree layer by layer to determine whether the target SMS is a spam SMS includes: Utilize the pinyin of the first word in the short message segmentation, and extract the first letter in the pinyin of the first word; Matching the first letter with the letter nodes in the second level of the dictionary tree to obtain the target letter node hit by the first letter; Traversing the third level of the dictionary tree to determine whether there is a target pinyin node that matches the pinyin of the first character; If the target pinyin node exists, traverse the fourth level corresponding to the target pinyin node to determine whether there is a first keyword node that matches the short message segmentation; If there is a first keyword node that matches the SMS segmentation, and the first keyword node that matches the SMS segmentation carries a preset tag, then the target SMS is determined to be a spam SMS; the preset tag is used to mark the spam keyword rule; If there is no pinyin node matching the pinyin of the first character in the third level of the dictionary tree, or if there is no first keyword node matching the SMS word segmentation in the fourth level corresponding to the target pinyin node, it is determined that the target SMS is not a spam SMS.

3. The method according to claim 2, characterized in that The first keyword node is associated with a plurality of keyword combinations, the plurality of keyword combinations are stored in a subtree structure of the dictionary tree, and each level of the subtree structure includes at least one keyword; The first level of the subtree structure corresponds to the first keyword node in the dictionary tree; The method further comprises: If the first keyword node matching the short message segmentation does not carry a preset tag, obtaining a second keyword located at the second level of the subtree structure; Determining whether there is a text message segmentation matching the second keyword in the segmentation set; If there is a text message segmentation matching the second keyword, detecting whether the matching second keyword carries a preset tag; If the matched second keyword carries a preset tag, determining that the target SMS is a spam SMS; Or, if there is no SMS segmentation matching the second keyword, the next SMS segmentation of the SMS segmentation is obtained from the segmentation set, and the step of matching the next SMS segmentation with the first keyword located at the first level of the subtree structure is repeated.

4. The method according to claim 3, characterized in that The method further comprises: If the matched second keyword does not carry the preset tag, detecting whether the matched second keyword has an associated third level in the subtree structure; If the third level associated with the matched second keyword exists in the subtree structure, extracting the third keyword from the third level associated with the matched second keyword; Determine whether there is a short message segmentation in the segmentation set that matches the third keyword; If there is a text message segmentation that matches the third keyword, it is detected whether the matched third keyword carries a preset tag. If the matched third keyword carries the preset tag, it is determined that the target text message is a spam text message.

5. The method according to claim 3, characterized in that: The method for constructing the dictionary tree includes: Splitting the at least one junk keyword rule to obtain a corresponding keyword combination; determining a first junk keyword in the keyword combination, and constructing a subtree structure using the keyword combination and the corresponding first junk keyword; Determine the pinyin and first letter of the first word of the first junk keyword in the keyword combination; Obtain an initial tree structure, and create an empty node in the initial tree structure; Create a letter node corresponding to the first letter of the first character of the first junk keyword with the empty node as the root node; Create a corresponding pinyin node under each letter node based on the pinyin of the first character in each first junk keyword; The subtree structure corresponding to the first junk keyword is added to the pinyin node corresponding to the initial tree structure to obtain the dictionary tree.

6. The method according to claim 5, characterized in that After obtaining the dictionary tree, it also includes: Calculate the proportion of keyword combinations under each letter node in the dictionary tree, and use the letter nodes with a proportion higher than a preset threshold as the first letter nodes and the letter nodes with a proportion lower than the preset threshold as the second letter nodes; The first junk keyword in the keyword combination under the first letter node is re-determined to adjust the keyword combination in the first letter node to the second letter node to obtain an updated dictionary tree.

7. The method according to claim 5, characterized in that After obtaining the dictionary tree, the method further includes: Get the newly added junk keyword rules corresponding to each letter node; Splitting the junk keyword rules to obtain new keyword combinations; The dictionary tree is updated using the newly added keyword combination.

8. A junk text message identification device, characterized in that: The device comprises: An acquisition module, used to acquire a target SMS to be processed and extract the SMS text content of the target SMS; A processing module, configured to perform word segmentation processing on the text content of the SMS to obtain a word segmentation set, wherein the word segmentation set includes a plurality of SMS word segments; An analysis module is used to extract the first letter of the first word in the SMS word segmentation in sequence, and perform layer-by-layer matching based on the first letter with a pre-constructed dictionary tree to determine whether the target SMS is a spam SMS; wherein the first level of the pre-constructed dictionary tree is an empty node, the second level of the dictionary tree includes multiple letter nodes, the third level of the dictionary tree includes at least one pinyin node associated with each letter node, and the fourth level of the dictionary tree includes at least one first keyword node associated with each pinyin node, and the pinyin represented by the pinyin node is the pinyin of the first character of the keyword represented by the first keyword node.

9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.