Malicious mail detection method and system based on semantic rule mining

Through the method based on semantic rule mining, a wide coverage of malicious email screening rules is generated and detected in combination with the graph diffusion algorithm. The problems of low efficiency and limited coverage in the existing technology are solved, and efficient and accurate malicious email detection is achieved.

CN120200814APending Publication Date: 2025-06-24GUANGDONG COREMAIL COMPUTER TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510372994.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing malicious email detection scheme relies on manual setting of screening rules, resulting in limited coverage and low detection efficiency, making it easy to miss malicious emails.

Method used

Using a method based on semantic rule mining, semantic labels are extracted and rule mining is performed to generate malicious email filtering rules with wide coverage and rich semantic granularity, and malicious email detection is performed in combination with the graph diffusion algorithm.

Benefits of technology

It realizes automatic and accurate detection of malicious emails in massive emails, reduces the rate of missed errors, and improves detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200814A_ABST
    Figure CN120200814A_ABST
Patent Text Reader

Abstract

The invention discloses a malicious mail detection method and system based on semantic rule mining, and the method comprises the steps: carrying out the rule mining of a plurality of semantic tags extracted from an existing malicious mail set based on a mail detection precision requirement, and generating a malicious mail screening rule; performing text matching on an existing malicious mail set and to-be-detected mails, and screening out a first threshold suspicious mail from the to-be-detected mails; and according to the malicious mail screening rule, calculating the label similarity between each suspicious mail and the existing malicious mail set through a graph diffusion algorithm, and screening out a plurality of malicious mails from the suspicious mails according to the label similarity. The malicious mail screening rule with wider coverage and richer semantic granularity is mined, malicious mails can be automatically and accurately detected in massive mails, and the false and missing judgment rate of malicious mail detection is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of email security detection, and specifically relates to a malicious email detection method and system based on semantic rule mining. Background Art

[0002] With the rapid development of the Internet and the wide application of email, malicious emails have become one of the most serious current network security threats. Malicious emails can bypass existing defense mechanisms to attack users, damage computers or steal data, causing damage to users' personal information and property security. Because the number of malicious emails exposed to the security model is scarce, it is difficult to use machine learning models to capture them in a timely manner.

[0003] The commonly used malicious email detection solution in the industry is a combination of an email classification model and rules, that is, through channels such as random inspection and user reporting, the email is reviewed to determine whether it is a malicious email. After confirming that it is a malicious email, the email administrator summarizes and combines the keywords, senders, header features, etc. of the malicious email through industry knowledge to obtain the corresponding machine learning model features, and then based on the obtained model features as screening rules, manually initiate a query to count the malicious emails. However, the screening rules set manually in the above solution are too precise, resulting in some malicious emails being easily missed; the second is that manually combining the rules leads to a huge search space, and a large amount of time and effort are required for verification and debugging to determine whether the combined screening rules have a wide coverage and can accurately detect malicious emails. Summary of the Invention

[0004] This application proposes a malicious email detection method and system based on semantic rule mining, which can mine malicious email screening rules with a wider coverage and richer semantic granularity, and can automatically and accurately detect malicious emails in a large number of emails, reducing the false omission rate of malicious email detection.

[0005] The first aspect of this application provides a malicious email detection method based on semantic rule mining, and the method includes:

[0006] Based on the email detection accuracy requirement, perform rule mining on several semantic tags extracted from the existing malicious email set to generate malicious email screening rules;

[0007] Perform text matching on the existing malicious email set and the emails to be detected, and screen out the first threshold number of suspicious emails from the emails to be detected;

[0008] According to the malicious email screening rules, calculate the label similarity between each suspicious email and the existing malicious email set through the graph diffusion algorithm, and screen out several malicious emails from the suspicious emails according to the label similarity.

[0009] The above solution first extracts semantic tags that may contain risk information from multiple existing malicious emails, and then combines these semantic tags multiple times through rule mining to generate malicious email screening rules with a wide coverage and strong applicability, which can accurately detect malicious emails in a large number of emails. Then, the existing malicious email set is textually matched with the emails to be detected, and the first threshold number of suspicious emails with a high text content match is selected as the data basis for subsequent email detection. Finally, according to the malicious email screening rules, the tag similarity between the suspicious emails and the existing malicious email set is calculated, and the suspicious emails with a higher tag similarity are selected as malicious emails to timely protect the user information security.

[0010] In a possible implementation method of the first aspect, based on the email detection accuracy requirement, rule mining is performed on several semantic tags extracted from the existing malicious email set to generate malicious email screening rules, specifically:

[0011] The text features of the existing malicious email set are respectively extracted and clustered through a preset first semantic model and a preset second semantic model, and several semantic tags and text clustering IDs are output;

[0012] When the email detection accuracy requirement is higher than the set average value, the semantic tags and the text clustering IDs are combined and dynamically screened through a rule tree visualization algorithm to generate malicious email screening rules;

[0013] When the email detection accuracy requirement is not higher than the set average value, the semantic tags and the text clustering IDs are combined through a rule tree integration algorithm to generate malicious email screening rules.

[0014] The above solution extracts multiple semantic tags respectively and clusters the email text to obtain corresponding clustering IDs, providing the function of semantic clustering for subsequent rule mining. If the detection accuracy requirement is high, the rule tree visualization algorithm is selected to dynamically screen the semantic tags through the provided keywords; if the detection accuracy requirement is not high, the rule tree integration algorithm is selected to directly generate malicious email screening rules, improving the efficiency of email detection.

[0015] In a possible implementation method of the first aspect, the semantic tags and the text clustering IDs are combined and dynamically screened through a rule tree visualization algorithm to generate malicious email screening rules, specifically:

[0016] The text features of the existing malicious email set are processed through a preset third semantic model to generate several dynamic semantic tags;

[0017] Using the preset decision tree model with the semantic tags as tree nodes, and combining the tree nodes based on the text clustering ID to generate a number of decision trees; wherein, the maximum depth of the decision tree is adjusted according to the mail detection accuracy requirement;

[0018] Calculating the proportion of malicious email features of each node in the decision tree according to the preset iterative tags and generating a corresponding co-occurrence graph of words;

[0019] Mining and combining the nodes according to the co-occurrence graph of words and the preset malicious email keyword rules to generate the malicious email screening rules.

[0020] The above solution uses a third semantic model to generate dynamic semantic tags, which can quickly add / delete new semantic and intent tags according to user needs to obtain more comprehensive tag data; then using the semantic tags as tree nodes, generating multiple decision trees according to the given maximum tree depth, where the maximum tree depth is determined by the mail detection accuracy requirement, which can further control the complexity of the algorithm and improve the calculation efficiency; then showing the proportion of malicious email features of each tree node according to the generated decision tree and performing co-occurrence analysis of words, calculating the frequency of occurrence of pairwise semantic tags in each email text, and combining the malicious email keyword rules for rule mining to generate malicious email screening rules with a wide coverage and strong applicability that can accurately detect malicious emails in a large number of emails.

[0021] In a possible implementation method of the first aspect, calculating the proportion of malicious email features of each node in the decision tree according to the preset iterative tags and generating a corresponding co-occurrence graph of words, specifically:

[0022] Calculating the proportion of malicious email features of each node in the decision tree according to the preset iterative tags, and selecting the nodes with the proportion of malicious email features exceeding the second threshold as gray nodes;

[0023] Dividing the existing malicious email set according to the gray nodes to obtain a number of email subsets;

[0024] Performing co-occurrence analysis of words on the gray nodes in each email subset to generate a co-occurrence graph of words corresponding to each email subset.

[0025] In a possible implementation method of the first aspect, combining the semantic tags and the text clustering ID through a rule tree integration algorithm to generate malicious email screening rules, specifically:

[0026] Using the preset decision tree model with the semantic tags as tree nodes, and combining the tree nodes based on the text clustering ID to generate a number of decision trees, and obtaining a number of tag combinations through the decision trees; wherein, the maximum depth of the decision tree is adjusted according to the mail detection accuracy requirement;

[0027] Record the occurrence times of each of the said tag combinations, and select the said malicious email screening rules that meet the preset malicious email feature conditions and whose occurrence times exceed the first threshold.

[0028] In a possible implementation method of the first aspect, perform text matching between the existing malicious email set and the emails to be detected, and screen out the first threshold number of suspicious emails from the said emails to be detected, specifically:

[0029] Decompose the existing malicious email set into several email texts, and then perform sparse matching between the said email texts and the emails to be detected within the third threshold number of days, and calculate the similarity scores between each of the said emails to be detected and the said email texts;

[0030] Convert the text content of the said emails to be detected whose similarity scores exceed the fourth threshold into dense text vectors, and perform vector calculations on the said dense text vectors and the said email texts, and find the first threshold number of suspicious emails from the said emails to be detected according to the vector calculation results.

[0031] The above solution performs text matching based on the email texts of the decomposed malicious email set. First, the similarity scores between texts are initially calculated through sparse matching, such as calculating word weights, and then vector calculations are performed through dense matching to find multiple similar emails that are close to each other, providing data support for subsequent malicious email detection.

[0032] In a possible implementation method of the first aspect, calculate the label similarity between each of the said suspicious emails and the existing malicious email set through the graph diffusion algorithm, and screen out several malicious emails from the said suspicious emails according to the said label similarity, specifically:

[0033] According to the text content of each of the said suspicious emails, determine the connection relationship between each semantic label through the malicious email screening rules;

[0034] Calculate the occurrence times of the said semantic labels in each of the said suspicious emails within the third threshold number of days, so as to determine the relationship weight of each of the said connection relationships;

[0035] Generate a corresponding bidirectional graph through the said connection relationship and the said relationship weight, and obtain a corresponding adjacency matrix through the said bidirectional graph;

[0036] Iterate the relationship weights in the adjacency matrix based on the preset label propagation coefficient. When the iteration end condition is met, output the said suspicious emails determined to be malicious emails; wherein, the said label propagation coefficient includes the maximum number of iteration rounds and the relationship weight threshold.

[0037] The above solution first determines the connection relationship between each semantic tag according to whether these semantic tags are included in each suspicious email, and then obtains the corresponding relationship weight based on the frequency of occurrence of this connection relationship within the last third threshold days, so as to construct an adjacency matrix for graph diffusion. Iterate the adjacency matrix through the graph diffusion algorithm and judge whether the relationship weight is greater than the set threshold during each iteration, and output the detected malicious emails after the iteration ends.

[0038] In a possible implementation method of the first aspect, the malicious email screening rule is specifically:

[0039] Mine rules from several screened malicious emails, and update the malicious email screening rule according to the mining results.

[0040] The above solution will conduct rule mining on each newly detected malicious email, select new semantic tags and tag combinations that may contain risks, so as to optimize the malicious email screening rule, expand the coverage of the screening rule, and reduce the false omission rate of malicious email detection.

[0041] The second aspect of this application provides a malicious email detection system based on semantic rule mining. The system includes: a screening rule generation module, a suspicious email screening module, and a malicious email detection module;

[0042] Among them, the screening rule generation module is used to conduct rule mining on several semantic tags extracted from the existing malicious email set based on the email detection accuracy requirement, and generate a malicious email screening rule;

[0043] The suspicious email screening module is used to perform text matching between the existing malicious email set and the email to be detected, and screen out the first threshold number of suspicious emails from the emails to be detected;

[0044] The malicious email detection module is used to calculate the label similarity between each suspicious email and the existing malicious email set through the graph diffusion algorithm according to the malicious email screening rule, and screen out several malicious emails from the suspicious emails according to the label similarity.

[0045] The third aspect of this application provides a terminal device. The device includes: a terminal device, including a processor and a memory. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of a malicious email detection method based on semantic rule mining according to any one of the embodiments of this application. Description of the Drawings

[0046] To more clearly illustrate the technical solutions of the present application, the accompanying drawings required for implementation will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0047] Figure 1 It is a schematic flowchart of a specific process of a malicious email detection method based on semantic rule mining provided by an embodiment of the present application;

[0048] Figure 2 It is a specific structural diagram of a malicious email detection system based on semantic rule mining provided by an embodiment of the present application;

[0049] Figure 3 An embodiment of the present application provides a structural diagram of a terminal device. Specific implementation manners

[0050] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0051] It should be understood that the step numbers used in the text are only for convenient description and are not intended to limit the execution order of the steps.

[0052] The first embodiment

[0053] The means of malicious attackers are endless. It is difficult to capture malicious emails in a timely manner only relying on existing defense mechanisms. Moreover, the existing malicious email screening rules are composed of keywords, senders, header features, and combinations of multiple features, with a low coverage rate and a narrow application range. This is because the rules set artificially are too precise and easy to be circumvented. However, setting the screening rules too coarsely will also cause misjudgments. Therefore, how to set a malicious email screening rule with a wide coverage rate and capable of accurately capturing the characteristics of malicious emails and perform accurate malicious email detection is the main research direction of the present application.

[0054] As Figure 1 shown, Figure 1 This is a schematic flowchart of a specific process of a malicious email detection method based on semantic rule mining provided by an embodiment of the present application. The malicious email detection method based on semantic rule mining in this embodiment includes steps S1 to S3, which are described in detail as follows:

[0055] Step S1: Based on the requirements for email detection accuracy, perform rule mining on several semantic tags extracted from the existing malicious email set to generate malicious email screening rules.

[0056] In the embodiment of the present application, first extract multiple semantic tags from the provided existing malicious email set. First, use the trained first semantic model to extract text features from the existing malicious email set and output several semantic tags; alternatively, use the trained second semantic model to perform text clustering on the existing malicious email set and output several text clustering IDs.

[0057] Specifically, the lightweight first semantic model can be used to analyze the email text online, classify it based on the text in the existing malicious email set, and output multiple semantic tags related to the risk of malicious emails. The first semantic model can meet the high-throughput performance requirements in an environment with limited computing resources, but it is difficult to consider the relationship between semantics and the word order of the text during the model analysis process, resulting in limited extraction effects of semantic tags. Using the second semantic model can perform deep clustering on the text of the existing malicious email set and output multiple text clustering IDs for subsequent screening rule mining, which can improve the coverage of malicious email screening rules. However, the disadvantage is that the data annotation cost of the second semantic model is relatively high.

[0058] Furthermore, the second semantic model can also perform offline analysis and online dense matching on the semantic tags, and distill the output result of the third semantic model into the first semantic model to improve the output effect of the first semantic model.

[0059] Optionally, the semantic tags can be table-like features as the input of the second semantic model; the first semantic model is based on the FastText model, and the FastText model can perform efficient text classification and word embedding, mainly used for processing large-scale text data, and has the characteristics of fast training and classification. The second semantic model is trained based on the Transformers model, can continuously improve the output effect of the first semantic model, and output text clustering IDs. More reliable screening rules can be extracted through the text clustering IDs.

[0060] Then, select the mining method of the malicious email screening rules according to the set requirements for email detection accuracy. If the requirement for detection accuracy is relatively high, use the rule tree visualization algorithm to combine and dynamically screen the semantic tags and the text clustering IDs to generate malicious email screening rules. Otherwise, use the rule tree integration algorithm to directly generate the corresponding malicious email screening rules through the semantic tags and the text clustering IDs.

[0061] The rule tree integration algorithm is end-to-end rule mining, which mainly combines the semantic tags output by the first semantic model and the text clustering IDs output by the second semantic model to generate malicious email screening rules that meet the malicious email feature conditions. Specifically: First, input the preset decision tree parameters, including the number of decision trees, the depth of each decision tree, the frequency threshold, and the accuracy threshold, etc. And select semantic features and non-semantic features from the provided malicious email feature templates, and input the selected features, decision tree parameters, and iteration labels into the trained decision tree model for processing to generate multiple decision trees. Among them, the maximum depth of the decision tree can be set according to the email detection accuracy requirement; if the maximum depth of the decision tree is larger, it means that the complexity of rule mining is greater.

[0062] Then, based on the text clustering ID, with the semantic tag as the tree node and the iteration label as the true value of the tree node, select the path with a predicted result of true through label combination and true value determination. The path is the combination of multiple semantic tags based on the text clustering ID. At the same time, through the input frequency threshold, the paths that appear more than the frequency threshold in all decision trees are used as high-frequency paths, and the high-frequency paths with a predicted result of true are screened and some unreasonable paths are deleted to find the high-frequency paths with a predicted result of true that meet the accuracy threshold as malicious email screening rules.

[0063] Exemplarily, for the obtained high-frequency path with a predicted result of true, if there are two tree nodes 'number of recipients > 3' and 'number of recipients < 2' among them, it means that there is a contradiction in this high-frequency path and it is unrealistic, and this high-frequency path should be deleted.

[0064] The rule tree visualization algorithm is semi-automatic rule mining, which mainly obtains malicious email screening rules based on keywords + semantic tags + other features through word co-occurrence analysis with the provided keywords and semantic tags, etc. Specifically: The data input into the trained decision tree model is basically the same, but the text features of the existing malicious email set are also processed through a preset third semantic model, and dynamic semantic tags are scattered and output quickly based on the existing semantic tags according to actual needs, and then the mined dynamic semantic tags are distilled into the first semantic model through the second semantic model to improve the comprehensiveness of the semantic tags.

[0065] Among them, the third semantic model is an offline generation model that can quickly add / delete new semantics and intent tags according to user needs and output some dynamic semantic tags that pose a risk of malicious emails. The dynamic semantic tags are new semantic tags with summarization ability, which can assist the second semantic model and the first semantic model to improve the data processing effect. However, the drawback is that the throughput of the third semantic model is relatively low and it requires a large amount of computing resources. Therefore, the third semantic model is only adopted when the email detection accuracy requirement is high and the attack methods of malicious emails iterate relatively fast, so as to improve the interception success rate of malicious emails.

[0066] Input the above data into the decision tree model, with semantic tags as tree nodes, and generate a decision tree with the same structure as that in the rule tree integration algorithm. For each tree node, display the proportion of malicious email features of each tree node in the form of a pie chart, and mark the node field features and the corresponding relationship weight thresholds. The proportion of malicious email features is the probability that the semantic tag corresponding to this tree node will cause the email to be a malicious email.

[0067] Then, divide the finite malicious email set into several email subsets, select the semantic tags with relatively high proportions of malicious email features, conduct co-occurrence analysis on these semantic tags based on the email subsets, and generate corresponding co-occurrence graphs. Mine and combine the tree nodes based on the co-occurrence graphs according to the provided malicious email keyword rules to generate the malicious email screening rules.

[0068] Furthermore, label propagation can be performed through other provided semantic features to generate more semantic tags to be detected as tree nodes, so as to expand the sample range.

[0069] Optionally, the decision tree model is constructed based on the GBDT model. The full name of GBDT is Gradient Boost Decision Tree, which is widely used in the risk control industry. In the case of small data volume and tabular data, it is widely considered to have stronger capabilities than deep models and has strong interpretability, making it very suitable for rule mining. At the same time, all data points will be very intuitively under a certain feature threshold segmentation (the intermediate result of GBDT). By controlling parameters such as the maximum tree depth, the complexity of the model can be constrained, thus making rule mining visualization and further co-occurrence analysis possible. This ensures the interactivity and interpretability of semantic rule mining. In addition, different from deep models, the decision tree model can show better model effects in the case of low data volume.

[0070] Step S2: Perform text matching on the existing malicious email set and the emails to be detected, and screen out the first threshold number of suspicious emails from the emails to be detected.

[0071] In the embodiment of the present application, first split the existing malicious email set into multiple email texts, and then calculate the similarity between the email to be detected and the existing malicious email set through sparse matching and dense matching in sequence to complete the preliminary screening of malicious emails.

[0072] First, perform sparse matching on the email text and the emails to be detected in the past n days. Using the email text as the query text, calculate the word weights between the email text and the emails to be detected, and then calculate the similarity score between the email text and the emails to be detected in combination with the inverse document frequency. Among them, the inverse document frequency is a measure used to evaluate the general importance of a word in a document set. Its basic idea is: if a word appears less frequently in the document set, then its importance in a specific document is higher; conversely, if a word appears more frequently in the document set, then its importance in a specific document is lower.

[0073] Then perform dense matching on the emails to be detected with relatively high similarity scores and the email text. First, convert the email text into a dense text vector, and perform vector calculation on the dense text vector and the email text. According to the vector calculation result, find the first threshold number of suspicious emails from the emails to be detected. These suspicious emails will be used as the data basis for the investigation of malicious emails in the next stage. Among them, the vector calculation can be maximum inner product calculation or minimum cosine distance calculation.

[0074] Step S3, according to the malicious email screening rule, calculate the label similarity between each of the suspicious emails and the existing malicious email set through the graph diffusion algorithm, and screen out several malicious emails from the suspicious emails according to the label similarity.

[0075] In the embodiment of the present application, first determine the connection relationship between each semantic label according to the text content of each suspicious email through the malicious email screening rule. For example, if a suspicious email contains "email address A" and "sender a", then there is a connection relationship between "email address A" and "sender a". Then calculate the number of occurrences of the semantic label in each suspicious email in the past n days to determine the relationship weight of each connection relationship. For example, if the number of times "email address A" and "sender a" appear in an email in the past n days is 5 times, then the corresponding relationship weight can be determined to be 5.

[0076] Exemplarily, the relationship weight can also be the frequency of the sender using a certain ip in the past n days, the frequency of the sender's email containing a certain url, the frequency of the sender's email having a certain attachment, the frequency of the sender's email having a certain attachment name, and so on.

[0077] Then set the label confidence of each suspicious email. For example, for semantic labels "Business Communication" and "Advertisement", set the confidence of "Business Communication" in a business communication email to 0.97 and that of "Advertisement" to 0.03.

[0078] Then given the label propagation coefficients, including the damping coefficient, the maximum number of iteration rounds, and the relationship weight threshold. Among them, the relationship weight threshold corresponding to each connection relationship is different and depends on the specific semantic label.

[0079] Generate a corresponding bipartite graph through the connection relationship and the relationship weight, and obtain the corresponding adjacency matrix through the bipartite graph. Then iterate the relationship weights in the adjacency matrix based on the label propagation coefficients. Stop iterating when the maximum number of iteration rounds is reached or when a certain connection weight is greater than the corresponding relationship weight threshold, and output the suspicious emails determined as malicious emails.

[0080] Exemplarily, if the frequency of a sender's email with a certain attachment within the recent n days is greater than 5, it indicates that the attachment is a potential candidate for a malicious email.

[0081] To further optimize the malicious email screening rules, rule mining is also performed on several screened malicious emails, and the malicious email screening rules are updated according to the mining results.

[0082] Implementing the embodiments of the present application has the following beneficial effects:

[0083] In the embodiments of the present application, first, semantic labels that may contain risk information are extracted from existing multiple malicious emails, and then these semantic labels are combined multiple times through rule mining to generate malicious email screening rules with a wide coverage, strong applicability, and the ability to accurately detect malicious emails in a large number of emails. Then, the existing malicious email set is textually matched with the emails to be detected, and the first threshold number of suspicious emails with a high text content match is selected as the data basis for subsequent email detection. Finally, the label similarity between the suspicious emails and the existing malicious email set is calculated according to the malicious email screening rules, and the suspicious emails with a relatively high label similarity are selected as malicious emails to timely protect the user information security.

[0084] Second Embodiment

[0085] Furthermore, to execute the malicious email detection system based on semantic rule mining corresponding to the above method embodiments to achieve the corresponding functions and technical effects, Figure 3 A structural diagram of a malicious email detection system based on semantic rule mining is provided. For ease of description, only parts related to this embodiment are shown. The malicious email detection system based on semantic rule mining provided by the embodiments of the present application includes:

[0086] The screening rule generation module 201 is configured to perform rule mining on a number of semantic tags extracted from an existing malicious email set based on the email detection accuracy requirement, and generate a malicious email screening rule.

[0087] In an embodiment of the present application, text features of the existing malicious email set are extracted and clustered respectively through a preset first semantic model and a preset second semantic model, and a number of semantic tags and text clustering IDs are output.

[0088] When the email detection accuracy requirement is higher than the set average value, the semantic tags and the text clustering IDs are combined and dynamically screened through a rule tree visualization algorithm to generate a malicious email screening rule.

[0089] When the email detection accuracy requirement is not higher than the set average value, the semantic tags and the text clustering IDs are combined through a rule tree integration algorithm to generate a malicious email screening rule.

[0090] The suspicious email screening module 202 is configured to perform text matching between an existing malicious email set and a to-be-detected email, and screen out the first threshold number of suspicious emails from the to-be-detected email.

[0091] In an embodiment of the present application, the existing malicious email set is first split into multiple email texts, and then the similarity between the to-be-detected email and the existing malicious email set is calculated sequentially through sparse matching and dense matching to complete the preliminary screening of malicious emails.

[0092] First, sparse matching is performed between the email text and the to-be-detected emails within the recent n days. Taking the email text as the query text, the word weights between the email text and the to-be-detected emails are calculated first, and then the similarity score between the email text and the to-be-detected emails is calculated in combination with the inverse document frequency. Among them, the inverse document frequency is a measure used to evaluate the general importance of a word in a document set. Its basic idea is: if a word appears less frequently in the document set, then its importance in a specific document is higher; conversely, if a word appears more frequently in the document set, then its importance in a specific document is lower.

[0093] Then, dense matching is performed on the to-be-detected emails with higher similarity scores and the email text. First, the email text is converted into a dense text vector, and vector calculation is performed on the dense text vector and the email text. According to the vector calculation result, the first threshold number of suspicious emails are found from the to-be-detected emails, and the suspicious emails will be used as the data basis for the next stage of malicious email investigation. Among them, the vector calculation can be maximum inner product calculation or minimum cosine distance calculation.

[0094] The malicious email detection module 203 is used to calculate the label similarity between each of the suspicious emails and the existing malicious email set through a graph diffusion algorithm according to the malicious email screening rule, and screen out several malicious emails from the suspicious emails according to the label similarity.

[0095] In the embodiment of the present application, according to the malicious email screening rule, the label similarity between each of the suspicious emails and the existing malicious email set is calculated through a graph diffusion algorithm, and several malicious emails are screened out from the suspicious emails according to the label similarity.

[0096] In some embodiments, the screening rule generation module 201 is specifically:

[0097] First, extract multiple semantic labels from the provided existing malicious email set. First, perform text feature extraction on the existing malicious email set through a trained first semantic model, and output several semantic labels; alternatively, perform text clustering on the existing malicious email set through a trained second semantic model, and output several text clustering IDs.

[0098] Specifically, the lightweight first semantic model can be used to analyze the email text online, classify it based on the text in the existing malicious email set, and output multiple semantic labels related to the malicious email risk. The first semantic model can meet the high-throughput performance requirements in an environment with limited computing resources, but it is difficult to consider the relationship between semantics and the word order of the text during the model analysis process, resulting in limited extraction effect of semantic labels. Using the second semantic model can perform deep clustering on the text of the existing malicious email set, output multiple text clustering IDs for subsequent screening rule mining, and can improve the coverage of the malicious email screening rule, but the disadvantage is that the data annotation cost of the second semantic model is relatively high.

[0099] Furthermore, the second semantic model can also perform offline analysis and online dense matching on the semantic labels, and distill the output result of the third semantic model into the first semantic model to improve the output effect of the first semantic model.

[0100] Optionally, the semantic label can be a table-like feature as the input of the second semantic model; the first semantic model is based on the FastText model, and the FastText model can perform efficient text classification and word embedding, mainly used to process large-scale text data, and has the characteristics of fast training and classification.

[0101] Then, select the mining method of the malicious email screening rule according to the set email detection accuracy requirement. If a high detection accuracy is required, the rule tree visualization algorithm is used to combine and dynamically screen the semantic tags and the text clustering IDs to generate a malicious email screening rule. Otherwise, the rule tree integration algorithm is used to directly generate the corresponding malicious email screening rule through the semantic tags and the text clustering IDs.

[0102] The rule tree integration algorithm is an end-to-end rule mining. It mainly combines the semantic tags output by the first semantic model and the text clustering IDs output by the second semantic model to generate a malicious email screening rule that meets the malicious email feature conditions. Specifically: First, input the preset decision tree parameters, including the number of decision trees, the depth of each decision tree, the frequency threshold, and the accuracy threshold, etc. And select semantic features and non-semantic features from the provided malicious email feature templates. Then input the selected features, decision tree parameters, and iteration labels into the trained decision tree model for processing to generate multiple decision trees. Among them, the maximum depth of the decision tree can be set according to the email detection accuracy requirement; if the maximum depth of the decision tree is larger, it means that the complexity of the rule mining is greater.

[0103] Then, based on the text clustering ID, with the semantic tag as the tree node and the iteration label as the true value of the tree node, select the path with the prediction result being true through label combination and true value determination. The path is the combination of multiple semantic tags based on the text clustering ID. At the same time, through the input frequency threshold, the paths that appear more than the frequency threshold in all decision trees are used as high-frequency paths. Screen the high-frequency paths with the prediction result being true and delete some unreasonable paths to find the high-frequency paths that meet the accuracy threshold and have the prediction result being true as the malicious email screening rule.

[0104] Exemplarily, for the obtained high-frequency path with the prediction result being true, if there are two tree nodes 'number of recipients > 3' and 'number of recipients < 2' in it, it means that there is a contradiction in this high-frequency path and it is unrealistic, and this high-frequency path should be deleted.

[0105] The rule tree visualization algorithm is a semi-automatic rule mining. It mainly obtains the malicious email screening rule based on keyword + semantic tag + other features through word co-occurrence analysis with the provided keywords and semantic tags, etc. Specifically: The data input into the trained decision tree model is basically the same, but the text features of the existing malicious email set are also processed by a preset third semantic model. According to actual needs, dynamic semantic tags are scattered and output on the basis of the existing semantic tags quickly. Then the mined dynamic semantic tags are distilled into the first semantic model through the second semantic model to improve the comprehensiveness of the semantic tags.

[0106] Among them, the third semantic model is an offline generation model that can quickly add / delete new semantics and intent tags according to user needs and output some dynamic semantic tags that pose a risk of malicious emails. The dynamic semantic tags are new semantic tags with summarization ability, which can assist the second semantic model and the first semantic model to improve the data processing effect. However, the drawback is that the throughput of the third semantic model is relatively low and it requires a large amount of computing resources. Therefore, the third semantic model is only adopted when the email detection accuracy requirement is high and the attack methods of malicious emails iterate relatively fast to improve the interception success rate of malicious emails.

[0107] Input the above data into the decision tree model, with semantic tags as tree nodes, and generate a decision tree with the same structure as that in the rule tree integration algorithm. For each tree node, display the proportion of malicious email features of each tree node in the form of a pie chart, and mark the node field features and the corresponding relationship weight thresholds. The proportion of malicious email features is the probability that the semantic tag corresponding to this tree node will cause the email to be a malicious email.

[0108] Then, divide the finite malicious email set into several email subsets, select the semantic tags with relatively high proportions of malicious email features, perform word co-occurrence analysis on the basis of the email subsets, generate the corresponding word co-occurrence graph, and mine and combine the tree nodes based on the provided malicious email keyword rules to generate the malicious email screening rules.

[0109] Furthermore, label propagation can also be performed through other provided semantic features to generate more semantic tags to be detected as tree nodes to expand the sample range.

[0110] Optionally, the decision tree model is constructed based on the GBDT model. GBDT is short for Gradient Boost Decision Tree, which is widely used in the risk control industry. In the case of small data volume and tabular data, it is widely considered to have stronger capabilities than deep models and has strong interpretability, making it very suitable for rule mining. At the same time, all data points will be very intuitively under a certain feature threshold segmentation (the intermediate result of GBDT). By controlling parameters such as the maximum tree depth, the complexity of the model can be constrained, making it possible for rule mining visualization and further word co-occurrence analysis. This ensures the interactivity and interpretability of semantic rule mining. In addition, different from deep models, the decision tree model can show better model effects in the case of low data volume.

[0111] In some embodiments, the malicious email detection module 203 is specifically:

[0112] First, determine the connection relationship between each semantic tag according to the text content of each said suspicious email through the malicious email screening rules. For example, if a suspicious email contains "Email address A" and "Sender a", there is a connection relationship between "Email address A" and "Sender a". Then calculate the number of occurrences of the said semantic tags in each said suspicious email within the last n days to determine the relationship weight of each said connection relationship. For example, if the number of times "Email address A" and "Sender a" appear simultaneously in an email within the last n days is 5 times, the corresponding relationship weight can be determined to be 5.

[0113] Exemplarily, the relationship weight can also be the frequency of the sender using a certain IP within the last n days, the frequency of a certain URL being included in the sender's email, the frequency of a certain attachment being carried in the sender's email, the frequency of a certain attachment name being carried in the sender's email, and so on.

[0114] Then set the label confidence of each suspicious email. For example, for the semantic tags "Business communication" and "Advertisement", in a business communication email, set "Business communication" to 0.97 and "Advertisement" to 0.03.

[0115] Then given the label propagation coefficients, including the damping coefficient, the maximum number of iteration rounds, and the relationship weight threshold. Among them, the relationship weight threshold corresponding to each connection relationship is different and depends on the specific semantic tag.

[0116] Generate a corresponding bipartite graph through the said connection relationship and the said relationship weight, and obtain the corresponding adjacency matrix through the bipartite graph. Then iterate the relationship weights in the adjacency matrix based on the label propagation coefficients. Stop the iteration when the maximum number of iteration rounds is reached or a certain connection weight is greater than the corresponding relationship weight threshold, and output the said suspicious emails determined to be malicious emails.

[0117] Exemplarily, if the frequency of a certain attachment being carried in the sender's email within the last n days is greater than 5, it indicates that the attachment is a potential malicious email candidate attachment.

[0118] To further optimize the malicious email screening rules, rule mining is also performed on several screened malicious emails, and the malicious email screening rules are updated according to the mining results.

[0119] Implementing the embodiments of the present application has the following beneficial effects:

[0120] In the embodiment of the present application, semantic tags that may contain risk information are first extracted from existing multiple malicious emails, and then these semantic tags are combined multiple times through rule mining to generate malicious email screening rules that have a wide coverage, strong applicability, and can accurately detect malicious emails in a large number of emails. Then, the existing malicious email set is textually matched with the emails to be detected, and the first threshold number of suspicious emails with a high text content match is selected as the data basis for subsequent email detection. Finally, according to the malicious email screening rules, the tag similarity between the suspicious emails and the existing malicious email set is calculated, and the suspicious emails with a higher tag similarity are selected as malicious emails to timely protect the user information security.

[0121] Further, Figure 3 The following is a structural diagram of a terminal device provided by an embodiment of the present application. As Figure 3 shown, the terminal device 3 of this embodiment includes: at least one processor 30 (only one is shown Figure 3 here), a memory 31, and a computer program 32 stored in the memory 31 and executable on the at least one processor. When the processor 30 executes the computer program 32, the steps of a malicious email detection method according to any one of the embodiments of the present application can be implemented.

[0122] The terminal device 3 can be a computing device such as a desktop computer, a cloud server, and a laptop computer. The computing device can include but is not limited to the processor 30 and the memory 31. Figure 3 This is only an example of the terminal device 3 and does not limit the terminal device 3. It may include more or fewer components than those shown in the figure.

[0123] The above specific embodiments have further elaborated on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above are only specific embodiments of the present application and are not used to limit the protection scope of the present application. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A malicious email detection method based on semantic rule mining, characterized in that: include: Based on the accuracy requirements of email detection, several semantic tags extracted from the existing malicious email set are mined to generate malicious email filtering rules; Performing text matching between the existing malicious email set and the emails to be detected, and filtering out suspicious emails of a first threshold value from the emails to be detected; According to the malicious email screening rule, the label similarity between each of the suspicious emails and the existing malicious email set is calculated by a graph diffusion algorithm, and a number of malicious emails are screened out from the suspicious emails according to the label similarity.

2. The malicious email detection method based on semantic rule mining according to claim 1 is characterized in that: Based on the requirements of email detection accuracy, rule mining is performed on several semantic tags extracted from the existing malicious email set to generate malicious email screening rules, which are specifically: Extracting and clustering text features of the existing malicious email set using a preset first semantic model and a preset second semantic model, and outputting a plurality of semantic labels and text clustering IDs; When the email detection accuracy requirement is higher than the set average value, the semantic label and the text cluster ID are combined and dynamically screened by a rule tree visualization algorithm to generate a malicious email screening rule; When the email detection accuracy requirement is not higher than a set average value, the semantic label and the text clustering ID are combined through a rule tree integration algorithm to generate a malicious email screening rule.

3. The malicious email detection method based on semantic rule mining according to claim 2 is characterized in that: The semantic tags and the text cluster IDs are combined and dynamically screened by the rule tree visualization algorithm to generate malicious email screening rules, which are specifically: Using a preset third semantic model to process text features of the existing malicious email set to generate a plurality of dynamic semantic tags; The semantic labels are used as tree nodes by a preset decision tree model, and the tree nodes are combined based on the text clustering ID to generate a plurality of decision trees; wherein the maximum depth of the decision tree is adjusted according to the email detection accuracy requirement; Calculate the malicious email feature ratio of each node in the decision tree according to the preset iteration label and generate a corresponding word co-occurrence graph; The nodes are mined and combined according to the word co-occurrence graph and preset malicious email keyword rules to generate the malicious email screening rules.

4. The malicious email detection method based on semantic rule mining according to claim 3 is characterized in that: The calculation of the malicious email feature ratio of each node in the decision tree and the generation of the corresponding word co-occurrence graph is specifically as follows: Calculating the malicious email feature ratio of each node in the decision tree, and selecting nodes whose malicious email feature ratio exceeds a second threshold as gray nodes; The existing malicious email set is segmented according to the gray nodes to obtain several email subsets; A word co-occurrence analysis is performed on the gray nodes in each email subset to generate a word co-occurrence graph corresponding to each email subset.

5. The malicious email detection method based on semantic rule mining according to claim 2 is characterized in that: The semantic tag and the text cluster ID are combined by the rule tree integration algorithm to generate a malicious email screening rule, which is specifically: The semantic labels are used as tree nodes by a preset decision tree model, and the tree nodes are combined based on the text clustering ID to generate a plurality of decision trees, and a plurality of label combinations are obtained through the decision trees; wherein the maximum depth of the decision tree is adjusted according to the accuracy requirement of the email detection; The number of occurrences of each of the label combinations is recorded, and the malicious email screening rule that meets the preset malicious email characteristic condition and whose number of occurrences exceeds a first threshold is selected.

6. The malicious email detection method based on semantic rule mining according to claim 1 is characterized in that: The text matching of the existing malicious email set with the emails to be detected is performed to filter out suspicious emails with a first threshold from the emails to be detected, specifically: Decomposing the existing malicious email set into a plurality of email texts, then sparsely matching the email texts with the emails to be detected within a third threshold number of days, and calculating a similarity score between each of the emails to be detected and the email texts; The text content of the to-be-detected email whose similarity score exceeds a fourth threshold is converted into a dense text vector, and vector calculation is performed on the dense text vector and the email text, and a first threshold number of suspicious emails are found from the to-be-detected emails according to the vector calculation result.

7. The malicious email detection method based on semantic rule mining according to claim 1 is characterized in that: The method of calculating the label similarity between each of the suspicious emails and the existing malicious email set by using a graph diffusion algorithm, and screening out a number of malicious emails from the suspicious emails according to the label similarity, is specifically as follows: According to the text content of each suspicious email, determining the connection relationship between each semantic tag through malicious email screening rules; Calculating the number of occurrences of the semantic tag in each of the suspicious emails within a third threshold number of days, thereby determining a relationship weight of each of the connection relationships; Generate a corresponding bidirectional graph through the connection relationship and the relationship weight, and obtain a corresponding adjacency matrix through the bidirectional graph; The relationship weights in the adjacency matrix are iterated based on a preset label propagation coefficient, and the suspicious email determined to be a malicious email is output when the iteration end condition is met; wherein the label propagation coefficient includes a maximum iteration round and a relationship weight threshold.

8. The malicious email detection method based on semantic rule mining according to any one of claims 1 to 7, characterized in that: The malicious email filtering rules are specifically as follows: Rule mining is performed on the filtered malicious emails, and the malicious email filtering rules are updated according to the mining results.

9. A malicious email detection system based on semantic rule mining, characterized in that: include: Screening rule generation module, suspicious email screening module and malicious email detection module; Among them, the screening rule generation module is used to perform rule mining on several semantic tags extracted from the existing malicious email set based on the email detection accuracy requirements to generate malicious email screening rules; The suspicious email screening module is used to perform text matching between the existing malicious email set and the emails to be detected, and screen out suspicious emails with a first threshold value from the emails to be detected; The malicious email detection module is used to calculate the label similarity between each of the suspicious emails and the existing malicious email set according to the malicious email screening rules through a graph diffusion algorithm, and screen out several malicious emails from the suspicious emails according to the label similarity.

10. A terminal device, characterized in that: It comprises a processor and a memory, the memory stores a computer program, and when the processor executes the computer program, it implements the steps of a malicious email detection method based on semantic rule mining as described in any one of claims 1 to 8.