Email-based risk prompt information generation method and apparatus, and medium

By detecting emails from multiple angles, high-comprehensive risk warning information is generated, and the problem of difficulty in generating effective risk warnings in the existing technology is solved, and fast and effective risk warnings and email security guarantees are achieved.

WO2025130083A1PCT designated stage expired Publication Date: 2025-06-26GUANGDONG COREMAIL COMPUTER TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/111701
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-08-13
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The prior art is difficult to generate effective comprehensive risk warning information based on email content, resulting in the problem of missing reminders by users.

Method used

Different detection methods are adopted from six perspectives: SPF verification, spelling of the sending domain name, unified resource locator, macro attachment, suspected impersonation name and money information, to generate extremely comprehensive risk warning information.

Benefits of technology

It realizes rapid and effective risk warnings, avoids the occurrence of missed reminders, and ensures the security of target emails. The prompts are calculated through specific data analysis and have high credibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024111701_26062025_PF_FP_ABST
    Figure CN2024111701_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention are an Email-based risk prompt information generation method and apparatus, and a medium. The method comprises: acquiring a target Email; performing SPF verification on the target Email, and generating a prompt word on the basis of a verification result; on the basis of a sending domain name, performing calculation to obtain a probability value, and combining the probability value with a confidence lower limit and a spelling feature to generate a prompt word; using a neural network model to perform segmentation and real number conversion on a uniform resource locator so as to generate a prompt word; performing feature analysis on a macro attachment, a suspected fraudulently-used name and money information to respectively generate corresponding prompt words; and generating risk prompt information on the basis of the prompt words. According to the Email-based risk prompt information generation method and apparatus, and the medium provided by the present invention, the target Email is comprehensively analyzed and detected from six different angles so as to generate comprehensive risk prompt information, and thus, the present invention can solve the problem of being difficult to generate effective comprehensive risk prompt information on the basis of Email content, and avoids the problem of prompt omission.
Need to check novelty before this filing date? Find Prior Art

Description

A method, device and medium for generating risk warning information based on email Technical Field

[0001] The present invention relates to the field of email technology, and in particular to a method, device and medium for generating risk warning information based on email. Background Art

[0002] As a communication method, email boasts a simple protocol, easy access to email addresses, and widespread use. Therefore, in addition to normal communication, it has also been targeted and abused by those with ulterior motives to deliver spam and phishing emails. Compared to standard spam, carefully crafted phishing emails are identical in content and format to regular emails, making it difficult for inexperienced readers to accurately identify the fraudulent content they contain. Furthermore, compared to fraud detection scenarios like cash scams and money transfers, email-based prevention solutions lack sufficient expertise in leveraging user agency and stimulating their awareness to reduce the likelihood of successful fraud. Existing technologies primarily alert users to the risks of email use by marking unfamiliar emails and providing pre-linked links with redirect pages.

[0003] However, the direct marking and provision of a transfer page method is rather mechanical, which can easily lead to users missing the mark and being deceived. Moreover, direct marking is to mark a number of abnormal points and remind users multiple times, resulting in a high frequency of reminders, a poor user experience and easy omission of reminders. Although the email client displays the real URL (Uniform Resource Locator) or [SPAM] tag, most users do not have sufficient knowledge to understand the composition of such characters and tags, resulting in the failure of the mark. In addition, the provision of a transfer page method can only determine whether the jump link is a domain name owned by the user himself, and the accuracy of this prompt is low.

[0004] Summary of the Invention

[0005] The present invention provides a method, device and medium for generating risk warning information based on emails, so as to solve the problem that it is difficult to generate effective comprehensive risk warning information based on email content and avoid missing reminders.

[0006] In order to solve the above problems, the present invention provides a method for generating risk warning information based on email, comprising:

[0007] Get the target email;

[0008] Performing an SPF check on the target email and generating a first prompt based on the check result;

[0009] Based on the spelling of the sending domain name in the target email, a probability value is calculated using a state transition probability method, and the probability value is combined with a preset confidence lower limit and the spelling characteristics of the target email to generate a second prompt;

[0010] Using a preset neural network model, segmenting and converting the uniform resource locator in the target email into a real number to generate a third prompt;

[0011] Performing feature analysis on the macro-containing attachments, suspected fraudulent names, and monetary information in the target email to generate a fourth prompt, a fifth prompt, and a sixth prompt, respectively;

[0012] Risk warning information is generated according to the first prompt, the second prompt, the third prompt, the fourth prompt, the fifth prompt, and the sixth prompt.

[0013] The present invention utilizes different detection methods from six perspectives: SPF verification, spelling of the sending domain name, Uniform Resource Locator (URL), attachments with macros, suspected fraudulent names, and monetary information. These methods generate corresponding warnings, thereby generating highly comprehensive risk warning information. This allows for rapid and effective risk warnings, ensuring the security of the target email. The probability value calculated using the state transition probability method reflects the degree of anomaly in the target email's email address, thus providing credibility support for the generation of the second warning. The use of a neural network model to segment the URL addresses addresses the issue of lengthy URLs, facilitating data analysis and accelerating the acquisition of the third warning.

[0014] Compared with the existing technology, this solution generates comprehensive risk warning information by conducting a comprehensive analysis and detection of the target email from six different angles, which can provide users with effective risk warnings at one time and avoid the occurrence of missed reminders. Moreover, since the prompts are obtained through specific data analysis and calculation, they have a higher credibility and can avoid ambiguous and non-targeted risk warning information, helping users avoid risks. Therefore, it can solve the problem of difficulty in generating effective comprehensive risk warning information based on email content and avoid missed reminders.

[0015] As a preferred solution, a probability value is calculated based on the spelling of the sending domain name in the target email using a state transition probability method. The probability value is combined with a preset confidence lower limit and the spelling characteristics of the target email to generate a second prompt, specifically:

[0016] Using a preset N value, traverse the sending domain name in the target email using the N-Gram algorithm to generate a character substring;

[0017] Combined with the preset character set, all possible N-Gram tuples are constructed according to the character substring to obtain a state transition probability matrix;

[0018] Splitting the sending domain name in the state transition probability matrix to obtain a plurality of tuples;

[0019] Calculating average state transition probability values ​​of the plurality of tuples to obtain the probability value;

[0020] The probability value is compared with the confidence lower limit, and the comparison result is combined with the spelling feature in the target email to generate a second prompt.

[0021] This preferred solution is to detect the sender's email address and then generate a second prompt based on it; starting from the spelling of the sender's domain name, the process of constructing all possible N-Gram tuples based on character substrings is the process of building a Markov chain. By constructing a Markov chain based on spelling to calculate the average state transition probability value as the probability value of the email address appearing, it can provide probabilistic credibility support for the generation of the second prompt.

[0022] As a preferred solution, the probability value is compared with a preset confidence lower limit, and the comparison result is combined with the spelling characteristics of the target email to generate a second prompt, specifically:

[0023] Comparing the probability value with a preset confidence lower limit to obtain a comparison result; wherein the confidence lower limit is obtained by calculating the probability values ​​of all domain names in a preset trusted domain name list;

[0024] Performing a standardization test on the spelling features of vowels and special symbols in the target email to obtain a test result;

[0025] The second prompt is generated according to the comparison result, the detection result, and the low-frequency communication top-level domain list of the mailbox where the target email is located.

[0026] In this preferred solution, since the lower limit of confidence is obtained by calculating the probability values ​​of all domain names in the preset trusted domain name list, it indicates the minimum value of the email credibility. Therefore, the probability value and the lower limit of confidence are compared, and the comparison result obtained can reflect the abnormality level of the email address of the current target email from an official perspective; by combining the character spelling feature detection results of the target email itself and the low-frequency communication top-level domain list, the email address structure of the target email itself and the email addresses with less communication can be considered, so that the second prompt is not only the result generated by analyzing the target email from an external perspective, but also includes a detailed consideration of internal details, making the second prompt more objective, accurate and effective.

[0027] As a preferred solution, the confidence lower limit is obtained by calculating the probability values ​​of all domain names in a preset trusted domain name list, specifically:

[0028] Obtain a list of websites ranked globally, and use partial listing information from the list to construct a list of trusted domain names;

[0029] Calculating probability values ​​of all domain names in the trusted domain name list to obtain a probability value set;

[0030] The minimum value in the probability value set is used as the confidence lower limit.

[0031] The trusted domain name list in this preferred solution is established based on the world's comprehensive ranking list and therefore has high reliability, thereby ensuring that the lower limit of confidence can become a comprehensive credibility measurement standard for comparison with the probability value to generate the second prompt.

[0032] As a preferred solution, a preset neural network model is used to segment and convert the uniform resource locator in the target email into a real number to generate a third prompt, specifically:

[0033] Using a preset neural network model, segmenting the uniform resource locator in the target email into a first character set;

[0034] Filtering out characters in the first character set that are not in the preset character set to obtain a second character set, and defining character identifiers of the second character set as unknown characters;

[0035] adding a preset beginning character and ending character to the beginning and end of the unknown character respectively to obtain a numbered list of the second character set;

[0036] The number list is transformed using a binary classification network structure in the neural network model to obtain an abnormality probability of the uniform resource locator, and the third prompt is generated according to the abnormality probability.

[0037] This preferred solution detects the Uniform Resource Locator (URL) contained in the email and then generates a third prompt based on it. By segmenting and filtering the URL in the target email, unknown characters are identified. By adding characters to the beginning and end of the unknown characters, the characters in the numbered list are clearly organized, facilitating faster conversion speeds during the transformation process, accelerating the process of determining the abnormality probability of the URL, and rapidly generating the third prompt. Compared to the sending domain name, the URL contained in the email is longer, making it unsuitable for simple state transition probabilities and character count analysis. Using a neural network model to convert the abnormality probability significantly reduces analysis time and improves the accuracy of the third prompt.

[0038] As a preferred solution, feature analysis is performed on the macro-containing attachment in the target email to generate a fourth prompt, specifically:

[0039] Obtaining the macro-containing attachment in the target email;

[0040] If a dangerous file in the focus list appears in the macro-attached attachment, generating the fourth prompt according to the dangerous file;

[0041] The focus list is established based on the prohibited attachment types for uploading to the mailbox where the target email is located and preset focus files.

[0042] This preferred solution detects attachments in emails and generates a fourth warning based on this information. By focusing on the list to determine whether macro-containing attachments are dangerous files, it can quickly determine whether macro-containing attachments in the target emails pose operational risks. The fourth warning is generated in a simple and quick manner.

[0043] As a preferred solution, feature analysis is performed on the suspected fraudulent name in the target email to generate a fifth prompt, specifically:

[0044] Obtaining the email header of the target email, and matching the email header with a preset dictionary tree to obtain a first matching result; wherein the dictionary tree is established based on a list of government agencies, banks, public security, procuratorial and judicial agencies, and the name of a permanent establishment of an enterprise;

[0045] Obtaining the entity name of the target email, and matching the entity name with the company list and the individual list in the sequence labeling algorithm to obtain a second matching result;

[0046] If the first matching result or the second matching result is a successful match, the sending email address of the target email is checked, and the fifth prompt is generated according to the checking result.

[0047] This preferred solution checks the email's subject name for forgery, and then generates the fifth prompt based on this. First, through matching, we can determine whether the target email successfully matches, that is, whether it constitutes a statement. Since a statement is a formal statement issued by an official body such as a government, enterprise, or organization, typically in response to a certain event or situation, expressing an official position and attitude, the email address from which the email is sent is generally official and formal. Therefore, after determining that it constitutes a statement, we directly check the sending email address to quickly determine whether the target email is forged, and then generate the fifth prompt.

[0048] As a preferred solution, the email address from which the target email is sent is checked, and the fifth prompt is generated according to the check result, specifically:

[0049] Performing category determination on the email address from which the target email originates, and obtaining the declared object of the target email;

[0050] If the declared object is a government agency or a public security, procuratorial or judicial agency, then check whether the top-level domain name of the target email belongs to a preset special top-level domain;

[0051] If the target of the declaration is a bank, invoice service provider or enterprise, check whether the domain name used in the target email belongs to a pre-collected list;

[0052] If the declared object is a permanent establishment of an enterprise, check whether the sender and recipient of the target email are in the same domain or subdomain;

[0053] If the declared object is a person's name, check whether the declared object is in the preset address book;

[0054] The fifth prompt is generated according to the inspection result.

[0055] In this preferred solution, since government agencies and public security, procuratorial, and judicial bodies often have unique top-level domains, direct checks can be performed on these domains to determine if any mailboxes are anomalous. Since permanent corporate entities often share affiliations, checks can be performed to determine if any mailboxes are anomalous by checking whether they are on the same domain or subdomain. Using pre-collected lists and address books, it's possible to quickly and directly determine if bank or individual mailboxes are anomalous. This solution provides specific methods for checking the sending email addresses of government agencies, banks, permanent corporate entities, and individuals. This highly targeted approach can accelerate the generation of the fifth prompt.

[0056] As a preferred solution, feature analysis is performed on the money information in the target email to generate a sixth prompt, specifically:

[0057] Obtaining text information of the target email;

[0058] Using a preset classifier to segment the text information to obtain a number of character strings;

[0059] Matching the plurality of character strings with a preset money vocabulary, and generating the sixth prompt according to the matching result;

[0060] The money word list is established based on preset financial fraud information. This preferred solution uses a classifier to segment the text information, which is equivalent to data segmentation and extraction of the original email text, so that the vocabulary in the obtained character strings can be easily compared and matched with the money word list, thereby accelerating the acquisition process of the sixth prompt.

[0061] The present invention also provides a device for generating risk warning information based on emails, comprising an email acquisition module, a first generation module, a second generation module, a third generation module, a fourth generation module and a summary module;

[0062] Wherein, the email acquisition module is used to acquire target emails;

[0063] The first generating module is configured to perform an SPF check on the target email and generate a first prompt according to the check result;

[0064] The second generating module is configured to calculate a probability value based on the spelling of the sending domain name in the target email using a state transition probability method, and generate a second prompt by combining the probability value with a preset confidence lower limit and spelling features in the target email;

[0065] The third generating module is configured to generate a third prompt by segmenting and converting the uniform resource locator in the target email into a real number using a preset neural network model;

[0066] The fourth generating module is configured to perform feature analysis on the macro-containing attachment, suspected fraudulent name, and monetary information in the target email, and generate a fourth prompt, a fifth prompt, and a sixth prompt, respectively;

[0067] The summarizing module is configured to generate risk warning information according to the first prompt, the second prompt, the third prompt, the fourth prompt, the fifth prompt, and the sixth prompt.

[0068] As a preferred solution, the second generation module includes a character unit, a matrix unit, a splitting unit, a probability value unit and a comparison unit;

[0069] The character unit is used to use a preset N value to traverse the sender domain name in the target email through the N-Gram algorithm to generate a character substring;

[0070] The matrix unit is used to construct all possible N-Gram tuples according to the character substring in combination with a preset character set to obtain a state transition probability matrix;

[0071] The splitting unit is used to split the sending domain name in the state transition probability matrix to obtain a plurality of tuples;

[0072] The probability value unit is used to calculate the average state transition probability value of the plurality of tuples to obtain the probability value;

[0073] The comparison unit is configured to compare the probability value with the confidence lower limit, and generate a second prompt based on the comparison result and the spelling feature in the target email.

[0074] As a preferred solution, the comparison unit includes a first subunit, a second subunit and a third subunit;

[0075] The first subunit is configured to compare the probability value with a preset confidence lower limit to obtain a comparison result; wherein the confidence lower limit is obtained by calculating the probability values ​​of all domain names in a preset trusted domain name list;

[0076] The second subunit is configured to perform a standardization detection on the spelling features of vowels and special symbols in the target email to obtain a detection result;

[0077] The third subunit is configured to generate the second prompt according to the comparison result, the detection result, and the low-frequency communication top-level domain list of the mailbox where the target email is located.

[0078] As a preferred solution, the confidence lower limit is obtained by calculating the probability values ​​of all domain names in a preset trusted domain name list, specifically:

[0079] Obtain a list of websites ranked globally, and use partial listing information from the list to construct a list of trusted domain names;

[0080] Calculating probability values ​​of all domain names in the trusted domain name list to obtain a probability value set;

[0081] The minimum value in the probability value set is used as the confidence lower limit.

[0082] As a preferred solution, the third generation module includes a first segmentation unit, a screening unit, a reconstruction unit and a transformation unit;

[0083] The first segmentation unit is configured to segment the uniform resource locator in the target email into a first character set using a preset neural network model;

[0084] The screening unit is configured to screen out characters in the first character set that are not in the preset character set to obtain a second character set, and define character identifiers of the second character set as unknown characters;

[0085] The reconstruction unit is configured to add a preset beginning character and an ending character to the beginning and the end of the unknown character, respectively, to obtain a number list of the second character set;

[0086] The transformation unit is used to use the binary classification network structure in the neural network model to transform the number list to obtain the abnormality probability of the uniform resource locator, and generate the third prompt according to the abnormality probability.

[0087] As a preferred solution, the fourth generating module includes a first acquiring unit and a first generating unit;

[0088] The first acquiring unit is configured to acquire the macro-bearing attachment in the target email;

[0089] The first generating unit is configured to generate the fourth prompt according to a dangerous file in a key concern list if the macro-attached attachment contains the dangerous file;

[0090] The focus list is established based on the prohibited attachment types for uploading to the mailbox where the target email is located and preset focus files.

[0091] As a preferred solution, the fourth generation module includes a second acquisition unit, a first matching unit and a checking unit;

[0092] The second acquisition unit is configured to acquire a mail header of the target mail, and match the mail header with a preset dictionary tree to obtain a first matching result; wherein the dictionary tree is established based on a list of government agencies, banks, public security, procuratorial and judicial agencies, and the names of permanent establishments of enterprises;

[0093] The first matching unit is configured to obtain an entity name of the target email, and match the entity name with a company list and an individual list in a sequence labeling algorithm to obtain a second matching result;

[0094] The checking unit is configured to check the sending email address of the target email if the first matching result or the second matching result is a successful match, and generate the fifth prompt according to the checking result.

[0095] As a preferred solution, the inspection unit includes a fourth subunit, a fifth subunit, a sixth subunit, a seventh subunit, an eighth subunit and a ninth subunit;

[0096] The fourth subunit is configured to determine the category of the email address from which the target email originates, and obtain the declared object of the target email.

[0097] The fifth subunit is configured to check whether the top-level domain name of the target email belongs to a preset special top-level domain if the declared object is a government agency or a public security, procuratorial or judicial agency;

[0098] The sixth subunit is configured to check whether the domain name used in the target email belongs to a pre-collected list if the declared object is a bank, invoice service provider or enterprise;

[0099] The seventh subunit is configured to check whether the sender and recipient of the target email are in the same domain or subdomain if the declared object is a permanent establishment of an enterprise;

[0100] The eighth subunit is configured to check whether the declared object is in a preset address book if the declared object is a person's name;

[0101] The ninth subunit is configured to generate the fifth prompt according to the inspection result.

[0102] As a preferred solution, the fourth generation module includes a third acquisition unit, a second segmentation unit and a second matching unit;

[0103] Wherein, the third obtaining unit is used to obtain text information of the target email;

[0104] The second segmentation unit is configured to segment the text information using a preset classifier to obtain a plurality of character strings;

[0105] The second matching unit is configured to match the plurality of character strings with a preset money vocabulary, and generate the sixth prompt according to the matching result;

[0106] The money vocabulary is established based on preset financial fraud information. The present invention also provides a storage medium having a computer program stored thereon, which is called and executed by a computer to implement the above-mentioned method for generating risk warning information based on email. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] FIG1 is a flow chart of a method for generating risk warning information based on emails provided by an embodiment of the present invention;

[0108] FIG2 is a schematic structural diagram of an email-based risk warning information generation device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0109] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0110] In the description of this application, it should be understood that the terms "first," "second," "third," ..., and "tenth" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first," "second," "third," ..., and "tenth" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "several" means two or more.

[0111] The email-based risk warning information generation method described in the embodiment of the present invention is mainly used when using an email mailbox and needs to perform risk detection on the received target email to generate prompt information to provide corresponding prompts to the user and improve the user's vigilance.

[0112] Example 1:

[0113] Referring to FIG1 , an embodiment of the present invention provides a method for generating risk warning information based on emails, including S1 to S6. The specific implementation steps are as follows:

[0114] S1. Get the target email.

[0115] Step S1 of the embodiment of the present invention is specifically as follows:

[0116] Get the target email from the mailbox system.

[0117] S2. Perform SPF verification on the target email and generate a first prompt based on the verification result.

[0118] Step S2 of the embodiment of the present invention is specifically as follows:

[0119] Use SPF (Sender Policy Framework), DKIM (DomainKeys Identified Mail) and DMARC (Domain-based Message Authentication, Reporting, and Conformance) to verify the target email and generate the first prompt based on the verification results.

[0120] In this embodiment, because SMTP does not verify the sender's email address, SPF can supplement the security of SMTP (Simple Mail Transfer Protocol) in the target email; however, the SPF protocol has many problems during use. For example, some domain names lack strict standard SPF configuration, which makes it possible for some emails that are not forged to fail SPF certification. Therefore, not all emails that fail SPF certification will pose serious risks, so letting users know the SPF verification results can help them make judgments; and, through DKIM and DMARC, SPF can be further assisted to obtain more complete prompts, providing users with more comprehensive prompt information.

[0121] S3. According to the spelling of the sending domain name in the target email, a probability value is calculated using a state transition probability method, and the probability value is combined with a preset confidence lower limit and the spelling characteristics of the target email to generate a second prompt.

[0122] In step S3 of the embodiment of the present invention, S3 includes S3.1 to S3.5, specifically:

[0123] S3.1. Convert the domain name in the target email to lowercase.

[0124] Using the preset N value, the N-Gram algorithm traverses the sending domain name in the target email and generates a character substring.

[0125] For example, assuming the target email's sending domain name is 123456.cn and N is 3, then after traversing the sending domain name using the N-Gram algorithm, the final substring list obtained is: "123, 234, 345, 456, 56., 6.c, .cn".

[0126] S3.2. Based on a preset character set, construct all possible N-Gram tuples according to the character substrings to obtain a state transition probability matrix. The character set is established based on RFC standards related to characters that can be used in the sending domain name, and its source can be RFC 1035, RFC 1123, RFC 2181, and RFC 5892.

[0127] Split the sending domain name in the state transition probability matrix to obtain several tuples;

[0128] The average state transition probability value of several tuples is calculated to obtain the probability value.

[0129] In this step S3.2, in order to reduce the matrix size, N=2 can be specified, thereby directly constructing all possible 2-gram tuples of size M, and then forming an initial state transition matrix of size M×M. Then, the probability of state transition between all 2-gram tuples is calculated to obtain the state transition probability matrix.

[0130] The specific implementation of step S3.2 is as follows:

[0131] Assuming that the corpus contains only one sentence, "I robot", by traversing the sending domain name in the target email through the N-Gram algorithm, we can find that when N = 3, the tuples after splitting are: "iro, rob, obo, bot". According to the order of front and back, we can know that the transfer object of the iro tuple is only the rob tuple, and the other tuples are similar. Therefore, the transfer probability values ​​of several tuples are shown in Table 1.

[0132] Table 1 Tuple transition probability value comparison table

[0133] For this application embodiment of the present invention, please refer to Table 1. Table 1 provides a tuple transition probability value comparison table, which is the transition probability values ​​of several tuples in the specific implementation example of step S3.2.

[0134] This embodiment starts with the spelling of the domain name of the email sending. The process of constructing all possible N-Gram tuples based on the character substring is the process of building a Markov chain. By building a Markov chain based on the spelling to calculate the average state transition probability value as the probability value of the email address appearing, it can provide probabilistic credibility support for the generation of the second prompt.

[0135] S3.3. Compare the probability value with a preset lower confidence limit to obtain a comparison result; wherein the lower confidence limit is obtained by calculating the probability values ​​of all domain names in the preset trusted domain name list;

[0136] The construction process of the lower confidence limit is as follows:

[0137] Obtain the Alexa Rank list of websites (the world's comprehensive ranking list), and use the list information of the top 20,000 in the Alexa Rank list to build a list of trusted domain names;

[0138] Calculate the probability values ​​of all domain names in the trusted domain name list to obtain a probability value set;

[0139] The minimum value in the probability value set is taken as the lower confidence limit.

[0140] The trusted domain name list in this embodiment is established based on the world's comprehensive ranking list and therefore has high reliability. This ensures that the lower limit of confidence can become a comprehensive credibility measurement standard for comparison with the probability value to generate the second prompt.

[0141] S3.4. Perform a standardization check on the spelling features of vowels and special symbols in the target email to obtain a test result.

[0142] Specifically, the number of vowels and special symbols (such as "-", "_", and ".") in the target email is counted. Based on the upper threshold calculated from historical statistical data, the domain names exceeding the upper threshold are identified and the detection result is obtained.

[0143] Among them, the detection results can show the number of symbols with low usage rates and the number of vowels that account for too low a proportion.

[0144] S3.5. Generate a second prompt based on the comparison result, the detection result, and the low-frequency communication top-level domain list of the mailbox where the target email is located.

[0145] The specific process of constructing the low-frequency communication top-level domain list is as follows:

[0146] We summarize pre-collected spam data and vendor security reports to create a list of top-level domains. The inclusion criteria for this list include low frequency of use, low registration cost, and high spam frequency.

[0147] Collect historical email data of the user in the mailbox where the target email is located, and calculate the top-level domains with high communication frequency. If the top-level domain of the sending mailbox is not on the list, it can be considered as a top-level domain with low communication frequency, and a list of low-frequency top-level domains is obtained;

[0148] A low-frequency communication top-level domain list is established based on the top-level domain list and the low-frequency communication top-level domain list.

[0149] In step S3 of this embodiment, it is difficult for each security expert or recipient to independently identify whether a sender's email address may have problems. Therefore, risk detection and reminder for email addresses are of great importance.

[0150] Since the lower confidence limit is calculated by calculating the probability values ​​of all domain names in the preset trusted domain name list, it indicates the minimum value of the email credibility. Therefore, comparing the probability value and the lower confidence limit can reflect the degree of abnormality of the current target email address from an official perspective.

[0151] In addition, since the lower threshold limit of the state transition probability calculation uses the minimum value, it is relatively conservative. Therefore, by combining the character spelling feature detection results of the target email itself and the low-frequency communication top-level domain list, the email address structure of the target email itself and the email addresses with less frequent communication can be taken into account. In addition, the second prompt can not only be the result of analyzing the target email from an external perspective, but also include a detailed consideration of internal details, making the second prompt more objective, accurate and effective.

[0152] S4. Using a preset neural network model, the third prompt is generated by segmenting and converting the uniform resource locator in the target email into a real number.

[0153] In step S4 of the embodiment of the present invention, S1 includes S4.1 to S4.4, specifically:

[0154] S4.1. Use a preset LSTM (Long Short-Term Memory) deep learning network structure to segment the URL (Uniform Resource Locator) in the target email into a first character set. The input processing of the LSTM deep learning network structure refers to the input word processing method of the large language model, and the network structure defines a large number of characters that can support segmentation.

[0155] S4.2. Filter out characters in the first character set that are not in the preset character set to obtain a second character set, and define character identifiers in the second character set as unknown characters;

[0156] Preset start characters and end characters are added to the head and tail of the unknown character respectively to obtain a number list of the second character set.

[0157] S4.3. Generate comprehensive auxiliary prompts based on the IP host, short URL, and external download link in the target email, specifically:

[0158] If the IP address of the host where the target email is located is not an intranet IP, the first auxiliary prompt will be output to indicate that the IP address of the target email is not an intranet IP.

[0159] If the URL in the target email ends with a browser download link, a second auxiliary prompt will be output to indicate that the target email points to an external download link. For example, https: / / www.123.com / abc.png is an image, so most browsers will not trigger an automatic download, so it is not an external download link. However, if the URL is https: / / www.123.com / abc.exe, most browsers will interpret it as a download link and pop up a save window, thus constituting a download link.

[0160] If the host location of the target email is matched with the domain name in the pre-obtained short URL service list through regular expression, a third auxiliary prompt is output to indicate that the target email points to the short URL;

[0161] A comprehensive auxiliary prompt is generated according to the first auxiliary prompt, the second auxiliary prompt and the third auxiliary prompt.

[0162] This embodiment generates comprehensive auxiliary prompts based on the differences in different scenarios such as using IP as the host bit, using short URLs, and pointing to external download links. Since spelling-based URL checking is not suitable in these scenarios, the LSTM network cannot be used to detect these URLs. Therefore, by taking corresponding detection measures for these three types of information, a comprehensive auxiliary prompt can be obtained to provide relevant reminders.

[0163] S4.4. Use the binary classification network structure in the LSTM deep learning network structure to perform a softmax transformation on the number list to obtain the abnormal probability of the Uniform Resource Locator and generate an initial prompt based on the abnormal probability;

[0164] The initial prompt and the comprehensive auxiliary prompt are combined to obtain a third prompt.

[0165] It should be noted that in the embodiments of the present invention, the combination of the first-level domain name and the top-level domain after excluding the top-level domain is considered the host location. For multiple URLs with the same host location, only the first URL that appears is checked. For example, the host location of 123.com.cn and 456.123.com.cn are both 123.com.cn, but 123.com and 123.cn are not the same host location. In addition, the softmax transform is a mathematical transformation of the original variable, which uses the softmax function. The softmax function is also called the normalized exponential function. It is a mathematical function that is generally used to convert a set of arbitrary real numbers into real numbers representing a probability distribution.

[0166] From a general perspective, step S4 of this embodiment, since spam phishing emails often achieve their goals by guiding users to click on corresponding URLs, prompting potentially risky URLs in target emails can provide a more refined and accurate prompting effect than using a forwarding page to forward all external links.

[0167] This embodiment detects the Uniform Resource Locator (URL) in an email and then generates a third prompt based on it. By segmenting and filtering the URL in the target email, unknown characters can be obtained. By adding characters at the beginning and end of the unknown characters, the characters in the numbered list can be clearly organized, facilitating faster conversion speeds during the transformation process, accelerating the process of obtaining the probability of anomalies in the URL, and quickly generating the third prompt.

[0168] Compared with the sending domain name, the uniform resource locator in the email is longer, so it is not suitable to be judged by simple state transition probability and number of characters. By using a neural network model to convert abnormal probabilities, it can greatly save analysis time and improve the accuracy of the third prompt.

[0169] S5. Perform feature analysis on the macro-containing attachments, suspected fraudulent names, and monetary information in the target email, and generate a fourth prompt, a fifth prompt, and a sixth prompt, respectively.

[0170] In step S5 of the embodiment of the present invention, S5 includes S5.1 to S5.3, specifically:

[0171] S5.1. Obtain the macro-containing attachments in the target email;

[0172] If a dangerous file on the focus list appears in the macro attachment, a fourth prompt will be generated based on the dangerous file;

[0173] The focus list is created based on the prohibited attachment types of the mailbox where the target email is located (for example, the list of attachment types prohibited from uploading in Outlook) and the preset focus files;

[0174] Among them, the files of focus include files that can be directly executed, such as ".exe", ".bat", ".ps" and ".sh", as well as macro files such as ".docm" and ".xlsm".

[0175] This embodiment detects attachments carried in emails and then generates a fourth prompt based on this. Because the aforementioned focused files are easily executed due to misoperation or misleading content in emails, for example, an Office file containing macros is generally safe even if it carries a macro virus when in read-only mode with macros disabled, but may be infected if the user enables macros. Therefore, by using the focused list to determine whether macro-containing attachments are dangerous files, it is possible to quickly determine whether macro-containing attachments in the target emails pose operational risks and provide timely reminders to users. The generation method of this fourth prompt is simple and quick.

[0176] S5.2. Obtain the email header of the target email and match the email header with a preset dictionary tree to obtain a first matching result; wherein the dictionary tree is constructed based on a list of government agencies, banks, public security, procuratorial and judicial agencies, and the names of permanent establishments of enterprises;

[0177] Obtain the entity name of the target email, match the entity name with the B-ORG (company or organization) list and the I-PER (individual) list in the sequence annotation algorithm, and obtain the second matching result;

[0178] If the first matching result or the second matching result is a successful match, the category of the email address from which the target email was sent is determined to obtain the declared object of the target email;

[0179] If the target of the claim is a government agency or a public security, procuratorial or judicial agency, check whether the top-level domain name of the target email belongs to the preset special top-level domain (such as ".gov" and ".mit");

[0180] If the target is a bank, invoice service provider, or enterprise, check whether the domain name used in the target email belongs to a pre-collected list;

[0181] If the target of the claim is a permanent establishment of an enterprise, check whether the sender and recipient of the target email are in the same domain or subdomain;

[0182] If the declared object is a person's name, check whether the declared object is in the preset address book;

[0183] A fifth prompt is generated according to the inspection result.

[0184] This embodiment detects whether the subject name of the email is forged, and then generates the fifth prompt based on this. Although protocols such as SPF have a certain degree of protection against forged email addresses, they cannot prevent the use of other people's names by exploiting the visual effects of email reading. Therefore, it is important to find and determine whether the subject name is forged by finding and judging the entity.

[0185] First, through matching, we can determine whether the target email is a match, that is, whether it constitutes a statement. A statement is a formal statement issued by an official body such as a government, enterprise, or organization. It is usually issued in response to an event or situation, expressing an official position and attitude. The sending email address is generally official and formal. Therefore, after determining that it constitutes a statement, we directly check the sending email address to quickly determine whether the target email is a forgery, thereby obtaining the fifth prompt. In addition, the generation of the second matching result utilizes the open source BiLSTM+CRF technology, which can provide a corresponding list for key targets for fast matching.

[0186] Furthermore, since the top-level domains of government, public security, procuratorial, and judicial organizations often have unique top-level domains, direct checks can be performed on these domains to determine if any mailboxes are anomalous. Since permanent corporate entities often share affiliations, checking for shared domains or subdomains can be used to determine if any mailboxes are anomalous. Using pre-collected lists and address books, it's possible to quickly and directly determine if bank or individual mailboxes are anomalous. This solution provides specific methods for checking the sending email addresses of government agencies, banks, permanent corporate entities, and individuals. This highly targeted approach can accelerate the generation of the fifth prompt.

[0187] S5.3. Obtain the text information of the target email;

[0188] Use the TextCNN (convolutional neural network) module as a classifier to segment text information and obtain several string tokens. The classifier controls its physical size and memory space by presetting the vocabulary.

[0189] Match several strings with the preset money word list, mask the tokens that are not in the money word list, and obtain the matching results;

[0190] generating a sixth prompt according to the matching result;

[0191] Among them, the money vocabulary is established based on preset financial fraud information.

[0192] In this embodiment, general email classification only determines whether an email is spam and does not care much about the specific type of spam. However, among all scams, those involving money are often the most sensitive. Therefore, it is necessary to provide additional reminders for messages involving money and potentially constituting fraud.

[0193] Using a classifier to segment the text information is equivalent to performing data segmentation and extraction on the original email text, so that the vocabulary in the obtained character strings can be easily compared and matched with the money word list, thereby accelerating the acquisition process of the sixth prompt.

[0194] S6. Generate risk warning information according to the first prompt, the second prompt, the third prompt, the fourth prompt, the fifth prompt, and the sixth prompt.

[0195] Step S6 of the embodiment of the present invention is specifically as follows:

[0196] Assign a unique ID and label to each of the first, second, third, fourth, fifth, and sixth prompts. Based on a pre-established multilingual text dictionary and corresponding placeholders for the output of each label, control the output of the first, second, third, fourth, fifth, and sixth prompts corresponding to each label by inputting language tag parameters. Combined information is then generated to obtain risk warning information.

[0197] Among them, the multilingual copywriting dictionary is written manually.

[0198] This embodiment controls the generation of prompts by formatting output to support multiple languages. Compared with the method of using machine translation, the prompts in this solution are written manually, and the language is more appropriate and natural, which can enhance the user experience.

[0199] In general, the embodiments of the present invention have the following beneficial effects:

[0200] The embodiment of the present invention adopts different detection methods from six perspectives: SPF verification, spelling of the sending domain name, Uniform Resource Locator, macro attachments, suspected fraudulent names, and monetary information, to obtain corresponding prompts, and then generates highly comprehensive risk warning information. This can quickly and effectively provide risk warnings at one time, avoid the occurrence of missed warnings, and ensure the information security of the target email.

[0201] Among them, a method based on entity recognition and knowledge base matching is used to identify potential mismatched identity claims to prompt possible impersonation emails, which covers a wider range and can reduce the risk of omissions; the entity recognition-based solution can find the organization name in a given text to discover potential entity impersonations, which is impossible to do through simple contact matching; and compared with the "one-size-fits-all" solution of prohibiting direct access to external links, the embodiment of the present invention uses a method that combines empirical knowledge and deep learning models to identify specific suspicious URLs, which can avoid the problem of users being numb and missing information reminders due to overly general prompts.

[0202] Example 2:

[0203] Referring to FIG2 , an embodiment of the present invention provides an apparatus for generating risk warning information based on emails, comprising an email acquisition module 10 , a first generation module 20 , a second generation module 30 , a third generation module 40 , a fourth generation module 50 , and a summary module 60 ;

[0204] Among them, the email acquisition module 10 is used to acquire the target email;

[0205] A first generating module 20 is configured to perform an SPF check on the target email and generate a first prompt based on the check result;

[0206] A second generating module 30 is configured to calculate a probability value based on the spelling of the sending domain name in the target email using a state transition probability method, and to generate a second prompt by combining the probability value with a preset confidence lower limit and spelling characteristics in the target email;

[0207] A third generating module 40 is configured to generate a third prompt by segmenting and converting the Uniform Resource Locator in the target email using a preset neural network model;

[0208] A fourth generating module 50 is configured to perform feature analysis on macro-containing attachments, suspected fraudulent names, and monetary information in the target email, and generate a fourth prompt, a fifth prompt, and a sixth prompt, respectively;

[0209] The summarizing module 60 is configured to generate risk warning information according to the first prompt, the second prompt, the third prompt, the fourth prompt, the fifth prompt, and the sixth prompt.

[0210] In one embodiment, the email acquisition module 10 specifically includes:

[0211] Get the target email from the mailbox system.

[0212] In one embodiment, the first generating module 20 is specifically:

[0213] Use SPF (Sender Policy Framework), DKIM (DomainKeys Identified Mail) and DMARC (Domain-based Message Authentication, Reporting, and Conformance) to verify the target email and generate the first prompt based on the verification results.

[0214] In this embodiment, because SMTP does not verify the sender's email address, SPF can supplement the security of SMTP (Simple Mail Transfer Protocol) in the target email; however, the SPF protocol has many problems during use. For example, some domain names lack strict standard SPF configuration, which makes it possible for some emails that are not forged to fail SPF certification. Therefore, not all emails that fail SPF certification will pose serious risks, so letting users know the SPF verification results can help them make judgments; and, through DKIM and DMARC, SPF can be further assisted to obtain more complete prompts, providing users with more comprehensive prompt information.

[0215] In one embodiment, the second generation module 30 includes a character unit, a matrix unit, a splitting unit, a probability value unit, a matrix unit, a first sub-unit, a second sub-unit, and a third sub-unit;

[0216] The character unit is used to convert the sending domain name in the target email into lowercase form;

[0217] The character unit is also used to use a preset N value to traverse the sending domain name in the target email through the N-Gram algorithm to generate a character substring.

[0218] For example, assuming the target email's sending domain name is 123456.cn and N is 3, then after traversing the sending domain name using the N-Gram algorithm, the final substring list obtained is: "123, 234, 345, 456, 56., 6.c, .cn".

[0219] A matrix unit is used to construct all possible N-Gram tuples according to character substrings in combination with a preset character set to obtain a state transition probability matrix; wherein the character set is established based on RFC standards related to characters that can be used in the sending domain name, and its source can be RFC 1035, RFC 1123, RFC 2181, and RFC 5892;

[0220] A splitting unit is used to split the sending domain name in the state transition probability matrix to obtain a number of tuples;

[0221] The probability value unit is used to calculate the average state transition probability value of a plurality of tuples to obtain a probability value.

[0222] In the matrix unit, in order to reduce the matrix size, N = 2 can be specified, thereby directly constructing all possible 2-gram tuples of size M, and then forming an initial state transition matrix of size M×M. Then, the probability of state transition between all 2-gram tuples is calculated to obtain the state transition probability matrix.

[0223] The specific implementation examples of the matrix unit, the splitting unit and the probability value unit are as follows:

[0224] Assuming that the corpus contains only one sentence, "I robot", by traversing the sending domain name in the target email through the N-Gram algorithm, we can find that when N = 3, the tuples after splitting are: "iro, rob, obo, bot". According to the order of front and back, we can know that the transfer object of the iro tuple is only the rob tuple, and the other tuples are similar. Therefore, the transfer probability values ​​of several tuples are shown in Table 1.

[0225] Table 1 Tuple transition probability value comparison table

[0226] This application embodiment of the present invention, please refer to Table 1, which provides a tuple transition probability value comparison table, which is the transition probability values ​​of several tuples in the specific implementation examples of the matrix unit, split unit and probability value unit.

[0227] This embodiment starts with the spelling of the domain name of the email sending. The process of constructing all possible N-Gram tuples based on the character substring is the process of building a Markov chain. By building a Markov chain based on the spelling to calculate the average state transition probability value as the probability value of the email address appearing, it can provide probabilistic credibility support for the generation of the second prompt.

[0228] The first subunit is configured to compare the probability value with a preset confidence lower limit to obtain a comparison result; wherein the confidence lower limit is obtained by calculating the probability values ​​of all domain names in a preset trusted domain name list;

[0229] The construction process of the lower confidence limit is as follows:

[0230] Obtain the Alexa Rank list of websites (the world's comprehensive ranking list), and use the list information of the top 20,000 in the Alexa Rank list to build a list of trusted domain names;

[0231] Calculate the probability values ​​of all domain names in the trusted domain name list to obtain a probability value set;

[0232] The minimum value in the probability value set is taken as the lower confidence limit.

[0233] The trusted domain name list in this embodiment is established based on the world's comprehensive ranking list and therefore has high reliability. This ensures that the lower limit of confidence can become a comprehensive credibility measurement standard for comparison with the probability value to generate the second prompt.

[0234] The second subunit is used to perform standardization detection on the character spelling features of vowels and special symbols in the target email to obtain a detection result.

[0235] Specifically, the number of vowels and special symbols (such as "-", "_", and ".") in the target email is counted. Based on the upper threshold calculated from historical statistical data, the domain names exceeding the upper threshold are identified and the detection result is obtained.

[0236] Among them, the detection results can show the number of symbols with low usage rates and the number of vowels that account for too low a proportion.

[0237] The third subunit is configured to generate a second prompt based on the comparison result, the detection result, and the low-frequency communication top-level domain list of the mailbox where the target email is located.

[0238] The specific process of constructing the low-frequency communication top-level domain list is as follows:

[0239] We summarize pre-collected spam data and vendor security reports to create a list of top-level domains. The inclusion criteria for this list include low frequency of use, low registration cost, and high spam frequency.

[0240] Collect historical email data of the user in the mailbox where the target email is located, and calculate the top-level domains with high communication frequency. If the top-level domain of the sending mailbox is not on the list, it can be considered as a top-level domain with low communication frequency, and a list of low-frequency top-level domains is obtained;

[0241] A low-frequency communication top-level domain list is established based on the top-level domain list and the low-frequency communication top-level domain list.

[0242] From the overall perspective, the second generation module 30 in this embodiment is difficult for each security expert or recipient to independently identify whether a sender's email address may have problems. Therefore, risk detection and reminder for email addresses are of great importance.

[0243] Since the lower confidence limit is calculated by calculating the probability values ​​of all domain names in the preset trusted domain name list, it indicates the minimum value of the email credibility. Therefore, comparing the probability value and the lower confidence limit can reflect the degree of abnormality of the current target email address from an official perspective.

[0244] In addition, since the lower threshold limit of the state transition probability calculation uses the minimum value, it is relatively conservative. Therefore, by combining the character spelling feature detection results of the target email itself and the low-frequency communication top-level domain list, the email address structure of the target email itself and the email addresses with less frequent communication can be taken into account. In addition, the second prompt can not only be the result of analyzing the target email from an external perspective, but also include a detailed consideration of internal details, making the second prompt more objective, accurate and effective.

[0245] In one embodiment, the third generation module 40 includes a first segmentation unit, a screening unit, a reconstruction unit, an auxiliary unit, and a transformation unit;

[0246] Among them, the first segmentation unit is used to use a preset LSTM (Long Short-Term Memory) deep learning network structure to segment the URL (Uniform Resource Locator) in the target email into a first character set; among them, the input processing of the LSTM deep learning network structure refers to the input word processing method of the large language model, and the network structure defines a large number of characters that can support segmentation.

[0247] A screening unit, configured to screen out characters in the first character set that are not in the preset character set, obtain a second character set, and define character identifiers in the second character set as unknown characters;

[0248] The reconstruction unit is used to add a preset beginning character and an ending character to the beginning and the end of the unknown character respectively, so as to obtain a number list of the second character set.

[0249] The auxiliary unit is used to generate comprehensive auxiliary prompts based on the IP host, short URL and external download link in the target email, specifically:

[0250] If the IP address of the host where the target email is located is not an intranet IP, the first auxiliary prompt will be output to indicate that the IP address of the target email is not an intranet IP.

[0251] If the URL in the target email ends with a browser download link, a second auxiliary prompt will be output to indicate that the target email points to an external download link. For example, https: / / www.123.com / abc.png is an image, so most browsers will not trigger an automatic download, so it is not an external download link. However, if the URL is https: / / www.123.com / abc.exe, most browsers will interpret it as a download link and pop up a save window, thus constituting a download link.

[0252] If the host location of the target email is matched with the domain name in the pre-obtained short URL service list through regular expression, a third auxiliary prompt is output to indicate that the target email points to the short URL;

[0253] A comprehensive auxiliary prompt is generated according to the first auxiliary prompt, the second auxiliary prompt and the third auxiliary prompt.

[0254] This embodiment generates comprehensive auxiliary prompts based on the differences in different scenarios such as using IP as the host bit, using short URLs, and pointing to external download links. Since spelling-based URL checking is not suitable in these scenarios, the LSTM network cannot be used to detect these URLs. Therefore, by taking corresponding detection measures for these three types of information, a comprehensive auxiliary prompt can be obtained to provide relevant reminders.

[0255] A transformation unit is used to perform a softmax transformation on the number list using a binary classification network structure in an LSTM deep learning network structure to obtain an abnormality probability of the uniform resource locator and generate an initial prompt based on the abnormality probability;

[0256] The transformation unit is further used to combine the initial prompt and the comprehensive auxiliary prompt to obtain a third prompt.

[0257] It should be noted that in the embodiments of the present invention, the combination of the first-level domain name and the top-level domain after excluding the top-level domain is considered the host location. For multiple URLs with the same host location, only the first URL that appears is checked. For example, the host location of 123.com.cn and 456.123.com.cn are both 123.com.cn, but 123.com and 123.cn are not the same host location. In addition, the softmax transform is a mathematical transformation of the original variable, which uses the softmax function. The softmax function is also called the normalized exponential function. It is a mathematical function that is generally used to convert a set of arbitrary real numbers into real numbers representing a probability distribution.

[0258] Overall, the third generation module 40 of this embodiment, since spam phishing emails often achieve their goals by guiding users to click on corresponding URLs, prompting potentially risky URLs in target emails can achieve a more refined and accurate prompting effect than using a forwarding page to forward all external links.

[0259] This embodiment detects the Uniform Resource Locator (URL) in an email and then generates a third prompt based on it. By segmenting and filtering the URL in the target email, unknown characters can be obtained. By adding characters at the beginning and end of the unknown characters, the characters in the numbered list can be clearly organized, facilitating faster conversion speeds during the transformation process, accelerating the process of obtaining the probability of anomalies in the URL, and quickly generating the third prompt.

[0260] Compared with the sending domain name, the uniform resource locator in the email is longer, so it is not suitable to be judged by simple state transition probability and number of characters. By using a neural network model to convert abnormal probabilities, analysis time can be greatly saved and the accuracy of the third prompt can be improved.

[0261] In one embodiment, the fourth generation module 50 includes a first acquisition unit, a first generation unit, a second acquisition unit, a first matching unit, a fourth subunit, a fifth subunit, a sixth subunit, a seventh subunit, an eighth subunit, a ninth subunit, a third acquisition unit, a second segmentation unit, and a second matching unit.

[0262] The first acquisition unit is used to acquire the macro-containing attachment in the target email;

[0263] The first generating unit is configured to generate a fourth prompt based on a dangerous file in a key concern list if the dangerous file is present in the macro-bearing attachment;

[0264] The focus list is created based on the prohibited attachment types of the mailbox where the target email is located (for example, the list of attachment types prohibited from uploading in Outlook) and the preset focus files;

[0265] Among them, the files of focus include files that can be directly executed, such as ".exe", ".bat", ".ps" and ".sh", as well as macro files such as ".docm" and ".xlsm".

[0266] This embodiment detects attachments carried in emails and then generates a fourth prompt based on this. Because the aforementioned focused files are easily executed due to misoperation or misleading content in emails, for example, an Office file containing macros is generally safe even if it carries a macro virus when in read-only mode with macros disabled, but may be infected if the user enables macros. Therefore, by using the focused list to determine whether macro-containing attachments are dangerous files, it is possible to quickly determine whether macro-containing attachments in the target emails pose operational risks and provide timely reminders to users. The generation method of this fourth prompt is simple and quick.

[0267] a second acquisition unit, configured to acquire a mail header of a target mail and match the mail header with a preset dictionary tree to obtain a first matching result; wherein the dictionary tree is established based on a list of government agencies, banks, public security, procuratorial and judicial agencies, and the names of permanent establishments of enterprises;

[0268] The first matching unit is used to obtain the entity name of the target email, match the entity name with the B-ORG (company or organization) list and the I-PER (individual) list in the sequence labeling algorithm, and obtain a second matching result;

[0269] The fourth subunit is configured to determine the category of the email address from which the target email was sent, and obtain the declared object of the target email, if the first matching result or the second matching result is a successful match;

[0270] The fifth subunit is used to check whether the top-level domain name of the target email belongs to a preset special top-level domain (such as ".gov" and ".mit") if the target of the declaration is a government agency or a public security, procuratorial or judicial agency;

[0271] The sixth subunit is used to check whether the domain name used in the target email belongs to a pre-collected list if the target is a bank, invoice service provider or enterprise;

[0272] The seventh subunit is used to check whether the sender and recipient of the target email are in the same domain or subdomain if the declared object is a permanent establishment of an enterprise;

[0273] The eighth subunit is used to check whether the declared object is in the preset address book if the declared object is a person's name;

[0274] The ninth subunit is configured to generate a fifth prompt according to the inspection result.

[0275] This embodiment detects whether the subject name of the email is forged, and then generates the fifth prompt based on this. Although protocols such as SPF have a certain degree of protection against forged email addresses, they cannot prevent the use of other people's names by exploiting the visual effects of email reading. Therefore, it is important to find and determine whether the subject name is forged by finding and judging the entity.

[0276] First, through matching, we can determine whether the target email is a match, that is, whether it constitutes a statement. A statement is a formal statement issued by an official body such as a government, enterprise, or organization. It is usually issued in response to an event or situation, expressing an official position and attitude. The sending email address is generally official and formal. Therefore, after determining that it constitutes a statement, we directly check the sending email address to quickly determine whether the target email is a forgery, thereby obtaining the fifth prompt. In addition, the generation of the second matching result utilizes the open source BiLSTM+CRF technology, which can provide a corresponding list for key targets for fast matching.

[0277] Furthermore, since the top-level domains of government, public security, procuratorial, and judicial organizations often have unique top-level domains, direct checks can be performed on these domains to determine if any mailboxes are anomalous. Since permanent corporate entities often share affiliations, checking for shared domains or subdomains can be used to determine if any mailboxes are anomalous. Using pre-collected lists and address books, it's possible to quickly and directly determine if bank or individual mailboxes are anomalous. This solution provides specific methods for checking the sending email addresses of government agencies, banks, permanent corporate entities, and individuals. This highly targeted approach can accelerate the generation of the fifth prompt.

[0278] A third acquiring unit, configured to acquire text information of a target email;

[0279] The second segmentation unit is used to segment the text information using the TextCNN (convolutional neural network) module as a classifier to obtain a number of string tokens; wherein the classifier controls its physical size and occupied memory space through a preset vocabulary;

[0280] The second matching unit is used to match the plurality of character strings with a preset money vocabulary, mask the tokens that are not in the money vocabulary, and obtain a matching result;

[0281] The second matching unit is further configured to generate a sixth prompt according to the matching result;

[0282] Among them, the money vocabulary is established based on preset financial fraud information.

[0283] In this embodiment, general email classification only determines whether an email is spam and does not care much about the specific type of spam. However, among all scams, those involving money are often the most sensitive. Therefore, it is necessary to provide additional reminders for messages involving money and potentially constituting fraud.

[0284] Using a classifier to segment the text information is equivalent to performing data segmentation and extraction on the original email text, so that the vocabulary in the obtained character strings can be easily compared and matched with the money word list, thereby accelerating the acquisition process of the sixth prompt.

[0285] In one embodiment, the summarizing module 60 specifically includes:

[0286] Assign a unique ID and label to each of the first, second, third, fourth, fifth, and sixth prompts. Based on a pre-established multilingual text dictionary and corresponding placeholders for the output of each label, control the output of the first, second, third, fourth, fifth, and sixth prompts corresponding to each label by inputting language tag parameters. Combined information is then generated to obtain risk warning information.

[0287] Among them, the multilingual copywriting dictionary is written manually.

[0288] This embodiment controls the generation of prompts by formatting output to support multiple languages. Compared with the method of using machine translation, the prompts in this solution are written manually, and the language is more appropriate and natural, which can enhance the user experience.

[0289] In general, the embodiments of the present invention have the following beneficial effects:

[0290] This device uses different detection methods from six perspectives: SPF verification, spelling of the sending domain name, uniform resource locator, macro attachments, suspected fraudulent names and financial information, to obtain corresponding prompts, and then generates highly comprehensive risk warning information. It can quickly and effectively issue risk warnings at one time, avoid missing warnings, and ensure the information security of the target email;

[0291] Among them, a method based on entity recognition and knowledge base matching is used to identify potential mismatched identity claims to prompt possible impersonation emails, which covers a wider range and can reduce the risk of omissions; the entity recognition-based solution can find the organization name in a given text to discover potential entity impersonations, which is impossible to do through simple contact matching; and compared with the "one-size-fits-all" solution of prohibiting direct access to external links, the embodiment of the present invention uses a method that combines empirical knowledge and deep learning models to identify specific suspicious URLs, which can avoid the problem of users being numb and missing information reminders due to overly general prompts.

[0292] Example 3:

[0293] An embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the method for generating risk warning information based on email;

[0294] Among them, if the method for generating risk warning information based on email is implemented in the form of a software functional unit and used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0295] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for generating risk warning information based on email, characterized in that: include: Get the target email; Performing SPF verification on the target email, and generating a first prompt according to the verification result; According to the spelling of the sending domain name in the target email, a probability value is calculated using a state transition probability method, and the probability value is combined with a preset confidence lower limit and the spelling feature in the target email to generate a second prompt; Using a preset neural network model, segmenting and converting the uniform resource locator in the target email into a real number to generate a third prompt; Performing feature analysis on the macro-bearing attachments, suspected fraudulent names, and monetary information in the target email, and generating a fourth prompt, a fifth prompt, and a sixth prompt respectively; Risk warning information is generated according to the first prompt, the second prompt, the third prompt, the fourth prompt, the fifth prompt, and the sixth prompt.

2. A method for generating risk warning information based on emails as claimed in claim 1, characterized in that: According to the spelling of the sending domain name in the target email, a probability value is calculated using a state transition probability method, and the probability value is combined with a preset confidence lower limit and the spelling features in the target email to generate a second prompt, specifically: Using a preset N value, traversing the sending domain name in the target email through the N-Gram algorithm to generate a character substring; In combination with a preset character set, all possible N-Gram tuples are constructed according to the character substrings to obtain a state transition probability matrix; Splitting the sending domain name in the state transition probability matrix to obtain a plurality of tuples; Calculating average state transition probability values ​​of the plurality of tuples to obtain the probability value; The probability value is compared with the confidence lower limit, and the comparison result is combined with the spelling feature in the target email to generate a second prompt.

3. A method for generating risk warning information based on emails as claimed in claim 2, characterized in that: The probability value is compared with a preset confidence lower limit, and the comparison result is combined with the spelling feature in the target email to generate a second prompt, which is specifically: Comparing the probability value with a preset confidence lower limit to obtain a comparison result; wherein the confidence lower limit is obtained by calculating the probability values ​​of all domain names in a preset trusted domain name list; Performing a standardization test on the character spelling features of vowels and special symbols in the target email to obtain a test result; The second prompt is generated according to the comparison result, the detection result and the low-frequency communication top-level domain list of the mailbox where the target email is located.

4. A method for generating risk warning information based on emails as claimed in claim 3, characterized in that: The confidence lower limit is obtained by calculating the probability values ​​of all domain names in the preset trusted domain name list, specifically: Obtain a world comprehensive ranking list of websites, and use part of the list information in the world comprehensive ranking list to build a trusted domain name list; Calculate the probability values ​​of all domain names in the trusted domain name list to obtain a probability value set; The minimum value in the probability value set is used as the confidence lower limit.

5. The method for generating risk warning information based on email according to claim 1, characterized in that: Using a preset neural network model, by segmenting and converting the uniform resource locator in the target email into a real number, a third prompt is generated, specifically: Using a preset neural network model, segmenting the uniform resource locator in the target email into a first character set; Filter out characters in the first character set that are not in the preset character set to obtain a second character set, and define character identifiers in the second character set as unknown characters; Adding a preset beginning character and an ending character to the beginning and the end of the unknown character respectively to obtain a number list of the second character set; The binary classification network structure in the neural network model is used to transform the number list to obtain the abnormal probability of the uniform resource locator, and the third prompt is generated according to the abnormal probability.

6. A method for generating risk warning information based on emails as claimed in claim 1, characterized in that: Perform feature analysis on the macro-bearing attachment in the target email to generate a fourth prompt, specifically: Obtaining the macro-containing attachment in the target email; If a dangerous file in the focus list appears in the macro-attached attachment, generating the fourth prompt according to the dangerous file; The focus list is established based on the prohibited attachment types for uploading to the mailbox where the target email is located and preset focus files.

7. The method for generating risk warning information based on email according to claim 1, characterized in that: Perform feature analysis on the suspected fraudulent name in the target email and generate a fifth prompt, specifically: Obtaining the email header of the target email, matching the email header with a preset dictionary tree to obtain a first matching result; wherein the dictionary tree is established based on a list of government agencies, banks, public security, procuratorial and judicial agencies, and the name of a permanent establishment of an enterprise; Obtaining the entity name of the target email, and matching the entity name with the company list and the individual list in the sequence labeling algorithm to obtain a second matching result; If the first matching result or the second matching result is a successful match, the email address from which the target email is sent is checked, and the fifth prompt is generated according to the checking result.

8. A method for generating risk warning information based on emails as claimed in claim 7, characterized in that: The email address from which the target email is sent is checked, and the fifth prompt is generated according to the check result, which is specifically: Determine the category of the email address from which the target email is sent, and obtain the declared object of the target email; If the object of the declaration is a government agency or a public security, procuratorial or judicial agency, then check whether the top-level domain name of the target email belongs to a preset special top-level domain; If the declared object is a bank, invoice service provider or enterprise, check whether the domain name used in the target email belongs to the pre-collected list; If the declared object is a permanent establishment of an enterprise, check whether the sender and the recipient of the target email are in the same domain or subdomain; If the declared object is a person's name, checking whether the declared object is in the preset address book; The fifth prompt is generated according to the inspection result.

9. The method for generating risk warning information based on email according to claim 1, characterized in that: Perform feature analysis on the money information in the target email to generate a sixth prompt, specifically: Obtaining text information of the target email; Using a preset classifier to segment the text information to obtain a number of character strings; Matching the plurality of character strings with a preset money word list, and generating the sixth prompt according to the matching result; Wherein, the money vocabulary is established based on preset financial fraud information.

10. A device for generating risk warning information based on email, characterized in that: It includes an email acquisition module, a first generation module, a second generation module, a third generation module, a fourth generation module and a summary module; Wherein, the email acquisition module is used to acquire the target email; The first generating module is used to perform SPF verification on the target email and generate a first prompt according to the verification result; The second generating module is used to calculate a probability value according to the spelling of the sending domain name in the target email using a state transition probability method, and generate a second prompt by combining the probability value with a preset confidence lower limit and the spelling feature in the target email; The third generating module is used to generate a third prompt by segmenting and converting the uniform resource locator in the target email into a real number using a preset neural network model; The fourth generating module is used to perform feature analysis on the macro-bearing attachments, suspected fraudulent names and money information in the target email, and generate a fourth prompt, a fifth prompt and a sixth prompt respectively; The summary module is used to generate risk warning information according to the first prompt word, the second prompt word, the third prompt word, the fourth prompt word, the fifth prompt word and the sixth prompt word.

11. The device for generating risk warning information based on email according to claim 10, characterized in that: The second generation module includes a character unit, a matrix unit, a splitting unit, a probability value unit and a comparison unit; The character unit is used to use a preset N value to traverse the sender domain name in the target email through an N-Gram algorithm to generate a character substring; The matrix unit is used to construct all possible N-Gram tuples according to the character substring in combination with a preset character set to obtain a state transition probability matrix; The splitting unit is used to split the sending domain name in the state transition probability matrix to obtain a plurality of tuples; The probability value unit is used to calculate the average state transition probability value of the plurality of tuples to obtain the probability value; The comparison unit is used to compare the probability value with the confidence lower limit, and generate a second prompt by combining the comparison result with the spelling feature in the target email.

12. The device for generating risk warning information based on email according to claim 11, characterized in that: The comparison unit includes a first subunit, a second subunit and a third subunit; The first subunit is used to compare the probability value with a preset confidence lower limit to obtain a comparison result; wherein the confidence lower limit is obtained by calculating the probability values ​​of all domain names in a preset trusted domain name list; The second subunit is used to perform a standardization detection on the character spelling features of the vowels and special symbols in the target email to obtain a detection result; The third subunit is used to generate the second prompt according to the comparison result, the detection result and the low-frequency communication top-level domain list of the mailbox where the target email is located.

13. The device for generating risk warning information based on email according to claim 12, characterized in that: The confidence lower limit is obtained by calculating the probability values ​​of all domain names in the preset trusted domain name list, specifically: Obtain a world comprehensive ranking list of websites, and use part of the list information in the world comprehensive ranking list to build a trusted domain name list; Calculate the probability values ​​of all domain names in the trusted domain name list to obtain a probability value set; The minimum value in the probability value set is used as the confidence lower limit.

14. The device for generating risk warning information based on email according to claim 10, characterized in that: The third generation module includes a first segmentation unit, a screening unit, a reconstruction unit and a transformation unit; The first segmentation unit is used to segment the uniform resource locator in the target email into a first character set using a preset neural network model; The screening unit is used to screen out characters in the first character set that are not in the preset character set to obtain a second character set, and define character identifiers of the second character set as unknown characters; The reconstruction unit is used to add a preset beginning character and an ending character to the beginning and the end of the unknown character respectively, to obtain a number list of the second character set; The transformation unit is used to use the binary classification network structure in the neural network model to transform the number list to obtain the abnormal probability of the uniform resource locator, and generate the third prompt according to the abnormal probability.

15. The device for generating risk warning information based on email according to claim 10, characterized in that: The fourth generating module includes a first acquiring unit and a first generating unit; Wherein, the first acquisition unit is used to acquire the macro-bearing attachment in the target email; The first generating unit is configured to generate the fourth prompt according to the dangerous file in the focus list if the dangerous file in the macro attachment appears; The focus list is established based on the prohibited attachment types for uploading to the mailbox where the target email is located and preset focus files.

16. The device for generating risk warning information based on email according to claim 10, characterized in that: The fourth generating module includes a second acquiring unit, a first matching unit and a checking unit; The second acquisition unit is used to acquire the email header of the target email, and match the email header with a preset dictionary tree to obtain a first matching result; wherein the dictionary tree is established based on a list of government agencies, banks, public security, procuratorial and judicial agencies, and the name of a permanent establishment of an enterprise; The first matching unit is used to obtain the entity name of the target email, match the entity name with the company list and the individual list in the sequence labeling algorithm, and obtain a second matching result; The checking unit is used to check the sending email address of the target email if the first matching result or the second matching result is a successful match, and generate the fifth prompt according to the checking result.

17. The device for generating risk warning information based on email according to claim 16, characterized in that: The inspection unit includes a fourth subunit, a fifth subunit, a sixth subunit, a seventh subunit, an eighth subunit and a ninth subunit; The fourth subunit is used to determine the category of the email address from which the target email is sent, and obtain the declared object of the target email; The fifth subunit is used to check whether the top-level domain name of the target email belongs to a preset special top-level domain if the declared object is a government agency or a public security, procuratorial or judicial agency; The sixth subunit is used to check whether the domain name used by the target email belongs to a pre-collected list if the declared object is a bank, an invoice service provider or an enterprise; The seventh subunit is used to check whether the sender and the recipient of the target email are in the same domain or subdomain if the declared object is a permanent establishment of an enterprise; The eighth subunit is used for checking whether the declared object is in a preset address book if the declared object is a person's name; The ninth subunit is used to generate the fifth prompt according to the inspection result.

18. The device for generating risk warning information based on email according to claim 10, characterized in that: The fourth generation module includes a third acquisition unit, a second segmentation unit and a second matching unit; Wherein, the third acquisition unit is used to acquire text information of the target email; The second segmentation unit is used to segment the text information using a preset classifier to obtain a plurality of character strings; The second matching unit is used to match the plurality of character strings with a preset money word list, and generate the sixth prompt according to the matching result; Wherein, the money vocabulary is established based on preset financial fraud information.

19. A storage medium, characterized in that: The storage medium stores a computer program, which is called and executed by a computer to implement any one of the email-based risk warning information generation methods according to claims 1 to 9.

Citation Information

Patent Citations

  • Signal sending behavior identification method and device, terminal equipment and storage medium

    CN116170226A

  • Scoring-based DMARC detection method, system and equipment and storage medium

    CN116866057A

  • Risk prompt information generation method and device based on mail and medium

    CN117749496A

  • Automatic Phishing Email Detection Based on Natural Language Processing Techniques

    US20160344770A1