Junk mail detection method and system and terminal equipment

By uploading the original email or image, combined with text recognition, anti-spam detection and language analysis models, the problems of insufficient flexibility and high misjudgment rate in the existing technology are solved, and more accurate and detailed detection results and better user experience are achieved.

CN120223370APending Publication Date: 2025-06-27GUANGDONG COREMAIL COMPUTER TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510290215.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art lacks flexibility in spam detection, which makes it easy to misjudgment that normal emails are spam, and the IP-based detection mechanism is unreasonable in the case of dynamic IP and shared servers.

Method used

By uploading the original email or image, text recognition and extraction technology is used, combined with anti-spam detection, address detection and anti-virus detection, email analysis logs are generated, and interpreted through language analysis models to provide risk judgment results and processing suggestions.

Benefits of technology

It realizes more accurate and detailed spam detection, reduces the misjudgment rate, and improves user experience and email security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223370A_ABST
    Figure CN120223370A_ABST
Patent Text Reader

Abstract

The invention discloses a junk mail detection method and system and terminal equipment, and the method comprises the steps: carrying out the text recognition and extraction of uploaded mail data according to a data type, and obtaining to-be-detected information; performing anti-spam detection, address detection and anti-virus detection on different contents of the to-be-detected information to obtain a mail analysis log; and inputting the cut mail analysis log into a preset language analysis model to interpret a mail detection result, and generating a risk judgment result of the mail analysis log and a corresponding mail processing suggestion. According to the method, the junk mail can be identified by uploading the original mail or the image, the mail detection result and related processing suggestions which are more convenient for the user to understand are generated, and the use experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of email security detection, and specifically relates to a method, system, and terminal device for detecting spam emails. Background Art

[0002] As email has become an important tool for daily communication between individuals and enterprises, spam emails not only occupy network resources but also pose threats to data security and user privacy. Spam emails often hide advertisements, phishing links, or malware. Attackers use these means to steal sensitive information, spread viruses, or even launch ransomware attacks, causing huge economic and reputational losses to enterprises and individuals. For users with email security detection needs, enterprise users rely on anti-spam service products purchased by the company, and individual users rely on the anti-spam services of cloud email platforms. Whether it is enterprise users or individual users, they cannot quickly and conveniently use third-party email detection services to conduct email security detection.

[0003] Currently, most in the industry identify spam emails by comparing the email HTML tags with pre-stored spam email HTML codes, and determine whether an email is spam by analyzing whether the email IP address has malicious attack behavior within a preset time. However, these methods are not flexible enough. When some enterprises use HTML styles to beautify promotional marketing emails, the detection method based on email HTML tags is prone to misjudgment. Moreover, when the sender uses dynamic IPs and shared servers, the same IP may send both normal emails and spam emails at the same time. For the detection and interception mechanism based on IP, whether it is judged as a normal IP or an abnormal IP, it is unreasonable. At this time, it is easy to misjudge normal emails as spam emails. Summary of the Invention

[0004] This application proposes a method, system, and terminal device for detecting spam emails. By uploading the original email or image, spam emails can be identified, and more user-friendly email detection results and relevant handling opinions can be generated, improving the user experience.

[0005] The first aspect of this application provides a method for detecting spam emails, and the method includes:

[0006] According to the data type, perform text recognition and extraction on the uploaded email data to obtain the information to be detected;

[0007] Perform anti-spam detection, address detection, and anti-virus detection on different contents of the information to be detected respectively to obtain an email analysis log;

[0008] Input the cropped email analysis log into a preset language analysis model to interpret the email detection results, and generate a risk determination result of the email analysis log and corresponding email handling suggestions.

[0009] The above solution first extracts the necessary text content according to the type of the uploaded email data, which is convenient for subsequent detection and processing of spam information. Moreover, adopting different data extraction methods can extract more precisely according to the text type, and quickly and efficiently obtain the key text information. Then, for different types of data in the information to be detected, such as the email body, email header, email address, and attachments, different types of detection methods are adopted to conduct a more complete and detailed detection of the email to be detected, and an email analysis log is obtained. Then, the email analysis log is trimmed to remove redundant fields, and the key information retaining the detection results is input into the language analysis model for interpretation. Through the context semantics, a risk determination result that enables users to understand why the email to be detected is determined as a normal / spam email and the corresponding email processing suggestions are obtained. Taking the email processing suggestions as a reference, users can correctly process spam emails, minimize the risks brought by spam emails, and improve the user experience.

[0010] In a possible implementation method of the first aspect, according to the data type, text recognition and extraction are performed on the uploaded email data to obtain the information to be detected, specifically:

[0011] According to the data type of the email data, corresponding data processing methods are respectively used to perform text recognition and extraction on the email header, email body, and attachments of the email data to obtain the information to be detected;

[0012] Wherein, the data type includes email originals, email images, and email header images.

[0013] The above solution extracts and identifies the data at different positions of the email data by using the corresponding method through the data type of the email data, improving the accuracy of data extraction.

[0014] In a possible implementation method of the first aspect, it further includes:

[0015] For the email data belonging to the email image, OCR text recognition is performed on the preprocessed email data to extract the initial language text;

[0016] The initial language text is corrected for error characters through a spelling correction algorithm to obtain the information to be detected of the email data.

[0017] In a possible implementation method of the first aspect, anti-spam detection, address detection, and anti-virus detection are respectively performed on different contents of the information to be detected to obtain an email analysis log, specifically:

[0018] Input the information to be detected into a preset spam detection model. Detect the email body and email header of the information to be detected through an anti-spam detection module, detect the email address of the information to be detected through a URL sandbox detection module, detect the attachments of the information to be detected through an anti-virus detection module, and finally output an email analysis log.

[0019] The above solution detects data at different positions of the information to be detected through a spam detection model. Specifically, it uses an anti-spam detection module to detect whether the email body and email header contain spam information, a URL sandbox detection module to detect the security of the email address, and an anti-virus detection module to detect whether the email attachment carries a virus. Through the above detections, it comprehensively determines whether the email is spam.

[0020] In a possible implementation method of the first aspect, when detecting the email body and email header of the information to be detected through an anti-spam detection module, specifically:

[0021] Extract the email body and email header from the information to be detected to obtain the first detection data;

[0022] Classify the first detection data through a preset label, and then perform data cleaning on the classified first detection data;

[0023] Use a word segmentation tool to perform text segmentation on the first detection data after data cleaning to obtain several words and phrases;

[0024] Perform semantic analysis based on context association on the words and phrases to obtain the email analysis log related to the email body and email header.

[0025] In a possible implementation method of the first aspect, when detecting the email address of the information to be detected through a URL sandbox detection module, specifically:

[0026] Extract the email address from the information to be detected, and perform domain name detection on the email address according to a preset domain name blacklist;

[0027] Perform structure detection on the email address after domain name detection; wherein, the structure detection includes detecting whether the length of the email address exceeds a first threshold, and detecting whether the email address contains garbled characters or preset special characters;

[0028] Detect whether there is an abnormal jump in the email address after structure detection through redirect link analysis.

[0029] In a possible implementation method of the first aspect, the cropped mail analysis log is input into a preset language analysis model for interpreting the mail detection result, generating a risk determination result of the mail analysis log and a corresponding mail processing suggestion, specifically:

[0030] According to a preset keyword field template, the mail analysis log is cropped to retain the mail body content, mail detection information, and mail behavior characteristics; wherein, the mail behavior characteristics include mail communication information and mail text length;

[0031] The cropped mail analysis log is input into the language analysis model for field analysis, and the mail analysis log is interpreted based on a preset standardized answer template, generating a risk determination result of the mail analysis log and a corresponding mail processing suggestion.

[0032] In the above solution, since ordinary users may not be able to understand the content of the output mail analysis log, the language analysis model is used to interpret the mail detection result, obtaining a more popular risk determination result for users to understand the reason why the mail is determined as an abnormal mail, and providing mail processing suggestions to assist users in disposing of spam, minimizing the risk brought by spam.

[0033] In a possible implementation method of the first aspect, the uploaded mail data is specifically:

[0034] Take a photo or screenshot of the mail to be detected to obtain mail data belonging to the mail image or the mail header image;

[0035] Preprocess the mail data, and then upload the preprocessed mail data;

[0036] Wherein, the preprocessing includes image denoising, image binarization, and image skew correction.

[0037] In the above solution, since the mail image may have problems such as unclear fonts, uneven illumination, or skewed content due to shooting light and angle problems during the shooting process, it is necessary to preprocess the image type data to improve the quality of the image.

[0038] The second aspect of the present application provides a spam detection system, which includes: a data extraction module, an information detection module, and a detection result interpretation module;

[0039] Among them, the data extraction module is used to perform text recognition and extraction on the uploaded mail data according to the data type to obtain the information to be detected;

[0040] The information detection module is used to perform anti-spam detection, address detection, and anti-virus detection on different contents of the information to be detected respectively to obtain a mail analysis log;

[0041] The detection result interpretation module is used to input the cropped email analysis log into a preset language analysis model for interpreting the email detection result, and generate a risk determination result of the email analysis log and corresponding email processing suggestions.

[0042] The third aspect of this application provides a terminal device, which includes: a terminal device, including a processor and a memory. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of a method for detecting spam according to any one of the embodiments of this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of this application, the drawings required for implementation will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0044] Figure 1 is a specific flowchart of a method for detecting spam provided by an embodiment of this application;

[0045] Figure 2 is a prompt diagram of the camera interface of a method for detecting spam provided by an embodiment of this application;

[0046] Figure 3 is a specific result diagram of a spam detection system provided by an embodiment of this application;

[0047] Figure 4 is a structural diagram of a terminal device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of this application belong to the scope of protection of this application.

[0049] It should be understood that the step numbers used in the text are only for convenient description and are not intended to limit the execution order of the steps.

[0050] The First Embodiment

[0051] With the wide application of email, spam has become an increasingly serious cyber security threat. Phishing emails, virus emails, etc. all fall into the category of spam. A phishing email refers to a fake email sent by impersonating a legitimate entity, usually aiming to deceive trusted recipients into leaking their personal sensitive information, clicking on malicious links or downloading malicious attachments. A virus email will spread the virus in a self-replicating manner when the user opens the email, causing damage to the computer system or data and affecting the normal operation of the computer. However, among the users with the need for email security detection currently, enterprise users rely on the anti-spam service products purchased by the company, and individual users rely on the anti-spam service of the cloud email platform. Whether it is enterprise users or individual users, they cannot quickly and conveniently use third-party email detection services to conduct email security detection. Therefore, how to quickly obtain a concise and clear email detection result and disposal opinions on how to handle spam by only uploading the email image through a third-party program, and improve the user's email security awareness education and user experience, is the main research direction of the embodiments of this application.

[0052] As Figure 1 shown, Figure 1 FIG. shows a specific flowchart of a method for detecting spam provided by an embodiment of this application. The method for detecting spam in this embodiment includes steps S1 to S3, which are described in detail as follows:

[0053] Step S1, according to the data type, perform text recognition and extraction on the uploaded email data to obtain the information to be detected.

[0054] The components of an email include the email header, the email body, and the attachment. The email header mainly includes the transmission information of the email, such as the sender, recipient, date, and subject, etc.; the email body is the main content part of the email, containing actual text, pictures, formatting information, or other data. Since all three of these parts may carry virus or phishing information, etc., the embodiments of this application respectively perform spam detection on the email header, the email body, and the attachment to obtain a complete detection result.

[0055] In the embodiments of this application, there are a total of three selectable ways to upload email data: uploading the original email, uploading the email image, and uploading the email header image. Because in the prior art, it generally relies on the anti-spam detection engine provided by an external platform to perform security detection on the received emails, while in the embodiments of this application, a flexible anti-spam detection service is provided. Users do not need to rely on email providers to conduct email anti-spam detection work, and can flexibly select the way to upload email information according to the usage habits of different platforms. The anti-spam detection service can be directly invoked in the form of a web page or a small program, enabling users to bypass processes such as enterprise procurement and manufacturer application docking to obtain email result interpretation and measure suggestions.

[0056] Exemplarily, the user can upload the email information to be detected through methods such as web pages, WeChat mini-programs, and Alipay mini-programs. For mobile devices such as mobile phones, the camera is used by default to take pictures of the email, obtaining the corresponding email image and email header image (taking a picture of the email header of the email to be detected to obtain the email header image). For PC devices, local pictures are obtained by default through screenshotting and other methods, and the local pictures are uploaded. In addition, the email to be detected can be saved and the email of the email to be detected can be uploaded.

[0057] Figure 2 A prompt diagram for the camera interface is provided. When taking a picture, the user places the content to be photographed within the viewfinder, and then preliminarily checks whether the taken picture is qualified. Generally, the projection method is used to determine whether the tilt angle of the picture is greater than 10°, and edge detection is used to determine whether the background noise is within a reasonable range, etc. If the conditions are not met, the user is prompted to take a picture / upload again; if the conditions are met, the uploaded image is preprocessed to reduce the cost of subsequent OCR text recognition.

[0058] Then, for the uploaded email information of different data types, different data extraction schemes are respectively adopted to extract the necessary information in the email information, obtaining the information to be detected.

[0059] The email content of the original email is complete, including all information such as the email header, email body, and attachments. Therefore, first, the original email is disassembled into three parts for these three types of email information, and the necessary information is identified and extracted. Then, the extracted information is converted into structured data information for subsequent anti-spam detection processing. Among them, since some spam emails store information in attachments, for the disassembly of attachments, it is necessary to use the correct disassembly method to extract the attachments according to the type format of the attachments (such as file extension, MIME type verification, file header signature verification, etc.). For document types, text information needs to be extracted, and for picture information, text needs to be extracted through OCR technology.

[0060] If the user uploads an email image, the image needs to be preprocessed first, specifically including image denoising, image binarization, background noise removal, and image skew correction.

[0061] Image denoising is to remove random noise in the email image by using filtering algorithms such as Gaussian filtering and median filtering.

[0062] Image binarization is to convert the email image into black and white to simplify the recognition process.

[0063] Removing background noise is because the captured photos may contain interference or cluttered elements in the background. Therefore, image segmentation or background removal methods are used to clean up this background noise, making the text content of the email image more prominent.

[0064] Image skew correction is because a poor shooting angle may cause the email image to be skewed or distorted. Therefore, methods such as perspective transformation and affine transformation are needed to correct the text in the email image to be horizontal or upright, ensuring that the OCR algorithm can accurately read the characters.

[0065] If the quality of the email image is poor, directly using the OCR algorithm for text recognition will incur a high processing cost. And because the picture quality is too low, performing OCR text recognition after high-level processing will cause recognition deviations. Therefore, it is necessary to preprocess the uploaded email image to improve the text quality and then use the OCR engine for text recognition and extraction.

[0066] Among them, OCR stands for Optical Character Recognition, also known as optical character recognition, which is a technology for extracting text information from images. Because in the embodiments of this application, the emails to be detected may have differences in multilingual texts and special character sets, it is necessary to train and adjust the OCR engine in advance using common languages and character sets to improve the accuracy of text recognition and extraction.

[0067] The text extraction results of the OCR engine are not always 100% accurate. Usually, character substitution, misspelled words, missing characters, etc. will occur. To correct these errors, for the already recognized text, spelling correction algorithms (such as dictionary-based correction, edit distance algorithms, language models, etc.) are used to detect and correct spelling mistakes. This part also needs to consider that the body of spam emails may use similar-style characters to confuse the detection engine. Therefore, the correction algorithm should not be excessive. When necessary, calculate the character mixing ratio as a subsequent anti-spam detection feature.

[0068] The email header image is obtained by taking a photo or screenshot of the email header of the email to be detected. The email header image is generally standard text. Therefore, as long as the uploaded email header image is corrected for text and character errors, the corresponding text information can be obtained.

[0069] By performing text recognition and extraction on the email data such as the uploaded original email, email image, and email header image as described above, the information to be detected that can be applied to anti-spam detection is obtained. Using different data extraction methods for extraction can more precisely extract according to the text type and quickly and efficiently obtain the key text information.

[0070] Step S2: Perform anti-spam detection, address detection, and anti-virus detection on different contents of the information to be detected to obtain an email analysis log.

[0071] In the embodiment of the present application, the detection of spam mainly includes three parts: anti-spam detection, address detection and anti-virus detection. These three detections are the core elements of anti-spam detection for emails. The optimized and trained spam detection model is used to perform anti-spam detection on the previously obtained information to be detected (including email headers and email body data), perform sandbox detection on the extracted email address, i.e., URL, and perform anti-virus detection on email attachments. Among them, the spam detection model includes an anti-spam detection module, a URL sandbox detection module and an anti-virus detection module.

[0072] Optionally, in the embodiment of the present application, the anti-spam detection module is fine-tuned based on BERT used for data processing, and optimized by cross-entropy loss function, and the optimizer selects Adam W. The test data is divided into multiple training data sets and verification data sets. During the module training process, the module performance is evaluated on the verification set for indicators such as accuracy, recall rate and F1 score after each epoch to check the generalization ability of the module, and finally an anti-spam detection module that meets the production application conditions is obtained.

[0073] After the information to be detected is input into the anti-spam detection module, the emails corresponding to the information to be detected are classified and predicted based on the machine learning results. The emails are labeled as normal emails or spam emails based on the extracted email body and email header information. If you want to further improve the detection accuracy, you can also subdivide the spam emails. Then the labeled information to be detected is cleaned. For example, for the email body in HTML format, HTML tags need to be removed and special characters need to be deleted. The cleaned data is divided into multiple words and phrases through a word segmentation tool. In order to ensure the contextual relevance of the word vector, a dynamic word vector method based on deep learning (BERT) is selected, and a large amount of unsupervised text data pre-training is used to process the words and phrases, capture more complex grammatical and semantic information, and obtain the email analysis log related to the email body and email header.

[0074] Optionally, you can use NLTK, spacy, etc. as word segmentation tools. NLTK stands for Natural Language Toolkit, and is a Python library that is good at processing text data. spacy is also a Python natural language processing library that can automatically complete various basic NLP tasks such as word segmentation, part-of-speech tagging, dependency parsing, named entity recognition, sentiment analysis, syntactic analysis, etc.

[0075] Input the information to be detected related to the email address into the URL sandbox detection module. This module detects the security of the email URL through a simulated and isolated environment, and can identify potential malicious behaviors such as phishing attacks, malware downloads, cross-site scripting attacks, etc. It is developed based on the open-source URL sandbox and analyzes various aspects of information such as the structure, domain name, page content, redirection behavior, file downloads, and network traffic of the URL to determine whether the URL is secure.

[0076] First, analyze the components of the email address, including the protocol (http, https), domain name, path, query parameters, etc. By analyzing this metadata, the URL sandbox detection module can initially identify some common malicious behavior patterns. The detection process specifically includes:

[0077] 1. Check the domain name: Verify whether the domain name matches known malicious domain names or domain names listed in the blacklist. Malicious links often use domain names that imitate legitimate websites.

[0078] 2. View the URL structure: Check for URLs that are too long, contain garbled characters or special characters. These features are common in phishing or malware propagation and are recorded as one of the suspicious features.

[0079] 3. Analyze link redirection: Some attack links will jump to the final malicious website through multiple redirections. The sandbox can analyze the redirection link and check for abnormal jumps. Record it as one of the suspicious features.

[0080] Secondly, for the URL landing page, parse and check its HTML tags and attributes. Conventional phishing URLs often prompt users to output and submit account and password information. Therefore, the embodiments of the present application focus on monitoring the following HTML tags:

[0081] (1) Username input box: Usually an input tag with type = "text" or type = "email";

[0082] (2) Password input box: Usually an input tag with type = "password";

[0083] (3) Confirm button: Usually an input tag with type = "submit" or <button>Label, type = "submit".

[0084] If the label feature exists in the email address and the anti-spam detection result of the corresponding email is determined to be phishing, then the URL is a phishing link.

[0085] If the email address does not have the above two major features, then it is necessary to analyze the traffic during URL access to detect whether there is abnormal external communication, such as connection to a malicious control server address, downloading an executable malicious attachment, etc.

[0086] Through the above process, the sandbox detection result of the email address can be obtained.

[0087] Input the information to be detected related to the email attachment into the anti-virus detection module to perform anti-virus detection on the attachment. Generally, if the user uploads the original email, it is necessary to perform anti-virus detection on its attachment.

[0088] The anti-virus detection module is docked with different anti-virus engine SDKs, which can improve the detection rate, that is, adopt a multi-engine mode to detect whether the email contains viruses. Among them, there are mainly two logical processing methods in the multi-engine mode: one is for the purpose of improving the detection rate. For example, as long as one engine detects a virus, it is determined that the email attachment contains a virus; the other is for the purpose of improving the accuracy rate. For example, it is necessary for two or more engines to detect a virus before it is determined that the email attachment contains a virus.

[0089] According to the above description, the spam detection model is used to detect the email body, email header, email address, and attachment respectively, and an email analysis log containing the detection results is obtained.

[0090] Step S3, input the cropped email analysis log into a preset language analysis model to interpret the email detection results, and generate the risk determination result of the email analysis log and the corresponding email processing suggestions.

[0091] In the embodiment of the present application, considering that users may not have a high level of security awareness of spam and do not have a comprehensive understanding of the security of emails. Therefore, when returning the email analysis log, a risk determination result that is easy to understand and the corresponding email processing suggestions will also be generated through the email result interpretation function to help users understand the reason why the email is determined to be an abnormal email and provide a reference for disposal suggestions.

[0092] There are a large number of fields defined in the email analysis log. If directly given to the large language model for interpretation, it will consume more tokens; the value mapping relationship of some fields cannot be understood by the untrained large language model either. Therefore, the following two steps are also required:

[0093] According to the preset keyword field template, the email analysis log is trimmed, and redundant intermediate fields are deleted, retaining the log data that needs to be analyzed and interpreted. For example, the email size, email hash information, complete email header information, etc. need to be retained.

[0094] To ensure that the subsequent language analysis model can understand the information of the interpreted email, generally the following fields need to be retained:

[0095] a. The main content of the email, enabling the model to understand the email content, including the sender, recipient, subject of the email, email body, attachment names, from information, and URL links, etc.

[0096] b. Email detection information, enabling the model to understand the reasons for email detection anomalies, including the detection results (labels and scores, etc.) output by the anti-spam detection module, the detection results output by the URL sandbox detection module, and the detection results of the anti-virus detection module.

[0097] c. Email behavior characteristics, enabling the model to conduct rational supplementary analysis in combination with the above points a and b, including the sending IP, sending location, CC address, reply address, xmailer information, recipient, and text length, etc.

[0098] Then, the trimmed email analysis log is input into the trained language analysis model for field analysis, and a preset standardized answer template is provided to the model for output reference, obtaining the risk determination result of the email analysis log and the corresponding email handling suggestions. Regarding the email handling suggestions, the embodiment of the present application also adds a prompt to guide the answer for the disposal plan of individual users, which enables users to well understand the disposal methods of various emails and minimizes the risks brought by spam emails. For example, executable attachments in emails should not be easily downloaded, and other communication methods need to be used to confirm with the email sender.

[0099] Through the text generation ability of the language analysis model, information interpretation of the email analysis log is carried out, making it possible for ordinary users to understand the reasons why an email is detected as normal / spam, and improving the user experience. For users, understanding the reasons for email determination can enhance personal email security awareness education and reduce the risk of being attacked.

[0100] Implementing the embodiment of the present application has the following beneficial effects:

[0101] First, according to the type of uploaded email data, the necessary text content is extracted to facilitate the subsequent detection and processing of spam information; and different data extraction methods can be used to extract text types more accurately, and key text information can be obtained quickly and efficiently. Then, different types of detection methods are used for different types of data in the information to be detected, such as email body, email header, email address, and attachments, to perform more complete and detailed detection on the emails to be detected, and obtain email analysis logs. Then, the email analysis logs are trimmed, redundant fields are removed, and the key information of the detection results is retained. Input into the language analysis model for interpretation, and through context semantics, the risk determination results and corresponding email processing suggestions that enable users to understand why the email to be detected is determined to be normal / spam are obtained. Using the email processing suggestions as a reference, users can correctly handle spam, minimize the risks brought by spam, and improve the user experience.

[0102] Second embodiment

[0103] Furthermore, in order to implement the spam detection system corresponding to the above method embodiment to achieve corresponding functions and technical effects, Figure 3 A structural diagram of a spam detection system is provided. For ease of description, only the parts related to this embodiment are shown. The spam detection system provided by the embodiment of the present application includes:

[0104] The data extraction module 201 is used to perform text recognition and extraction on the uploaded email data according to the data type to obtain the information to be detected.

[0105] In an embodiment of the present application, a photo or screenshot of the mail to be inspected is taken to obtain mail data belonging to the mail image or the mail header image; the mail data is preprocessed, and then the preprocessed mail data is uploaded; wherein the preprocessing includes image denoising, image binarization and image tilt correction.

[0106] Then, according to the data type of the email data, corresponding data processing methods are used to perform text recognition and extraction on the email header, email body and attachments of the email data to obtain information to be detected;

[0107] The data types include mail originals, mail images and mail header images.

[0108] The information detection module 202 is used to perform anti-spam detection, address detection and anti-virus detection on different contents of the information to be detected, and obtain an email analysis log.

[0109] In the embodiments of the present application, the information to be detected is input into a preset spam detection model. The anti-spam detection module detects the email body and email header of the information to be detected, the URL sandbox detection module detects the email address of the information to be detected, and the anti-virus detection module detects the attachments of the information to be detected. Finally, an email analysis log is output.

[0110] The detection result interpretation module 203 is configured to input the cropped email analysis log into a preset language analysis model for interpreting the email detection result, and generate a risk determination result of the email analysis log and corresponding email processing suggestions.

[0111] In the embodiments of the present application, considering that users may not have a high level of security awareness regarding spam and do not have a comprehensive understanding of the security of emails. Therefore, when returning the email analysis log, a risk determination result that is easy to understand and corresponding email processing suggestions will also be generated through the email result interpretation function, helping users understand the reasons why the email is determined to be an abnormal email and providing reference for disposal suggestions.

[0112] There are a large number of fields defined for development in the email analysis log. If directly given to a large language model for interpretation, it will consume a large number of tokens; the value mapping relationships of some fields cannot be understood by an untrained large language model either. Therefore, the following two steps need to be carried out:

[0113] According to a preset keyword field template, the email analysis log is cropped, redundant intermediate fields are deleted, and the log data that needs to be analyzed and interpreted is retained. For example, the email size, email hash information, complete email header information, etc. need to be retained.

[0114] To ensure that the subsequent language analysis model can understand the information of the interpreted email, generally the following fields need to be retained:

[0115] d. The main content of the email, enabling the model to understand the email content, including the sender, recipient, subject of the email, email body, attachment name, from information, and URL link, etc.

[0116] e. Email detection information, enabling the model to understand the reasons for email detection anomalies, including the detection results (tags and scores, etc.) output by the anti-spam detection module, the detection results output by the URL sandbox detection module, and the detection results of the anti-virus detection module.

[0117] f. Email behavior characteristics, enabling the model to conduct a reasonable supplementary analysis in combination with the above points a and b, including the sending IP, sending location, CC address, reply address, xmailer information, recipient, and text length, etc.

[0118] Then, the cropped email analysis log is input into the trained language analysis model for field analysis, and a preset standardized answer template is provided to the model for output reference, obtaining the risk determination result of the email analysis log and the corresponding email processing suggestions. For the email processing suggestions, the embodiment of the present application also adds a prompt to guide the answer to the disposal plan of individual users, enabling users to well understand the disposal methods of various emails and minimizing the risks brought by spam. For example, executable attachments in emails should not be easily downloaded, and other communication methods need to be used to confirm with the email sender.

[0119] Through the text generation ability of the language analysis model, information interpretation of the email analysis log is carried out, making it possible for ordinary users to understand the reasons why emails are detected as normal / spam, and improving the user experience. For users, understanding the reasons for email determination can enhance personal email security awareness education and reduce the risk of being attacked.

[0120] In some embodiments, the data extraction module 201 is specifically:

[0121] The components of an email include the email header, the email body, and the attachment. The email header mainly includes the transmission information of the email, such as the sender, recipient, date, and subject, etc.; the email body is the main content part of the email, containing actual text, pictures, formatting information, or other data. Since all three parts may carry viruses or phishing information, etc., the embodiment of the present application respectively performs spam detection on the email header, the email body, and the attachment to obtain the complete detection result.

[0122] In the embodiment of the present application, there are three optional ways to upload email data: uploading the original email, uploading the email image, and uploading the email header image. Since in the prior art, it generally relies on the anti-spam detection engine provided by an external platform to perform security detection on the received emails, while in the embodiment of the present application, a flexible callable spam detection service is provided. Users do not need to rely on email providers to perform email anti-spam detection work, and can flexibly select the email information upload method according to the usage habits of different platforms. The spam detection service can be directly called in the form of a web page or a small program, enabling users to bypass processes such as enterprise procurement and manufacturer application docking to obtain email result interpretation and measure suggestions.

[0123] Exemplarily, the user can upload the email information to be detected through methods such as web pages, WeChat mini-programs, and Alipay mini-programs. For mobile devices such as mobile phones, the camera is defaultly used to take pictures of the emails, obtaining the corresponding email images and email header images (taking a picture of the email header of the email to be detected will obtain the email header image). For PC devices, local pictures are defaultly obtained through screenshotting and other methods and then uploaded. Additionally, the email to be detected can be saved and the email of the email to be detected can be uploaded.

[0124] Then, for the uploaded email information of different data types, different data extraction schemes are respectively adopted to extract the necessary information in the email information, obtaining the information to be detected.

[0125] The email content of the original email is complete, including all information such as the email header, email body, and attachments. Therefore, first, the original email is disassembled into three parts for these three types of email information, and the necessary information is identified and extracted. Then, the extracted information is converted into structured data information for subsequent anti-spam detection processing. Among them, since some spam emails store information in attachments, the disassembly of attachments needs to be based on the type format of the attachments (such as file extension, MIME type verification, file header signature verification, etc.), and the attachments are extracted using the correct disassembly method. For document types, text information needs to be extracted, and for image information, text needs to be extracted through OCR technology.

[0126] If the user uploads an email image, the image needs to be preprocessed first, specifically including image denoising, image binarization, background noise removal, and image skew correction.

[0127] Image denoising is to remove random noise in the email image by using filtering algorithms such as Gaussian filtering and median filtering.

[0128] Image binarization is to convert the email image into black and white to simplify the recognition process.

[0129] Background noise removal is because the taken photos may contain interference or clutter elements in the background. Therefore, image segmentation or background removal methods are used to clean up this background noise, making the text content of the email image more prominent.

[0130] Image skew correction is because poor shooting angles may cause the email image to be skewed or distorted. Therefore, methods such as perspective transformation and affine transformation are needed to correct the text in the email image to be horizontal or upright, ensuring that the OCR algorithm can accurately read the characters.

[0131] If the quality of the email image is poor, directly using the OCR algorithm for text recognition will incur a high processing cost, and because the image quality is too low, performing OCR text recognition after high-level processing will cause recognition deviation. Therefore, it is necessary to preprocess the uploaded email image to improve the text quality and then use the OCR algorithm for text recognition and extraction.

[0132] Among them, OCR stands for Optical Character Recognition, also known as optical character recognition, which is a technology for extracting text information from images. Since in the embodiments of this application, the emails to be detected may have differences in multi-language texts and special character sets, it is necessary to train and adjust the OCR engine in advance using common languages and character sets to improve the accuracy of text recognition and extraction.

[0133] The text extraction results of the OCR engine are not always 100% accurate, and usually there will be character replacements, typos, missing characters, etc. To correct these errors, for the recognized text, spelling correction algorithms (such as dictionary-based correction, edit distance algorithms, language models, etc.) are used to detect and correct spelling mistakes. This part also needs to consider that the body of spam emails may use similar-style characters to confuse the detection engine. Therefore, the correction algorithm should not be excessive, and when necessary, calculate the character mixing ratio as a subsequent anti-spam detection feature.

[0134] The email header image is obtained by taking a photo or screenshot of the email header of the email to be detected. The email header image is generally standard text, so as long as the uploaded email header image is corrected for text and character errors, the corresponding text information can be obtained.

[0135] By performing text recognition and extraction on the email data such as the uploaded original email, email image, and email header image as described above, the information to be detected that can be applied to anti-spam detection is obtained. Using different data extraction methods for extraction can more accurately extract according to the text type and quickly and efficiently obtain the key text information.

[0136] In some embodiments, the information detection module 202 is specifically:

[0137] In the embodiments of this application, the detection of spam emails mainly includes three parts: anti-spam detection, address detection, and anti-virus detection. These three detections are the core elements of email anti-spam detection. An optimized and trained spam email detection model is used to perform anti-spam detection on the previously obtained information to be detected (including email headers and email body data), perform sandbox detection on the extracted email addresses (i.e., URLs), and perform anti-virus detection on email attachments. Among them, the spam email detection model includes an anti-spam detection module, a URL sandbox detection module, and an anti-virus detection module.

[0138] Optionally, in the embodiments of the present application, the anti-spam detection module is fine-tuned based on BERT used for data processing, optimized by the cross-entropy loss function, and the optimizer is selected as AdamW. The test data is divided into multiple training data sets and validation data sets. During the module training process, after each epoch, the module performance is evaluated on the validation set for metrics such as accuracy, recall, and F1 score to check the generalization ability of the module, and finally an anti-spam detection module that meets the production application conditions is obtained.

[0139] After the information to be detected is input into the anti-spam detection module, the email corresponding to the information to be detected is classified and predicted through the machine learning result. The email is labeled as a normal email or a spam email according to the extracted email body and email header information. If you want to further improve the detection accuracy, the spam can also be subdivided. Then, the text cleaning is performed on the labeled information to be detected. For example, for the email body in HTML format, the HTML tags need to be removed and special characters need to be deleted. The text after text cleaning is segmented into multiple words and phrases through a tokenization tool. To ensure the context relevance of the word vectors, a deep learning-based dynamic word vector method (BERT) is selected, and a large amount of unsupervised text data is used for pre-training to process the words and phrases, capture more complex syntax and semantic information, and obtain the email analysis log related to the email body and email header.

[0140] Optionally, NLTK, spacy, etc. can be used as the tokenization tool. NLTK, short for Natural Language Toolkit, is a Python library that is good at processing text data. Spacy is also a Python natural language processing library that can automatically complete various basic NLP tasks such as tokenization, part-of-speech tagging, dependency parsing, named entity recognition, sentiment analysis, and syntactic analysis.

[0141] The information to be detected related to the email address is input into the URL sandbox detection module. This module detects the security of the email URL through a simulated and isolated environment and can identify potential malicious behaviors such as phishing attacks, malware downloads, cross-site scripting attacks, etc. Based on the open-source URL sandbox, secondary development is carried out to judge whether the URL is safe by analyzing various aspects of information such as the structure, domain name, page content, redirection behavior, file download, and network traffic of the URL.

[0142] First, the components of the email address are parsed, including the protocol (http, https), domain name, path, query parameters, etc. By analyzing this metadata, the URL sandbox detection module can initially identify some common malicious behavior patterns. The detection process specifically includes:

[0143] 4. Check the domain name: Verify whether the domain name matches known malicious domain names or blacklisted domain names. Malicious links often use domain names that imitate legitimate websites.

[0144] 5. Examine the URL structure: Check for URLs that are overly long, contain garbled characters or special characters, which are common features in phishing or malware distribution and are recorded as one of the suspicious features.

[0145] 6. Analyze link redirection: Some attack links will jump to the final malicious website through multiple redirections. The sandbox can analyze the redirection link and check for abnormal jumps, which are recorded as one of the suspicious features.

[0146] Secondly, for the URL landing page, parse and check its HTML tags and attributes. Conventional phishing URLs often prompt users to output and submit account and password information. Therefore, the embodiments of this application focus on monitoring the following HTML tags:

[0147] (1) Username input box: Usually an input tag with type = "text" or type = "email";

[0148] (2) Password input box: Usually an input tag with type = "password";

[0149] (3) Confirm button: Usually an input tag with type = "submit" or< / button> <button>Label, type = "submit".

[0150] If the email address has this label feature and the anti-spam detection result of the corresponding email is determined to be phishing, then this URL is a phishing link.

[0151] If the email address does not have the above two features, then the traffic during URL access needs to be analyzed to detect whether there is abnormal external communication, such as connections to malicious control server addresses, downloading executable malicious attachments, etc.

[0152] Through the above process, the sandbox detection result of the email address can be obtained.

[0153] Input the information to be detected related to the email attachment into the anti-virus detection module to perform anti-virus detection on the attachment. Generally, if the user uploads the original email, it is necessary to perform anti-virus detection on its attachment.

[0154] The anti-virus detection module docks different anti-virus engine SDKs, which can improve the detection rate, that is, adopt a multi-engine mode to detect whether the email contains viruses. Among them, there are mainly two logical processing methods in the multi-engine mode: one is for the purpose of improving the detection rate. For example, as long as one engine detects a virus, it is determined that the email attachment contains a virus; the other is for the purpose of improving the accuracy. For example, it is necessary for two or more engines to detect a virus to determine that the email attachment contains a virus.

[0155] According to the above description, the spam detection model is used to detect the email body, email header, email address, and attachment respectively, and an email analysis log containing the detection results is obtained.

[0156] Implementing the embodiments of the present application has the following beneficial effects:

[0157] First, according to the type of the uploaded email data, the necessary text content is extracted to facilitate subsequent detection and processing of spam information; moreover, different data extraction methods can be used for extraction to more accurately extract according to the text type and quickly and efficiently obtain the key text information. Then, different types of detection methods are respectively adopted for different types of data in the information to be detected, such as the email body, email header, email address, and attachment, etc., to perform a more complete and detailed detection on the email to be detected, and an email analysis log is obtained. Then, the email analysis log is trimmed to remove redundant fields, and the key information retaining the detection results is input into the language analysis model for interpretation. Through the context semantics, a risk determination result that enables the user to understand why the email to be detected is determined to be normal / spam and the corresponding email processing suggestions are obtained. Using the email processing suggestions as a reference, the user can correctly handle spam, minimize the risk brought by spam, and improve the user experience.

[0158] Further, Figure 4 FIG. 3 is a structural diagram of a terminal device provided by an embodiment of the present application. As Figure 4 shown, the terminal device 3 of this embodiment includes: at least one processor 30 (only one is shown Figure 4 herein), a memory 31, and a computer program 32 stored in the memory 31 and executable on the at least one processor. When the processor 30 executes the computer program 32, the steps of a spam detection method described in any one of the embodiments of the present application can be implemented.

[0159] The terminal device 3 may be a computing device such as a desktop computer, a cloud server, and a laptop computer. The computing device may include, but is not limited to, the processor 30 and the memory 31. Figure 4 This is only an example of the terminal device 3 and does not limit the terminal device 3. It may include more or fewer components than shown in the figure.

[0160] The above specific embodiments further elaborate on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above are only specific embodiments of the present application and are not used to limit the protection scope of the present application. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.< / button>

Claims

1. A method for detecting spam, characterized in that: include: According to the data type, the uploaded email data is subjected to text recognition and extraction to obtain the information to be detected; Perform anti-spam detection, address detection and anti-virus detection on different contents of the information to be detected, and obtain the email analysis log; The trimmed email analysis log is input into a preset language analysis model to interpret the email detection result, and generate the risk determination result of the email analysis log and the corresponding email processing suggestion.

2. The method for detecting spam according to claim 1, characterized in that: According to the data type, text recognition and extraction are performed on the uploaded email data to obtain the information to be detected, specifically: According to the data type of the email data, respectively use corresponding data processing methods to perform text recognition and extraction on the email header, email body and attachments of the email data to obtain information to be detected; The data types include mail originals, mail images and mail header images.

3. The method for detecting spam according to claim 2, characterized in that: Also includes: For the mail data belonging to the mail image, performing OCR text recognition on the pre-processed mail data to extract the original language text; The initial language text is corrected for incorrect characters by a spelling correction algorithm to obtain the information to be detected of the mail data.

4. The method for detecting spam according to claim 1, characterized in that: The different contents of the information to be detected are respectively subjected to anti-spam detection, address detection and anti-virus detection to obtain an email analysis log, which is specifically: The information to be detected is input into a preset spam detection model, the email body and email header of the information to be detected are detected by the anti-spam detection module, the email address of the information to be detected is detected by the URL sandbox detection module, the attachment of the information to be detected is detected by the anti-virus detection module, and finally the email analysis log is output.

5. The method for detecting spam according to claim 4, characterized in that: The anti-spam detection module detects the email body and email header of the information to be detected, specifically: Extracting the mail body and mail header from the information to be detected to obtain first detection data; Classifying the first detection data by using preset labels, and then performing data cleaning on the classified first detection data; Using a word segmentation tool to segment the first detection data after data cleaning to obtain a number of words and phrases; The words and phrases are subjected to context-related semantic analysis to obtain the email analysis log related to the email body and email header.

6. The method for detecting spam according to claim 4, characterized in that: The URL sandbox detection module is used to detect the email address of the information to be detected, specifically: Extracting an email address from the information to be detected, and performing domain name detection on the email address according to a preset domain name blacklist; Performing a structure detection on the email address after the domain name detection; wherein the structure detection includes detecting whether the length of the email address exceeds a first threshold, and detecting whether the email address contains garbled characters or preset special characters; The email address after the structure detection is analyzed through redirection links to detect whether there is any abnormal jump.

7. The method for detecting spam according to claim 1, characterized in that: The trimmed email analysis log is input into a preset language analysis model to interpret the email detection result, and generate the risk determination result of the email analysis log and the corresponding email processing suggestion, specifically: According to the preset key field template, the email analysis log is trimmed to retain the email body content, email detection information and email behavior characteristics; wherein the email behavior characteristics include email exchange information and email text length; The trimmed email analysis log is input into the language analysis model for field analysis, and the email analysis log is interpreted based on a preset standardized answer template to generate a risk determination result of the email analysis log and a corresponding email processing suggestion.

8. The spam detection method according to any one of claims 1 to 3, characterized in that: The uploaded email data is specifically: Taking a photo or screenshot of the mail to be inspected, and obtaining mail data belonging to the mail image or the mail header image; Preprocessing the mail data, and then uploading the preprocessed mail data; Wherein, the preprocessing includes image denoising, image binarization and image tilt correction.

9. A spam detection system, characterized in that: include: Data extraction module, information detection module and detection result interpretation module; The data extraction module is used to perform text recognition and extraction on the uploaded email data according to the data type to obtain the information to be detected; The information detection module is used to perform anti-spam detection, address detection and anti-virus detection on different contents of the information to be detected, and obtain the email analysis log; The detection result interpretation module is used to input the trimmed email analysis log into a preset language analysis model to interpret the email detection result, and generate a risk determination result of the email analysis log and a corresponding email processing suggestion.

10. A terminal device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a spam detection method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Abnormal mail detection method and device, computer equipment and storage medium

    CN121619130A