Personalized junk email filtering method and system, device, and medium

By training the recommendation model and calculating the similarity between the user vector and the email vector, personalized spam filtering is realized, solving the problem of high misjudgment rate in the existing technology, and improving the accuracy and user experience of email filtering.

WO2025138872A1PCT designated stage expired Publication Date: 2025-07-03GUANGDONG COREMAIL COMPUTER TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/111772
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-08-13
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing email anti-spam technology cannot effectively distinguish the needs of different users, resulting in a high misjudgment rate and a low efficiency mechanism that relies on manual feedback, which affects the email sending and receiving experience.

Method used

By obtaining the historical log information of the recipient, training the recommendation model, calculating user vector information, comparing it with the email-related vectors, modifying and judging according to the similarity threshold, building a whitelist to reduce misjudgment.

Benefits of technology

It improves the accuracy of spam filtering, reduces misjudgments, and improves the personalized processing capabilities of email filtering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024111772_03072025_PF_FP_ABST
    Figure CN2024111772_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of information filtering, and in particular to a personalized junk email filtering method and system, a device, and a medium. The method comprises: acquiring historical log information of a recipient, and inputting the historical log information into a recommendation model for training; acquiring a user mailbox address, and inputting the user mailbox address into the recommendation model to obtain user vector information; acquiring an email related vector, and comparing the user vector information with the email related vector to obtain a vector similarity; and acquiring a similarity threshold, and rejudging an email when the vector similarity is greater than the similarity threshold. Thus, misjudgments during email filtering are reduced, and the accuracy of junk email filtering is improved.
Need to check novelty before this filing date? Find Prior Art

Description

A personalized spam filtering method, system, device and medium Technical Field

[0001] The present application relates to the field of information filtering technology, and in particular to a personalized spam filtering method, system, device and medium. Background Art

[0002] Current email anti-spam technology primarily filters globally, ignoring the differences between different users. For example, some users may find emails soliciting papers, requesting manuscripts, or reviewing manuscripts harassing, while the intended users may find them useful. The difference lies in whether the subject matter of the email is relevant to the recipient's major, or whether the conference's level matches the recipient's academic level. For example, emails from non-top conferences may be harassing to recipients at key universities, but useful to recipients at less prestigious universities. If a user-based personalized filtering engine were to refilter emails identified as spam by the global engine, since it can learn the recipient's email semantics or habits, such misjudgments could be corrected based on this personalized filtering engine.

[0003] Any machine learning algorithm can potentially misjudge spam. Anti-spam systems typically rely heavily on human feedback to detect misjudgments. Only after an email is misjudged and the recipient submits a complaint can further misjudgments be prevented by adding sender restrictions or improving the model's training sample extraction. However, in reality, due to user unfamiliarity with feedback, most misjudgments go unanswered. This results in ineffective resolution of misjudgments in email filtering, impacting the user experience. This issue remains to be addressed.

[0004] Summary of the Invention

[0005] In order to reduce misjudgments in email filtering and improve the accuracy of spam filtering, this application provides a personalized spam filtering method, system, device, and medium, which adopt the following technical solutions:

[0006] In a first aspect, the present application provides a personalized spam filtering method, comprising:

[0007] Obtain the recipient's historical log information and input the historical log information into the recommendation model for training;

[0008] Obtain the user's email address and input it into the recommendation model to obtain the user vector information;

[0009] Obtain email-related vectors, compare user vector information with email-related vectors, and obtain vector similarity;

[0010] Get the similarity threshold and change the email judgment when the vector similarity is greater than the similarity threshold.

[0011] Preferably, it also includes:

[0012] Obtain the sender information of emails whose vector similarity is greater than the similarity threshold, and build a whitelist based on the sender information of emails.

[0013] Preferably, it also includes:

[0014] Get the text content of the email, segment the text content to get a number of word information, and build a recommendation model based on the several word information to filter spam.

[0015] Preferably, it also includes:

[0016] Obtain the time information of email reception and input the time information of email reception into the recommendation model to filter spam emails.

[0017] Preferably, it also includes:

[0018] Obtain the number of recurrences of the email and input the information into the recommendation model to filter out spam emails.

[0019] In a second aspect, the present application provides a personalized spam filtering system, comprising:

[0020] The training module is used to obtain the recipient's historical log information and input the historical log information into the recommendation model for training;

[0021] User vector processing module, used to obtain the user's email address, input the user's email address into the recommendation model, and obtain user vector information;

[0022] The comparison module is used to obtain email-related vectors, compare user vector information with email-related vectors, and obtain vector similarity;

[0023] The re-judgment module is used to obtain a similarity threshold and re-judge the email when the vector similarity is greater than the similarity threshold.

[0024] Preferably, it also includes:

[0025] The whitelist module is used to obtain the sender information of emails whose vector similarity is greater than the similarity threshold, and build a whitelist based on the sender information of emails.

[0026] Preferably, it also includes:

[0027] The spam filtering module is used to obtain the text content of the email, segment the text content into words, obtain a number of word information, and build a recommendation model based on the multiple word information to filter spam.

[0028] In a third aspect, the present application provides a personalized spam filtering device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the personalized spam filtering method as described above.

[0029] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the personalized spam filtering method as described above when running.

[0030] In summary, compared with the prior art, the technical solution provided by this application has at least the following beneficial effects:

[0031] This application trains the recommendation model through historical log information, obtains the user vector information in the recommendation model by extracting the user email address, and then obtains the relevant vectors of emails judged as spam and compares them with the user vector information. The comparison obtains the vector similarity, and the spam is rejudged based on the quantitative relationship between the vector similarity and the similarity threshold, thereby reducing the misjudgment of email filtering and improving the accuracy of spam filtering. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] FIG1 is a flow chart of a personalized spam filtering method according to an embodiment of the present application.

[0033] FIG2 is a module diagram of a personalized spam filtering system according to an embodiment of the present application.

[0034] Explanation of the accompanying symbols: 1. Training module; 2. User vector processing module; 3. Comparison module; 4. Rejudgment module; 5. Whitelist module; 6. Spam filtering module. DETAILED DESCRIPTION

[0035] The present application is further described in detail below with reference to FIG. 1 and FIG. 2 . The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to be limiting.

[0036] When receiving emails, different users have different requirements for the content of the email information. Some users may think that emails soliciting papers or requesting manuscripts for review are spam, while others may find them useful. The difference lies in whether the subject of such emails is relevant to the recipient's major, or whether the level of the conference matches the recipient's academic level. Some emails from non-top conferences may be harassment to recipients from key universities, but may be useful to recipients from non-key universities. Another example is advertising emails, e-commerce promotions, or promotions for foreign trade companies. Some users happen to have related needs in their lives and therefore consider the emails to be normal emails, while some users happen to have no needs and therefore consider them to be spam.

[0037] Any machine learning algorithm may always have misjudgment problems when judging spam, and the misjudgment detection mechanism of general anti-spam systems is very dependent on manual feedback. On the basis of misjudgment, these misjudged senders may be added at any time. For example, a customer has launched a new OA system, which has led to the development of sending new notification letters. A recruitment website has suddenly added a type of notification email to notify users of a certain type of matter. A new academic conference will send a mass academic conference email, etc. Therefore, in order to deal with these misjudgments, it is necessary to find the misjudged senders at a faster speed and quickly correct the judgment errors of the machine learning model by adding sender restriction information. However, if you want to rely on user feedback to find these misjudged senders, you usually have to wait for several weeks to accumulate enough user feedback for us to discover these misjudgments. In order to solve the misjudgment problem in this scenario, the solution of this application is proposed.

[0038] 1 , the present application relates to a personalized spam filtering method, specifically comprising:

[0039] Step S1: Obtain the recipient's historical log information and input the historical log information into the recommendation model for training;

[0040] Step S2: Obtain the user's email address, input the user's email address into the recommendation model, and obtain user vector information;

[0041] Step S3: Obtain email-related vectors, compare the user vector information with the email-related vectors, and obtain vector similarity;

[0042] Step S4: Obtain a similarity threshold, and re-judge the email if the vector similarity is greater than the similarity threshold.

[0043] Specifically, a global spam classification system is used to classify emails. Historical log information is then used to train the recommendation model. This historical log information is extracted from all emails sent to each recipient, including the time the email was received, the number of times the email was repeated, and the text content of the emails. The user vectors of the recommendation model, trained using information related to all the recipient's emails, are very similar to the email vectors of the normal emails that the recipient frequently receives. The user vector information in the recommendation model is then extracted using the user's email address. The vectors associated with emails deemed spam are then compared with the user vector information to determine vector similarity. The spam classification is then revised based on the quantitative relationship between the vector similarity and the similarity threshold, reducing misjudgments in email filtering and improving spam filtering accuracy.

[0044] The embodiment of the present application introduces a recommendation system into the anti-spam scenario, changing the determination of whether an email is spam to the determination of whether the degree to which the email is recommended by the recipient is high enough, so as to achieve personalized anti-spam filtering.

[0045] For each user, the recommendation system calculates a vector based on their historical information, such as search results and the titles of items they clicked on. For each item in the system, a vector is also calculated based on historical information, such as which users clicked on it, the item's title, and its description. By comparing the similarity between these two vectors, the system can assess the degree to which the item is recommended to the user.

[0046] As one implementation method, the text content of the email is obtained, the text content is segmented to obtain a plurality of word information, and a recommendation model is constructed based on the plurality of word information to screen spam emails.

[0047] Specifically, in the email anti-spam scenario, the email text is first broken down into its word information. For example, if the email content is "This is a test," it is segmented into three words: "this," "a," and "test." These words are then mapped to "products" in the recommendation system. Therefore, receiving an email with "This is a test" is equivalent to the user "clicking" on the three products: "this," "a," and "test." These "product" words are actually "clicked" by multiple different users. This allows a conventional recommendation system to be trained to generate a model. The model takes the user and email text as input and outputs a recommendation score. If the recommendation score exceeds a threshold, it is considered that the user is likely to want to receive the email.

[0048] As one implementation method, it also includes:

[0049] Obtain the time information of email reception and input the time information of email reception into the recommendation model to filter spam emails.

[0050] Specifically, recommendation algorithms can accept inputs other than discrete variables like text. For example, the time an email was received. By including the time an email was received as an input feature, the same email content, sent to the same recipient during the day, can be classified as normal, but sent at midnight as spam.

[0051] As one implementation method, it also includes:

[0052] Obtain the number of recurrences of the email and input the information into the recommendation model to filter out spam emails.

[0053] Specifically, current recommendation algorithms accept inputs not only discrete variables but also continuous variables. For example, the number of times an email appears repeatedly. By including this as an input feature, we can identify the same email content sent to the same recipient as normal if the number of repetitions is small, but classify it as spam if it is sent 1,000 times.

[0054] By using each user's email sending and receiving records in the cloud, and capturing samples of both legitimate and spam emails, a unique model can be trained to support personalized models for hundreds of millions of users. After filtering emails using the global anti-spam filtering engine, the anti-spam system uses this personalized model again to determine the degree of match between the email and the recipient. As long as the match exceeds a threshold, the email is likely legitimate.

[0055] After classifying spam emails, the recipient's historical log information is obtained and fed into the recommendation model for training. Specifically, each recipient's log information, including the time the email was received, the number of times the email was repeated, and the text content of the email, is extracted to construct a vector associated with the email. Using these vectors as training samples and running the training program, a corresponding recommendation model is trained. The user vectors in this model will be very similar to the vectors of frequently received legitimate emails and inconsistent with the vectors of frequently received spam emails.

[0056] Obtain the user's email address and input it into the recommendation model to obtain the user vector information. The user email address is the recipient's email address. After the model is trained, the user vector corresponding to each user is already stored in the model. Simply input the user's email address to retrieve the user's corresponding user vector information from the model.

[0057] The recommendation system obtains the email-related vector and compares the user vector with the email-related vector to determine the vector similarity. Specifically, the spam email's text content, the time the email was received, and the number of times the email was repeated are passed as parameters to the recommendation system model. The vector associated with the email is calculated and then compared with the user vector to determine the similarity between the recipient vector and the email-related vector. A pre-set similarity threshold is then used. If the similarity exceeds the pre-selected threshold, the email is reclassified as a legitimate email and delivered to the user's inbox.

[0058] As one implementation method, the email sender information whose vector similarity is greater than a similarity threshold is obtained, and a whitelist is constructed according to the email sender information.

[0059] Specifically, the revised sending emails and sending domain names are recorded in real time, and the number of revisions for each sending email and sending domain name is recorded. There are several thresholds for the number of revisions, and the embodiment of the present application is set with T1 to T4. When the number of revisions of the sending email exceeds the threshold T1, an audit notification is triggered; if the number of revisions of the sending domain name exceeds the threshold T2, an audit notification is triggered; if the number of revisions of the sending email exceeds the threshold T3, and it has not yet been reviewed, the restriction information is automatically added first; if the number of revisions of the sending domain name exceeds the threshold T4, and it has not yet been reviewed, the sender restriction information is automatically added first; after receiving the manual review notification, the restriction information maintenance personnel determines whether the sender restriction information needs to be added. If it is considered necessary to add it and the system has not automatically added it, the restriction information is manually added; after receiving the manual review notification, the maintenance personnel determines whether the sender restriction information needs to be added. If it is considered unnecessary to add it, check whether the restriction information has been automatically added. If the restriction information has been automatically added, the corresponding restriction information is manually deleted.

[0060] 2 , a personalized spam filtering system is provided for an embodiment of the present application. The system includes:

[0061] Training module 1 is used to obtain the recipient's historical log information and input the historical log information into the recommendation model for training;

[0062] User vector processing module 2, used to obtain the user's email address, input the user's email address into the recommendation model, and obtain user vector information;

[0063] Comparison module 3 is used to obtain email-related vectors, compare user vector information with email-related vectors, and obtain vector similarity;

[0064] The judgment re-judgment module 4 is used to obtain a similarity threshold and to re-judge the email when the vector similarity is greater than the similarity threshold.

[0065] Specifically, the system further includes a whitelist module 5 for obtaining email sender information whose vector similarity is greater than a similarity threshold, and constructing a whitelist based on the email sender information.

[0066] The system also includes a spam filtering module 6, which is configured to obtain the text content of an email, segment the text content to obtain a plurality of word information, and construct a recommendation model based on the plurality of word information to filter out spam emails. The system is also configured to obtain the time information of email receipt and input the time information of email receipt into the recommendation model to filter out spam emails. The system is also configured to obtain the number of recurrences of an email and input the number of recurrences of an email into the recommendation model to filter out spam emails.

[0067] An embodiment of the present application provides a personalized spam filtering device, including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to execute the personalized spam filtering method as described above.

[0068] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the aforementioned personalized spam filtering method when running.

[0069] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and products can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0070] In several embodiments provided in this application, it should be understood that the disclosed methods, systems, devices, and program products may be implemented in other ways.

[0071] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0072] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A personalized spam filtering method, characterized in that, Including: Obtain the historical log information of the recipient, and input the historical log information into the recommendation model for training; Obtain the user's email address, input the user's email address into the recommendation model, and obtain the user vector information; Obtain the email-related vector, compare the user vector information with the email-related vector, and obtain the vector similarity; Obtain the similarity threshold, and rejudge the email when the vector similarity is greater than the similarity threshold.

2. The personalized spam filtering method according to claim 1, wherein Also including: Obtain the email sender information with vector similarity greater than the similarity threshold, and construct a whitelist according to the email sender information.

3. The personalized spam filtering method according to claim 1, wherein Also including: Obtain the text content of the email, perform word segmentation on the text content to obtain several word information, and construct a recommendation model according to the several word information for spam screening.

4. The personalized spam filtering method according to claim 3, wherein Also including: Obtain the time information when the email is received, and input the time information when the email is received into the recommendation model for spam screening.

5. The personalized spam filtering method according to claim 3, characterized in that, Also including: Obtain the repeated occurrence times information of the email, and input the repeated occurrence times information of the email into the recommendation model for spam screening.

6. A personalized spam filtering system, characterized in that, Including: A training module, configured to obtain the historical log information of the recipient, and input the historical log information into the recommendation model for training; A user vector processing module, configured to obtain the user's email address, input the user's email address into the recommendation model, and obtain the user vector information; A comparison module, configured to obtain the email-related vector, compare the user vector information with the email-related vector, and obtain the vector similarity; A rejudgment module, configured to obtain the similarity threshold, and rejudge the email when the vector similarity is greater than the similarity threshold.

7. The personalized spam filtering system according to claim 6, wherein Also including: A whitelist module, configured to obtain the email sender information with vector similarity greater than the similarity threshold, and construct a whitelist according to the email sender information.

8. The personalized spam filtering system according to claim 6, wherein, Also including: A spam screening module, configured to obtain the text content of the email, perform word segmentation on the text content to obtain several word information, and construct a recommendation model according to the several word information for spam screening.

9. A personalized spam filtering device, characterized in that, Including a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the personalized spam filtering method according to any one of claims 1-5.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program is configured to execute the personalized spam filtering method according to any one of claims 1-5 when running.

Citation Information

Patent Citations

  • A FastText algorithm-based high-precision intelligent misjudgment prevention method and device for a mail system

    CN109831373A

  • Method and device for identifying junk mail, server and storage medium

    CN110213152A

  • Mail management method and device, electronic equipment and computer readable storage medium

    CN116957527A

  • Personalized spam filtering method, system and equipment and medium

    CN117834579A

  • Method and system for determining a spam prediction error parameter

    US20220109649A1