A risk assessment system and method for mailbox security
By constructing email analysis features and setting cycles, a personalized self-learning and updating mechanism is formed based on sender characteristics, solving the problem of difficult to identify trusted email risks in the existing technology, realizing effective risk assessment and adaptive protection of trusted emails, and ensuring enterprise information security.
Patent Information
- Application Number
- CN202411498259.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-10-25
AI Technical Summary
The prior art is difficult to effectively identify and prevent email risks from trusted bodies. Traditional email security detection methods cannot accurately identify email risks from trusted bodies, and the self-learning frequency can lead to false alarms or phishing emails.
By constructing email analysis features, detecting historical email databases, setting data cycles and verification cycles, forming a personalized self-learning and update mechanism based on sender characteristics, feeding it back to the data port for self-learning and update, and keeping update records to the data cloud.
It realizes an effective risk assessment of trustworthy emails, reduces the risk of email fraud, ensures the security of corporate information, and avoids the hidden dangers of false alarms and phishing emails.
Smart Images

Figure CN119402242B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mailbox security risk assessment, and specifically to a risk assessment system and method for mailbox security. Background Art
[0002] In the information age, email has become an important tool for enterprise communication. However, with the continuous upgrading of email security attack means, the email attack means have evolved from widely sending thousands of phishing emails on the Internet to sending phishing emails targeting specific individuals in combination with other network attacks. Among current technical means, most software can identify risks for emails from unknown users. However, for emails from trusted entities (such as partners, subsidiaries, other companies within a group, and internal employees), it is often easy to deceive the detection software. Traditional email security detection means are difficult to identify the email risks of trusted entities. In emerging email security detection technologies, they often rely only on the characteristics of emails for identification, such as common sending time periods, common cities for sending emails, common language types, common email signatures, frequently sent email recipients, frequently sent attachment types, frequently discussed email topics, and common sending frequencies, etc. Then, self-learning of the feature library is carried out through the method of big data portraits. However, during the self-learning process, too fast a learning frequency will lead to data conflicts, accelerated fitting, and a large number of false alarms; too slow a learning frequency will give some phishing emails an opportunity. Therefore, how to provide a self-learning frequency according to different email characteristics is one of the problems that need to be solved currently. Summary of the Invention
[0003] The purpose of the present invention is to provide a risk assessment system and method for mailbox security to solve the problems raised in the prior art.
[0004] To achieve the above purpose, the present invention provides the following technical solution: A risk assessment method for mailbox security, the method includes the following steps:
[0005] S1. Construct email analysis features, detect the historical email database, and perform data sorting on the historical emails in the
[0006] historical email database based on the email analysis features;
[0007] S2. Set a cycle for the historical email database, the cycle includes a data cycle and a verification cycle,
[0008] the data cycle is used to extract email analysis feature data, and the verification cycle is used to extract the false positive rate and the false negative rate;
[0009] S3. Construct a mailbox security risk assessment model, form a sender personalized self-learning update mechanism based on the sender characteristics, and feedback it to the data port;
[0010] S4. The data port performs corresponding self-learning updates based on the ID information of different senders, and retains the update records in the data cloud.
[0011] According to the above technical solution, in step S1, the email analysis features include: the sender's email address, the sender's IP address, the email sending time, the email subject, the email language, and the email attachment format.
[0012] Among them, the language of the email is determined by reading the Unicode codes of all characters in the email body, counting the number of Chinese characters within the Chinese Unicode range, calculating whether the number of Chinese characters exceeds 50% of the total number of characters. If it exceeds 50%, it is considered that the email body uses Chinese; otherwise, it is considered that the email body does not use Chinese.
[0013] The data collation means: grouping and processing according to the same sender, forming a historical email sequence in chronological order, and forming a data analysis group of historical emails. The data analysis group of historical emails is denoted as [A 0 、B 0 、C 0 、D 0 、E 0 , where A 0 refers to the sender's IP address; B 0 refers to the corresponding email sending time; C 0 refers to the corresponding email subject; D 0 refers to the corresponding email language; E 0 refers to the corresponding attachment format; the attachment format includes documents, pictures, audio, and video.
[0014] According to the above technical solution, in step S2, a data period T 0 and a verification period T 1 are set for any sender;
[0015] The verification period T 1 is adjacent to the data period T 0 and is the latter time period of the data period T 0 ;
[0016] Collect the data analysis group of historical emails in the data period T 0 , collect the fixed learning update duration in the data period T 0 ; collect the misjudgment rate and the wrong judgment rate of emails within the verification period T 1 ;
[0017] The misjudgment rate refers to the number of normal emails judged as warning files within the verification period T 1The proportion of the total number of emails from the same sender within; the misjudgment rate refers to the proportion of the number of spam emails misjudged as normal files in the verification period T 1 The proportion of the total number of emails from the same sender within; the warning file refers to the file determined to feedback an alarm to the administrator port;
[0018] Based on the data analysis group of historical emails, form the data period T 0 The following training data set: [A 1 、B 1 、C 1 、D 1 、E 1 , where A 1 refers to the number of IP addresses of the senders in the data period T 0 ; B 1 refers to the number of differences in the sending time of the corresponding emails in the data period T 0 ; C 1 refers to the number of different topics of the corresponding emails in the data period T 0 ; D 1 refers to the number of non-Chinese languages of the corresponding emails in the data period T 0 ; E 1 refers to the number of attachment format types of the corresponding emails in the data period T 0 ;
[0019] Among them, B 1 is processed using central tendency. Taking 24 hours per day as a range interval, the median is used to separate the data. Set a fixed value t, and the data outside the range of [K - t, K + t] is recorded as difference data, where K refers to the median of the sending time of the emails;
[0020] Construct the training data of the current sender for any sender, including the training data set in the data period T 0 : [A 1 、B 1 、C 1 、D 1 、E 1 ; the fixed update time period of the current sender; the sum of the misjudgment rate and the misclassification rate of the emails within the verification period T 1
[0021] According to the above technical solution, in step S3, the construction of the mailbox security risk assessment model includes:
[0022] Continuously change the position of the data period T 0 on the time axis to form several groups of training data of the current sender, denoted as the first data group of the current sender;
[0023] Formed under a fixed update time period, the current sender's verification period T 1 The sum of the false positive rate and the wrong positive rate of the emails and the data cycle T 0 The function fitting relationship between the training data sets below:
[0024] H=x 1 A 1 +x 2 B 1 +x 3 C 1 +x 4 D 1 +x 5 E 1 +ω
[0025] Among them, H refers to the verification period T of the current sender 1 The sum of the misjudgment rate and the wrong judgment rate of the internal mail; ω refers to the error term; x 1 、x 2 、x 3 、x 4 、x 5 Represent the coefficients of function fitting respectively;
[0026] The function fitting relationship is calculated for all senders, and a representative data coordinate is taken for each sender. The representative data coordinate takes the fixed update time period value of the sender as the ordinate and the most recent data period T of the current sender as the ordinate. 0 [A 1 , B 1 , C 1 , D 1 、E 1 ] as the basic data, substitute the function fitting relationship of each sender, and the sum of the misjudgment rate and the wrong judgment rate is used as the horizontal axis; form a number of coordinate scatter points, and use the linear equation fitting to form:
[0027] p=k 1 *H 0 +b
[0028] Where p refers to the fixed update time period of the sender; H 0 Refers to the function fitting relationship corresponding to the sender, using the current sender's most recent data period T 0 [A 1 , B 1 , C 1 , D 1 、E 1 ] is the sum of the misjudgment rate and the wrong judgment rate formed as the basic data; k 1 Represents the coefficient value, k 1 >0; b represents a constant term;
[0029] Set a fixed update time period threshold p 0 ; When H 0 takes 0, if p does not satisfy p≥p 0 >0, then take p = p 0 Output; When H 0 takes 0, if p satisfies p≥p 0 >0, then take p as output;
[0030] The data port receives the update time period p and performs corresponding self - learning updates.
[0031] A risk assessment system for mailbox security, which includes a mail feature processing module, a cycle analysis module, a mailbox security risk assessment module, and a self - learning update module;
[0032] The mail feature processing module is used to construct mail analysis features, detect the historical mail database, and organize the data of historical mails in the historical mail database based on the mail analysis features; the cycle analysis module is used to set a cycle for the historical mail database, and the cycle includes a data cycle and a verification cycle. The data cycle is used to extract mail analysis feature data, and the verification cycle is used to extract the false positive rate and the false negative rate; the mailbox security risk assessment module is used to construct a mailbox security risk assessment model, form a sender - personalized self - learning update mechanism based on the sender's features, and feedback it to the data port; the self - learning update module performs corresponding self - learning updates based on the ID information of different senders and retains the update records in the data cloud;
[0033] The output end of the mail feature processing module is connected to the input end of the cycle analysis module; the output end of the cycle analysis module is connected to the input end of the mailbox security risk assessment module; the output end of the mailbox security risk assessment module is connected to the input end of the self - learning update module.
[0034] According to the above technical solution, the mail feature processing module includes a mail analysis feature extraction unit and a data organization unit;
[0035] The mail analysis feature extraction unit is used to construct mail analysis features, and the mail analysis features include: the sender's mail address, the sender's IP address, the mail sending time, the mail subject, the mail language, and the mail attachment format; the data organization unit is used to detect the historical mail database and organize the data of historical mails in the historical mail database based on the mail analysis features;
[0036] The output end of the mail analysis feature extraction unit is connected to the input end of the data organization unit.
[0037] According to the above technical solution, the cycle analysis module includes a data cycle unit and a verification cycle unit;
[0038] The data cycle unit is used to construct a data cycle and extract email analysis feature data; the verification cycle unit is used to construct a verification cycle and extract the false positive rate and false negative rate;
[0039] The verification cycle is adjacent to the data cycle and is the next time period after the data cycle.
[0040] According to the above technical solution, the mailbox security risk assessment module includes a model analysis unit and a data feedback unit;
[0041] The model analysis unit is used to construct a mailbox security risk assessment model and form a personalized self-learning update mechanism for the sender based on the sender's characteristics; the data feedback unit obtains the self-learning update time period based on the self-learning update mechanism and feeds it back to the data port;
[0042] The output end of the model analysis unit is connected to the input end of the data feedback unit.
[0043] According to the above technical solution, the self-learning update module includes a data port unit and a data cloud;
[0044] The data port unit is used to query the ID information of the sender and perform corresponding self-learning updates based on the ID information of different senders; the data cloud is used to retain the update records;
[0045] The output end of the data port unit is connected to the data cloud.
[0046] Compared with the prior art, the beneficial effects of the present invention are as follows: By analyzing long-term historical emails, the present invention extracts the long-term historical behavior characteristics of the email sender and uses them as the historical portrait of the email address, and matches them with new emails from the sender's email address to discover abnormal risk behaviors inconsistent with historical behaviors, and timely send alarms to enterprise IT administrators and recipient employees. It solves the problem that traditional email security detection technologies only analyze the reputation of "single emails" and cannot discover risk emails from trusted entities. Through the implementation of the present invention, enterprises can effectively detect risk emails from trusted partners, subsidiaries, other companies in the group, and internal employees, implement an adaptive self-learning update mechanism, reduce the risk of email fraud, and ensure enterprise information security. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic flowchart of a risk assessment method for mailbox security according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0049] Embodiment: As Figure 1 shown, the present invention provides a risk assessment method for mailbox security, and the method includes: constructing mail analysis features, detecting a historical mail database, and sorting the historical mails in the historical mail database based on the mail analysis features;
[0050] The mail analysis features include: the mail address of the sender, the IP address of the sender, the sending time of the mail, the subject of the mail, the language of the mail, and the attachment format of the mail;
[0051] Among them, the language of the mail is determined by reading the Unicode codes of all characters in the mail body, counting the number of Chinese characters within the Chinese Unicode range, calculating whether the number of Chinese characters exceeds 50% of the total number of characters. If it exceeds 50%, it is considered that the mail body uses Chinese; otherwise, it is considered that the mail body does not use Chinese;
[0052] The data sorting means: grouping and processing according to the same sender, forming a historical mail sequence in chronological order, and forming a data analysis group of historical mails. The data analysis group of historical mails is denoted as [A 0 、B 0 、C 0 、D 0 、E 0 , where A 0 refers to the IP address of the sender; B 0 refers to the sending time of the corresponding mail; C 0 refers to the subject of the corresponding mail; D 0 refers to the language of the corresponding mail; E 0 refers to the attachment format of the corresponding mail; the attachment format includes documents, pictures, audio, and video.
[0053] Set a cycle for the historical mail database, and the cycle includes a data cycle and a verification cycle.
[0054] The data cycle is used to extract mail analysis feature data, and the verification cycle is used to extract the false positive rate and the false negative rate;
[0055] Set a data cycle T 0 and a verification cycle T 1 for any sender;
[0056] The verification period T 1 is adjacent to the data period T 0 and is the subsequent time period of the data period T 0 ;
[0057] The data analysis group of historical emails in the data collection period T 0 collects the fixed learning and update duration in the data collection period T 0 ; The misjudgment rate and the wrong judgment rate of emails within the collection verification period T 1 are collected;
[0058] The misjudgment rate refers to the proportion of the number of normal emails determined as warning files within the verification period T 1 to the total number of emails of the same sender; The wrong judgment rate refers to the proportion of the number of spam emails determined as normal files within the verification period T 1 to the total number of emails of the same sender; The warning file refers to a file determined to feedback an alarm to the administrator port;
[0059] Based on the data analysis group of historical emails, a training data set in the data period T 0 is formed: [A 1 , B 1 , C 1 , D 1 , E 1 , where A 1 refers to the number of IP addresses of senders in the data period T 0 ; B 1 refers to the number of differences in the sending times of the corresponding emails in the data period T 0 ; C 1 refers to the number of different subjects of the corresponding emails in the data period T 0 ; D 1 refers to the number of non-Chinese languages of the corresponding emails in the data period T 0 ; E 1 refers to the number of attachment format types of the corresponding emails in the data period T 0 ;
[0060] Among them, B 1 is processed using the central tendency. Taking every 24 hours of a day as a range interval, the median is used to separate the data, and a fixed value t is set. The data outside the range of [K - t, K + t] is recorded as difference data, where K refers to the median of the sending time of the emails;
[0061] For any sender, the training data of the current sender is constructed, including the training data set in the data period T 0 : [A 1 , B 1 , C 1 , D1 , E 1 ; The fixed update time period of the current sender; The verification period T 1 The sum of the misjudgment rate and the wrong judgment rate of the emails within it.
[0062] Construct a mailbox security risk assessment model, form a self-learning update mechanism personalized for the sender based on the sender's characteristics, and feedback it to the data port;
[0063] The construction of the mailbox security risk assessment model includes:
[0064] Continuously change the position of the data period T 0 on the time axis to form several sets of training data for the current sender, denoted as the first data set of the current sender;
[0065] Form the verification period T of the current sender under the fixed update time period 1 The sum of the misjudgment rate and the wrong judgment rate of the emails within it and the function fitting relationship between the data period T 0 and the training data set under it:
[0066] H = x 1 A 1 + x 2 B 1 + x 3 C 1 + x 4 D 1 + x 5 E 1 + ω
[0067] Among them, H refers to the sum of the misjudgment rate and the wrong judgment rate of the emails within the verification period T of the current sender 1 ; ω refers to the error term; x 1 , x 2 , x 3 , x 4 , x 5 respectively represent the coefficients of the function fitting;
[0068] Calculate the respective function fitting relationships for all senders, take the representative data coordinates for each sender, and use the fixed update time period value of the sender as the ordinate, and use the [A 0 , 1 , B 1 , C 1 , D 1 , E 1 under the most recent data period T of the current sender as the basic data, substitute it into the respective function fitting relationships of each sender, and use the sum of the misjudgment rate and the wrong judgment rate formed as the abscissa; form several coordinate scatter points, and use a linear equation to fit and form:
[0069] p = k 1 *H 0 + b
[0070] Wherein, p refers to the fixed update time period value of the sender; H 0 refers to the sum of the misjudgment rate and the wrong judgment rate formed by [A 0 , B 1 , C 1 , D 1 , E 1 as the basic data under the function fitting relationship corresponding to the sender, using the most recent data period T of the current sender 1 ; k 1 represents the coefficient value, k 1 > 0; b represents the constant term;
[0071] Set the fixed update time period threshold p 0 ; When H 0 takes 0, if p does not satisfy p ≥ p 0 > 0, then take p = p 0 for output; When H 0 takes 0, if p satisfies p ≥ p 0 > 0, then take p for output;
[0072] The data port receives the update time period p, and the data port performs corresponding self-learning updates based on the ID information of different senders, and retains the update records to the data cloud.
[0073] In this embodiment, a risk assessment system for mailbox security is further provided. The system includes a mail feature processing module, a period analysis module, a mailbox security risk assessment module, and a self-learning update module;
[0074] The mail feature processing module is used to construct mail analysis features, detect the historical mail database, and sort out the data of the historical mails in the historical mail database based on the mail analysis features; the period analysis module is used to set a period for the historical mail database, and the period includes a data period and a verification period. The data period is used to extract mail analysis feature data, and the verification period is used to extract the misjudgment rate and the wrong judgment rate; the mailbox security risk assessment module is used to construct a mailbox security risk assessment model, form a sender-personalized self-learning update mechanism based on the sender features, and feedback it to the data port; the self-learning update module performs corresponding self-learning updates based on the ID information of different senders, and retains the update records to the data cloud;
[0075] The output end of the mail feature processing module is connected to the input end of the period analysis module; the output end of the period analysis module is connected to the input end of the mailbox security risk assessment module; the output end of the mailbox security risk assessment module is connected to the input end of the self-learning update module.
[0076] The mail feature processing module includes a mail analysis feature extraction unit and a data collation unit;
[0077] The mail analysis feature extraction unit is used to construct mail analysis features, which include: the sender's mail address, the sender's IP address, the mail sending time, the mail subject, the mail language, and the mail attachment format; the data collation unit is used to detect the historical mail database and collate the historical mails in the historical mail database based on the mail analysis features.
[0078] The output end of the mail analysis feature extraction unit is connected to the input end of the data collation unit.
[0079] The period analysis module includes a data period unit and a verification period unit;
[0080] The data period unit is used to construct a data period and extract mail analysis feature data; the verification period unit is used to construct a verification period and extract the false positive rate and the false negative rate.
[0081] The verification period is adjacent to the data period and is the subsequent time period of the data period.
[0082] The mailbox security risk assessment module includes a model analysis unit and a data feedback unit;
[0083] The model analysis unit is used to construct a mailbox security risk assessment model and form a sender-personalized self-learning update mechanism based on the sender's features; the data feedback unit obtains the self-learning update time period based on the self-learning update mechanism and feeds it back to the data port.
[0084] The output end of the model analysis unit is connected to the input end of the data feedback unit.
[0085] The self-learning update module includes a data port unit and a data cloud;
[0086] The data port unit is used to query the ID information of the sender and perform corresponding self-learning updates based on the ID information of different senders; the data cloud is used to retain the update records.
[0087] The output end of the data port unit is connected to the data cloud.
[0088] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in all respects, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. A risk assessment method for mailbox security, characterized by: The method comprises the following steps: S1. Construct email analysis features, detect the historical email database, and sort the historical emails in the historical email database based on the email analysis features; S2. Setting a cycle for the historical email database, wherein the cycle includes a data cycle and a verification cycle. The data cycle is used to extract email analysis feature data, and the verification cycle is used to extract the false positive rate and the wrong positive rate; S3. Build a mailbox security risk assessment model, form a sender-specific self-learning update mechanism based on sender characteristics, and feed it back to the data port; S4, the data port performs corresponding self-learning updates based on the ID information of different senders, and retains the update records to the data cloud; In step S3, the construction of the mailbox security risk assessment model includes: The position of the data period T0 on the time axis is continuously changed to form a number of training data of the current sender, which are recorded as the first data group of the current sender; The function fitting relationship between the sum of the misjudgment rate and the wrong judgment rate of the emails in the verification period T1 of the current sender under a fixed update time period and the training data set under the data period T0 is formed: H=x1A1+x2B1+x3C1+x4D1+x5E1+ω Among them, H refers to the sum of the false positive rate and the wrong positive rate of the emails in the verification period T1 of the current sender; ω refers to the error term; x1, x2, x3, x4, and x5 represent the coefficients of the function fitting respectively; The function fitting relationship of each sender is calculated, and the representative data coordinates are taken for each sender. The representative data coordinates use the fixed update time period value of the sender as the vertical coordinate, and use [A1, B1, C1, D1, E1] of the current sender's most recent data period T0 as the basic data. Substitute the function fitting relationship of each sender, and the sum of the misjudgment rate and the wrong judgment rate is used as the horizontal coordinate; a number of coordinate scatter points are formed, and the linear equation is used for fitting: p=k1*H0+b Among them, p refers to the fixed update time period value of the sender; H0 refers to the sum of the misjudgment rate and the wrong judgment rate formed by the function fitting relationship corresponding to the sender, using [A1, B1, C1, D1, E1] of the current sender in the most recent data period T0 as the basic data; k1 represents the coefficient value, k1>0; b represents the constant term; Set a fixed update time period threshold p0; when H0 is 0, if p does not satisfy p≥p0>0, then take p=p0 for output; when H0 is 0, if p satisfies p≥p0>0, then take p for output; The data port receives the update time period p and performs the corresponding self-learning update; Among them, A1 refers to the number of sender IP addresses in data period T0; B1 refers to the number of differences in the sending time of the corresponding emails in data period T0; C1 refers to the number of different subjects of the corresponding emails in data period T0; D1 refers to the number of non-Chinese languages of the corresponding emails in data period T0; E1 refers to the number of attachment format types in data period T0.
2. A risk assessment method for mailbox security according to claim 1, characterized in that: In step S1, the email analysis features include: the sender's email address, the sender's IP address, the email sending time, the email subject, the email language, and the email attachment format; The language of the email is determined by reading the Unicode codes of all characters in the email body, counting the number of Chinese characters within the Chinese Unicode range, and calculating whether the number of Chinese characters exceeds 50% of the number of all characters. If it exceeds 50%, it is considered that the email body uses Chinese; otherwise, it is considered that the email body does not use Chinese. The data sorting refers to: grouping according to the same sender, forming a historical email sequence in chronological order, and forming a data analysis group of historical emails, the data analysis group of historical emails is recorded as [A0, B0, C0, D0, E0], where A0 refers to the sender's IP address; B0 refers to the sending time of the corresponding email; C0 refers to the subject of the corresponding email; D0 refers to the language of the corresponding email; E0 refers to the corresponding attachment format; the attachment format includes documents, pictures, audio and video.
3. A risk assessment method for mailbox security according to claim 2, characterized in that: In step S2, a data period T0 and a verification period T1 are established for any sender; The verification period T1 is adjacent to the data period T0 and is a period after the data period T0; The data analysis group for historical emails in the data collection period T0, the fixed learning update duration in the data collection period T0; the misjudgment rate and wrong judgment rate of emails in the collection verification period T1; The false positive rate refers to the ratio of the number of normal emails judged as warning files to the total number of emails from the same sender during the verification period T1; the false positive rate refers to the ratio of the number of spam emails judged as normal files to the total number of emails from the same sender during the verification period T1; the warning file refers to the file judged as a file that feeds back an alarm to the administrator port; The data analysis group based on historical emails forms the training data set under data period T0: [A1, B1, C1, D1, E1]; Among them, B1 uses the central tendency to process, taking 24 hours a day as a range interval, using the median to separate the data, setting a fixed value t, and recording the data outside the range of [Kt, K+t] as difference data, where K refers to the median of the email sending time; For any sender, the training data of the current sender is constructed, including the training data set under the data period T0: [A1, B1, C1, D1, E1]; the fixed update time period of the current sender; and the sum of the false positive rate and the wrong positive rate of the emails in the verification period T1.
4. A risk assessment system for mailbox security, using the risk assessment method for mailbox security as claimed in claim 1, characterized in that: The system includes an email feature processing module, a cycle analysis module, an email security risk assessment module and a self-learning update module; The email feature processing module is used to construct email analysis features, detect the historical email database, and organize the data of historical emails in the historical email database based on the email analysis features; the cycle analysis module is used to set a cycle for the historical email database, and the cycle includes a data cycle and a verification cycle. The data cycle is used to extract email analysis feature data, and the verification cycle is used to extract the false positive rate and the wrong positive rate; the mailbox security risk assessment module is used to construct a mailbox security risk assessment model, and form a sender-personalized self-learning update mechanism based on the sender's characteristics, and feed it back to the data port; the self-learning update module performs corresponding self-learning updates based on the ID information of different senders, and retains the update records to the data cloud; The output end of the email feature processing module is connected to the input end of the periodic analysis module; the output end of the periodic analysis module is connected to the input end of the mailbox security risk assessment module; the output end of the mailbox security risk assessment module is connected to the input end of the self-learning update module.
5. A risk assessment system for mailbox security according to claim 4, characterized in that: The email feature processing module includes an email analysis feature extraction unit and a data sorting unit; The email analysis feature extraction unit is used to construct email analysis features, which include: the sender's email address, the sender's IP address, the email sending time, the email subject, the email language, and the email attachment format; the data sorting unit is used to detect the historical email database and sort the historical emails in the historical email database based on the email analysis features; The output end of the email analysis feature extraction unit is connected to the input end of the data sorting unit.
6. A risk assessment system for mailbox security according to claim 5, characterized in that: The cycle analysis module includes a data cycle unit and a verification cycle unit; The data cycle unit is used to construct a data cycle and extract email analysis feature data; the verification cycle unit is used to construct a verification cycle and extract the false positive rate and the wrong positive rate; The verification period is adjacent to the data period and is a time period after the data period.
7. A risk assessment system for mailbox security according to claim 5, characterized in that: The mailbox security risk assessment module includes a model analysis unit and a data feedback unit; The model analysis unit is used to construct a mailbox security risk assessment model and form a sender-personalized self-learning update mechanism based on sender characteristics; the data feedback unit obtains a self-learning update time period based on the self-learning update mechanism and feeds it back to the data port; The output end of the model analysis unit is connected to the input end of the data feedback unit.
8. A risk assessment system for mailbox security according to claim 5, characterized in that: The self-learning update module includes a data port unit and a data cloud; The data port unit is used to query the sender's ID information and perform corresponding self-learning updates based on the ID information of different senders; the data cloud is used to retain update records; The output end of the data port unit is connected to the data cloud.
Citation Information
Patent Citations
Learning rate adjusting method, device and equipment and readable storage medium
CN111353867A
Mail security detection device, method and equipment and storage medium
CN117768142A