Email Classification Method, Device and Electronic Equipment
By using pre-trained machine learning models to segment, group and normalize the URLs in the email, the problem of slow and low efficiency of email classification in the existing technology is solved, and the effect of quickly identifying and marking phishing emails is achieved.
Patent Information
- Application Number
- CN202210408239.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-04-19
AI Technical Summary
In the prior art, the classification of received mail is slow, inefficient, and it is difficult to effectively identify and mark phishing mail.
The pre-trained machine learning model is used to process word segmentation, group and normalize the URLs in the email, and the array is synthesized according to the location order, and finally classified it to quickly identify and mark phishing emails.
Parallel processing improves the efficiency of email classification, can quickly identify and mark phishing emails, and improves user security and efficiency.
Smart Images

Figure CN114742532B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of risk identification, and in particular, to a method, device, and electronic device for classifying emails. Background Art
[0002] Phishing emails usually contain emails that induce users to reply with personal private information (such as ID numbers, bank card passwords), or include website links that may disclose personal private information. Thus, in order to prevent users from replying with personal private information to phishing emails or clicking on the website links in phishing emails, when receiving an email, it is necessary to analyze the content of the email to classify the received email. Thus, when an email is classified as a phishing email, it can be marked to prompt the user.
[0003] Currently, the speed of classifying received emails is slow and the efficiency is low. Summary of the Invention
[0004] This application provides a method, device, and electronic device for classifying emails to solve the problem of slow speed and low efficiency in classifying received emails.
[0005] In a first aspect, this application provides a method for classifying emails, which is applied to a server. The method provided by this application includes:
[0006] Obtain the website addresses included in the emails to be identified;
[0007] Perform word segmentation on the website addresses based on a pre-trained machine learning model to obtain each first character in the website addresses. The machine learning model is obtained by inputting a training sample set composed of multiple website addresses marked with a first identifier and multiple website addresses carrying a second identifier into a network to be trained. The first identifier is used to indicate the existence of risks, and the second identifier is used to indicate the non-existence of risks;
[0008] Convert each first character in the website addresses based on the machine learning model to obtain a first array;
[0009] Group the first array based on the machine learning model to obtain N second arrays, and record the positional order between the N second arrays, where N is an integer greater than or equal to 2;
[0010] Perform normalization processing on the N second arrays in parallel based on the machine learning model to obtain the N normalized second arrays;
[0011] Based on the machine learning model, synthesize the N normalized second arrays into a normalized first array according to the recorded positional order between the N second arrays;
[0012] Classify the normalized first array based on a machine learning model, and output the classification result of the email carrying the website address.
[0013] The email classification method provided by this application can convert each first character in the website address based on a machine learning model to obtain a first array; group the first array based on the machine learning model to obtain N second arrays, and record the positional order among the N second arrays, where N is an integer greater than or equal to 2; perform normalization processing on the N second arrays in parallel based on the machine learning model to obtain N normalized second arrays. Since the normalization processing of the N second arrays is performed in parallel, the efficiency is high. Furthermore, based on the machine learning model, according to the recorded positional order among the N second arrays, the N normalized second arrays are combined into a normalized first array. In this way, the normalized first array can be classified based on the machine learning model, and the classification result of the email carrying the website address can be output. Thus, the efficiency of obtaining the classification result is also high.
[0014] In a second aspect, this application provides an email classification device, which is applied to a server. The device provided by this application includes:
[0015] An information acquisition unit, configured to acquire the website address included in the email to be recognized;
[0016] A word segmentation processing unit, configured to perform word segmentation processing on the website address based on a pre-trained machine learning model to obtain each first character in the website address, where the machine learning model is obtained by inputting a training sample set composed of multiple website addresses marked with a first identifier and multiple website addresses carrying a second identifier into a network to be trained, where the first identifier is used to indicate the existence of risks, and the second identifier is used to indicate the non-existence of risks;
[0017] A data conversion unit, configured to convert each first character in the website address based on a machine learning model to obtain a first array;
[0018] A data grouping unit, configured to group the first array based on the machine learning model to obtain N second arrays, and record the positional order among the N second arrays, where N is an integer greater than or equal to 2;
[0019] A normalization unit, configured to perform normalization processing on the N second arrays in parallel based on a machine learning model to obtain N normalized second arrays;
[0020] A data synthesis unit, configured to combine the N normalized second arrays into a normalized first array based on the machine learning model according to the recorded positional order among the N second arrays;
[0021] A data classification unit for classifying the normalized first array based on a machine learning model and outputting a classification result of the email carrying the website address.
[0022] In a third aspect, the present application further provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device is caused to execute the method provided in the first aspect of the present application.
[0023] In a fourth aspect, the present application further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the computer is caused to execute the method provided in the first aspect of the present application.
[0024] In a fifth aspect, the present application further provides a computer program product including a computer program. When the computer program is run, the computer is caused to execute the method provided in the first aspect of the present application.
[0025] In addition, the technical effects of the solutions provided in the second, third, fourth, and fifth aspects of the present application can refer to the technical effects of the email classification method provided in the first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0027] Figure 1 It is a flowchart of the email classification method provided in an embodiment of the present application;
[0028] Figure 2 It is an interaction schematic diagram between a server and a terminal device provided in an embodiment of the present application;
[0029] Figure 3 It is an architecture schematic diagram of the machine learning model provided in an embodiment of the present application;
[0030] Figure 4 is Figure 1 a specific flowchart of S104 in;
[0031] Figure 5 It is a functional module structure schematic diagram of the email classification device provided in an embodiment of the present application;
[0032] Figure 6 is Figure 5 a structure schematic diagram of a subunit of the data management module in;
[0033] Figure 7 is Figure 5 a structure schematic diagram of a subunit of the email recognition module in;
[0034] Figure 8 This is a circuit connection block diagram of an electronic device provided by an embodiment of the present application.
[0035] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0036] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application.
[0037] The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the drawings.
[0038] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other implementation manners of the present application. The present application is intended to cover any variations, uses, or adaptive changes of the present application, and these variations, uses, or adaptive changes follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary.
[0039] It should be understood that the present application is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
[0040] Please refer to Figure 1 , an embodiment of the present application provides a method for classifying emails, which is applied to a server 100. As Figure 2 shown, the server 100 is communicatively connected to the terminal device 200. Among them, the terminal device 200 can be a computer. As Figure 1 shown, the method provided by the present application includes:
[0041] S101: The server 100 inputs a training sample set composed of a plurality of URLs marked with a first identifier and a plurality of URLs carrying a second identifier into a network to be trained for training, and obtains a machine learning model.
[0042] Among them, multiple URLs marked with the first identifier can be the content of emails in a pre-stored first sample set and a second sample set. The first identifier is used to indicate the existence of risks, and the second identifier is used to indicate the non-existence of risks.
[0043] It should be noted that the first sample set can be called a public data set, and the emails in the first sample set are obtained from websites such as PhishTank, Millersmiles, and MalwarePatrol. The websites of PhishTank, Millersmiles, and MalwarePatrol include emails with URLs where fraud and personal privacy leaks exist. The second sample set can be called a private data set, including phishing emails confirmed in the enterprise's history or newly added phishing emails by users.
[0044] Among them, a part of the emails in the training sample set can be used as training samples, and the other part can be used as verification samples. The server can input multiple training samples into the network to be trained for training to obtain a machine learning model. Furthermore, input multiple verification samples into the machine learning model and output a classification result. When the accuracy rate of the classification result is lower than the preset ratio, continue to train the machine learning model until the accuracy rate of the classification result reaches the preset ratio (such as 95%).
[0045] In some embodiments, the network to be trained described above can be a transformer network.
[0046] S102: The server 100 obtains the URLs included in the email to be recognized.
[0047] Exemplarily, after receiving the email, the server 100 extracts the URL from the email. The URL can be referred to as a uniform resource locator (URL). For example, the URL can be "https: / / www.elgoog.com / ", and of course, it can also be other URLs, which are not limited here.
[0048] S103: The server 100 performs word segmentation processing on the URL based on the pre-trained machine learning model to obtain each first character in the URL.
[0049] It can be understood that the machine learning model is obtained by inputting a training sample set composed of multiple URLs marked with the first identifier and multiple URLs carrying the second identifier into the network to be trained, where the first identifier is used to indicate the existence of risks and the second identifier is used to indicate the non-existence of risks.
[0050] For example, when the website address is "https: / / www.elgoog.com / ", the respective first characters obtained by tokenizing the website address "https: / / www.elgoog.com / " are "h", "t", "t", "p", "s", ":", " / ", " / ", "w", "w", "w", ".", "e", "l", "g", "o", "o", "g", ".", "c", "o", "m", and " / ".
[0051] It should be noted that, as Figure 3 shown, the machine learning model includes a preprocessing layer, and the preprocessing layer of the machine learning model can execute S103.
[0052] S104: The server 100 converts each first character in the website address based on the machine learning model to obtain a first array. Among them, each element in the first array is an integer constant.
[0053] Exemplarily, as Figure 4 shown, S104 may include:[[]]
[0054] S401: The server 100 determines the length of the website address based on the machine learning model.
[0055] For example, when the website address is "https: / / www.elgoog.com / ", the server 100 determines that the length of the website address is 23.
[0056] S402: The server 100 processes the website address according to the length of the website address so that the length of the website address is equal to the preset length.
[0057] For example, when the website address is "https: / / www.elgoog.com / " and the preset length is 512, zeros can be appended to the end of "https: / / www.elgoog.com / " to obtain "https: / / www.elgoog.com / 000...000". Among them, the length of "https: / / www.elgoog.com / 000...000" is 512.
[0058] In some other embodiments, when the preset length is 512 and it is determined that the length of the website address is greater than 512, the first character at the end of the website address can be truncated so that the length of the website address is equal to 512.
[0059] It should be noted that the preprocessing layer of the machine learning model can also execute S104.
[0060] S403: The server 100 determines one by one whether the first character of the processed URL is in the preset word list. If so, execute S404; if not, execute S405.
[0061] Among them, the preset word list includes the mapping relationship between some first characters and integer constants. When the first character in the URL is in the preset word list, the integer constant corresponding to the first character can be found from the preset word list; conversely, the integer constant corresponding to the first character cannot be found.
[0062] S404: The server 100 converts the first character into the integer constant corresponding to the first character in the preset word list.
[0063] For example, the server 100 can convert the first character "h" into the integer constant "15"; for another example, the server 100 can convert the first character "t" into the integer constant "2". Exemplarily, when all the first characters in the processed "https: / / www.elgoog.com / 000...000" are converted, the processed URL "https: / / www.elgoog.com / 000...000" is converted into [15; 2; 2; 13; 9; 31; 4; 4; 33; 33; 33; 18; 3; 17; 26; 5; 5; 26; 18; 12; 5; 16; 4; 0; 0; 0;...; 0; 0; 0]. Among them, [15; 2; 2; 13; 9; 31; 4; 4; 33; 33; 33; 18; 3; 17; 26; 5; 5; 26; 18; 12; 5; 16; 4; 0; 0; 0;...; 0; 0; 0] can be called the first array.
[0064] S405: The server 100 converts the first character into a target character.
[0065] For example, when the first character is not in the preset word list, the server 100 converts the first character into the target character " <oom>”。
[0066] Understandably, the first character is converted into the integer constant corresponding to it in the preset vocabulary table to facilitate processing by the machine learning model.
[0067] It should be noted that still as Figure 3 shown, the machine learning model further includes an input layer, and the input layer is used to receive the first array from the preprocessing layer.
[0068] S105: The server 100 groups the first array based on the machine learning model to obtain N second arrays, and records the position order among the N second arrays, where N is an integer greater than or equal to 2.
[0069] Exemplarily, when N = 4, the first array [15; 2; 2; 13; 9; 31; 4; 4; 33; 33; 33; 18; 3; 17; 26; 5; 5; 26; 18; 12; 5; 16; 4; 0; 0; 0;...; 0; 0; 0] with a length of 512 can be evenly divided into the second array A [15; 2; 2; 13; 9; 31; 4; 4; 33; 33; 33; 18; 3; 17; 26; 5; 5; 26; 18; 12; 5; 16; 4; 0; 0; 0;...; 0; 0; 0]; the second array B [0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0;...; 0; 0; 0]; the second array C [0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0;...; 0; 0; 0]; the second array D [0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0; 0;...; 0; 0; 0]. The server 100 records the position order among the 4 second arrays, which is the second array A, the second data B, the second array C, and the second array D in sequence.
[0070] Of course, N can also be equal to integers such as 2, 3, 5, etc., which is not limited here.
[0071] It should be noted that still as Figure 3 shown, the machine learning model further includes a grouping layer and a positioning layer. The grouping layer is used to group the first array to obtain N second arrays. The positioning layer is used to record the position order among the N second arrays.
[0072] S106: Based on the machine learning model, perform normalization processing on the N second arrays in parallel to obtain the N normalized second arrays.
[0073] For example, when N = 4, since the normalization process is performed on 4 second arrays in parallel, the efficiency is high.
[0074] It should be noted that, still as Figure 3 shown, the machine learning model further includes a multi-head attention layer, and the multi-head attention layer is used to execute S106. The multi-head attention layer may include N attention heads ( Figure 3 in this case, 4 attention heads), and each attention head is used to perform normalization processing on a second array. That is, N multi-head attention layers perform normalization processing on N second arrays in parallel to obtain N normalized second arrays.
[0075] It should be noted that the model dimension d model of each attention head, the key dimension d k and the value dimension d v satisfy the relationship: d k = d v = d model / 4 = 128.
[0076] S107: Based on the machine learning model, according to the position order among the recorded N second arrays, the N normalized second arrays are combined into a normalized first array.
[0077] It should be noted that, still as Figure 3 shown, the machine learning model further includes a data merging layer, and the data merging layer is used to execute 107.
[0078] S108: Based on the machine learning model, classify the normalized first array and output the classification result of the email carrying the website.
[0079] Among them, the classification result is used to indicate whether the email is a phishing email or not. For example, the classification result can be the binary number "0", which is used to indicate that the email is a phishing email; the classification result can also be the binary number "1", which is used to indicate that the email is not a phishing email. For another example, the classification result can be the English word "false", which is used to indicate that the email is a phishing email; the classification result can also be the English word "true", which is used to indicate that the email is not a phishing email.
[0080] It should be noted that, still as Figure 3 shown, the machine learning model further includes a forward feedback network layer, and the forward feedback network layer is used to classify the normalized first array, and then use the softmax function to output the classification result of the email carrying the website. Exemplarily, the size of the forward feedback network layer can be 128.
[0081] In summary, the email classification method provided by the embodiments of the present application can convert each first character in the website address based on a machine learning model to obtain a first array; group the first array based on the machine learning model to obtain N second arrays, and record the positional order among the N second arrays, where N is an integer greater than or equal to 2; perform normalization processing on the N second arrays in parallel based on the machine learning model to obtain the N normalized second arrays. Since the normalization processing of the N second arrays is performed in parallel, the efficiency is high. Furthermore, based on the machine learning model, according to the recorded positional order among the N second arrays, the N normalized second arrays are combined into a normalized first array. In this way, the normalized first array can be classified based on the machine learning model to output the classification result of the email carrying the website address. Thus, the efficiency of obtaining the classification result is also high.
[0082] In addition, after the above S108, the method provided by the embodiments of the present application may further include:
[0083] S109: When the classification result indicates that the email is a phishing email, the server 100 sends a prompt message to the terminal device 200 for display, and the prompt message is used to indicate that the email is a phishing email.
[0084] For example, when the terminal device 200 displays the home page of the email application, the prompt message may be displayed on one side of the received emails in the email directory on the home page of the email application. The prompt message may be, but is not limited to, a text message such as "This email may be a phishing email".
[0085] S110: The server 100 responds to the user from the terminal device 200 marking a first identifier or a second identifier for the email carrying the website address.
[0086] The user can browse the content of the email on the terminal device 200. When the user discovers that the email carries risky content, the user can mark the first identifier for the email; when the user discovers that the email does not carry risky content, the user can mark the second identifier for the email.
[0087] S111: The server 100 adds the email marked with the first identifier or the second identifier to the training sample set.
[0088] For example, the server 100 can add the email marked with the first identifier or the second identifier to the above-mentioned private dataset. In this way, the training sample set can be continuously updated, making the reliability of the machine learning model obtained based on the updated training sample set higher later.
[0089] In addition, the server 100 can also determine whether the classification result is correct according to the classification result of the email carrying the website address and the first identifier or the second identifier marked by the user for the email carrying the website address.
[0090] For example, when the classification result is used to indicate that an email is risky, and the second identifier (used to indicate that the email is not risky) marked by the user is carried in the email with a website address, the server 100 determines that the classification result is incorrect. For another example, when the classification result is used to indicate that an email is risky, and the first identifier (used to indicate that the email is risky) marked by the user is carried in the email with a website address, the server 100 determines that the classification result is correct. The server 100 can also count the proportion of correct classifications, the proportion of incorrect classifications, and the recall rate of the machine learning model. The proportion of correct classifications, the proportion of incorrect classifications, and the recall rate are sent to the terminal device 200 for display. In this way, the maintainer of the machine learning model can adjust the machine learning model based on the proportion of correct classifications and the proportion of incorrect classifications.
[0091] The embodiment of the present application also provides an email classification device 500, which is applied to the server 100. It should be noted that the basic principle and the technical effects generated by the email classification device 500 provided in this embodiment are the same as those in the above embodiments. For a brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above embodiments. As Figure 5 shown, the device provided by the present application includes a data management module 502, a model management module, and an email recognition module 503. Among them,
[0092] The data management module 502 is used to store a training sample set composed of a plurality of website addresses marked with the first identifier and a plurality of website addresses carrying the second identifier, and input the training sample set into the network to be trained to obtain a machine learning model.
[0093] As Figure 6 shown, specifically, the data management module 502 may include a public data set unit 601, a private data set unit 602, and a data update unit 603.
[0094] Among them, the public data set unit 601 is used to store the first sample set, and the private data set unit 602 is used to store the second sample set. The data update unit 603 is used to send a prompt message to the terminal device 200 for display when the classification result indicates that the email is a phishing email, where the prompt message is used to indicate that the email is a phishing email; in response to the user from the terminal device 200 marking the first identifier or the second identifier for the email with a website address; adding the email marked with the first identifier or the second identifier to the training sample set.
[0095] The model management module is used to configure and update the machine learning model.
[0096] Exemplarily, the model management module is specifically configured to determine whether the classification result is correct according to the classification result of the email carrying the website address, the first identifier or the second identifier marked by the user for the email carrying the website address; count the correct classification ratio, the incorrect classification ratio, and the recall rate of the machine learning model; and send the correct classification ratio, the incorrect classification ratio, and the recall rate to the terminal device 200 for display.
[0097] The email recognition module 503 is configured to classify the obtained emails.
[0098] Specifically, as Figure 7 shown, the email recognition module 503 includes: an information acquisition unit 701, a word segmentation processing unit 702, a data conversion unit 703, a data grouping unit 704, a normalization unit 705, a data synthesis unit 706, and a data classification unit 707. Among them,
[0099] The information acquisition unit 701 is configured to acquire the website address included in the email to be recognized.
[0100] The word segmentation processing unit 702 is configured to perform word segmentation processing on the website address based on a pre-trained machine learning model to obtain each first character in the website address. The machine learning model is obtained by inputting a training sample set composed of multiple website addresses marked with the first identifier and multiple website addresses carrying the second identifier into the network to be trained. The first identifier is used to indicate the existence of risk, and the second identifier is used to indicate the non-existence of risk.
[0101] The data conversion unit 703 is configured to convert each first character in the website address based on the machine learning model to obtain a first array, where each element in the first array is an integer constant.
[0102] Specifically, the data conversion unit 703 is specifically configured to determine the length of the website address based on the machine learning model. According to the length of the website address, the website address is processed so that the length of the website address is equal to the preset length. When the first character of the processed website address is in the preset word list, the first character is converted into the integer constant corresponding to the first character in the preset word list. When the first character of the processed website address is not in the preset word list, the first character is converted into the target character.
[0103] The data grouping unit 704 is configured to group the first array by the machine learning model to obtain N second arrays, and record the position order between the N second arrays, where N is an integer greater than or equal to 2.
[0104] The normalization unit 705 is configured to perform parallel normalization processing on the N second arrays based on the machine learning model to obtain the normalized N second arrays.
[0105] Specifically, the normalization unit 705 is specifically configured to perform normalization processing on N second arrays in parallel by N multi-head attention layers to obtain N normalized second arrays, where any one of the multi-head attention layers performs normalization processing on one second array.
[0106] The data synthesis unit 706 is configured to synthesize the N normalized second arrays into a normalized first array based on the machine learning model according to the position order among the recorded N second arrays.
[0107] The data classification unit 707 is configured to classify the normalized first array based on the machine learning model and output the classification result of the email carrying the website.
[0108] Figure 8 is a block diagram of an electronic device shown according to an exemplary embodiment. The electronic device is a server 800. The server 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, an input / output (I / O) interface 812, and a communication component 816.
[0109] The processing component 802 generally controls the overall operation of the server 800. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components.
[0110] The memory 804 is configured to store various types of data to support the operation of the server 800. Examples of these data include instructions for any application or method operating on the server 800, a training sample set, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0111] The power supply component 806 provides power for various components of the server 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the server 800.
[0112] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, and the peripheral interface module may be a keyboard, a click wheel, a button, etc.
[0113] The communication component 816 is configured to facilitate communication between the server 800 and other devices in a wired or wireless manner.
[0114] In an exemplary embodiment, the server 800 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0115] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the above instructions can be executed by a processor 820 of the server 800 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc. When the instructions in the storage medium are executed by a processor of the terminal device, the electronic device can execute Figure 1 the mail classification method shown.
[0116] The embodiments of the present application also provide a computer program product including a computer program, and when the computer program is executed by a processor, it is like Figure 1 the mail classification method shown.
[0117] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary.
[0118] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.< / oom>
Claims
1. A method for classifying emails, characterized in that, Applied to a server, the method includes: Obtain the URLs included in the emails to be recognized; Perform word segmentation on the URLs based on a pre-trained machine learning model to obtain each first character in the URLs. Among them, the machine learning model is obtained by inputting a training sample set composed of multiple URLs marked with a first identifier and multiple URLs carrying a second identifier into a network to be trained. The first identifier is used to indicate the existence of risks, and the second identifier is used to indicate the non-existence of risks; Based on the machine learning model, convert each first character in the URLs to obtain a first array, including: Based on the machine learning model, determine the length of the URLs; According to the length of the URLs, process the URLs so that the length of the URLs is equal to a preset length; When the first character of the processed URLs is in the preset vocabulary, convert the first character to the integer constant corresponding to the first character in the preset vocabulary; When the first character of the processed URLs is not in the preset vocabulary, convert the first character to a target character; Based on the machine learning model, group the first array to obtain N second arrays, and record the positional order between the N second arrays, where N is an integer greater than or equal to 2; Based on the machine learning model, perform normalization processing on the N second arrays in parallel to obtain N normalized second arrays; Based on the machine learning model, according to the recorded positional order between the N second arrays, combine the N normalized second arrays into a normalized first array; Based on the machine learning model, classify the normalized first array and output the classification result of the email carrying the URLs.
2. The method according to claim 1, characterized in that, The machine learning model includes N multi-head attention layers. The step of based on the machine learning model, performing normalization processing on the N second arrays in parallel to obtain N normalized second arrays includes: The N multi-head attention layers perform normalization processing on the N second arrays in parallel to obtain N normalized second arrays, where any one of the multi-head attention layers performs normalization processing on one of the second arrays.
3. The method according to claim 1, characterized in that, After the step of based on the machine learning model, classifying the normalized first array and outputting the classification result of the email carrying the URLs, the method further includes: When the classification result indicates that the email is a phishing email, the server sends a prompt message to the terminal device for display, where the prompt message is used to indicate that the email is a phishing email; Respond to the user from the terminal device marking the first identifier or the second identifier for the email carrying the URLs; Add the emails marked with the first identifier or the second identifier to the training sample set.
4. The method according to claim 3, characterized in that, After adding the URLs included in the emails marked with the first identifier or the second identifier to the training sample set, the method further includes: According to the classification result of the email carrying the URLs and the first identifier or the second identifier marked by the user for the email carrying the URLs, determine whether the classification result is correct; Statistically calculate the proportion of correct classifications, the proportion of incorrect classifications, and the recall rate of the machine learning model; Send the proportion of correct classifications, the proportion of incorrect classifications, and the recall rate to a terminal device for display.
5. The method according to claim 1, characterized in that, Before obtaining the URLs included in the email to be recognized, the method further includes: Input a training sample set composed of the multiple URLs marked with the first identifier and the multiple URLs carrying the second identifier into a network to be trained for training to obtain the machine learning model.
6. An email classification device, characterized in that, Applied to a server, the device includes: An information acquisition unit, configured to acquire the URLs included in the email to be recognized; A word segmentation processing unit, configured to perform word segmentation processing on the URLs based on a pre-trained machine learning model to obtain each first character in the URLs, where the machine learning model is obtained by inputting a training sample set composed of multiple URLs marked with the first identifier and multiple URLs carrying the second identifier into a network to be trained, and where the first identifier is used to indicate the existence of risks and the second identifier is used to indicate the non-existence of risks; A data conversion unit, configured to convert each first character in the URLs based on the machine learning model to obtain a first array; The data conversion unit is specifically configured to determine the length of the URL based on the machine learning model; process the URL according to the length of the URL so that the length of the URL is equal to a preset length; when the first character of the processed URL is in a preset word list, convert the first character to the integer constant corresponding to the first character in the preset word list; when the first character of the processed URL is not in the preset word list, convert the first character to a target character; A data grouping unit, configured to group the first array by the machine learning model to obtain N second arrays, and record the positional order between the N second arrays, where N is an integer greater than or equal to 2; A normalization unit, configured to perform normalization processing on the N second arrays in parallel based on the machine learning model to obtain N normalized second arrays; A data synthesis unit, configured to synthesize the N normalized second arrays into a normalized first array based on the positional order between the N second arrays recorded by the machine learning model; A data classification unit, configured to classify the normalized first array based on the machine learning model and output a classification result of the email carrying the URL.
7. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the electronic device is caused to execute the method according to any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the computer is caused to execute the method according to any one of claims 1 to 5.
9. A computer program product, characterized in that, Including a computer program, when the computer program is run, the computer is caused to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Phishing website distinguishing method and device based on deep learning
CN110365691A
Mail processing method, mail processing device, electronic equipment and storage medium
CN113343682A