Multi-language junk mail detection method and system and terminal equipment
By using the semantic recognition model in multilingual spam detection to extract text features and combine multi-dimensional features, the problem of accuracy and inefficiency of multilingual spam detection in the prior art is solved, and more efficient multilingual spam detection is achieved.
Patent Information
- Application Number
- CN202510094816.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-27
AI Technical Summary
It is difficult for the existing technology to build a high-precision multilingual spam detection model, especially due to the diversity of language types and the difficulty of obtaining email data, the model has insufficient accuracy and generalization capabilities in different language environments.
By obtaining the email data to be detected, a preset semantic recognition model is entered to extract text features and convert them into text matrix, and the text vector is then determined. The text vector is then inputted to the pretrained spam detection model for spam detection based on multidimensional features.
It improves the accuracy and efficiency of multilingual spam detection, and can more accurately identify spam in different locale environments.
Smart Images

Figure CN120050255A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer network communication, and particularly to a method, system, and terminal device for detecting multi-language spam. Background Art
[0002] With the rapid development of Internet technology, email has become an indispensable communication bridge. However, the popularization of email is accompanied by a problem that cannot be ignored - the rampant spread of spam. Spam not only seriously disrupts the normal order of information exchange, but also often contains various advertising promotions, fraud information, and even malicious links, posing a great threat to the personal information security and property security of users. In addition, with the in-depth development of globalization and the booming rise of cross-border trade, spam has begun to be spread in multiple languages, which undoubtedly further upgrades the complexity of spam detection.
[0003] Currently, the detection strategies for spam generally integrate multiple technologies such as content filtering, behavior analysis, and reputation scoring, and train a learning model through a large-scale email dataset to achieve effective identification of spam. However, due to the diversity of language types and the difficulty of obtaining email data corresponding to each language, constructing a high-precision multi-language spam detection model has become a daunting challenge. At the same time, the scarcity of email data in different languages directly affects the accuracy and generalization ability of the model in different language environments.
[0004] Application Content
[0005] This application provides a method, system, and terminal device for detecting multi-language spam to improve the accuracy and efficiency of multi-language spam detection.
[0006] In a first aspect, this application provides a method for detecting multi-language spam, including:
[0007] Obtain the email data to be detected;
[0008] Input the email data into a preset semantic recognition model to convert the first text features extracted from the email data into a text matrix, and determine the corresponding first text vector based on the text matrix, where the semantic recognition model is obtained by training with a first email training set containing different languages;
[0009] Input the first text vector into a preset spam detection model to determine the probability value that the email data belongs to spam based on the first text vector, and determine spam based on the probability value. The spam detection model is obtained by training based on multi-dimensional features and a second text vector in a second email training set. The second text vector is determined by the semantic recognition model inferring the second text features in the second email training set. The multi-dimensional features include the second text features.
[0010] In the embodiments of the present application, by obtaining the email data to be detected, it is convenient to input it into the semantic recognition model later to deeply understand the content corresponding to the email data. By inputting the email data into the preset semantic recognition model, the content corresponding to the email data in different languages can be accurately and quickly understood to identify the text features of emails in different languages and convert them into text matrices. Based on the text matrix, the corresponding first text vector is determined, and multi-language texts with similar semantics can be mapped into similar first text vectors, thereby extracting the semantic features of multiple languages, which is convenient for subsequent spam detection. By pre-training the spam detection model with multi-dimensional features in advance, the trained model can quickly and accurately perform spam detection based on the first text vector. Compared with the prior art, the present application can improve the accuracy and efficiency of multi-language spam detection.
[0011] Further, the conversion of the first text features extracted from the email data into a text matrix is specifically as follows:
[0012] Perform sub-word segmentation on the email data to obtain a number of first text features;
[0013] Perform vectorization representation on each of the first text features to obtain a vectorization result, and determine the corresponding text matrix based on the vectorization result.
[0014] In this way, by performing sub-word segmentation on the email data, the email data can be converted into simple first text features, and the content corresponding to the email data in different languages can be accurately and quickly understood.
[0015] Further, the determination of the corresponding first text vector based on the text matrix is specifically as follows:
[0016] Perform convolution processing on adjacent word vectors in the text matrix to obtain a convolution matrix, and perform pooling processing on the convolution matrix to obtain a pooling matrix;
[0017] Construct a position encoding matrix, and fuse the pooling matrix and the position encoding matrix to obtain a position information matrix;
[0018] Extract the context features from the position information matrix to obtain an attention matrix, and determine the corresponding first text vector based on the attention matrix.
[0019] By determining the corresponding first text vector based on the text matrix in this way, multilingual texts with similar semantics can be mapped into similar first text vectors, thereby extracting the semantic features of multiple languages, which is convenient for subsequent spam detection.
[0020] Furthermore, the calculation formula for the convolution process is specifically:
[0021] y j = f(∑ i w i x ij + b);
[0022] In the formula, y j is the convolution result of the j-th submatrix, and the y values of each submatrix form a convolution matrix; f is the activation function; w i represents the weight at the i-th position in the convolution kernel; x ij is the position of the i-th data in the j-th submatrix of the text matrix; b is the bias.
[0023] Furthermore, the semantic recognition model is obtained by training based on the first email training set containing different languages, specifically:
[0024] Obtain the initial email data, and translate the third text feature of the initial email data to obtain the first email training set containing different languages, where the third text feature includes the first subject feature and the first body feature;
[0025] Concatenate all the second subject features and second body features in the first email training set into a fourth text feature, encode the fourth text feature to obtain encoded information, and train the initial semantic recognition model based on the encoded information;
[0026] Determine the paired emails in the first email training set, and input the paired emails into the initial semantic recognition model to obtain the corresponding vector group;
[0027] Calculate the similarity of each vector in the vector group to obtain the similarity calculation result, and optimize the initial semantic recognition model based on the similarity calculation result to determine the semantic recognition model.
[0028] In this way, by constructing the first email training set containing different languages for training the semantic recognition model and training the semantic recognition model through contrast learning of paired emails, the semantic recognition model can make the vectors mapped by emails with different text features but the same semantics in different languages as close as possible.
[0029] Further, the spam detection model includes an embedding layer, a fully connected layer, and an activation function. Determining the probability value that the email data belongs to spam based on the first text vector is specifically as follows:
[0030] After receiving the first text vector, the embedding layer converts the first text vector into a corresponding feature vector;
[0031] The fully connected layer transforms the feature vector according to preset model parameters to obtain a prediction result;
[0032] The activation function determines the probability value of spam based on the prediction result.
[0033] In this way, by determining the probability value of spam, spam can be detected accurately and quickly.
[0034] Further, the spam detection model is obtained by training based on multi-dimensional features and a second text vector in a second email training set, specifically as follows:
[0035] Obtain a second email training set containing Chinese and English email data, and extract multi-dimensional features in the second email training set. Among them, the multi-dimensional features include second text features, attachment features, reputation features, and metadata features;
[0036] Input the second text feature into the semantic recognition model to obtain a second text vector;
[0037] Input the second text vector, the attachment feature, the reputation feature, and the metadata feature into the spam detection model to respectively obtain corresponding third text vectors, attachment vectors, reputation vectors, and metadata vectors;
[0038] Fuse the third text vector, the attachment vector, the reputation vector, and the metadata vector to obtain a feature interaction vector, and determine an output matrix based on the feature interaction vector;
[0039] Train the spam detection model based on the output matrix.
[0040] In this way, by pre-training the spam detection model with multi-dimensional features in advance, the trained model can quickly and accurately detect spam based on the first text vector.
[0041] Further, the calculation formula for determining the output matrix based on the feature interaction vector is specifically as follows:
[0042] y = b + wx T +xVVT x T ;
[0043] where y is an output matrix of size M×M; b is a bias; w is a weight; x is a matrix formed by stacking the feature interaction vectors into a matrix of size M×N, where M is the number of features and N is the vector length; V is a feature interaction vector of length N; and T is the transpose.
[0044] In a second aspect, the present application provides a multi - language spam detection system, including: an acquisition module, an identification module, and a detection module;
[0045] The acquisition module is configured to acquire mail data to be detected;
[0046] The identification module is configured to input the mail data into a preset semantic recognition model to convert the first text features extracted from the mail data into a text matrix, and determine a corresponding first text vector based on the text matrix, where the semantic recognition model is obtained by training according to a first mail training set including different languages;
[0047] The detection module is configured to input the first text vector into a preset spam detection model to determine a probability value that the mail data belongs to spam based on the first text vector, and determine spam based on the probability value, where the spam detection model is obtained by training according to multi - dimensional features and a second text vector in a second mail training set, the second text vector is inferred and determined by the semantic recognition model for the second text features in the second mail training set, and the multi - dimensional features include the second text features.
[0048] By acquiring the mail data to be detected in the embodiments of the present application, it is convenient to input the data into the semantic recognition model later to deeply understand the content corresponding to the mail data; by inputting the mail data into the preset semantic recognition model, it is possible to accurately and quickly understand the content corresponding to the mail data in different languages, identify the text features of different - language mails and convert them into a text matrix; by determining the corresponding first text vector based on the text matrix, multi - language texts with similar semantics can be mapped into similar first text vectors, thereby extracting multi - language semantic features, which is convenient for subsequent spam detection; by pre - training the spam detection model with multi - dimensional features in advance, the trained model can quickly and accurately perform spam detection based on the first text vector. Compared with the prior art, the present application can improve the accuracy and efficiency of multi - language spam detection.
[0049] In a third aspect, the present application further provides a terminal device, including: one or more processors; a memory coupled to the processors for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method for detecting multi-language spam as described in the present application. Description of the Drawings
[0050] Figure 1 is a schematic flowchart of an embodiment of the method for detecting multi-language spam provided by the present application;
[0051] Figure 2 is Figure 1 a schematic flowchart of step S102 in
[0052] Figure 3 is a schematic structural diagram of the semantic recognition model provided by the present application;
[0053] Figure 4 is Figure 2 a schematic flowchart of step S202 in
[0054] Figure 5 is a schematic diagram of the training process of the multi-semantic recognition model provided by the present application;
[0055] Figure 6 is a schematic diagram of the training process of the spam detection model provided by the present application;
[0056] Figure 7 is a schematic structural diagram of an embodiment of the multi-language spam detection system provided by the present application;
[0057] Figure 8 is a schematic structural diagram of an embodiment of the terminal device provided by the present application. Detailed Embodiments
[0058] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0059] It should be understood that the step numbers used in the text are only for convenient description and do not limit the execution order of the steps.
[0060] It should be understood that the terms used in the specification of the present invention are merely for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0061] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0062] The term "and / or" refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0063] At present, email has become an indispensable communication bridge, but the problem of spam is severe, which not only interferes with normal communication but also may threaten the information security of users. Under globalization, the multi-language spread of spam has intensified, resulting in a further upgrade in the difficulty of spam detection. The current spam detection strategies combine technologies such as content filtering, behavior analysis, and reputation scoring to identify spam through big data training models. However, the diversity of language types and the difficulty of obtaining corresponding email data in different languages make it difficult to accurately detect multi-language spam.
[0064] Next, the nouns involved in this application are analyzed:
[0065] Subword Tokenization is a technique that splits text into smaller units (usually subwords or character n-grams), which helps to handle out-of-vocabulary (OOV) words and morphologically rich languages. Through subword tokenization, email data can be decomposed into a series of first text features (here, it can be understood as a sequence of subwords or characters).
[0066] The masked language model is a natural language processing method. The idea is to randomly replace some words in the input text with masks, and then let the model predict the masked words according to the context, so that the model can learn context information and language structure, and thus obtain a vector that can represent the text meaning.
[0067] Contrastive learning is a self-supervised learning method that can learn the representation of data by comparing the similarities and differences between different samples. The core idea is to make similar samples as close as possible in the feature space, while making dissimilar samples as far away as possible.
[0068] Based on this, the embodiments of this application provide a method, system and medium for detecting multi-language spam, which can improve the accuracy and efficiency of multi-language spam detection.
[0069] A method, system, and medium for detecting multilingual spam provided by an embodiment of the present application will be described in detail through the following embodiments. First, the method for detecting multilingual spam in the embodiments of the present application will be described.
[0070] The method for detecting multilingual spam provided by the embodiment of the present application relates to the field of computer network communication. The method for detecting multilingual spam provided by the embodiment of the present application can be applied to a terminal, a server, or software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing a method for detecting multilingual spam, etc., but is not limited to the above forms.
[0071] The present application can also be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0072] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.
[0073] Embodiment 1
[0074] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the multi-language spam detection method provided by the present application, including steps S101 to S103;
[0075] Step S101, obtain the email data to be detected;
[0076] In some embodiments, the email data to be detected can be obtained through, but not limited to, an email server or an email client. Among them, the email data can be emails in languages such as Chinese, English, the Indic language family, the Iranian language family, other Germanic language families except English, the Romance language family, or the Slavic language family.
[0077] In some embodiments, the email data can include, but not be limited to, the email subject, the body, attachments, and metadata, etc.
[0078] Step S102, input the email data into a preset semantic recognition model to convert the first text features extracted from the email data into a text matrix, and determine the corresponding first text vector based on the text matrix, where the semantic recognition model is obtained by training according to a first email training set containing different languages;
[0079] Please refer to Figure 2 In some embodiments, step S102 can include, but not be limited to, steps S201 to S202;
[0080] Step S201, input the email data into a preset semantic recognition model to convert the first text features extracted from the email data into a text matrix;
[0081] In some embodiments, after obtaining the email data, it is input into a preset semantic recognition model, where the semantic recognition model is formed by concatenating an embedding model, a convolutional neural network, and a Transformer in sequence. After the semantic recognition model receives the email data, the embedding model converts the first text features extracted from the email data into a text matrix, and then the convolutional neural network and the Transformer determine the corresponding first text vector based on the text matrix. The structural schematic diagram of the semantic recognition model provided in this application is as shown in Figure 3 shown.
[0082] In some embodiments, the conversion of the first text features extracted from the email data into a text matrix includes: performing sub-word segmentation on the email data to obtain a number of first text features; vectorizing each of the first text features to obtain a vectorization result, and determining the corresponding text matrix based on the vectorization result. Specifically, first, after the semantic recognition model receives the email data, it performs sub-word segmentation to decompose the email data into a series of first text features (which can be understood as sub-words or character sequences), where the first text features include a subject feature and a body feature; second, each first text feature is encoded with a word embedding model and vectorized to obtain a vectorization result. Once each first text feature is converted into a vectorization result, a structured text matrix can be constructed based on these vectorization results.
[0083] In this way, by performing sub-word segmentation on the email data, the email data can be converted into simple first text features, and the content corresponding to the email data in different languages can be accurately and quickly understood.
[0084] Step S202, determining the corresponding first text vector based on the text matrix.
[0085] Please refer to Figure 4 , in some embodiments, step S202 may but is not limited to including steps S401 to S403;
[0086] Step S401, performing convolution processing on adjacent word vectors in the text matrix to obtain a convolution matrix, and performing pooling processing on the convolution matrix to obtain a pooling matrix;
[0087] In some embodiments, after obtaining the text matrix through the word embedding model, the convolutional neural network then processes the text matrix, where the convolutional neural network is composed of a convolutional layer and a pooling layer.
[0088] In some embodiments, convolutional processing is performed on adjacent word vectors in the text matrix to obtain a convolutional matrix. Specifically, the convolutional layer performs a convolutional operation on each sub-matrix in the text matrix, that is, adjacent word vectors, through one or more convolutional kernels, and shares the weights between each sub-matrix to extract the local spatial features of each sub-matrix, that is, the convolutional result. Furthermore, the convolutional results of each sub-matrix are combined to obtain a convolutional matrix. This process can construct the correlation features of adjacent words.
[0089] In some embodiments, the calculation formula for the convolutional processing is specifically:
[0090] y j = f(∑ i w i x ij + b);
[0091] In the formula, y j is the convolutional result of the j-th sub-matrix, and the y values of each sub-matrix form a convolutional matrix; f is an activation function; w i represents the weight at the i-th position in the convolutional kernel; x ij is the position of the i-th data in the j-th sub-matrix of the text matrix; b is a bias.
[0092] In some embodiments, pooling processing is performed on the convolutional matrix to obtain a pooling matrix. Specifically, after obtaining the convolutional matrix, the pooling layer performs pooling processing on each sub-matrix in the convolutional matrix to obtain the pooling results of each sub-matrix, and combines the pooling results of each sub-matrix to obtain a compressed matrix, that is, a pooling matrix. This process can improve the inference efficiency of the subsequent semantic recognition model.
[0093] In some embodiments, the maximum pooling formula is specifically:
[0094] z j = max i y ij ;
[0095] In the formula, y ij represents the data at the i-th position in the j-th sub-matrix of the convolutional matrix, and z j is the maximum pooling result of the j-th sub-matrix, and the maximum pooling results of each sub-matrix form a pooling matrix.
[0096] Step S402: Construct a position encoding matrix, and fuse the pooling matrix and the position encoding matrix to obtain a position information matrix;
[0097] In some embodiments, after obtaining the pooling matrix through a convolutional neural network, the pooling matrix is then processed by a Transformer. The Transformer consists of two main parts: an encoder and a decoder. Only the encoding part of the Transformer is used in this step. The purpose of this step is to further construct a matrix containing semantics based on the pooling matrix, so as to facilitate the semantic recognition model to map the first text feature into a first text vector.
[0098] In some embodiments, a positional encoding matrix is constructed. Specifically, when the output of the convolutional neural network, that is, the pooling matrix, is input into the Transformer, a positional encoding (PE) matrix of the same size as the pooling matrix will be constructed.
[0099] In some embodiments, the formula for constructing the positional encoding matrix is:
[0100]
[0101] In the formula, pos represents the position of the first dimension of the input matrix (i.e., the position of the row vector); d represents the size of the second dimension of the matrix (i.e., the size of the row vector); 2i and 2i + 1 respectively represent the even and odd positions of the data points in the row vector (2i, 2i + 1 ≤ d).
[0102] In some embodiments, the output of the convolutional neural network, that is, the pooling matrix and the positional encoding matrix, are fused to obtain a position information matrix, so as to introduce position information in text processing. The fusion method can be but is not limited to simple addition, concatenation, or more complex transformations.
[0103] Step S403: Extract the context features from the position information matrix to obtain an attention matrix, and determine the corresponding first text vector based on the attention matrix.
[0104] In some embodiments, extracting the context features from the position information matrix to obtain an attention matrix is specifically as follows: When the position information matrix is obtained, it is input into the Transformer encoder again. Since the Transformer encoder is composed of multiple identical layers stacked together, each layer is mainly composed of a multi-head attention mechanism and a feed-forward network. Therefore, when one layer of the Transformer encoder is completed, the encoded data after the operation is then input into multiple layers in sequence to obtain an output matrix with position information and context information. That is, after the multi-head attention mechanism receives the position information matrix, it will convert it into three matrices: query (Q), key (K), and value (V). Then, multiple self-attention mechanisms extract the context correlation features in the text in sequence to output the attention matrix.
[0105] In some embodiments, the relevant formula of the self-attention mechanism is as follows:
[0106]
[0107] In the formula, Q, K, and V are three matrices of query (Q), key (K), and value (V). The sizes of the three matrices are the same as those of the input matrix, that is, the position information matrix; d is the size of the second dimension of the input matrix, that is, the position information matrix.
[0108] In some embodiments, determining the corresponding first text vector based on the attention matrix specifically includes: expanding the attention matrix and outputting, by a fully connected layer, an email text vector with features such as text association information, position information, and context information, that is, the first text vector.
[0109] It should be noted that the convolutional neural network adopted by the multi-language semantic model of emails in the present invention includes two convolutional layers and one pooling layer. The encoder of the Transformer adopts a double-layer structure, and the fewer number of layers enables faster training and inference.
[0110] In this way, by determining the corresponding first text vector based on the text matrix, multi-language texts with similar semantics can be mapped into similar first text vectors, and then the semantic features of multi-languages can be extracted, which is convenient for subsequent spam detection.
[0111] In some embodiments, the semantic recognition model is obtained by training based on a first email training set including different languages, including: obtaining initial email data, and translating the third text features of the initial email data to obtain the first email training set including different languages, where the third text features include a first subject feature and a first body feature; splicing all the second subject features and second body features in the first email training set into a fourth text feature, encoding the fourth text feature to obtain encoded information, and training an initial semantic recognition model based on the encoded information; determining paired emails in the first email training set, inputting the paired emails into the initial semantic recognition model to obtain a corresponding vector group; calculating the similarity of each vector in the vector group to obtain a similarity calculation result, and optimizing the initial semantic recognition model based on the similarity calculation result to determine the semantic recognition model. The schematic diagram of the training process of the multi-semantic recognition model is as Figure 5 shown.
[0112] It should be noted that the training of the semantic recognition model includes two methods - masked language model and contrastive learning.
[0113] In some embodiments, initial email data is obtained, and the third text feature of the initial email data is translated to obtain the first email training set containing different languages. Specifically, first, a number of initial email data containing Chinese and English are collected, and the third text feature in the initial email data is extracted, where the third text feature includes a first subject feature and a first body feature. Second, the third text feature is translated through a large language model, that is, the email data is translated into email texts in different languages, and the email texts in multiple language versions are combined to obtain the first email training set containing different languages.
[0114] It should be noted that when the third text feature is extracted, duplicate removal is required according to the first subject feature and the first body feature of the email data. Then, by setting the unicode encoding range of common Chinese characters, dirty data (such as non-Chinese characters) in the email data is filtered out. Then, the fasttext language detection model is used to detect the email data whose text is Chinese or English, and the collected data is stored in the text database to obtain the third text feature.
[0115] It should be noted that the translation of the third text feature through a large language model (Large Language Model, LLM) is specifically as follows: First, a translation prompt template is set. Since it is necessary to maintain the corresponding email specifications in different languages during translation, when designing the prompt, it is necessary to make the large model maintain the style and language specifications of the email during translation. The designed prompt is roughly as follows: "You are an expert in the field of emails proficient in various languages. Now you need to translate the subject and body of the email into {lang}. Please keep the translated email in line with the email specifications. Your translation result must only include the subject and the body. The first line is the translated subject, and the following lines are the translated body. The email is as follows: Subject: {subject}, Body: {content}" where {lang} in the prompt is a language parameter, and its value can be "Japanese", "Korean", etc., {subject} is the email subject, and {content} is the email body.
[0116] It should be noted that when the large model performs translation, if the email data is in Chinese, it needs to be translated into languages such as English, the Indic language family, the Iranian language family, other Germanic language families except English, the Romance language family, or the Slavic language family. If the email data is in English, it needs to be translated into languages such as Chinese, the Indic language family, the Iranian language family, other Germanic language families except English, the Romance language family, or the Slavic language family. Designing the prompt to construct email texts in multiple languages with Chinese and English email texts can improve the problem of the lack of multi-language training data in the field of emails.
[0117] In some embodiments, all the second topic features and second body features in the first email training set are concatenated into a fourth text feature, and the fourth text feature is encoded to obtain encoded information, and an initial semantic recognition model is trained based on the encoded information. Specifically: First, read the first email training set, and concatenate the second topic features and second body features in the first email training set to obtain a long text, that is, the fourth text feature. Subsequently, a sub-word segmentation strategy is used to encode the vocabulary in the fourth text feature to obtain multiple tokens, that is, the encoded information. Part of the tokens in the fourth text feature are replaced with masks, and the original replaced tokens are saved. The fourth text feature tokens and the replaced tokens are input into a multilingual semantic recognition model. During training, a cross-entropy loss function is used to guide the model to learn. The context information of the fourth text feature tokens is used in the form of multi-classification to predict the replaced tokens, so as to train the initial semantic recognition model. At this time, the initial semantic recognition model has the ability of text understanding and can output vectors with preliminary text representativeness.
[0118] In some embodiments, paired emails in the first email training set are determined, and the paired emails are input into the initial semantic recognition model to obtain a corresponding vector group. Specifically: First, read the first email training set, and randomly pair every two emails. If the semantics of the two emails are the same, they are marked as 1. If the semantics of the two emails are different, they are marked as 0. Subsequently, the paired two emails, that is, the paired email boxes, are simultaneously input into the initial semantic recognition model to obtain a corresponding vector group, that is, two different vectors, and they are converted into unit vectors.
[0119] In some embodiments, the similarity of each vector in the vector group is calculated to obtain a similarity calculation result, and the initial semantic recognition model is optimized based on the similarity calculation result to determine a semantic recognition model. Specifically: Calculate the dot product of the two output unit vectors to measure the similarity of the two vectors. Finally, a cross-entropy loss function is used to guide binary classification training, so that text vectors with similar semantics are as close as possible in the feature space, while text vectors with different semantics are as far away as possible. Through the training of contrastive learning, the semantic recognition model has the ability to recognize semantics in different languages and can map text with similar semantics to similar vectors.
[0120] In this way, by constructing a first email training set containing different languages for training the semantic recognition model, and through the contrastive learning of paired emails to train the semantic recognition model, it can be made that the semantic recognition model can make the vectors mapped by emails with different languages but the same semantics as close as possible.
[0121] It should be noted that the first or the second does not represent the order of precedence. It is only used to distinguish the same nouns from different sources, and the first or the second can be understood as a name. It can be understood that the first email training set is different from the second email training set. Among them, the first email training set includes email data in multiple types of languages and is used to train the semantic recognition model, enabling the semantic recognition model to support email texts in multiple languages. The second email training set only contains two basic languages, Chinese and English, and is used to train the spam detection model.
[0122] In this way, by combining the convolutional neural network and the Transformer with fewer layers, the semantic recognition model can ensure better extraction of semantic information while improving the training and inference efficiency. At the same time, a constructed multi-language email dataset is adopted and trained by the contrastive learning method of the masked language model, so that the vectors mapped by emails with different languages but the same semantics are as close as possible.
[0123] Step S103: Input the first text vector into a preset spam detection model to determine the probability value that the email data belongs to spam based on the first text vector, and determine spam based on the probability value. Among them, the spam detection model is obtained by training based on the multi-dimensional features and the second text vector in the second email training set. The second text vector is determined by the semantic recognition model through reasoning on the second text features in the second email training set. The multi-dimensional features include the second text features.
[0124] In some embodiments, the spam detection model is composed of structures such as an embedding model, a convolutional neural network, and a recurrent neural network. Before building the model, multi-dimensional features such as text features, attachment features, reputation features, and metadata features are constructed from the second email training set. Then, corresponding network structures are used for these multi-dimensional features to build the model. Next, the vectors output by each feature are fused by the second-order cross-feature method. Finally, the fully connected layer maps the fused vector to the detection probabilities of normal emails and spam, thereby obtaining the spam detection result.
[0125] In some embodiments, the spam detection model includes an embedding layer, a fully-connected layer, and an activation function. Determining the probability value that the email data belongs to spam based on the first text vector is specifically as follows: After receiving the first text vector, the embedding layer converts the first text vector into a corresponding feature vector; the fully-connected layer transforms the feature vector according to preset model parameters to obtain a prediction result; the activation function determines the probability value of spam based on the prediction result. Specifically, first, when the spam detection model receives the first text vector, it will convert the first text vector into a corresponding low-dimensional and dense feature vector through the embedding layer to capture the semantic relationship between words. Subsequently, the fully-connected layer performs a linear transformation on the feature vector output by the embedding layer according to the preset model parameters to obtain a prediction result. Finally, the activation function converts the prediction result output by the fully-connected layer into a probability value, which helps to map the linear prediction result into the interval [0, 1], thereby representing the probability that the email belongs to spam.
[0126] In some embodiments, the preset model parameters include a weight matrix and a bias vector, where each row of the weight matrix corresponds to the connection weight between an element in the feature vector and an element in the prediction result. Multiply the feature vector by the weight matrix and add the bias vector to obtain the prediction result.
[0127] It should be noted that the Sigmoid function is usually selected as the activation function of the output layer.
[0128] In this way, by determining the probability value of spam, spam can be detected accurately and quickly.
[0129] In some embodiments, determining spam based on the probability value is specifically as follows. After determining the probability value, a threshold can be set. If the probability value exceeds this threshold, the email is marked as spam; otherwise, it is regarded as a normal email.
[0130] In some embodiments, the spam detection model is obtained by training based on multi-dimensional features and second text vectors in a second email training set. Specifically: Obtain a second email training set containing Chinese and English email data, and extract the multi-dimensional features in the second email training set. Among them, the multi-dimensional features include second text features, attachment features, reputation features, and metadata features; Input the second text features into the semantic recognition model to obtain second text vectors; Input the second text vectors, the attachment features, the reputation features, and the metadata features into the spam detection model to respectively obtain corresponding third text vectors, attachment vectors, reputation vectors, and metadata vectors; Fuse the third text vectors, the attachment vectors, the reputation vectors, and the metadata vectors to obtain a feature interaction vector, and determine an output matrix based on the feature interaction vector; Train the spam detection model based on the output matrix. The schematic diagram of the training process of the spam detection model is as shown in Figure 6 shown.
[0131] In some embodiments, to obtain a second email training set containing Chinese and English email data and extract the multi-dimensional features in the second email training set, specifically: First, collect a number of Chinese email data and English email data and integrate them into a second email training set. Second, extract the subject content and body content from the email bodies in the second email training set to obtain second text features. Since some phishing emails or ordinary spam emails may attach some executable files or web page files, and some emails for normal communication may attach text files, the attachment type is also a feature that can assist in distinguishing the email type. Therefore, it is also necessary to check whether the email contains attachments. If it contains attachments, extract features such as the type, size, and whether there is malicious code of the attachments to obtain attachment features. Evaluate the reputation of the sender based on the sender's historical behavior (such as the frequency of sending spam emails, the number of reports, etc.) to obtain reputation features. Finally, extract the metadata of the email, such as the sender address, recipient address, sending time, email subject, etc., to obtain metadata features. It should be noted that the extraction order of these features can be adjusted, and this application does not make any restrictions.
[0132] It should be noted that the reputation features include domain name reputation and IP reputation. The higher the reputation value, the higher the normal degree of the email. If the sender and recipient have communicated with each other in a certain period of time, the sending and receiving domain names and IPs will increase their reputations. Count the number of days when the domain name, sub-domain name, and IP address appear respectively for the emails that have communicated with each other in a certain period of time as the reputation value.
[0133] It should be noted that the metadata features include the html text of the email and the text information of the email sending client. Although these texts do not belong to natural language, they can also represent specific meanings. For example, if some spam emails are sent using automated tools, the email sending client may contain information about the tool.
[0134] In some embodiments, the second text feature is input into the semantic recognition model to obtain a second text vector. Specifically, after the second text feature is extracted from the second email training set, it is input into the trained semantic recognition model to obtain the second text vector corresponding to the second text feature through the semantic recognition model.
[0135] In some embodiments, the second text vector, the attachment feature, the reputation feature, and the metadata feature are input into the spam detection model to respectively obtain corresponding third text vectors, attachment vectors, reputation vectors, and metadata vectors. Specifically, after the second text vector, the attachment feature, the reputation feature, and the metadata feature are input into the spam detection model, the spam detection model will respectively use corresponding network structures to process these features to respectively obtain corresponding third text vectors, attachment vectors, reputation vectors, and metadata vectors.
[0136] It should be noted that the text feature network structure is composed of fully connected layers, and the third text vector can be obtained by directly inputting the second text vector into the network structure of the spam detection model. The attachment feature network structure is composed of an embedding layer and fully connected layers. Multiple attachment types in an email can be directly input into the embedding layer to obtain an output matrix, and the output matrix is then unfolded and input into the fully connected layer to obtain the attachment vector. The reputation feature network is only constructed by fully connected layers and activation functions, and the reputation feature can be directly input into the network to obtain the reputation vector. The metadata feature network is composed of an embedding layer, a convolutional neural network, a recurrent neural network, and fully connected layers. Before the metadata text is input into this network, the text is split into multiple tokens according to text punctuation or other delimiters, and the tokens are numerically encoded and then input into the embedding. At this time, each token will be mapped to a vector, and each vector is combined to obtain a metadata matrix; then, the metadata matrix is sequentially output to the convolutional neural network and the recurrent neural network, where the convolutional neural network compresses the metadata matrix and extracts the correlation information of adjacent data in the text; then, the recurrent neural network is used to receive the data output by the convolutional neural network to extract the context information of the text; finally, the fully connected layer outputs the metadata vector.
[0137] In some embodiments, the third text vector, the attachment vector, the reputation vector, and the metadata vector are fused to obtain a feature interaction vector, and an output matrix is determined based on the feature interaction vector. Specifically, after obtaining the third text vector, the attachment vector, the reputation vector, and the metadata vector, the four vectors need to be fused to obtain a feature interaction vector, and then the output matrix is determined through the operation of the second-order feature interaction formula.
[0138] In some embodiments, the calculation formula for determining the output matrix based on the feature interaction vector is specifically:
[0139] y = b + wx T + xVV T x T ;
[0140] In the formula, y is an output matrix of size M×M; b is a bias; w is a weight; x is a matrix formed by stacking the feature interaction vectors with a size of M×N, where M is the number of features and N is the vector length; V is a feature interaction vector of length N; T is the transpose.
[0141] In some embodiments, the spam detection model is trained based on the output matrix. Specifically, the cross-entropy loss function is used to guide the training of binary classification to adjust the parameters of the spam detection model.
[0142] In this way, the spam detection model combines lightweight neural network structures such as convolutional neural networks and recurrent neural networks to ensure efficient training and inference efficiency. At the same time, through various email features and multi-language text vectors output by the semantic recognition model for training, the model can achieve multi-language email detection.
[0143] In the embodiments of the present application, by obtaining the email data to be detected, it is convenient to input it into the semantic recognition model later to deeply understand the content corresponding to the email data; by inputting the email data into the preset semantic recognition model, the content corresponding to the email data in different languages can be accurately and quickly understood to identify the text features of different language emails and convert them into text matrices; based on the text matrix, the corresponding first text vector can be determined, and multi-language texts with similar semantics can be mapped into similar first text vectors, thereby extracting multi-language semantic features, which is convenient for subsequent spam detection; by pre-training the spam detection model with multi-dimensional features in advance, the trained model can quickly and accurately perform spam detection based on the first text vector. Compared with the prior art, the present application can improve the accuracy and efficiency of multi-language spam detection.
[0144] Embodiment 2
[0145] Please refer toFigure 3 , Figure 3 is a schematic structural diagram of an embodiment of the multi - language spam detection system provided by this application, including: an acquisition module 100, an identification module 200, and a detection module 300;
[0146] The acquisition module 100 is used to acquire the mail data to be detected;
[0147] The identification module 200 is used to input the mail data into a preset semantic recognition model to convert the first text features extracted from the mail data into a text matrix, and determine the corresponding first text vector based on the text matrix, where the semantic recognition model is obtained by training according to a first mail training set containing different languages;
[0148] The detection module 300 is used to input the first text vector into a preset spam detection model to determine the probability value that the mail data belongs to spam based on the first text vector, and determine spam based on the probability value, where the spam detection model is obtained by training according to the multi - dimensional features and the second text vector in the second mail training set, the second text vector is inferred and determined by the semantic recognition model for the second text features in the second mail training set, and the multi - dimensional features include the second text features.
[0149] Regarding the information interaction, execution process, etc. among the modules in the above - mentioned multi - language spam detection system, since they are based on the same concept as the embodiment of the multi - language spam detection method in the first aspect of the present invention, the achieved technical effects are basically the same. For specific content, refer to the description in the first embodiment of the method of the present invention, and details will not be repeated here.
[0150] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the method of this embodiment.
[0151] Please refer to Figure 8 , Figure 8 which schematically shows the hardware structure of a terminal device in another embodiment. The terminal device includes:
[0152] The processor 801 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0153] The memory 802 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 802 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 802 and are called by the processor 801 to execute the large model-based conversation risk assessment method of the embodiments of the present application;
[0154] The input / output interface 803 is used to implement information input and output;
[0155] The communication interface 804 is used to implement communication interaction between this device and other devices, and can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0156] The bus 805 transmits information between the various components of the device (such as the processor 801, the memory 802, the input / output interface 803, and the communication interface 804);
[0157] Among them, the processor 801, the memory 802, the input / output interface 803, and the communication interface 804 achieve communication connections with each other inside the device through the bus 805.
[0158] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the multi-language spam detection method as described in Embodiment 1 above.
[0159] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above various methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0160] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above description is only for the specific embodiments of the present application and is not used to limit the protection scope of the present application.
[0161] It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A method for detecting multilingual spam, characterized in that: include: Get the email data to be tested; Inputting the email data into a preset semantic recognition model to convert the first text feature extracted from the email data into a text matrix, and determining a corresponding first text vector based on the text matrix, wherein the semantic recognition model is obtained by training based on a first email training set containing different languages; The first text vector is input into a preset spam detection model to determine a probability value that the email data belongs to spam based on the first text vector, and to determine spam based on the probability value, wherein the spam detection model is obtained by training based on multi-dimensional features and a second text vector in a second email training set, and the second text vector is determined by inferring a second text feature in the second email training set based on the semantic recognition model, and the multi-dimensional feature includes the second text feature.
2. The multilingual spam detection method according to claim 1, characterized in that: The first text feature extracted from the email data is converted into a text matrix, specifically: Performing subword segmentation on the email data to obtain a plurality of first text features; Each of the first text features is vectorized to obtain a vectorization result, and a corresponding text matrix is determined based on the vectorization result.
3. The multilingual spam detection method according to claim 1, characterized in that: The determining of the corresponding first text vector based on the text matrix is specifically: Performing convolution processing on adjacent word vectors in the text matrix to obtain a convolution matrix, and performing pooling processing on the convolution matrix to obtain a pooling matrix; Constructing a position coding matrix, and fusing the pooling matrix and the position coding matrix to obtain a position information matrix; Context features in the position information matrix are extracted to obtain an attention matrix, and a corresponding first text vector is determined based on the attention matrix.
4. The multilingual spam detection method according to claim 3, characterized in that: The calculation formula of the convolution processing is specifically: y j =f(∑ i w i x ij +b); In the formula, y j The convolution result of the jth submatrix, the y value of each submatrix constitutes the convolution matrix; f is the activation function; w i Represents the weight of the i-th position in the convolution kernel; x ij is the position of the ith data in the jth submatrix in the text matrix; b is the bias.
5. The multilingual spam detection method according to claim 1, characterized in that: The semantic recognition model is obtained by training a first email training set containing different languages, specifically: Acquire initial email data, and translate the third text feature of the initial email data to obtain the first email training set containing different languages, wherein the third text feature includes a first subject feature and a first body feature; splicing all the second subject features and the second text features in the first email training set into a fourth text feature, encoding the fourth text feature to obtain encoding information, and training an initial semantic recognition model based on the encoding information; Determine paired emails in the first email training set, and input the paired emails into the initial semantic recognition model to obtain corresponding vector groups; A similarity calculation is performed on each vector in the vector group to obtain a similarity calculation result, and the initial semantic recognition model is optimized based on the similarity calculation result to determine a semantic recognition model.
6. The multilingual spam detection method according to claim 1, characterized in that: The spam detection model includes an embedding layer, a fully connected layer, and an activation function. The probability value of the email data belonging to spam is determined based on the first text vector, specifically: After receiving the first text vector, the embedding layer converts the first text vector into a corresponding feature vector; The fully connected layer transforms the feature vector according to preset model parameters to obtain a prediction result; The activation function determines a probability value of spam based on the prediction result.
7. The multilingual spam detection method according to claim 1, characterized in that: The spam detection model is obtained by training based on the multi-dimensional features and the second text vector in the second email training set, specifically: Obtaining a second email training set containing Chinese and English email data, and extracting multi-dimensional features from the second email training set, wherein the multi-dimensional features include a second text feature, an attachment feature, a reputation feature, and a metadata feature; Inputting the second text feature into the semantic recognition model to obtain a second text vector; Inputting the second text vector, the attachment feature, the reputation feature, and the metadata feature into the spam detection model to obtain corresponding third text vectors, attachment vectors, reputation vectors, and metadata vectors, respectively; fusing the third text vector, the attachment vector, the reputation vector and the metadata vector to obtain a feature interaction vector, and determining an output matrix based on the feature interaction vector; The spam detection model is trained based on the output matrix.
8. The multilingual spam detection method according to claim 1, characterized in that: The calculation formula for determining the output matrix based on the feature interaction vector is specifically: y=b+wx T +xVV T x T ; Where y is an output matrix of size M×M; b is a bias; w is a weight; x is the feature interaction vector stacked into a matrix of size M×N, where M is the number of features and N is the vector length; V is a feature interaction vector of length N; and T is the transpose.
9. A multi-language spam detection system, characterized in that: include: Acquisition module, recognition module and detection module; The acquisition module is used to acquire the mail data to be detected; The recognition module is used to input the email data into a preset semantic recognition model to convert the first text feature extracted from the email data into a text matrix, and determine the corresponding first text vector based on the text matrix, wherein the semantic recognition model is obtained by training based on a first email training set containing different languages; The detection module is used to input the first text vector into a preset spam detection model to determine the probability value of the email data belonging to spam based on the first text vector, and determine spam based on the probability value, wherein the spam detection model is obtained by training based on the multi-dimensional features and the second text vector in the second email training set, the second text vector is determined by inferring the second text feature in the second email training set based on the semantic recognition model, and the multi-dimensional features include the second text feature.
10. A terminal device, characterized in that: include: one or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the multilingual spam detection method according to any one of claims 1 to 8.
Citation Information
Cited By
Policy game and large language model-based phishing mail detection method, apparatus and device, and medium
CN121056235A
Reusable mail block generation and suggestion
US20260113294A1