Malicious domain name generation method and apparatus, and device and medium
By encoding and extracting the semantics of the initial malicious domain name, and combining Gaussian mixture classification and domain name generation network, potential malicious domain names with strong relevance to the initial malicious domain name are generated. This solves the problems of timeliness and poor relevance in existing technologies and improves the quality of malicious domain names.
Patent Information
- Application Number
- PCT/CN2024/105188
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-10
- Filing Date
- 2024-07-12
- Publication Date
- 2025-11-13
AI Technical Summary
Existing technologies suffer from poor timeliness in mining potential malicious domains, and the generated potential malicious domains have poor correlation with the original malicious domains, resulting in poor quality of the mined potential malicious domains.
By character encoding the initial malicious domain name, extracting semantic features using a pre-built attention mechanism, determining the target category based on the cluster center of Gaussian mixture categories, and generating potential malicious domain names in conjunction with the domain name generation network, the Gaussian mixture categories are clustered and generated using updated Gaussian distribution parameters.
It improves the quality of potential malicious domain names, enhances the ability to identify and respond to them, and generates potential malicious domain names that are highly correlated with the initial malicious domain names, are timely, and do not rely on historical data.
Smart Images

Figure CN2024105188_13112025_PF_FP_ABST
Abstract
Description
Methods, devices, equipment, and media for generating malicious domain names Technical Field
[0001] This disclosure relates to the field of network security technology, and in particular to a method, apparatus, device, and medium for generating malicious domain names. Background Technology
[0002] Malicious domains are the domains of malicious websites, and detecting them is crucial for cybersecurity. Malicious domain discovery involves identifying more potential malicious domains based on existing ones. By uncovering more potential malicious domains, we can detect potential threats early and take timely defensive measures.
[0003] In related technologies, the discovery of malicious domain names often relies on the domain's inherent characteristics and historical data. This results in poor timeliness of the discovered potential malicious domain names, and the generated potential malicious domain names have poor correlation with the original malicious domain names, easily leading to the generation of essentially normal domain names. Therefore, the quality of the ultimately discovered potential malicious domain names is relatively poor.
[0004] Summary of the Invention
[0005] The main objective of this disclosure is to provide a method, apparatus, device, and medium for generating malicious domain names, which can improve the quality of the discovered potential malicious domain names.
[0006] To achieve the above objectives, a first aspect of this disclosure provides a method for generating malicious domain names, comprising:
[0007] Obtain the initial malicious domain name and encode each character in the initial malicious domain name to obtain the initial domain name characteristics;
[0008] Semantic features are obtained by semantically extracting the initial domain name features based on a pre-built attention mechanism;
[0009] Based on the cluster centers of multiple preset Gaussian mixture categories, the distribution probability of the initial malicious domain name belonging to different Gaussian mixture categories is determined, and the target category to which the initial malicious domain name belongs is determined from the multiple Gaussian mixture categories according to the magnitude relationship of each distribution probability.
[0010] The semantic features are input into the domain name generation network, and the clustering features of the target category are extracted. The clustering features are then used to guide the domain name generation network to generate potentially malicious domain names.
[0011] The Gaussian mixture category is obtained by updating the initial Gaussian mixture category according to the updated Gaussian distribution parameters. The updated Gaussian distribution parameters are obtained by first determining multiple initial Gaussian mixture categories for multiple sample domain name features, then calculating the posterior probability of each sample domain name feature belonging to each initial Gaussian mixture category, and updating the current Gaussian distribution parameters according to the posterior probability.
[0012] In some embodiments, the Gaussian mixture category is determined through the following steps:
[0013] Multiple sample domain name features are obtained, and Gaussian mixture clustering is performed on the multiple sample domain name features to obtain multiple initial Gaussian mixture categories;
[0014] Determine the current Gaussian distribution parameters for each initial Gaussian mixture category, and determine the probability density function of each sample domain name feature under each initial Gaussian mixture category, wherein the Gaussian distribution parameters include the mean vector, covariance matrix, and mixing coefficients;
[0015] Based on the current mean vector, covariance matrix, mixing coefficients, and probability density function, calculate the posterior probability that each sample domain name feature belongs to each initial Gaussian mixture category;
[0016] The current Gaussian distribution parameters are updated based on the posterior probability, and the initial Gaussian mixture category is updated based on the updated Gaussian distribution parameters to obtain the updated Gaussian mixture category.
[0017] In some embodiments, calculating the posterior probability of each sample domain name feature belonging to each initial Gaussian mixture category based on the current mean vector, the covariance matrix, the mixing coefficients, and the probability density function includes:
[0018] The feature x of each sample domain name is calculated according to the following formula. n The posterior probability γ(Z) of belonging to the k-th initial Gaussian mixture category nk ):
[0019] Wherein, N(x) in formula (1) n |μ k ,Σ k ) is the sample domain name feature x n The probability density function under the k-th initial Gaussian mixture category, π k It is the mixing coefficient of the k-th initial Gaussian mixture category, μ k It is the mean vector of the k-th initial Gaussian mixture class, Σk It is the covariance matrix of the k-th initial Gaussian mixture class;
[0020] Updating the current Gaussian distribution parameters based on the posterior probability includes:
[0021] According to the posterior probability γ(Z) nk The formula for updating the current Gaussian distribution parameters is as follows:
[0022] Among them, N in formulas (2) to (4) K is the total number of sample domain name features belonging to the k-th initial Gaussian mixture category, and N is the total number of sample domain name features.
[0023] In some embodiments, obtaining multiple sample domain name features includes:
[0024] Obtain multiple sample malicious domain names;
[0025] For each of the sample malicious domain names, extract the first character of the top-level domain portion and the second character of the subdomain portion of the sample malicious domain name;
[0026] The first character feature is extracted from the first character, and the second character feature is extracted from the second character;
[0027] The first character feature and the second character feature corresponding to each malicious domain name of the sample are combined to obtain multiple sample domain name features.
[0028] In some embodiments, extracting a first character feature from the first character and extracting a second character feature from the second character includes:
[0029] One-hot encoding is performed on the first character to obtain the feature of the first character;
[0030] The second character is identified by counting the substrings that meet a preset length. The second character is then encoded based on the frequency of occurrence of each substring under multiple malicious domain names in the sample, thus obtaining the second character feature.
[0031] In some embodiments, performing Gaussian mixture clustering on multiple sample domain name features to obtain multiple initial Gaussian mixture categories includes:
[0032] Determine the initial number of categories, and perform Gaussian mixture clustering on the features of multiple sample domain names under the number of categories to obtain multiple corresponding sub-Gaussian mixture categories;
[0033] Calculate the error value between the domain name features of each sample in each sub-Gaussian mixture category and the corresponding cluster center;
[0034] Gradually increase the number of categories, and re-perform Gaussian mixture clustering with the corresponding number of categories, and calculate the error value of each sub-Gaussian mixture category with the corresponding number of categories;
[0035] The minimum target error value is determined from the error values under different numbers of categories, and the number of categories corresponding to the target error value is determined as the target number of categories;
[0036] Gaussian mixture clustering is performed on the features of multiple sample domain names under the target number of categories to obtain multiple initial Gaussian mixture categories.
[0037] In some embodiments, the type of the covariance matrix is determined by the following steps:
[0038] Based on the Gaussian distribution parameters calculated from different covariance matrix types, Gaussian mixture clustering is performed on multiple sample domain name features to obtain multiple test Gaussian mixture categories;
[0039] Determine the clustering effect of the corresponding test Gaussian mixture categories under different covariance matrix types, and determine the covariance matrix as a diagonal covariance matrix based on the clustering effect.
[0040] In some embodiments, the semantic extraction of the initial domain name features based on a pre-built attention mechanism to obtain semantic features includes:
[0041] The initial domain name features are input into a sequentially connected multi-layer attention network;
[0042] In each layer of the attention network, the current input data is bidirectionally encoded through an attention mechanism to obtain a corresponding attention weight matrix. The hidden state is obtained and output based on the attention weight matrix and the current input data.
[0043] Semantic features are obtained through the output of the attention network in the last layer.
[0044] In some embodiments, the step of inputting the semantic features into the domain name generation network, extracting the clustering features of the target category, and using the clustering features to guide the domain name generation network to generate potentially malicious domain names includes:
[0045] Extract the clustering features of the target category;
[0046] The initial domain name features, the clustering features, and the semantic features are input into the domain name generation network. An attention mechanism based on the clustering features and the semantic features is used to diffuse the initial domain name features to generate potential malicious domain names.
[0047] In some embodiments, after guiding the domain name generation network to generate potentially malicious domain names using the clustering features, the method further includes:
[0048] Obtain the domain name to be tested;
[0049] Based on the potential malicious domain name, the domain name to be detected is subjected to malicious detection, and the corresponding domain name detection results are obtained.
[0050] To achieve the above objectives, a second aspect of this disclosure provides an apparatus for generating malicious domain names, comprising:
[0051] An encoding module is used to obtain an initial malicious domain name and encode each character in the initial malicious domain name to obtain the initial domain name characteristics;
[0052] The semantic extraction module is used to extract semantic features from the initial domain name features based on a pre-built attention mechanism to obtain semantic features;
[0053] The clustering module is used to determine the distribution probability of the initial malicious domain name belonging to different Gaussian mixture categories based on the cluster centers of multiple preset Gaussian mixture categories, and to determine the target category to which the initial malicious domain name belongs from the multiple Gaussian mixture categories based on the magnitude relationship of each distribution probability.
[0054] The domain name generation module is used to input the semantic features into the domain name generation network, extract the clustering features of the target category, and use the clustering features to guide the domain name generation network to generate potentially malicious domain names;
[0055] The Gaussian mixture category is obtained by updating the initial Gaussian mixture category according to the updated Gaussian distribution parameters. The updated Gaussian distribution parameters are obtained by first determining multiple initial Gaussian mixture categories for multiple sample domain name features, then calculating the posterior probability of each sample domain name feature belonging to each initial Gaussian mixture category, and updating the current Gaussian distribution parameters according to the posterior probability.
[0056] To achieve the above objectives, a third aspect of this disclosure provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the malicious domain name generation method described in the first aspect embodiment above.
[0057] To achieve the above objectives, a fourth aspect of this disclosure provides a storage medium, which is a computer-readable storage medium storing a computer program that, when executed by a processor, implements the malicious domain name generation method described in the first aspect of the embodiment.
[0058] The malicious domain name generation method, apparatus, device, and medium disclosed in this embodiment can be applied to a malicious domain name generation apparatus. By executing the malicious domain name generation method, each character in the initial malicious domain name is first encoded to obtain initial domain name features. Semantic extraction of the initial domain name features is performed based on a pre-built attention mechanism, which significantly improves the accuracy and efficiency of semantic extraction, enhancing the ability to identify and respond to malicious domain names, and obtaining semantic features. Next, based on the cluster centers of multiple preset Gaussian mixture categories, the distribution probability of the initial malicious domain name belonging to different Gaussian mixture categories is determined. The target category to which the initial malicious domain name belongs is determined from the multiple Gaussian mixture categories based on the magnitude of each distribution probability. The semantic features are input into a domain name generation network, and the clustering features of the target category are extracted. The Gaussian mixture category is obtained by updating the initial Gaussian mixture category based on the updated Gaussian distribution parameters. The updated Gaussian distribution parameters are obtained by first determining multiple initial Gaussian mixture categories for multiple sample domain name features, then calculating the posterior probability of each sample domain name feature belonging to each initial Gaussian mixture category, and updating the current Gaussian distribution parameters based on the posterior probability. In this way, each Gaussian mixture category can accurately cluster similar domain names, and the clustering features of the target category can well represent the data characteristics of the Gaussian distribution. Subsequently, the clustering features can be used to guide the domain name generation network to generate similar potential malicious domain names within the range indicated by the target category. The obtained potential malicious domain names have a stronger correlation with the initial malicious domain name, and since the potential malicious domain names are generated based on the currently input initial malicious domain name, they do not need to rely on historical data, so the timeliness is better. Therefore, the quality of the ultimately mined potential malicious domain names is better. Attached Figure Description
[0059] Figure 1 is a flowchart illustrating the method for generating malicious domain names provided in this embodiment of the present disclosure;
[0060] Figure 2 is a flowchart illustrating the process for determining Gaussian mixture categories provided in an embodiment of this disclosure;
[0061] Figure 3 is a flowchart illustrating the further components of step S201 in Figure 2;
[0062] Figure 4 is a flowchart illustrating the further components of step S303 in Figure 3;
[0063] Figure 5 is a schematic diagram of another process that further includes step S201 in Figure 2;
[0064] Figure 6 is a flowchart illustrating the process for determining the type of covariance matrix provided in an embodiment of this disclosure;
[0065] Figure 7 is a flowchart illustrating the further components of step S102 in Figure 1;
[0066] Figure 8 is a schematic diagram of the encoder provided in an embodiment of this disclosure;
[0067] Figure 9 is a flowchart illustrating the further components of step S104 in Figure 1;
[0068] Figure 10 is a schematic diagram of the domain name generation network training process provided in an embodiment of this disclosure;
[0069] Figure 11 is a flowchart illustrating the process following step S104 in Figure 1.
[0070] Figure 12 is a schematic diagram of the functional modules of the malicious domain name generation device provided in the embodiments of this disclosure;
[0071] Figure 13 is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0072] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this disclosure.
[0073] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used herein is for the purpose of describing embodiments of this disclosure only and is not intended to be limiting of this disclosure.
[0074] First, let's analyze some of the terms used in this disclosure:
[0075] Artificial Intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.
[0076] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0077] A Gaussian Mixture Model (GMM) is a clustering method based on a probability density function. It assumes that each cluster is a mixture of multiple Gaussian distributions. The goal of a GMM is to estimate the model parameters by maximizing the likelihood function, including the mean, variance, and mixing coefficients of each Gaussian distribution, as well as the probability that a data point belongs to each cluster.
[0078] In related technologies, the discovery of malicious domain names often relies on the domain's inherent characteristics and historical data. This results in poor timeliness of the discovered potential malicious domain names, insufficient speed in processing newly generated unknown malicious domain names, a small detection range, and an inability to accommodate and discover a large number of potential malicious domain names. Furthermore, the generated potential malicious domain names have poor correlation with the original malicious domain names, and are prone to generating essentially normal domain names. Therefore, the quality of the ultimately discovered potential malicious domain names is poor.
[0079] Furthermore, related technologies also rely on patterns in domain name character characteristics to identify legitimate and malicious domains. However, this approach is ineffective for detecting domains composed of English words, and the continuous emergence of new malicious website domain families leads to insufficient training data and low recognition rates.
[0080] Based on this, the present disclosure provides a method, apparatus, device, and medium for generating malicious domain names, which can improve the quality of the discovered potential malicious domain names.
[0081] The method for generating malicious domain names in this disclosure can be illustrated through the following embodiments.
[0082] The embodiments disclosed herein can acquire and process relevant data based on artificial intelligence technology.
[0083] The malicious domain name generation method provided in this disclosure can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the malicious domain name generation method, but is not limited to the above forms.
[0084] This disclosure can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This disclosure can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and so on that perform a specific task or implement a specific abstract data type. This disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0085] It should be noted that in all specific embodiments of this disclosure, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. For example, when obtaining domain name-related information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this disclosure require obtaining sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to a confirmation page. Only after obtaining the user's separate permission or consent is the necessary user-related data for the proper functioning of embodiments of this disclosure obtained.
[0086] Figure 1 is an optional flowchart of a method for generating malicious domain names provided in this disclosure embodiment. The method in Figure 1 may include, but is not limited to, steps S101 to S104.
[0087] Step S101: Obtain the initial malicious domain name and encode each character in the initial malicious domain name to obtain the initial domain name characteristics;
[0088] Step S102: Semantic features are extracted from the initial domain name features based on the pre-built attention mechanism to obtain semantic features;
[0089] Step S103: Based on the cluster centers of multiple preset Gaussian mixture categories, determine the distribution probability of the initial malicious domain name belonging to different Gaussian mixture categories, and determine the target category to which the initial malicious domain name belongs from multiple Gaussian mixture categories according to the magnitude relationship of each distribution probability.
[0090] Step S104: Input semantic features into the domain name generation network and extract clustering features of the target category. Use the clustering features to guide the domain name generation network to generate potential malicious domain names.
[0091] The Gaussian mixture category is obtained by updating the initial Gaussian mixture category based on the updated Gaussian distribution parameters. The updated Gaussian distribution parameters are obtained by first determining multiple initial Gaussian mixture categories for multiple sample domain name features, then calculating the posterior probability of each sample domain name feature belonging to each initial Gaussian mixture category, and updating the current Gaussian distribution parameters based on the posterior probability.
[0092] Regarding step S101 above, the initial malicious domain name is the domain name of any malicious website. This website exhibits certain malicious behaviors that may violate laws and regulations and impact network security. Therefore, the domain name of this website is defined as malicious. The initial malicious domain name is input data used to mine potential malicious domain names. Furthermore, the initial malicious domain name can come from an initial domain name dataset, which contains multiple malicious domain names, referred to as multiple data points. Therefore, at least one of these datasets can be selected as the initial malicious domain name for mining potential malicious domain names; that is, the initial malicious domain name can be a data point within the dataset. Therefore, this embodiment of the disclosure can perform mining based on multiple initial malicious domain names, without specific limitations.
[0093] It should be noted that the embodiments of this disclosure can collect relevant network traffic data and employ related packet capture technologies, such as the Data Plane Development Kit (DPDK). These embodiments can focus on the traffic of website information in certain specific malicious domains and construct a dataset focused on this through in-depth analysis. However, in the process of acquiring data, all relevant laws and ethical guidelines must be followed to ensure user privacy and data security. The resulting dataset is directly related to relevant network activities and therefore has direct practical applicability. It not only highlights the importance of fine-grained monitoring of network traffic but also provides a reliable data foundation for in-depth research on network security and identification of potential threats.
[0094] After obtaining the initial malicious domain name, it is necessary to encode each character in the initial malicious domain name to obtain the initial domain name characteristics. Exemplarily, in terms of encoding, this embodiment first constructs a character encoding table by combining all top-level domains appearing in the dataset, including the initial malicious domain name, as well as individual characters and null characters generated after splitting subdomains (such as second-level domains), and assigns a number to each character. Given that the top-level domains of most malicious website domains are unique and indivisible, they need to be treated as a whole during encoding. Then, the characters in the domain names are mapped one-to-one with the numbers in the table, resulting in initial domain name vectors of varying lengths. To ensure that all domain name vectors have a uniform length, null characters can be used to supplement them, ensuring that each vector has the same length and is composed of character encodings. Therefore, this embodiment of the disclosure, by encoding the domain names in the dataset, has advantages such as unified characteristics, preservation of key information, simplified processing, convenient comparison and matching, reduced data sparsity, and easy expansion. Finally, after encoding the current initial malicious domain name, the initial domain name characteristics can be obtained.
[0095] Regarding step S102 above, the attention mechanism is an important concept in deep learning, mimicking the attention mechanism in the human visual system. When processing large amounts of information, the human visual system selectively focuses on certain important information while ignoring others. Similarly, in deep learning, the attention mechanism enables the model to automatically learn and focus on the most task-relevant parts of the input data. However, in the context of malicious domain name detection, the initial domain name features are obtained by encoding the individual characters in the domain name. While these features contain basic information about the domain name, they may not directly reflect its semantic meaning or potential malice. Therefore, further processing is needed to extract more meaningful semantic features.
[0096] Based on this, the semantic extraction in this embodiment is performed using a pre-built attention mechanism. Essentially, it learns a weighting strategy based on the initial domain name features. This strategy assigns different weights to different characters or character combinations according to their importance in the domain name; for example, important characters or character combinations receive higher weights, thus playing a greater role in subsequent processing. The resulting semantic features not only contain basic domain name information but also incorporate the semantic relationships between characters or character combinations and their contribution to the domain's malice, providing richer and more meaningful information for subsequent processing and analysis, and helping to improve the ability to identify and respond to malicious domain names.
[0097] Furthermore, the attention mechanism can calculate weights using one or more neural network layers and apply these weights to the initial domain name features; this disclosure does not impose specific limitations. In this way, it is possible to learn which characters or character combinations are more relevant to malicious domain names, thereby extracting more representative semantic features.
[0098] Regarding step S103 above, in this embodiment of the disclosure, the preset multiple Gaussian mixture categories are different clusters of malicious domain name characteristics. These clusters can be pre-set based on historical data or expert knowledge, or they can be obtained by clustering different malicious domain names in the dataset. Therefore, each cluster center represents a specific malicious behavior pattern or feature set.
[0099] When given an initial malicious domain name, its feature vector, obtained after encoding and semantic extraction, is input into a Gaussian mixture model. The model calculates the probability that the feature vector belongs to each Gaussian mixture category. For example, the final distribution probability can be obtained by calculating the distance (such as Mahalanobis distance) or likelihood between the feature vector and each Gaussian distribution, and then combining the weights of each distribution. The specific method for calculating the distribution probability is not limited here.
[0100] Subsequently, based on the magnitude of these probability distributions, the Gaussian mixture category to which the initial malicious domain name most likely belongs, i.e., the target category, can be determined. Therefore, embodiments of this disclosure can cluster malicious domain names with similar characteristics or behavioral patterns together, and extract representative clustering features for each category. These clustering features not only help in understanding the commonalities and differences of malicious domain names, but also provide guidance for the generation of subsequent potential malicious domain names.
[0101] It should be noted that the Gaussian mixture model assumes that all data points are a mixture of a finite number of Gaussian distributions (i.e., normal distributions). In this embodiment, the Gaussian mixture category is a different clustering of domain name features. Each cluster represents a specific malicious behavior pattern or feature set. Therefore, iterative optimization is needed to gradually improve the clustering effect.
[0102] In refining the clustering results, the first step is to predetermine some initial Gaussian mixture categories based on the existing sample domain name features. Next, for each sample domain name feature, the probability of it belonging to each initial Gaussian mixture category is calculated. These probabilities, called posterior probabilities, reflect the likelihood of a sample feature belonging to a particular category. Based on the posterior probabilities, the parameters of each Gaussian distribution can be updated, including the mean vector and covariance matrix. For example, the mean of each distribution is updated to a weighted average of the sample features belonging to that distribution, and the covariance matrix is adjusted accordingly based on the distribution of the sample features. As the Gaussian distribution parameters are updated, the original Gaussian mixture categories also change. Some previously similar categories may merge, while some previously significantly different categories may further differentiate. After multiple iterations, the Gaussian mixture categories gradually stabilize, and each category more accurately reflects the distribution and clustering of domain name features.
[0103] Through the above process, Gaussian mixture categories are constructed and updated. This model not only considers the statistical distribution of domain name features but also gradually improves the accuracy and efficiency of clustering through iterative optimization. During clustering, Gaussian mixture clustering assigns data points to the clusters with the highest probabilities, rather than rigidly assigning data points to a specific cluster as in K-Means. Therefore, using such Gaussian mixture categories to guide the generation process of a domain name generation network can generate new domain names that are similar to the initial malicious domain name and possess potential maliciousness, thereby improving the accuracy and efficiency of malicious domain name detection.
[0104] Regarding step S104 above, the domain name generation network is a deep learning model that, after training, can learn the rules and patterns of domain name generation. During this process, the network attempts to generate new domain names based on the input semantic features. Therefore, semantic features can be input into the domain name generation network. Furthermore, this embodiment of the disclosure also needs to extract clustering features of the target category. These clustering features are obtained based on a Gaussian mixture model and represent the common features and patterns of domain names within the target category. By extracting clustering features, the characteristics of the target category can be grasped more accurately, providing guidance for subsequent domain name generation.
[0105] By leveraging clustering features, domain name generation networks can be guided to generate domain names within the scope indicated by the target category. Specifically, clustering features can be used as conditions or constraints for the domain name generation network, ensuring that the generated domain names match the characteristics of the target category. In this way, the network can generate new domain names that are similar to the initial malicious domain name and possess potential maliciousness based on the clustering features of the target category. Since the potentially malicious domain names are generated based on the currently input initial malicious domain name and guided by the clustering features of the target category, their correlation with the initial malicious domain name is stronger. Furthermore, by combining deep learning and clustering analysis techniques, this generation method does not rely on historical data but generates domain names in real time based on the features of the current input, thus offering better timeliness and improving the accuracy and efficiency of malicious domain name generation.
[0106] In summary, through steps S101 to S104, this embodiment of the present disclosure obtains initial domain name features by first encoding each character in the initial malicious domain name, and then extracting semantics from the initial domain name features based on a pre-built attention mechanism. This significantly improves the accuracy and efficiency of semantic extraction, enhances the ability to identify and respond to malicious domain names, and obtains semantic features. Next, based on the cluster centers of multiple preset Gaussian mixture categories, the distribution probability of the initial malicious domain name belonging to different Gaussian mixture categories is determined. Based on the magnitude relationship of each distribution probability, the target category to which the initial malicious domain name belongs is determined from the multiple Gaussian mixture categories. The semantic features are input into the domain name generation network, and the clustering features of the target category are extracted. The Gaussian mixture category is obtained by updating the initial Gaussian mixture category based on the updated Gaussian distribution parameters. The updated Gaussian distribution parameters are obtained by first determining multiple initial Gaussian mixture categories of multiple sample domain name features, then calculating the posterior probability of each sample domain name feature belonging to each initial Gaussian mixture category, and updating the current Gaussian distribution parameters based on the posterior probability. In this way, each Gaussian mixture category can accurately cluster similar domain names, and the clustering features of the target category can well represent the data characteristics of the Gaussian distribution. Subsequently, the clustering features can be used to guide the domain name generation network to generate similar potential malicious domain names within the range indicated by the target category. The obtained potential malicious domain names have a stronger correlation with the initial malicious domain name, and since the potential malicious domain names are generated based on the currently input initial malicious domain name, they do not need to rely on historical data, so the timeliness is better. Therefore, the quality of the ultimately mined potential malicious domain names is better.
[0107] Referring to Figure 2, in some embodiments, the Gaussian mixture category is determined through the following steps, which may include steps S201 to S204:
[0108] Step S201: Obtain multiple sample domain name features and perform Gaussian mixture clustering on the multiple sample domain name features to obtain multiple initial Gaussian mixture categories;
[0109] Step S202: Determine the current Gaussian distribution parameters of each initial Gaussian mixture category, and determine the probability density function of each sample domain name feature under each initial Gaussian mixture category, wherein the Gaussian distribution parameters include the mean vector, covariance matrix and mixing coefficients;
[0110] Step S203: Calculate the posterior probability of each sample domain name feature belonging to each initial Gaussian mixture category based on the current mean vector, covariance matrix, mixing coefficients, and probability density function.
[0111] Step S204: Update the current Gaussian distribution parameters according to the posterior probability, and update the initial Gaussian mixture category according to the updated Gaussian distribution parameters to obtain the updated Gaussian mixture category.
[0112] Steps S201 to S204 above represent a series of key steps in determining and updating Gaussian mixture categories using the Gaussian mixture model. The purpose of these steps is to obtain a Gaussian mixture model that accurately reflects the data distribution by clustering the features of sample domain names, so that it can be used subsequently in the malicious domain name generation process.
[0113] Specifically, in this embodiment, multiple sample domain name features are first collected. These features can be used as data points for clustering, and can be obtained by encoding other domain names in the initial domain name dataset containing the initial malicious domain. Then, Gaussian mixture clustering is performed on the multiple sample domain name features to group similar features together, forming initial Gaussian mixture categories.
[0114] Next, the Gaussian distribution parameters include the mean vector, covariance matrix, and mixing coefficients. The mean vector represents the center position of the Gaussian distribution, the covariance matrix describes the dispersion of the data, and the mixing coefficients determine the weight of each Gaussian distribution in the mixture model. Therefore, embodiments of this disclosure need to determine these parameters for each initial Gaussian mixture category and calculate the probability density function of each sample domain name feature under these categories, where the probability density function describes the probability that the sample feature belongs to a certain Gaussian distribution.
[0115] Subsequently, based on the current mean vector, covariance matrix, mixing coefficients, and probability density function, the posterior probability of each sample domain name feature belonging to each initial Gaussian mixture category is calculated. The posterior probability refers to the probability that a sample feature belongs to a certain Gaussian mixture category given the sample features and other information. This step is achieved by calculating the probability density of the sample feature under each Gaussian distribution and combining it with the mixing coefficients to obtain the posterior probability of the sample feature belonging to each initial Gaussian mixture category.
[0116] Finally, the current Gaussian distribution parameters are updated based on the posterior probability, and the initial Gaussian mixture class is updated based on the updated Gaussian distribution parameters, resulting in the updated Gaussian mixture class. In this step, the calculated posterior probability is used to update the Gaussian distribution parameters. Specifically, the mean vector, covariance matrix, and mixing coefficients are updated by maximizing the likelihood function of the data or minimizing a certain loss function. As the parameters are updated, the initial Gaussian mixture class also changes, gradually approaching the true data distribution. After multiple iterations, the obtained Gaussian mixture class can more accurately reflect the clustering of the sample domain name features.
[0117] In summary, through this series of steps, the Gaussian mixture category is determined and updated, providing a foundation for its subsequent application in malicious domain name generation. This Gaussian mixture model can capture the complex distribution of domain name features, improving the accuracy and efficiency of malicious domain name detection and generation.
[0118] Furthermore, unlike the K-means algorithm, the GMM algorithm does not require pre-specifying the number of clusters; instead, it discovers clusters based on the probability density between data points. It assumes the data is a mixture of multiple Gaussian distributions and uses the Expectation-Maximization (EM) algorithm to estimate the parameters of these distributions. The specific process is as follows:
[0119] Initialization steps: First, randomly select a set of initial Gaussian distribution parameters, including the mean vector (μ) and covariance matrix (Σ). Then, calculate the probability that each data point belongs to each distribution. This process can be calculated using the probability density function (PDF) of the Gaussian distribution; in some embodiments, the probability density function can be used as the distribution probability in the above embodiments. The formula is shown below:
[0120] In formula (5), x represents the feature vector of the data point, such as the domain name feature of the sample, μ represents the cluster center, Σ is the covariance matrix of the Gaussian mixture categories, d is the dimension of the data point, and |Σ| represents the determinant of the covariance matrix. By calculating the probability that each data point belongs to each distribution, it can be used for subsequent expectation steps.
[0121] The expected step (E-step): In some embodiments, the posterior probability of each sample domain name feature belonging to each initial Gaussian mixture class is calculated based on the current mean vector, covariance matrix, mixing coefficients, and probability density function, including:
[0122] The domain name feature x for each sample is calculated using the following formula. n The posterior probability γ(Z) of belonging to the k-th initial Gaussian mixture class nk ):
[0123] Wherein, N(x) in formula (1) n |μ k ,Σ k ) is the feature x of the sample domain name n The probability density function under the k-th initial Gaussian mixture category, π k It is the mixing coefficient of the k-th initial Gaussian mixture class, μ k Σk is the mean vector of the k-th initial Gaussian mixture class, Σk is the covariance matrix of the k-th initial Gaussian mixture class, and the denominator is the sum of the mixture probabilities of all Gaussian distributions for the data points. Calculate the posterior probability γ(Z). nk This will result in a matrix where rows represent data points, columns represent Gaussian distributions, and each element represents the posterior probability that a data point belongs to the corresponding Gaussian distribution.
[0124] The maximization step (M-step) involves updating the Gaussian distribution parameters based on the posterior probabilities calculated in the expectation step. Specifically, the updated Gaussian distribution parameters include the mean μk, the covariance matrix Σk, and the mixing coefficients πk.
[0125] In some embodiments, updating the current Gaussian distribution parameters based on the posterior probability includes:
[0126] Based on the posterior probability γ(Z) nk The formula for updating the current Gaussian distribution parameters is as follows:
[0127] Among them, N in formulas (2) to (4) K is the total number of sample domain name features belonging to the k-th initial Gaussian mixture category, and N is the total number of sample domain name features. This is achieved by utilizing the posterior probability γ(Z). nk We update each Gaussian distribution parameter by weighted averaging of the data points to maximize the log-likelihood function.
[0128] In the expectation-maximization process (expectation step and maximization step), the expectation step calculates the posterior probability of each data point belonging to each distribution. This is done by calculating the probability of each data point on each distribution and normalizing it to obtain the posterior probability. In the maximization step, the model parameters are updated based on the posterior probabilities calculated in the expectation step, including updating parameters such as the distribution's mean vector, covariance matrix, and mixing coefficients, to maximize the log-likelihood function. The expectation and maximization steps are repeated until the parameters converge or the maximum number of iterations is reached, thus obtaining the optimal parameter estimates of the model and achieving clustering or parameter estimation of the data. This method assigns data points to each category with a certain probability, and is therefore a soft clustering method with advantages in handling data ambiguity or multi-class classification.
[0129] Furthermore, other improvements can be made to the above algorithm. For example, distributed computing and parallel algorithms can be introduced for clustering to fully utilize computing resources and accelerate the clustering process, especially when dealing with large-scale datasets. This allows the model to analyze and process large amounts of malicious domain name data more quickly, providing more efficient support for the model's clustering task. In addition, this embodiment improves the robustness and accuracy of the clustering algorithm by enhancing the model's ability to handle outliers. Traditional parameter estimation methods are sensitive to outliers, which may cause model parameters to deviate from reality. By employing the more robust parameter estimation method M-estimation, this embodiment can more effectively identify and handle outliers, enabling the model to exhibit better robustness to various data conditions. This makes the model more suitable for various complex data scenarios in the real world, thereby improving its practicality and reliability.
[0130] Referring to Figure 3, in some embodiments, obtaining multiple sample domain name features in step S201 may include steps S301 to S304:
[0131] Step S301: Obtain multiple sample malicious domain names;
[0132] Step S302: For each sample malicious domain name, extract the first character of the top-level domain part and the second character of the subdomain part of the sample malicious domain name;
[0133] Step S303: Extract the first character feature from the first character and extract the second character feature from the second character;
[0134] Step S304: Combine the first character features and second character features corresponding to each malicious domain name sample to obtain multiple sample domain name features.
[0135] This embodiment of the disclosure is optimized for the characteristics of domain name data. An improved initialization method is used in the process of obtaining sample domain name features, that is, before clustering, making it more suitable for processing domain name data.
[0136] In steps S301 to S304 above, the domain name may include a top-level domain and a subdomain. The subdomain may include a second-level domain or more levels of domain. The top-level domain typically indicates the category or organizational nature of the domain, while the subdomain reflects the specific website name or function. In this embodiment, multiple sample malicious domains can be obtained, and targeted feature extraction can be performed on these sample malicious domains. First, for each sample malicious domain, the first character of the top-level domain and the second character of the subdomain are extracted. Subsequently, by extracting the features of these two parts separately, the semantic information of different parts of the domain name can be better captured, thereby more accurately identifying the malicious domain.
[0137] After extracting characters related to the top-level domain and subdomains, feature extraction is required. Specifically, first-character features are extracted from the first character, and second-character features are extracted from the second character. The extraction process can convert characters into numerical values, encodings, calculate character statistics (such as frequency, length, etc.), or perform other forms of feature extraction. The obtained features will be used for subsequent processing to capture potential malicious patterns in the domain.
[0138] Finally, the first and second character features corresponding to each malicious domain name are combined to obtain multiple sample domain name features. The combination process involves fusing the first and second character features, assigning corresponding feature weights to them during the fusion process, and then weighting and fusing them based on these weights to obtain the final sample domain name features. In this way, after feature extraction, each malicious domain name can ultimately yield its corresponding sample domain name features, enabling more accurate identification of malicious domain names in subsequent steps.
[0139] Referring to Figure 4, in some embodiments, the extraction of the first character feature from the first character and the extraction of the second character feature from the second character in step S303 above may include steps S401 to S402:
[0140] Step S401: Perform one-hot encoding on the first character to obtain the feature of the first character;
[0141] Step S402: Count the substrings in the second character that meet the preset length, and encode the second character according to the occurrence frequency of each substring under multiple malicious domain names to obtain the second character feature.
[0142] In steps S401 to S402 above, for the feature representation of the top-level domain part, this embodiment of the disclosure adopts one-hot encoding, encoding the first character extracted from each top-level domain part into a binary vector to obtain the first character feature. For example, when the top-level domain appears, the vector at the corresponding position is set to 1, and other positions are set to 0, and the length of the vector is equal to the total number of top-level domains.
[0143] For the extraction of features from the subdomain portion, an n-gram feature representation method was adopted. This method considers all substrings of the second character that satisfy a preset length of n. Therefore, the frequency of occurrence of substrings of length n is taken into account, and low-frequency substrings are removed, retaining only the remaining substrings to construct a substring list. Then, a vector of corresponding length is created based on the length of the substring list, where each vector component represents the frequency of occurrence of the corresponding substring. This forms the n-gram feature vector of the subdomain portion, which is the feature of the second character.
[0144] Finally, the extracted top-level domain feature vectors and second-level domain feature vectors are combined to form a complete domain feature vector, which is the sample domain feature.
[0145] Referring to Figure 5, in some embodiments, the Gaussian mixture clustering of multiple sample domain name features in step S201 above to obtain multiple initial Gaussian mixture categories may include steps S501 to S505:
[0146] Step S501: Determine the initial number of categories, and perform Gaussian mixture clustering on the features of multiple sample domain names under the number of categories to obtain multiple corresponding sub-Gaussian mixture categories;
[0147] Step S502: Calculate the error value between the domain name features of each sample in each sub-Gaussian mixture category and the corresponding cluster center;
[0148] Step S503: Gradually increase the number of categories and re-perform Gaussian mixture clustering with the corresponding number of categories, and calculate the error value of each sub-Gaussian mixture category with the corresponding number of categories;
[0149] Step S504: Determine the minimum target error value from the error values under different categories, and determine the number of categories corresponding to the target error value as the target category number;
[0150] Step S505: Perform Gaussian mixture clustering on the features of multiple sample domain names under the target category number to obtain multiple initial Gaussian mixture categories.
[0151] In steps S501 to S505 above, to ensure the effectiveness of Gaussian mixture clustering, it is necessary to first determine the number of Gaussian distributions. First, the initial number of categories needs to be determined, such as 1 or 2. Under this number of categories, Gaussian mixture clustering is performed on the features of multiple sample domain names to obtain the corresponding number of sub-Gaussian mixture categories. An upper limit can also be set for the number of categories in this process.
[0152] Next, the error value between the domain name features of each sample in each sub-Gaussian mixture category and its corresponding cluster center is calculated. That is, for the clustering results with each number of categories, an evaluation metric for clustering effectiveness is calculated, such as the sum of squared errors, and used as the error value. This represents the sum of squared distances between each data point and its corresponding cluster center. Subsequently, the number of categories is gradually increased, and Gaussian mixture clustering is performed again with the corresponding number of categories, calculating the error value for each sub-Gaussian mixture category with the corresponding number of categories.
[0153] After obtaining the error values for different numbers of clusters, a curve can be plotted with the number of clusters as the X-axis and the corresponding error value as the Y-axis to show how the error value changes with the number of clusters. In the curve, the point where the rate of decrease in error value slows significantly is identified as the target point. To the left of the target point, the error value decreases rapidly with increasing number of clusters; to the right of the target point, the rate of decrease slows significantly with further increases in the number of clusters. This shows that as the number of clusters increases, the improvement in clustering effect decreases rapidly at first, then gradually levels off. This transition point represents the optimal number of clusters. Therefore, the number of clusters corresponding to the target point is the optimal Gaussian distribution number, meaning the minimum error value is the target error value. The number of clusters corresponding to the target error value is then determined as the target number of clusters.
[0154] Referring to Figure 6, in some embodiments, the type of the covariance matrix is determined by the following steps, which may include steps S601 to S602:
[0155] Step S601: Based on the Gaussian distribution parameters calculated from different covariance matrix types, Gaussian mixture clustering is performed on the features of multiple sample domain names to obtain multiple test Gaussian mixture categories;
[0156] Step S602: Determine the clustering effect of the corresponding test Gaussian mixture categories under different covariance matrix types, and determine the covariance matrix as a diagonal covariance matrix based on the clustering effect.
[0157] In steps S601 to S602 above, this embodiment of the disclosure needs to optimize the type of covariance matrix in the Gaussian mixture clustering algorithm. Considering the characteristics of domain name data, different types of covariance matrices are tried, including diagonal covariance matrices, to better reflect the relationship between data.
[0158] Specifically, this embodiment attempts to use different covariance matrix types for Gaussian mixture clustering. A covariance matrix is a mathematical tool describing the correlation between multidimensional random variables; in a Gaussian mixture model, it determines the shape and orientation of each Gaussian distribution. Different covariance matrix types include diagonal covariance matrices and full covariance matrices. A diagonal covariance matrix implies that the features are independent of each other, while a full covariance matrix considers the correlation between features. Therefore, based on each covariance matrix type, the corresponding Gaussian distribution parameters (including the mean vector, covariance matrix, and mixing coefficients) are calculated. Then, these parameters are used to perform Gaussian mixture clustering on the features of multiple sample domain names, resulting in multiple test Gaussian mixture categories. Each test Gaussian mixture category is a clustering result obtained based on a specific covariance matrix type.
[0159] Subsequently, this embodiment of the disclosure needs to evaluate the clustering effect of the test Gaussian mixture categories obtained under different covariance matrix types. This can be accomplished, for example, through some clustering effect evaluation metrics, such as the silhouette coefficient and the Calinski-Harabasz Index. These metrics can help quantify the quality of the clustering results. By comparing the clustering effects under different covariance matrix types, the best-performing covariance matrix type can be selected. In practical applications, the diagonal covariance matrix showed good clustering effects in the tests, better considering the characteristics of domain name data and better reflecting the relationships between data. Therefore, the diagonal covariance matrix was ultimately selected as the final covariance matrix type.
[0160] Once the type of covariance matrix is determined, it can be used to further adjust and optimize the Gaussian mixture model to improve the ability to identify and respond to malicious domain names, thereby effectively protecting network security.
[0161] Referring to Figure 7, in some embodiments, the semantic extraction of the initial domain name features based on a pre-built attention mechanism in step S102 to obtain semantic features may include steps S701 to S703:
[0162] Step S701: Input the initial domain name features into a multi-layer attention network that is connected sequentially;
[0163] Step S702: In each layer of the attention network, the current input data is bidirectionally encoded through the attention mechanism to obtain the corresponding attention weight matrix. The hidden state is obtained and output based on the attention weight matrix and the current input data.
[0164] Step S703: Obtain semantic features through the output of the last layer of the attention network.
[0165] In steps S701 to S703 above, the attention mechanism constructed in this embodiment includes a Bidirectional Encoder Representations from Transformers (BERT), which comprises a series of interconnected multi-layer attention networks. The initial domain name features are re-encoded using BERT to extract key features of the domain name. As shown in Figure 8, each BERT encoder (Trm) includes a multi-head self-attention mechanism and a feedforward neural network, followed by a residual and a normalization module. The input to each attention network layer is EN, and the output is TN. Residual connections help propagate gradients better, thus mitigating the vanishing gradient problem; while layer normalization helps reduce internal covariate bias and improve training speed.
[0166] After the initial domain name features are input into a series of interconnected multi-layer attention networks, each layer of the attention network performs bidirectional encoding on the current input data through an attention mechanism to obtain a corresponding attention weight matrix. Based on the attention weight matrix and the current input data, the hidden state is obtained and output. Specifically, the encoder input first passes through a BERT model, which contains multiple layers of Transformer encoders. The BERT model focuses on different parts of the domain name sequence through a self-attention mechanism, allowing the encoder to consider the context of the entire sentence when encoding specific words. Next, the semantic features are obtained through the output of the last layer of the attention network (such as TN), that is, the output of the last layer of the BERT model, i.e., the output of the last time step, is selected as the feature vector of the domain name.
[0167] Therefore, this embodiment of the disclosure maintains semantic consistency with the domain name generation network while acquiring a global understanding and abstract high-level feature representation of the entire domain name sequence. In this way, it provides more meaningful prior knowledge to the domain name generation network, helping to promote semantic consistency in the generation process and improve the quality of domain name generation.
[0168] Referring to Figure 9, in some embodiments, step S104 above, which involves inputting semantic features into the domain name generation network and extracting clustering features of the target category, and using these clustering features to guide the domain name generation network in generating potentially malicious domain names, may include steps S801 to S802:
[0169] Step S801: Extract clustering features of the target category;
[0170] Step S802: Input the initial domain name features, clustering features, and semantic features into the domain name generation network, and use the attention mechanism built based on the clustering features and semantic features to perform diffusion processing on the initial domain name features to generate potential malicious domain names.
[0171] In steps S801 to S802 above, the domain name generation network can be obtained through pre-training. After training, the domain name generation network can receive initial domain name features, clustering features, and semantic features as input, and use an attention mechanism built based on clustering features and semantic features to perform diffusion processing on the initial domain name features. In this way, the domain name generation network can combine clustering features and semantic features to generate potential malicious domain names, and the obtained potential malicious domain names are more correlated with the initial malicious domain names.
[0172] It should be noted that the domain name generation network can be a diffusion model, and diffusion models can also be used for text generation tasks. The diffusion model constructs a corresponding multi-layered structure and incorporates an attention mechanism based on clustering and semantic features. This attention mechanism helps the model pay more attention to clustering and semantic features when generating potentially malicious domain names, allocating more "attention" or weight to more important parts. Therefore, the initial domain name features can be diffused based on this attention mechanism, and the final model output can obtain potentially malicious domain names.
[0173] Furthermore, the domain name generation network can also be a SeqGAN, as shown in Figure 10. It includes a generator capable of receiving BERT-encoded feature vectors and a discriminator. The domain name generation network can be pre-trained adversarially, and the trained generator is used to generate potentially malicious domain names. Specifically:
[0174] The generator consists of a generator capable of receiving BERT-encoded feature vectors and a discriminator. The generator incorporates an attention mechanism and an LSTM layer. In tasks involving generating sequential data (such as text), the attention mechanism allows the domain name generation network to focus on different parts of the input sequence as it generates each element, thus utilizing the information of the input sequence more effectively. This mechanism helps the domain name generation network better understand the long-term dependencies and global structure of the input sequence, thereby improving the accuracy and fluency of the generation. In this domain name generation network, the introduction of the attention mechanism helps the generator better capture the dependencies and semantic relevance between characters, resulting in more accurate, similar, and diverse domain names. The Long Short-Term Memory (LSTM) layer is a variant of the Recurrent Neural Network (RNN) specifically designed to process sequential data and better capture long-term dependencies within the sequence. Compared to traditional RNNs, LSTMs introduce three gating units: an input gate, a forget gate, and an output gate, as well as a memory unit. These units work together to allow the network to better preserve and update information when processing long sequences. In text generation tasks, LSTM layers can better capture long-range dependencies in text, avoiding the gradient vanishing or exploding problems of traditional RNNs, and are therefore widely used in text generation tasks. In this domain name generation network, using LSTM layers helps the generator better model long-term dependencies in domain names, thus generating more semantically coherent and diverse domain names. Furthermore, the domain name generation network introduces a self-attention mechanism, which can better capture dependencies and semantic relevance between characters, enhance the ability to model long dependencies, improve global context awareness, and enhance the expressive power and flexibility of the domain name generation network through a multi-head attention mechanism. This allows the generator to consider longer-term contextual information when generating domain names, thereby generating more accurate, similar, and diverse domain names.
[0175] To ensure the accuracy of the discriminator, the domain name generation network employs a two-layer convolutional neural network with an added attention mechanism. After the generator generates potentially malicious domain names, it transforms the domain name vectors. The discriminator receives this domain name vector data as input, passes through an embedding layer, and then sequentially through a convolutional pooling combination layer, a terminal convolutional layer, a two-layer Highway network layer, and a fully connected layer, ultimately outputting the domain name classification result. The terminal convolutional layer contains a context layer for the attention layer. Details of the specific network structure will be described in detail later.
[0176] The generator is designed based on LSTM layers and consists of three main parts, with the LSTM layers introducing a self-attention mechanism. The input to the domain name generation network includes the semantic features extracted from the last self-attention layer after BERT encoding, and a pseudo-random seed. The pseudo-random seed is an initial random vector used to introduce randomness, thereby increasing the diversity of the generated text. In the domain name generation network, the role of the pseudo-random seed is to ensure that the domain name generation network produces different outputs each time, rather than strictly depending on the input data. The generator's task is to generate text sequences similar to real domain names, with the output dimension matching the size of the domain name character list. The main goal of this generator is to generate domain name text with a certain degree of variability while preserving the semantic features of the input domain names by introducing pseudo-randomness. Throughout the generation process, the pseudo-random seed plays a crucial role, introducing some initial uncertainty and variation to guide the generation process. This challenges the discriminator in trying to distinguish between real and fictitious generated domain name vectors, thus improving the generator's robustness and adversarial capabilities.
[0177] The generator's primary task is to generate a sequence as similar as possible to real, initial malicious domain name vectors, making it difficult for the discriminator to distinguish between real and fictitious input domain name vectors. In this process, the generator fully utilizes information from clustering to guide the generation of potential malicious domain names, ensuring that the generated potential malicious domain names are similar to real domain names in a specific cluster to some extent. First, the generator starts with a pseudo-random initial character number, which is then transformed into a meaningful embedding vector through an embedding layer. This process aims to introduce semantic information into the generated domain name sequence, enabling the generated domain names to have a certain degree of semantic coherence. This embedding vector provides the generator with a good starting point, allowing it to gradually adjust parameters during training to generate more plausible domain names. Next, the embedding vector is concatenated with semantic features obtained from BERT encoding, forming a fixed-size and meaningful feature sequence output. These features include not only structural and syntactic considerations but also higher-level semantic information, providing the generator with richer context. The introduction of Transformer encoding plays a crucial role in handling long-distance dependencies, effectively capturing the complex relationships between characters in the domain name and improving the overall coherence of the generated domain name sequence. In the LSTM layer, the unit structure of the Long Short-Term Memory (LSTM) network allows the network to better capture long-term dependencies in the sequence. Through a gating mechanism, the LSTM network can selectively remember or forget information when processing the input sequence, which helps the generator better understand and retain the semantic and structural information in the domain name sequence. Subsequently, the fully connected layer receives the input from the LSTM layer, containing prediction information for the next batch of character numbers. This design takes into account the dynamic process of domain name generation, converting it into a vector the size of the domain name character table for output. Through the use of the LogSoftMax function, the output is converted into a log probability distribution, ultimately generating a highly interpretable domain name vector. The entire generation process is perfected through continuous learning and parameter tuning iterations, enabling the generator to gradually improve the quality of the generated domain names and more accurately simulate the structure and characteristics of real domain names. This iterative process is key to the generator's self-improvement, aiming to better deceive the discriminator and generate more realistic potentially malicious domain names.
[0178] The discriminator's task is to distinguish between genuine and fake input domain name vectors. Its structure includes an embedding layer, a convolutional pooling combination layer (consisting of 30 2×2 filters and 15 3×3 filters), a terminal convolutional layer, a custom attention module, two layers of a high-speed network, and a fully connected layer. This combination of layers enables the discriminator to effectively extract features from the input domain name vectors and make a true / false judgment. First, a batch of domain name vectors passes through the embedding layer to obtain an embedding tensor, the shape of which is batch size × maximum domain name length × convolutional layer dimension. Next, the embedding tensor passes through the convolutional pooling combination layer to extract the character features of the domain names and output the corresponding feature tensor.
[0179] To enhance the discriminator's ability to extract input features, this domain name generation network adds a terminal convolutional layer after the convolutional pooling combination layer. A custom `attention_3d_block` function is used to implement the attention module, converting the feature tensor, which originally lacked hidden states, into a hidden state tensor with sequence information. A custom lambda function is then used to extract the hidden state at the last moment from the hidden state tensor. Next, the `dot` function is used to perform a dot product operation with the attention score vector in a specific dimension [2,1], resulting in a tensor representing the attention weights. In the subsequent Activation module, to allow the discriminator to focus on the key features of the domain name vector, the attention weights need to be normalized. This domain name generation network uses the softmax function to transform all attention weights into a probability distribution, placing them between [0,1] and summing them to 1, thus achieving the allocation of importance for the input features. Finally, in a specific dimension [1,1], the `dot` function is used again to perform a dot product operation, weighting and summing the result with the input hidden state to generate a context vector that integrates the importance of different positions, providing the discriminator with a more comprehensive semantic understanding.
[0180] To address the vanishing or exploding gradient problem that may arise from increasing the depth of the domain name generation network, a two-layer highway network is introduced. These highway network layers adaptively select the amount and path of information transmission, thus avoiding the adverse performance impact of gradient problems. Finally, the probability of the domain name vector is output through a fully connected layer. When the probability is greater than a preset threshold (e.g., 0.5), the input domain name vector is determined to be a vector of potentially malicious domain names; when the probability is less than the preset threshold, the potentially malicious domain names generated by the generator are deemed unacceptable.
[0181] Please refer to Figure 11. In some embodiments, after step S104 uses clustering features to guide the domain name generation network to generate potentially malicious domain names, it may also include steps S901 to S902:
[0182] Step S901: Obtain the domain name to be tested;
[0183] Step S902: Perform malicious detection on the domain to be detected based on the potential malicious domain name, and obtain the corresponding domain name detection results.
[0184] In steps S901 to S902 above, the domain name to be detected can be one submitted by a user, captured by network traffic monitoring, or a suspected malicious domain name detected by other security systems. Since the security of the domain name to be detected is unknown, subsequent security testing is required.
[0185] After obtaining the potential malicious domains in the above embodiments, since the potential malicious domains are generated based on the initial malicious domains in the dataset, that is, generated from any data point in the dataset, there can be multiple potential malicious domains generated in the end. There are various ways to apply potential malicious domains. For example, multiple potential malicious domains can be used to directly detect malicious domains, such as comparing the domain to be detected with potential malicious domains to check for similarities or common features, and obtaining corresponding domain detection results. If the domain to be detected and the potential malicious domain are highly similar in structure, character combination, semantics, etc., then it is very likely also a malicious domain. Furthermore, machine learning or deep learning models can be used to classify the domain to be detected and determine whether it is a malicious domain. Specifically, this includes training a malicious domain identification model using multiple generated potential malicious domains. After training, the identification model can accurately identify malicious behavior in the domain. Therefore, after inputting the domain to be detected into the identification model, the corresponding domain detection results can be accurately obtained.
[0186] Furthermore, the domain name detection results can be binary, indicating whether the domain name is malicious, or multi-classified, indicating which type of malicious behavior the domain name belongs to, or include more detailed scoring or probability information for subsequent security analysis or response measures. Therefore, in this embodiment, the generated potentially malicious domain names are applied to actual malicious domain name detection tasks, thereby enhancing network security protection capabilities and enabling timely detection and response to potential network threats.
[0187] Please refer to Figure 12. This embodiment of the disclosure also provides a malicious domain name generation apparatus, which can implement the above-described malicious domain name generation method. The malicious domain name generation apparatus includes:
[0188] Encoding module 1201 is used to obtain the initial malicious domain name and encode each character in the initial malicious domain name to obtain the initial domain name characteristics;
[0189] The semantic extraction module 1202 is used to extract semantic features from the initial domain name features based on a pre-built attention mechanism to obtain semantic features;
[0190] The clustering module 1203 is used to determine the distribution probability of the initial malicious domain name belonging to different Gaussian mixture categories based on the cluster centers of multiple preset Gaussian mixture categories, and to determine the target category to which the initial malicious domain name belongs from multiple Gaussian mixture categories based on the magnitude relationship of each distribution probability.
[0191] The domain name generation module 1204 is used to input semantic features into the domain name generation network and extract clustering features of the target category, and use the clustering features to guide the domain name generation network to generate potentially malicious domain names.
[0192] The Gaussian mixture category is obtained by updating the initial Gaussian mixture category based on the updated Gaussian distribution parameters. The updated Gaussian distribution parameters are obtained by first determining multiple initial Gaussian mixture categories for multiple sample domain name features, then calculating the posterior probability of each sample domain name feature belonging to each initial Gaussian mixture category, and updating the current Gaussian distribution parameters based on the posterior probability.
[0193] In summary, the malicious domain name generation device executes a malicious domain name generation method. First, it encodes each character in the initial malicious domain name to obtain initial domain name features. Based on a pre-built attention mechanism, it performs semantic extraction on the initial domain name features, which significantly improves the accuracy and efficiency of semantic extraction, enhancing the ability to identify and respond to malicious domain names, and obtaining semantic features. Next, based on the cluster centers of multiple preset Gaussian mixture categories, it determines the distribution probability of the initial malicious domain name belonging to different Gaussian mixture categories. Based on the magnitude of each distribution probability, it determines the target category to which the initial malicious domain name belongs from the multiple Gaussian mixture categories. The semantic features are input into the domain name generation network, and the clustering features of the target category are extracted. The Gaussian mixture category is obtained by updating the initial Gaussian mixture category based on the updated Gaussian distribution parameters. The updated Gaussian distribution parameters are obtained by first determining multiple initial Gaussian mixture categories for multiple sample domain name features, then calculating the posterior probability of each sample domain name feature belonging to each initial Gaussian mixture category, and updating the current Gaussian distribution parameters based on the posterior probability. In this way, each Gaussian mixture category can accurately cluster similar domain names, and the clustering features of the target category can well express the data features of the Gaussian distribution. Subsequently, by utilizing clustering features, the domain name generation network can be guided to generate similar potential malicious domain names within the range indicated by the target category. The resulting potential malicious domain names are more correlated with the initial malicious domain name. Furthermore, since the potential malicious domain names are generated based on the currently input initial malicious domain name and do not rely on historical data, they are more timely. Therefore, the quality of the final discovered potential malicious domain names is better.
[0194] The specific implementation of the malicious domain name generation device is basically the same as the specific embodiment of the malicious domain name generation method described above, and will not be repeated here. Subject to meeting the requirements of the embodiments of this disclosure, the malicious domain name generation device may also be equipped with other functional modules to implement the malicious domain name generation method in the above embodiments.
[0195] This disclosure also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned method for generating malicious domain names. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0196] Please refer to Figure 13, which illustrates the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0197] The processor 1301 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this disclosure.
[0198] The memory 1302 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1302 can store operating devices and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1302 and is called and executed by the processor 1301 to execute the malicious domain name generation method of the embodiments of this disclosure.
[0199] The input / output interface 1303 is used to implement information input and output;
[0200] The communication interface 1304 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0201] Bus 1305 transmits information between various components of the device (e.g., processor 1301, memory 1302, input / output interface 1303, and communication interface 1304);
[0202] The processor 1301, memory 1302, input / output interface 1303 and communication interface 1304 are connected to each other within the device via bus 1305.
[0203] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for generating malicious domain names.
[0204] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0205] The embodiments described in this disclosure are for the purpose of more clearly illustrating the technical solutions of this disclosure and do not constitute a limitation on the technical solutions provided by this disclosure. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by this disclosure are also applicable to similar technical problems.
[0206] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this disclosure, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0207] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0208] Those skilled in the art will understand that all or some of the steps, apparatuses, or functional modules / units in the methods disclosed above can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0209] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0210] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0211] In the several embodiments provided in this disclosure, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0212] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0213] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0214] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0215] The preferred embodiments of the present disclosure have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present disclosure. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of the present disclosure shall be within the scope of the claims of the present disclosure.
Claims
1. A method for generating malicious domain names, characterized in that, include: Obtain the initial malicious domain name and encode each character in the initial malicious domain name to obtain the initial domain name characteristics; Semantic features are obtained by semantically extracting the initial domain name features based on a pre-built attention mechanism; Based on the cluster centers of multiple preset Gaussian mixture categories, the distribution probability of the initial malicious domain name belonging to different Gaussian mixture categories is determined, and the target category to which the initial malicious domain name belongs is determined from the multiple Gaussian mixture categories according to the magnitude relationship of each distribution probability. The semantic features are input into the domain name generation network, and the clustering features of the target category are extracted. The clustering features are then used to guide the domain name generation network to generate potentially malicious domain names. The Gaussian mixture category is obtained by updating the initial Gaussian mixture category according to the updated Gaussian distribution parameters. The updated Gaussian distribution parameters are obtained by first determining multiple initial Gaussian mixture categories for multiple sample domain name features, then calculating the posterior probability of each sample domain name feature belonging to each initial Gaussian mixture category, and updating the current Gaussian distribution parameters according to the posterior probability.
2. The method for generating malicious domain names according to claim 1, characterized in that, The Gaussian mixture category is determined through the following steps: Multiple sample domain name features are obtained, and Gaussian mixture clustering is performed on the multiple sample domain name features to obtain multiple initial Gaussian mixture categories; Determine the current Gaussian distribution parameters for each initial Gaussian mixture category, and determine the probability density function of each sample domain name feature under each initial Gaussian mixture category, wherein the Gaussian distribution parameters include the mean vector, covariance matrix, and mixing coefficients; Based on the current mean vector, covariance matrix, mixing coefficients, and probability density function, calculate the posterior probability that each sample domain name feature belongs to each initial Gaussian mixture category; The current Gaussian distribution parameters are updated based on the posterior probability, and the initial Gaussian mixture category is updated based on the updated Gaussian distribution parameters to obtain the updated Gaussian mixture category.
3. The method for generating malicious domain names according to claim 2, characterized in that, The step of calculating the posterior probability of each sample domain name feature belonging to each initial Gaussian mixture category based on the current mean vector, covariance matrix, mixing coefficients, and probability density function includes: The feature x of each sample domain name is calculated according to the following formula. n The posterior probability γ(Z) of belonging to the k-th initial Gaussian mixture category nk ): Wherein, N(x) in formula (1) n |μ k ,Σ k ) is the sample domain name feature x n The probability density function under the k-th initial Gaussian mixture category, π k It is the mixing coefficient of the k-th initial Gaussian mixture category, μ k It is the mean vector of the k-th initial Gaussian mixture class, Σ k It is the covariance matrix of the k-th initial Gaussian mixture class; Updating the current Gaussian distribution parameters based on the posterior probability includes: According to the posterior probability γ(Z) nk The formula for updating the current Gaussian distribution parameters is as follows: Among them, N in formulas (2) to (4) K is the total number of sample domain name features belonging to the k-th initial Gaussian mixture category, and N is the total number of sample domain name features.
4. The method for generating malicious domain names according to claim 2, characterized in that, The acquisition of multiple sample domain name features includes: Obtain multiple sample malicious domain names; For each of the sample malicious domain names, extract the first character of the top-level domain portion and the second character of the subdomain portion of the sample malicious domain name; The first character feature is extracted from the first character, and the second character feature is extracted from the second character; The first character feature and the second character feature corresponding to each malicious domain name of the sample are combined to obtain multiple sample domain name features.
5. The method for generating malicious domain names according to claim 4, characterized in that, The step of extracting the first character feature from the first character and extracting the second character feature from the second character includes: One-hot encoding is performed on the first character to obtain the feature of the first character; The second character is identified by counting the substrings that meet a preset length. The second character is then encoded based on the frequency of occurrence of each substring under multiple malicious domain names in the sample, thus obtaining the second character feature.
6. The method for generating malicious domain names according to claim 2, characterized in that, The Gaussian mixture clustering of multiple sample domain name features yields multiple initial Gaussian mixture categories, including: Determine the initial number of categories, and perform Gaussian mixture clustering on the features of multiple sample domain names under the number of categories to obtain multiple corresponding sub-Gaussian mixture categories; Calculate the error value between the domain name features of each sample in each sub-Gaussian mixture category and the corresponding cluster center; Gradually increase the number of categories, and re-perform Gaussian mixture clustering with the corresponding number of categories, and calculate the error value of each sub-Gaussian mixture category with the corresponding number of categories; The minimum target error value is determined from the error values under different numbers of categories, and the number of categories corresponding to the target error value is determined as the target number of categories; Gaussian mixture clustering is performed on the features of multiple sample domain names under the target number of categories to obtain multiple initial Gaussian mixture categories.
7. The method for generating malicious domain names according to claim 2, characterized in that, The type of the covariance matrix is determined through the following steps: Based on the Gaussian distribution parameters calculated from different covariance matrix types, Gaussian mixture clustering is performed on multiple sample domain name features to obtain multiple test Gaussian mixture categories; Determine the clustering effect of the corresponding test Gaussian mixture categories under different covariance matrix types, and determine the covariance matrix as a diagonal covariance matrix based on the clustering effect.
8. The method for generating malicious domain names according to claim 1, characterized in that, The semantic features obtained by performing semantic extraction on the initial domain name features based on the pre-built attention mechanism include: The initial domain name features are input into a sequentially connected multi-layer attention network; In each layer of the attention network, the current input data is bidirectionally encoded through an attention mechanism to obtain a corresponding attention weight matrix. The hidden state is obtained and output based on the attention weight matrix and the current input data. Semantic features are obtained through the output of the attention network in the last layer.
9. The method for generating malicious domain names according to claim 1, characterized in that, The step of inputting the semantic features into the domain name generation network, extracting the clustering features of the target category, and using the clustering features to guide the domain name generation network to generate potentially malicious domain names includes: Extract the clustering features of the target category; The initial domain name features, the clustering features, and the semantic features are input into the domain name generation network. An attention mechanism based on the clustering features and the semantic features is used to diffuse the initial domain name features to generate potential malicious domain names.
10. The method for generating malicious domain names according to claim 1, characterized in that, After using the clustering features to guide the domain name generation network to generate potentially malicious domain names, the method further includes: Obtain the domain name to be tested; Based on the potential malicious domain name, the domain name to be detected is maliciously detected, and the corresponding domain name detection results are obtained.
11. A device for generating malicious domain names, characterized in that, include: An encoding module is used to obtain an initial malicious domain name and encode each character in the initial malicious domain name to obtain the initial domain name characteristics; The semantic extraction module is used to extract semantic features from the initial domain name features based on a pre-built attention mechanism to obtain semantic features; The clustering module is used to determine the distribution probability of the initial malicious domain name belonging to different Gaussian mixture categories based on the cluster centers of multiple preset Gaussian mixture categories, and to determine the target category to which the initial malicious domain name belongs from the multiple Gaussian mixture categories based on the magnitude relationship of each distribution probability. The domain name generation module is used to input the semantic features into the domain name generation network, extract the clustering features of the target category, and use the clustering features to guide the domain name generation network to generate potentially malicious domain names; The Gaussian mixture category is obtained by updating the initial Gaussian mixture category according to the updated Gaussian distribution parameters. The updated Gaussian distribution parameters are obtained by first determining multiple initial Gaussian mixture categories for multiple sample domain name features, then calculating the posterior probability of each sample domain name feature belonging to each initial Gaussian mixture category, and updating the current Gaussian distribution parameters according to the posterior probability.
12. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method for generating malicious domain names as described in any one of claims 1 to 10.
13. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for generating malicious domain names as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Domain name detection method and system for adversarial network
CN110545284A
Malicious domain name training data generation method based on generative adversarial network model
CN113190846A
Method and device for identifying domain name
CN115473726A
Domain name processing method, device and equipment
CN117376307A
Malicious domain name generation method and device, equipment and medium
CN118138382A
Cited By
Malicious domain name detection system fusing multi-modal embedding and dynamic weight Mama
CN122226480A