A bad domain name identification method and device and related equipment
Patent Information
- Application Number
- CN202611156833.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-31
- Publication Date
- 2026-09-25
AI Technical Summary
然而,上述不良域名识别方法存在一定不足
[0032]本申请在上述各方面提供的实现方式的基础上,还可以进行进一步组合以提供更多实现方式。
Smart Images

Figure CN122824481A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network security technology, specifically to a method, apparatus, and related equipment for identifying malicious domain names. Background Technology
[0002] With the rapid development of internet services, domain names have become a crucial entry point for accessing network resources and locating network services. At the same time, some domain names may be used in cybersecurity incidents such as online fraud, malware propagation, phishing attacks, illegal redirects, and data theft; these domain names are generally referred to as malicious domain names. Identifying malicious domain names is a vital part of network security protection, and the results can be used for risk alerts, access blocking, security auditing, and threat analysis.
[0003] Some methods for identifying malicious domain names typically involve blacklist matching, rule matching, manual feature filtering based on experience, or machine learning models. However, these methods have certain limitations. When dealing with domains with similar characteristics, the accuracy of malicious domain name identification can be low. Furthermore, when these methods are applied to environments with limited computing or storage resources, the identification model may negatively impact identification efficiency.
[0004] Therefore, there is an urgent need for a method that can improve the accuracy of identifying malicious domain names while reducing the impact of computing resources, storage resources, and response time on the operation of the malicious domain name identification model. Summary of the Invention
[0005] This application provides a method for identifying malicious domain names, which improves the accuracy of malicious domain name identification by utilizing domain name resolution data corresponding to domain names at multiple time points, and reduces storage overhead, computational overhead, and inference latency in the malicious domain name identification process by using a student model with a small number of parameters. Furthermore, this application also provides a corresponding malicious domain name identification device, computing equipment, computer-readable storage medium, and computer program product.
[0006] Firstly, this application provides a method for identifying malicious domain names, which can be executed by a malicious domain name identification device or a computing device with data processing capabilities. Specifically, the malicious domain name identification device acquires the domain name to be identified and the domain name resolution data corresponding to the domain name at multiple time points. Then, the malicious domain name identification device uses a temporal neural network to fuse features of the domain name resolution data corresponding to the domain name at multiple time points and the temporal order relationship of the multiple time points to obtain fused features. Next, the malicious domain name identification device uses a student model to infer the fused features to obtain an identification result, which is used to indicate whether the domain name to be identified is a malicious domain name. The student model is obtained by using training samples and knowledge distillation training of the initial student model based on the teacher model. The training samples include sample fusion features and training labels. The sample fusion features are obtained by fusing features of the domain name resolution data corresponding to the sample domain name at multiple training time points and the temporal order relationship of the multiple training time points. The training labels are used to indicate whether the sample domain name is a malicious domain name. The number of parameters in the student model is less than the number of parameters in the teacher model.
[0007] Thus, by fusing features of the domain name resolution data corresponding to the domain name to be identified at multiple time points and the temporal order relationship at multiple time points, the fused features can simultaneously reflect the domain name resolution information corresponding to each time point and the temporal correlation information between different time points. Compared to identification based solely on the domain name resolution data corresponding to the domain name to be identified at a single time point, the fused features can characterize the domain name resolution data and its changing patterns at multiple time points from a temporal dimension, thereby helping to identify the changing characteristics reflected in the domain name resolution data at different time points and improving the accuracy of the identification results. Furthermore, the student model is obtained by training the teacher model based on training samples through knowledge distillation. Through knowledge distillation, the identification knowledge learned by the teacher model based on the fused features and training labels of the samples can be transferred to the student model with fewer parameters, enabling the student model to have the ability to identify malicious domain names while reducing the model's storage usage and inference computation. This reduces the storage and computational resources required for deploying the malicious domain name identification model, improves the inference efficiency of malicious domain name identification, and enhances the applicability of this method in scenarios with limited computational resources.
[0008] In one possible implementation, the malicious domain name identification device utilizes a temporal neural network to extract features from the domain name resolution data corresponding to the domain name to be identified at multiple time points, obtaining a first feature. Then, the temporal neural network is used to extract features from the temporal order relationship at multiple time points, obtaining a second feature. Finally, the first and second features are fused to obtain a fused feature. The temporal order relationship at multiple time points can be represented by time stamps, sequence indices, relative order information, or time position codes corresponding to each time point. Thus, the fused feature can take into account both the information contained in the domain name resolution data itself and the temporal order information between multiple time points, enhancing the fused feature's ability to represent the data and temporal order relationship of the domain name to be identified at multiple time points, thereby helping to improve the accuracy of malicious domain name identification.
[0009] In one possible implementation, before the student model infers from the fused features, the method further includes training processes for both the teacher and student models. First, multiple labeled training domain name samples are acquired. Each training domain name sample includes the training domain name and its corresponding domain name resolution data at multiple training time points. The labels represent whether the corresponding training domain name is a normal or malicious domain name. Then, a domain name knowledge graph is constructed based on the multiple labeled training domain name samples. The domain name knowledge graph includes multiple domain name nodes and associated nodes. Domain name nodes indicate training domain names, and associated nodes include at least one of Internet Protocol (IP) address nodes, Autonomous System (AS) nodes, domain name registrant nodes, geographic location nodes, and certificate nodes. Next, based on the domain name resolution data of the multiple training domain names at multiple training time points, the node features corresponding to the multiple nodes in the domain name knowledge graph at different training time points are determined. Using the node features corresponding to the multiple nodes at different training time points and the labels corresponding to the multiple training domain name samples, the initial teacher model is trained to obtain the teacher model. Finally, the initial student model is trained using knowledge distillation based on the teacher model to obtain the student model. The initial student model has fewer parameters than the initial teacher model.
[0010] Thus, by constructing a domain knowledge graph including domain nodes and their associated nodes based on multiple labeled training domain samples, the relationship between training domains and associated nodes can be incorporated into the training process of the teacher model. This allows the teacher model to learn features by combining the association information between training domains and their associated nodes. Compared to relying solely on the information of the domain nodes themselves, this enriches the feature information used to distinguish between normal and malicious domains, enhancing the teacher model's ability to represent differences between different categories of domains. Furthermore, by determining the node features corresponding to multiple nodes in the domain knowledge graph at different training time points, the teacher model can comprehensively utilize the feature information of the same training domain and its associated nodes at different training time points, reducing the recognition bias caused by limited information when using node features from only a single training time point. Simultaneously, using the labels corresponding to the training domain samples to supervise the initial teacher model ensures that the node association information and node feature information learned by the teacher model correspond to the true category of the training domain, reducing the possibility of inconsistency between the teacher model's output and the actual category of the training domain, thereby improving the accuracy of the teacher model in distinguishing between normal and malicious domains. Finally, by using the teacher model to train the initial student model with a small number of parameters through knowledge distillation, the knowledge of identifying malicious domain names learned by the teacher model can be transferred to the student model, enabling the student model to have the corresponding ability to identify malicious domain names while reducing the number of parameters.
[0011] In one possible implementation, the associated nodes include IP address nodes, autonomous system nodes, and certificate nodes. The domain name knowledge graph is an undirected weighted graph. The domain name knowledge graph includes a first edge connecting domain name nodes and IP address nodes, a second edge connecting domain name nodes and certificate nodes, and a third edge connecting IP address nodes and autonomous system nodes. The weight of the first edge is greater than the weight of the second edge, and the weight of the second edge is greater than the weight of the third edge. Thus, by setting connecting edges between domain name nodes and IP address nodes, between domain name nodes and certificate nodes, and between IP address nodes and autonomous system nodes, the domain name knowledge graph can be used to represent the association relationships between corresponding nodes. By decreasing the weights of the first, second, and third edges sequentially, the relative importance of different types of association relationships can be distinguished, thereby helping to improve the domain name knowledge graph's ability to represent the association relationships between different nodes.
[0012] In one possible implementation, the process of training the initial teacher model includes: mapping the node features corresponding to different types of nodes in the domain knowledge graph at different training time points to the same feature space, obtaining node feature representations corresponding to different types of nodes. Then, based on the node identifiers corresponding to the node feature representations and the training time points, the node feature representations corresponding to different types of nodes for the same training domain at different training time points are arranged in chronological order to determine the node time series corresponding to different training domains. Finally, the initial teacher model is trained using the domain knowledge graph, the node time series corresponding to different training domains, and the labels corresponding to the training domain samples to obtain the teacher model. In this way, by mapping the node features of different types of nodes to the same feature space, the feature expression forms of different types of nodes can be unified, enabling the domain nodes corresponding to the same training domain and their different types of associated nodes to be jointly organized and processed. Furthermore, by determining the training domain corresponding to each node feature representation based on the node identifier, and arranging the node feature representations of different types of nodes corresponding to the same training domain in chronological order based on the training time points, the feature information originally scattered across different nodes and different training time points can be integrated into a node time series corresponding to the training domain. Therefore, the teacher model can combine complementary information provided by different types of nodes to learn the changes and correlations of the characteristics of the same training domain name and its associated nodes over time, thereby improving the accuracy of distinguishing between normal and bad domain names.
[0013] In one possible implementation, the temporal neural network includes a temporal convolutional network and a long short-term memory (LSTM) network. Based on this, the process of training the initial teacher model includes: convolving the time series of nodes corresponding to different training domains using the temporal convolutional network to obtain the time series output features corresponding to different training domains. Then, temporally fusing the time series output features corresponding to different training domains using the LTM network to obtain fused features corresponding to different training domains. Finally, using the domain knowledge graph, the fused features corresponding to different training domains, and the labels corresponding to the training domain samples, the initial teacher model is trained to obtain the teacher model. Thus, by convolving the node time series using the temporal convolutional network, local feature changes between adjacent or nearby training time points can be extracted; by temporally fusing the time series output features using the LTM network, feature information over a longer time range can be associated, learning long-term dependencies between different training time points. Therefore, the obtained fused features can take into account both local changes and overall evolution patterns of node features, enabling the teacher model to identify abnormal changes occurring in a short period and abnormal features that persist across multiple training time points, thereby improving the accuracy of distinguishing between normal and malicious domains.
[0014] In one possible implementation, the temporal convolutional network comprises multiple stacked temporal convolutional layers. These stacked layers perform dilated causal convolution on the time series of nodes corresponding to different training domains, yielding the output features for each time series. Thus, by performing dilated causal convolution on multiple stacked temporal convolutional layers, the temporal correlation range of node features can be expanded without significantly increasing computational cost, enabling the identification of feature change patterns across longer time periods. Simultaneously, the node features at each training time point are determined solely by the current and previous node features, preventing the node features at subsequent training time points from negatively influencing the feature representations of earlier time points.
[0015] In one possible implementation, the process of training an initial student model using knowledge distillation based on a teacher model includes: using both the teacher model and the initial student model to infer the fusion features corresponding to different training domain names, obtaining the output results of the teacher model and the initial student model; comparing the differences between the output results of the teacher model and the initial student model to obtain a distillation loss value; determining a classification loss value based on the output results of the initial student model and the labels corresponding to multiple training domain name samples, whereby the classification loss value characterizes the degree of difference between the output results of the initial student model and the labels; and adjusting the model parameters of the initial student model based on the distillation loss value and the classification loss value to obtain the student model. Thus, the distillation loss value can characterize the difference between the output results of the teacher model and the output results of the initial student model, guiding the output results of the initial student model to approach the output results of the teacher model, enabling the student model to learn the teacher model's judgment information on different categories. The classification loss value can characterize the difference between the output results of the initial student model and the training labels, guiding the output results of the student model to match the true categories of the training domain names, reducing the possibility that the student model only learns the output of the teacher model and deviates from the true categories.
[0016] In one possible implementation, before the student model infers from the fused features, the method further includes: acquiring newly added training domain name samples; and performing incremental distillation on the student model using the newly added training domain name samples to obtain an updated student model. The updated student model uses the same hyperparameters as the original student model, including the distillation temperature coefficient, distillation loss weights, and classification loss weights. After completing the incremental distillation, the updated student model is used to infer from the fused features to obtain the recognition result. Thus, by using newly added training domain name samples for incremental distillation, the student model can continuously learn the features of problematic domain names contained in the newly added training domain name samples, achieving adaptive iterative updates of the student model and improving its adaptability to newly added problematic domain name samples and dynamically changing problematic domain names. Furthermore, by keeping the distillation temperature coefficient, distillation loss weights, and classification loss weights unchanged, the original knowledge distillation training configuration can be used to update the student model without retraining based on all training domain name samples, thereby reducing the training overhead required for model updates and improving the iterative efficiency of the problematic domain name recognition model.
[0017] In one possible implementation, the identification result includes a predicted probability that the domain name to be identified is a malicious domain name, or it includes a classification result that characterizes whether the domain name to be identified is a normal domain name or a malicious domain name. Thus, the predicted probability can characterize the likelihood that the domain name to be identified belongs to a malicious domain name, giving the identification result a more granular level of distinction. The classification result can directly characterize whether the domain name to be identified is a normal domain name or a malicious domain name, facilitating a quick and clear identification conclusion.
[0018] Secondly, this application provides a device for identifying malicious domain names, comprising an acquisition module, a fusion module, and an inference module. The acquisition module acquires the domain name to be identified and its corresponding domain name resolution data at multiple time points. The fusion module uses a temporal neural network to fuse the domain name resolution data at multiple time points and the temporal order of these time points to obtain fused features. The inference module uses a student model to infer the fused features and obtain an identification result, which indicates whether the domain name to be identified is a malicious domain name. The student model is trained on a teacher model using knowledge distillation based on training samples. The training samples include sample fusion features and training labels. The sample fusion features are obtained by fusing the domain name resolution data at multiple training time points and the temporal order of these time points. The training labels indicate whether the sample domain name is a malicious domain name. The number of parameters in the student model is less than that in the teacher model.
[0019] In one possible implementation, the fusion module can utilize a temporal neural network to extract features from the domain name resolution data corresponding to the domain name to be identified at multiple time points, obtaining a first feature. Then, the fusion module uses the temporal neural network to extract features from the temporal sequence relationship at multiple time points, obtaining a second feature. Finally, the fusion module fuses the first feature and the second feature to obtain a fused feature.
[0020] In one possible implementation, the device further includes a knowledge graph construction module, a feature determination module, a teacher training module, and a distillation training module. Before inferring the fused features using the student model, the acquisition module is further used to acquire multiple labeled training domain name samples. These training domain name samples include training domain names and domain name resolution data corresponding to the training domain names at multiple training time points. The labels are used to characterize whether the corresponding training domain name is a normal domain name or a malicious domain name. The knowledge graph construction module is used to construct a domain name knowledge graph based on the multiple labeled training domain name samples. The domain name knowledge graph includes multiple domain name nodes and associated nodes of the domain name nodes. The domain name nodes are used to indicate training domain names, and the associated nodes include at least one of Internet Protocol (IP) address nodes, Autonomous System (AS) nodes, domain name registrant nodes, geographic location nodes, and certificate nodes. The feature determination module is used to determine the node features corresponding to multiple nodes in the domain name knowledge graph at different training time points based on the domain name resolution data corresponding to the multiple training domain names at multiple training time points. The teacher training module is used to train the initial teacher model using the node features corresponding to multiple nodes in the domain name knowledge graph at different training time points and the labels corresponding to the multiple training domain name samples to obtain the teacher model. The distillation training module is used to perform knowledge distillation training on the initial student model based on the teacher model to obtain the student model. The number of parameters in the initial student model is less than the number of parameters in the initial teacher model.
[0021] In one possible implementation, the associated nodes include IP address nodes, autonomous system nodes, and certificate nodes, and the domain name knowledge graph is an undirected weighted graph. The domain name knowledge graph includes a first edge connecting domain name nodes and IP address nodes, a second edge connecting domain name nodes and certificate nodes, and a third edge connecting IP address nodes and autonomous system nodes. The weight of the first edge is greater than the weight of the second edge, and the weight of the second edge is greater than the weight of the third edge.
[0022] In one possible implementation, the teacher training module maps the node features corresponding to different types of nodes in the domain knowledge graph at different training time points to the same feature space, obtaining node feature representations corresponding to different types of nodes. Next, based on the node identifiers corresponding to the node feature representations and the training time points, the teacher training module arranges the node feature representations corresponding to different types of nodes for the same training domain at different training time points in chronological order, determining the node time series corresponding to different training domains. Finally, the teacher training module uses the domain knowledge graph, the node time series corresponding to different training domains, and the labels corresponding to the training domain samples to train the initial teacher model, obtaining the teacher model.
[0023] In one possible implementation, the temporal neural network includes a temporal convolutional network and a long short-term memory network. The teacher training module performs convolutional processing on the time series of nodes corresponding to different training domains based on the temporal convolutional network to obtain the time series output features corresponding to different training domains. Next, the teacher training module performs temporal fusion on the time series output features corresponding to different training domains based on the long short-term memory network to obtain fused features corresponding to different training domains. Finally, the teacher training module uses the domain knowledge graph, the fused features corresponding to different training domains, and the labels corresponding to the training domain samples to train the initial teacher model, thus obtaining the teacher model.
[0024] In one possible implementation, the temporal convolutional network comprises multiple stacked temporal convolutional layers. The teacher training module, based on these stacked temporal convolutional layers, performs dilated causal convolution processing on the time series of nodes corresponding to different training domains to obtain the time series output features corresponding to different training domains.
[0025] In one possible implementation, the distillation training module uses the teacher model and the initial student model to infer the fusion features corresponding to different training domain names, respectively, to obtain the output results of the teacher model and the initial student model. The difference between the teacher model output and the initial student model output is compared to obtain the distillation loss value. Then, based on the initial student model output and the labels corresponding to multiple training domain name samples, the distillation training module determines the classification loss value, which characterizes the degree of difference between the initial student model's output and the labels. Finally, based on the distillation loss value and the classification loss value, the distillation training module adjusts the model parameters of the initial student model to obtain the student model.
[0026] In one possible implementation, the apparatus further includes an incremental distillation module. The acquisition module is further configured to acquire newly added training domain name samples before inferring the fused features using the student model. The incremental distillation module is used to perform incremental distillation on the student model using the newly added training domain name samples to obtain an updated student model. The hyperparameters of the updated student model are the same as those of the original student model, including distillation temperature coefficient, distillation loss weights, and classification loss weights. The inference module is used to infer the fused features using the updated student model to obtain the recognition result.
[0027] In one possible implementation, the identification result includes a predicted probability that the domain name to be identified is a bad domain name, or a classification result used to characterize whether the domain name to be identified is a normal domain name or a bad domain name.
[0028] The malicious domain name identification device provided in the second aspect corresponds to the malicious domain name identification method provided in the first aspect. Therefore, the technical effects of any implementation method in the second aspect can be found in the relevant descriptions of the technical effects of the corresponding implementation methods in the first aspect, and will not be repeated here.
[0029] Thirdly, this application provides a computing device including a memory and a processor. The processor executes a computer program stored in the memory, causing the computing device to perform the malicious domain name identification method of the first aspect or any possible implementation thereof. The memory may be integrated into the processor or separate from the processor. The computing device may also include a bus, through which the processor and the memory can communicate.
[0030] Fourthly, this application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the malicious domain name identification method of the first aspect or any possible implementation thereof.
[0031] Fifthly, this application provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the malicious domain name identification method of the first aspect or any possible implementation thereof.
[0032] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments provided in this application. For those skilled in the art, other drawings can be obtained based on these drawings.
[0034] Figure 1 A schematic diagram of the structure of a malicious domain name identification system provided in this application; Figure 2 A flowchart illustrating a method for identifying malicious domain names provided in this application; Figure 3 A flowchart illustrating the process of obtaining a student model provided in this application; Figure 4 A schematic diagram illustrating the attribute information of various nodes in the domain name knowledge graph provided in this application; Figure 5 A schematic diagram illustrating the processing procedure of a temporal convolutional network and a long short-term memory network provided in this application; Figure 6 A schematic diagram of an output layer knowledge distillation training process provided in this application; Figure 7 A schematic diagram of a hierarchical knowledge distillation training process provided in this application; Figure 8 A schematic diagram of a malicious domain name identification device provided in this application; Figure 9 This is a schematic diagram of the hardware structure of a computing device provided in this application. Detailed Implementation
[0035] The terms "first," "second," etc., used in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application.
[0036] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. The drawings are used to illustrate data flow, processing relationships, and functional division of labor, and do not limit the components to be deployed in separate devices.
[0037] To make the application scenarios of this application and the data interaction relationships between its components clearer, the following will first explain the process of identifying malicious domain names from the system level.
[0038] See Figure 1 , Figure 1 A schematic diagram of a malicious domain name identification system is shown. The system includes a domain name data source 100, a sample database 200, a feature fusion device 300, a model training device 400, and a domain name identification device 500, which is used to output the identification results.
[0039] Domain data source 100 is used to provide domain name data to be identified to feature fusion device 300. The domain name data to be identified may include the domain name to be identified, domain name resolution data corresponding to the domain name at multiple time points, and time stamps corresponding to each set of domain name resolution data.
[0040] In one possible implementation, domain name resolution data may include data directly obtained through the domain name resolution process, as well as associated data obtained based on domain name resolution relationships. Specifically, data directly obtained through the domain name resolution process may include one or more of the following: whether the domain name is resolvable, the IP address resolved by the domain name, and the type of the resolution record. Associated data obtained based on domain name resolution relationships may include one or more of the following: the autonomous system to which the IP address belongs, the certificate used by the domain name, the domain name registrant information, geographical location information, whether the domain name is reachable, and the collection time corresponding to each data point. The domain name data source 100 may be a domain name resolution monitoring platform, a domain name data collection server, a domain name resolution database, or other data sources capable of providing historical domain name resolution records and associated data.
[0041] In practical applications, multiple time points can be determined according to a fixed collection cycle. For example, the domain data source 100 can collect domain name resolution data corresponding to the domain name to be identified every six hours or every day. In another possible implementation, multiple time points can also be determined based on the moment when domain-related information changes. For example, when the IP address resolved to by the domain name to be identified, certificate association, domain reachability status, or domain resolvability status changes, the domain data source 100 can determine the corresponding change moment as the new time point. In addition, multiple time points can also be selected from existing historical resolution records. The number of time points and the time interval between adjacent time points can be set according to the observation period, the number of historical records, and the input length of the temporal neural network.
[0042] A time stamp is used to distinguish the time position corresponding to each set of domain name resolution data. The time stamp can include one or more of the following: collection timestamp, sequence index, relative time interval, or time position encoding. For example, the domain name data source 100 can obtain the domain name resolution data corresponding to the domain name to be identified at the first, second, and third time points, and configure corresponding timestamps or sequence indices for each of the three sets of domain name resolution data, so that the feature fusion device 300 can identify the temporal relationship between the acquired domain name resolution data.
[0043] The sample database 200 stores multiple labeled training domain name samples and provides these samples to the feature fusion device 300. Each training domain name sample may include the training domain name, the domain name resolution data corresponding to the training domain name at multiple training time points, and the corresponding label. The label is used to characterize whether the training domain name is a normal domain name or a malicious domain name.
[0044] In one possible implementation, labeled training domain name samples can be divided into a training subset and a validation subset. The training subset is used to adjust the model parameters of the teacher model or student model, while the validation subset is used to evaluate the model's recognition performance during training and can be used to determine when to terminate model training. Unlabeled domain name data can be used as domain name data to be identified and inferred by the trained student model, without participating in supervised training based on real labels.
[0045] In practical applications, the domain name data source 100 and the sample database 200 can be implemented by the same data management system, database server, or data platform. This data management system can logically divide the stored data according to whether the data is tagged and its purpose. Data tagged with normal or malicious domain names can be provided as training domain name samples to the feature fusion device 300; data without tags can be provided as domain name data to be identified to the feature fusion device 300. Therefore, Figure 1 The domain name data source 100 and sample database 200 are shown separately, mainly to distinguish their functions in terms of data usage and processing stage, and do not imply that they must be implemented by different physical devices. In another possible implementation, the domain name data source 100 and sample database 200 can also be deployed in different data acquisition devices and data storage devices.
[0046] The feature fusion device 300 is used to process training domain name samples or domain name data to be identified. For training domain name samples, the feature fusion device 300 can process the data corresponding to the training domain name at multiple training time points and their temporal order relationship to obtain the fused features corresponding to the training domain name, and send the fused features corresponding to the training domain name to the model training device 400. For domain name data to be identified, the feature fusion device 300 can use the same attribute definition, encoding method and temporal processing method as in the training stage to obtain the fused features corresponding to the domain name to be identified, and send the fused features to the domain name identification device 500.
[0047] The model training device 400 is used to obtain teacher and student models based on training domain name samples. The model training device 400 can receive the fused features corresponding to the training domain names sent by the feature fusion device 300, and obtain the labels corresponding to the training domain name samples and the domain name knowledge graph used during training. The model training device 400 uses the fused features corresponding to the training domain names, the domain name knowledge graph, and the corresponding labels to train the initial teacher model, thus obtaining the teacher model.
[0048] In one possible implementation, the teacher model and the student model can have the same or similar input and output formats. For example, both the teacher and student models take the fusion features corresponding to the training domain name as input and output the predicted probabilities of the training domain name being a normal domain name and a bad domain name. The number of network layers, hidden units, convolutional channels, intermediate feature dimensions, or classification layer size of the student model can be smaller than those of the teacher model, thus reducing the number of parameters in the student model. After completing knowledge distillation training, the model training device 400 can send the trained student model to the domain name recognition device 500.
[0049] The domain name recognition device 500 is used to receive the fusion features corresponding to the domain name to be recognized sent by the feature fusion device 300, and the student model provided by the model training device 400, and to use the student model to reason about the fusion features corresponding to the domain name to be recognized to obtain the recognition result.
[0050] For example, the domain name recognition device 500 can input the fusion features corresponding to the domain name to be identified into a student model, and the student model can output the predicted probabilities of the domain name to be identified being a normal domain name and the domain name to be identified being a malicious domain name. The domain name recognition device 500 can output the predicted probability of the domain name to be identified being a malicious domain name as the recognition result, or it can compare the predicted probability with a preset classification threshold to obtain a classification result used to characterize whether the domain name to be identified is a normal domain name or a malicious domain name. For example, when the student model outputs a predicted probability of 0.82 for the domain name to be identified as a malicious domain name, the domain name recognition device 500 can directly output 0.82 as the recognition result. When the preset classification threshold is 0.5, the domain name recognition device 500 can also output the classification result of the domain name to be identified as a malicious domain name based on the fact that 0.82 is greater than 0.5.
[0051] In some possible implementations, the feature fusion device 300, the model training device 400, and the domain name recognition device 500 can be set up separately or integrated through functional modules. The feature fusion device 300 can be used as an independent device to interact with the model training device 400 and the domain name recognition device 500 respectively; it can also be integrated into the model training device 400 to process the domain name resolution data and their temporal order relationship corresponding to the training domain name at multiple training time points to obtain fused features for teacher model training and knowledge distillation training; or it can be integrated into the domain name recognition device 500 to process the domain name resolution data and their temporal order relationship corresponding to the domain name to be recognized at multiple time points to obtain fused features for student model inference. The model training device 400 and the domain name recognition device 500 can be deployed on different computing devices such as the model training server and the model inference server, respectively. The model training device 400 sends the trained student model to the domain name recognition device 500. Alternatively, the model training device 400 can be integrated as a functional module into the domain name recognition device 500, enabling the domain name recognition device 500 to simultaneously perform feature fusion, teacher model training, student model knowledge distillation training, and domain name inference. Therefore... Figure 1 The feature fusion device 300, model training device 400, and domain name recognition device 500 shown in the figure are mainly used to distinguish the feature fusion function, model training function, and model inference function, and do not limit the above functions to be implemented by mutually independent physical devices.
[0052] In practical applications, domain names typically serve as the entry point for terminals to access network resources in enterprise network egress points, recursive domain name resolution servers, or domain name security monitoring platforms. The same domain name may correspond to different resolution states and resolution addresses at different points in time. After these resolution results are continuously recorded, a resolution behavior trajectory of the domain name over a period of time can be formed.
[0053] For example, domain A resolves to IP address A1 continuously from day one to day three. On day four, due to server migration, it switches to IP address A2 and remains stable thereafter. Domain B resolves to IP addresses B1, B2, and B3 sequentially over several hours, with successes and failures alternating between adjacent time points. If only a single point in time is observed, both domain A and domain B may be in a normal resolution state, and the difference between them is not obvious. However, by combining the resolution data from multiple time points and their temporal relationships, we can further observe the duration of the resolved addresses, the switching process, and the changes in the resolution status, thereby obtaining dynamic information that cannot be presented at a single point in time.
[0054] Traditional methods for identifying malicious domains rely on blacklists, fixed rules, or domain characteristics at a single point in time. However, for new domains not yet on the blacklist, or malicious domains deliberately displaying normal DNS resolution at the current time, the DNS information at a single point in time may be quite similar to that of normal domains, making it difficult to establish sufficient criteria for classification. Historical DNS behavior is not directly equivalent to malicious domains; for example, content delivery networks or disaster recovery systems may also change their DNS addresses. Therefore, simply setting a fixed threshold based on the number of address changes or DNS failures may misclassify normal domains as malicious ones.
[0055] On the other hand, to learn more complex domain name behavior patterns from a large number of historical samples, models with a large number of parameters can be used for training and recognition. However, in enterprise gateways, regional recursive resolution nodes, or edge monitoring nodes, the number of domain name requests is large. If each deployment node directly runs a model with a large number of parameters, it will consume a lot of storage and computing resources and increase the inference time of a single domain name recognition. Therefore, there is still room for improvement in the existing solution regarding the sufficiency of recognition evidence and the model deployment overhead.
[0056] Based on this, Figure 1 In the illustrated malicious domain name identification system, taking the identification of a malicious domain name as an example, the domain name data source 100 can provide the domain name to be identified and the domain name resolution data corresponding to the domain name at multiple time points to the feature fusion device 300. The feature fusion device 300 can use a temporal neural network to perform feature fusion on the domain name resolution data corresponding to multiple time points and the temporal order relationship of multiple time points to obtain the fused features corresponding to the domain name to be identified. The domain name identification device 500 can call the student model to infer the fused features to obtain the identification result used to indicate whether the domain name to be identified is a malicious domain name. The student model can be obtained by training the initial student model through knowledge distillation using the teacher model based on training samples. The training samples include the sample fusion features and training labels corresponding to the sample domain names. The sample fusion features can be obtained by performing feature fusion on the domain name resolution data corresponding to the sample domain names at multiple training time points and the temporal order relationship of multiple training time points. The training labels can be used to indicate whether the sample domain name is a malicious domain name. The number of parameters in the student model is less than the number of parameters in the teacher model.
[0057] Therefore, this application embodiment does not rely solely on the domain name resolution data corresponding to the domain name to be identified at a single point in time for identification. Instead, it combines the domain name resolution data corresponding to the domain name to be identified at multiple points in time and the temporal order of these points to form a fusion feature that reflects the changes in the domain name resolution behavior over time. Even if normal and malicious domain names have similar resolution performance at a certain point in time, the malicious domain name identification system can distinguish them by combining the resolution information at different points in time and their sequential relationships, thereby providing supplementary judgment criteria in the time dimension for malicious domain name identification and improving the accuracy of malicious domain name identification. Simultaneously, this application embodiment uses a student model trained through knowledge distillation in the actual identification stage, instead of directly running a teacher model with a large number of parameters. This reduces the number of model parameters loaded and calculated in the inference stage, lowering model storage overhead, computational overhead, and inference latency. Therefore, the above method can balance the accuracy of malicious domain name identification with model operating efficiency and improve its applicability to application scenarios such as DNS security monitoring platforms, enterprise security gateways, and edge monitoring nodes.
[0058] Understandable, Figure 1 The structure of the malicious domain name identification system shown is merely an exemplary structure, used to illustrate the data interaction relationship between the components in this application embodiment, and does not constitute a limitation on the scope of protection of this application. For example, other possible malicious domain name identification systems may also include domain name resolution data preprocessing equipment, domain name knowledge graph construction equipment, node feature determination equipment, identification result display equipment, and data storage equipment, etc. This application embodiment does not limit the number, deployment location, or connection method of the components in the system.
[0059] In clear Figure 1 Based on the system architecture shown, see [link / reference]. Figure 2 , Figure 2 This diagram illustrates a method for identifying malicious domain names. This method can be applied to... Figure 1 The aforementioned malicious domain name identification system can also be applied to other applicable malicious domain name identification systems. The following section will focus on its application in... Figure 1 The following is an example of a malicious domain name identification system. Figure 2 The method shown includes steps S201 to S203.
[0060] S201: Feature fusion device 300 acquires the domain name to be identified and the domain name resolution data corresponding to the domain name at multiple time points.
[0061] For example, domain name resolution data can include data obtained directly through the domain name resolution process, or related data obtained further based on the resolution relationship. For instance, domain name resolution data can include one or more of the following: whether the domain name is resolvable, the IP address resolved to by the domain name, the autonomous system to which the IP address belongs, the certificate used by the domain name, the domain name registrant information, geographical location information, whether the domain name is reachable, and the time information corresponding to each piece of data.
[0062] In practical applications, the domain name to be identified can be provided by the domain name data source 100. Each piece of domain name resolution data in the domain name data source 100 is associated with a corresponding time point, enabling subsequent processing to distinguish data from different time points. For example, if the domain name to be identified is query-a.example, the feature fusion device 300 can obtain the resolution data corresponding to the domain name at the first, second, and third time points. At the first time point, the domain name resolves to the first IP address and uses the first certificate; at the second time point, the domain name resolves to the second IP address; at the third time point, the certificate authority and reachability status of the domain name change. The data from these multiple time points collectively reflect dynamic information such as IP address switching, certificate changes, and changes in domain name reachability status.
[0063] S202: The feature fusion device 300 uses a temporal neural network to fuse the domain name resolution data corresponding to the domain name to be identified at multiple time points and the temporal order relationship of multiple time points to obtain fused features.
[0064] In one possible implementation, the temporal neural network may include a data feature extraction branch and a temporal relationship extraction branch. The data feature extraction branch encodes and performs temporal processing on the domain name resolution data corresponding to each time point to obtain a first feature. The first feature is used to characterize the node attributes, relationships, and state change information contained in the domain name resolution data itself. The temporal relationship extraction branch can encode a temporal relationship input based on time identifiers, sequential indices, relative time intervals, or time positions, and extracts features from the temporal relationship input to obtain a second feature. The second feature is used to characterize the sequential order and interval relationships between multiple time points. For example, the first, second, and third time points can be encoded using sequential indices 1, 2, and 3, respectively; when the time intervals between adjacent time points are not the same, the temporal relationship extraction branch can also encode the time intervals between adjacent time points into the second feature.
[0065] The first feature and the second feature can respectively characterize the information contained in the domain name resolution data itself and the time sequence information between multiple time points. In one possible implementation, the feature fusion device 300 can fuse the first feature and the second feature through feature concatenation, weighted summation, or gated fusion to obtain a fused feature.
[0066] For example, the first feature can characterize the IP address type, certificate attributes, domain name reachability, and domain name resolvability at multiple time points, while the second feature can characterize the sequence and time intervals between multiple time points. After fusing the two types of features, the resulting fused feature can both retain the resolution status at each time point and characterize the chronological order of different state changes, thus distinguishing between different change processes such as "certificate change occurs first, then IP address switch" and "IP address switch occurs first, then certificate change."
[0067] In another possible implementation, the feature fusion device 300 can first arrange the domain name resolution data corresponding to the domain name to be identified at multiple time points in chronological order, forming a domain name resolution time series, and then use a temporal neural network to perform feature extraction and fusion processing on the domain name resolution time series. This implementation organizes the domain name resolution data and the chronological order relationship in the same sequence input, and can also obtain fused features that take into account both data content and chronological order.
[0068] S203: The domain name recognition device 500 uses a student model to reason about the fused features and obtain the recognition result.
[0069] The student model may include a classification network and an output layer. The classification network receives fused features and generates category representations, while the output layer outputs the probability that the domain name to be identified is a normal domain name and the probability that the domain name to be identified is a malicious domain name, based on the category representations. The domain name identification device 500 can directly output the probability that the domain name to be identified is a malicious domain name, or it can convert the probability into a classification result of normal or malicious domain name according to a preset classification rule. For example, when the probability of a malicious domain name output by the student model is 0.87, the identification result can be expressed as the probability of a malicious domain name is 0.87, or it can be expressed as "malicious domain name".
[0070] Figure 2 This explains how the student model is used in the domain name recognition stage. To further illustrate the process of obtaining the student model, the following section combines... Figures 3 to 7 This paper explains the relationship between training domain name sample acquisition, domain name knowledge graph construction, node temporal feature processing, teacher model training, and knowledge distillation training.
[0071] In one possible implementation, Figure 3 The process of obtaining the student model shown can be performed during execution Figure 2 Execute before step S203 in the process.
[0072] See Figure 3 , Figure 3 This diagram illustrates a process for obtaining a student model. This process can be carried out by... Figure 1The feature fusion device 300 and the model training device 400 work together, or they can be executed by a computing device that integrates feature processing and model training functions. Figure 3 The process shown includes steps S301 to S305.
[0073] S301: Feature fusion device 300 acquires multiple labeled training domain name samples.
[0074] Each training domain sample can include the training domain name, the domain name resolution data corresponding to the training domain name at multiple training time points, and corresponding labels. Training domain name samples can be derived from confirmed categories of normal domain name data, malicious domain name data, and their historical resolution data. Labels are used to characterize whether the corresponding training domain name is a normal or malicious domain name. For example, normal domain name samples can be labeled with label 0, and malicious domain name samples can be labeled with label 1; labels can also use one-hot encoding or other encoding methods that can distinguish between normal and malicious domain names.
[0075] For example, for a training domain name sample-a.example that has been identified as a malicious domain name, the feature fusion device 300 can obtain the training domain name at the training time point. , and The corresponding domain name resolution data is then configured, and malicious domain name tags are assigned to them. At the training time point... The training domain name resolves to the first IP address and uses the first certificate; at the training time point The training domain name resolves to the second IP address and continues to use the first certificate; at the training time point The training domain name continues to resolve to the second IP address, but the certificate used is changed to the second certificate. The above data can provide a data foundation for subsequent construction of a domain name knowledge graph, determination of node features, and extraction of temporal change features.
[0076] In one possible implementation, the labeled training domain name samples can be divided into a training subset and a validation subset. The training subset can be used to adjust the model parameters of the teacher model and the initial student model, while the validation subset can be used to evaluate the recognition performance of the model during training and to determine the end time of teacher model training or knowledge distillation training.
[0077] In practical applications, after acquiring training domain name samples, the feature fusion device 300 can preprocess the training domain name samples. Preprocessing can include one or more of the following: data format unification, category attribute encoding, Boolean attribute encoding, time attribute conversion, and numerical attribute normalization. For example, Boolean type attributes such as whether a domain name is reachable and whether it is resolvable can be encoded using 0 and 1; category type attributes such as registrar, affiliated organization, certificate authority, and country can be represented by dictionary indexing, one-hot encoding, or embedding vectors; time type attributes such as registration time, expiration time, certificate issuance time, and certificate expiration time can be converted into timestamps, or the registration duration, remaining validity period, relative time interval, or normalized value can be calculated based on the corresponding time; continuous numerical type attributes can be normalized or standardized according to a preset value range.
[0078] The training and domain name recognition phases can use the same attribute definitions, encoding rules, and feature processing parameters to ensure that the training domain name samples and the domain name data to be recognized have a consistent data representation. For example, the categorical attribute encoding dictionary, numerical normalization parameters, and time attribute transformation rules used in the training phase can be continued in the domain name recognition phase.
[0079] S302: Feature fusion device 300 constructs a domain name knowledge graph based on multiple labeled training domain name samples.
[0080] The domain name knowledge graph includes multiple domain name nodes and associated nodes. Domain name nodes indicate training domain names, and associated nodes can include at least one of IP address nodes, autonomous system nodes, domain name registrant nodes, geographic location nodes, and certificate nodes. The feature fusion device 300 can establish connections between different types of nodes based on the domain name resolution relationships, certificate usage relationships, autonomous system affiliation relationships, and other associations contained in the training domain name samples. For example, when a training domain name resolves to an IP address, the feature fusion device 300 can establish a connection between the corresponding domain name node and the IP address node; when a training domain name uses a certificate, the feature fusion device 300 can establish a connection between the corresponding domain name node and the certificate node; when a network segment corresponding to an IP address belongs to an autonomous system, the feature fusion device 300 can establish a connection between the corresponding IP address node and the autonomous system node. Taking a training domain name as an example, the training domain name sample-a.example at the training time point... Resolve to IP address and use certificates IP address The corresponding network segment belongs to the autonomous system. The feature fusion device 300 can establish the domain name node and IP address node corresponding to the training domain name sample-a.example. The connection between them, the domain name node and the certificate node The connections between them, and the IP address nodes With autonomous system nodes The connections between them. Therefore, the training domain name at the training time point. The corresponding domain name resolution relationships and associations can be recorded in the domain name knowledge graph.
[0081] In one possible implementation, the associated nodes include IP address nodes, autonomous system nodes, and certificate nodes, and the domain name knowledge graph can be an undirected weighted graph. Domain name nodes are connected to IP address nodes via a first edge, to certificate nodes via a second edge, and to autonomous system nodes via a third edge. The weight of the first edge is greater than the weight of the second edge, and the weight of the second edge is greater than the weight of the third edge. For example, the weights of the first, second, and third edges can be set to 3, 2, and 1, respectively. Different edge weights can be used to distinguish the relative importance of different types of connections; the aforementioned values are merely examples and do not limit the actual weight values used. The weights corresponding to the first, second, and third edges can be stored as attributes of the corresponding connection edges in the domain name knowledge graph.
[0082] Furthermore, domain name registrants and geographical locations can be set as independent nodes. For example, the feature fusion device 300 can establish a connection between a domain name node and its corresponding domain name registrant node, or between an IP address node and its corresponding geographical location node. In another possible implementation, domain name registration information and geographical location information can also be stored as attributes of other nodes. For example, domain name registrars and domain name registrants can be attributes of domain name nodes, the country and province of an IP address can be attributes of IP address nodes, and the autonomous system's region can be attributes of autonomous system nodes.
[0083] In practical applications, multiple training domains can share the same associated node. For example, when training domains... and training domain name All resolved to IP addresses At that time, the domain name knowledge graph can be used to build training domain names separately. Corresponding domain name nodes and IP address nodes The first side between them, and the training domain name Corresponding domain name nodes and IP address nodes The first side between them. Thus, the domain knowledge graph can record situations where multiple training domains share the same IP address node. Similarly, when multiple training domains use the same certificate at the same or different training points, these training domains can share the corresponding certificate node; when the IP addresses resolved by multiple training domains belong to the same autonomous system, these training domains can be indirectly associated with the same autonomous system node through their respective associated IP address nodes. By sharing associated nodes, the domain knowledge graph can preserve the indirect relationships formed between different training domains, enabling subsequent model training processes to combine the node attributes of individual training domains with the association information between multiple training domains for feature learning.
[0084] Through step S302, the feature fusion device 300 can determine the node types and the connection relationships between nodes in the domain name knowledge graph. When the domain name knowledge graph is an undirected weighted graph, the feature fusion device 300 can also determine the weights corresponding to different types of connection relationships. To further illustrate the attribute information that various types of nodes in the domain name knowledge graph can contain, the following section combines... Figure 4 Please provide an explanation.
[0085] See Figure 4 , Figure 4 It displays the attribute information of various nodes in the domain knowledge graph. Figure 4 Taking domain name nodes, IP address nodes, autonomous system nodes, and certificate nodes as examples, the document illustrates the attribute names, meanings, and lengths that each type of node can include. These node attributes can be used in subsequent step S303 to determine the node features corresponding to each node at different training time points.
[0086] The attributes of a domain node can include one or more of the following: domain name, registration time, expiration time, registrar, registrant, domain status, domain reachability, and domain resolvability. The domain name can be used to distinguish different domain nodes; the registration time, expiration time, registrar, and registrant can be used to represent the domain's registration information; and the domain status, domain reachability, and domain resolvability can be used to represent the domain's operational status at the corresponding training time point.
[0087] For example, some attribute values of the same domain name node may differ at different training time points. For instance, the training domain name at different training time points... In a parseable state, at the training time point It is still in a parseable state at the training time point. It becomes unresolvable. The feature fusion device 300 can determine the domain name node at each training time point based on the attribute values collected at each training time point. , and The corresponding node features.
[0088] It should be noted that the situation where a training domain name resolves to different IP addresses at different training time points can be represented by the connection relationship between the domain name node and different IP address nodes. For example, the training domain name at training time point... The first IP address was resolved at the training time point. When the second IP address is resolved, the domain knowledge graph can record the association between the domain node and the first IP address node and the second IP address node at the corresponding training time points.
[0089] An IP address node's attributes may include one or more of the following: country of origin, country abbreviation, province, organization, whether it is a name server IP address, whether it corresponds to an A record, and whether it is a recursive server IP address. These attributes can be used to characterize the geographical location of the IP address, the organization to which it belongs, and the purpose of the IP address in the domain name resolution process.
[0090] The attributes of an autonomous system node can include the region information corresponding to the autonomous system. In addition, autonomous system nodes can also include attributes such as autonomous system number, autonomous system name, or operating entity to represent the identity and affiliation information of the corresponding autonomous system.
[0091] The attributes of a certificate node can include one or more of the following: the algorithm used, the certificate authority, the country of the certificate authority, the issuance date, and the expiration date. These attributes can be used to characterize the encryption algorithm used by the corresponding certificate, the certificate's origin, and its validity period. For example, when a training domain uses different certificates at different training points, the domain knowledge graph can record the associations between domain nodes and different certificate nodes at different training points. For instance, a training domain at a training point... and Using the first certificate, at the training time point When using a second certificate, the domain knowledge graph can record the relationships between domain nodes and both the first and second certificate nodes. Each certificate node carries attributes such as its algorithm, certificate authority, issuance date, and expiration date.
[0092] Figure 4 The listed node attributes and attribute lengths are examples of one implementation. In other implementations, the feature fusion device 300 can select based on the available data types and the input format of the model. Figure 4 The listed attributes can be supplemented with additional attributes that can represent domain name nodes and associated nodes. Furthermore, the attribute length can be set according to the data format and encoding method of the corresponding attribute, and is not limited to a fixed length. Figure 4The specific values shown.
[0093] In practical applications, the feature fusion device 300 can perform the preprocessing as described in step S301. Figure 4 The different types of attributes shown are encoded or numerically converted to obtain attribute data that can be used to determine node characteristics.
[0094] After clarifying the attributes that various nodes in the domain name knowledge graph can contain, the processing flow can return to... Figure 3 Continue with step S303 to determine the node features corresponding to each node at different training time points.
[0095] S303: Feature fusion device 300 determines the node features of multiple nodes in the domain name knowledge graph at different training time points based on the domain name resolution data corresponding to multiple training domain names at multiple training time points.
[0096] Specifically, the feature fusion device 300 can form the original node features corresponding to each node at different training time points based on the attribute values of each node at the corresponding training time points, and determine the set of associated nodes corresponding to each training domain name at the corresponding training time point based on the effective connection relationships in the domain name knowledge graph at each training time point.
[0097] In practical applications, for node types nodes The feature fusion device 300 can fuse nodes according to a preset attribute order. At training time points The corresponding attribute values are encoded or converted into numerical values and then organized into nodes. At training time points The corresponding original node features. These original node features can be represented as:
[0098] in, Indicates the node type. Indicates the node identifier. Indicates node type The corresponding number of attributes, Represents a node At training time points The corresponding number Item attribute value, Represents a node At training time points The corresponding original node features.
[0099] For example, for domain name nodes, the feature fusion device 300 can form corresponding domain name node features according to a preset order of attributes such as registration time, expiration time, registrar, registrant, domain name status, domain name reachability, and domain name resolvability; for IP address nodes, autonomous system nodes, and certificate nodes, node features can also be formed according to their respective attribute order.
[0100] Some node attributes can remain unchanged at different training points in time, while others can change over time. For example, the domain name string can remain constant, while attribute values such as domain status, domain reachability, and domain resolvability can change. For certificate nodes, attributes such as certificate authority, issuance date, and expiration date typically remain unchanged after the certificate is generated; when a training domain changes its certificate, the domain knowledge graph can characterize the certificate change by associating the training domain with different certificate nodes at different training points in time.
[0101] As an implementation example, the set of associated nodes can include the domain name node corresponding to the training domain name, the IP address node and certificate node associated with the domain name node, and the autonomous system node associated through the IP address node. Taking the aforementioned training domain name sample-a.example as an example, at the training time point... The feature fusion device 300 can determine the domain name node corresponding to the training domain name, and at the training time point. The node characteristics corresponding to the associated IP address nodes, certificate nodes, and autonomous system nodes. At the training time point. When the IP address resolved by the training domain name changes while the associated certificate remains unchanged, the feature fusion device 300 can adjust the feature based on the training time point. Valid connection relationships determine the node characteristics corresponding to the changed IP address node, the unchanged certificate node, and the autonomous system node associated with the changed IP address node.
[0102] Therefore, the differences in the set of node features corresponding to the training domain at different training time points can originate from changes in the attribute values of the same node, or from changes in the nodes associated with the training domain. The feature fusion device 300 can establish associations between each node feature and its corresponding node identifier, node type, training domain identifier, and training time point, so that the model training device 400 can distinguish the node features corresponding to different node types, different training domains, and different training time points.
[0103] Subsequently, step S304 can be executed to train the initial teacher model using the node features corresponding to multiple nodes at different training time points and the labels corresponding to the training domain name samples.
[0104] S304: The model training device 400 uses the node features corresponding to multiple nodes in the domain knowledge graph at different training time points, as well as the labels corresponding to multiple training domain samples, to train the initial teacher model and obtain the teacher model.
[0105] Before training the initial teacher model, the model training device 400 can... Figure 4 The original node features of different types of nodes are mapped to the same feature space, and the mapped node features are organized according to the training time order. Since the attribute types and original feature dimensions corresponding to different types of nodes may be different, the model training device 400 can unify the feature representation of different types of nodes through feature mapping.
[0106] In one possible implementation, for node type nodes At training time points The corresponding original node features can be used by the model training device 400 to utilize node types. The corresponding feature mapping matrix is processed to obtain the mapped node feature representation:
[0107] in, Indicates the node type. Indicates the node identifier. Represents a node At training time points The corresponding original node features, Indicates node type The corresponding feature mapping matrix, Indicates the bias term. This represents the activation function. This represents the node feature representation after mapping to the same feature space. Through the above processing, the node feature representations of different types of nodes can have the same feature dimension, so that they can jointly participate in subsequent feature organization and temporal processing.
[0108] In another possible implementation, the model training device 400 can also perform feature mapping through fully connected layers, embedding layers, or other feature transformation networks. When the domain knowledge graph also includes domain registrant nodes or geographic location nodes, the model training device 400 can also set corresponding feature mapping parameters for the respective node types.
[0109] After completing the feature mapping of different types of nodes, the model training device 400 can obtain the training domain name determined in step S303. At training time points The corresponding domain name nodes and associated nodes are analyzed, and the node feature representations after the above node mapping are fused. Specifically, when a node type is at a certain training time point... When there is only one node, the model training device 400 can directly use the node feature representation after mapping that node as the corresponding type feature representation. When a node type corresponds to multiple nodes, the model training device 400 can obtain the corresponding type feature representation through summation, averaging, pooling, or weighted aggregation. Among them, the weights used in weighted aggregation can be determined according to the weights of the connection edges between nodes, or they can be learned during model training.
[0110] After obtaining the type feature representations corresponding to different node types, the model training device 400 can fuse the type feature representations corresponding to domain name nodes, IP address nodes, autonomous system nodes, and certificate nodes through methods such as concatenation, summation, averaging, or weighted fusion to obtain the training domain name. At training time points The corresponding static features of the nodes. The above feature fusion methods can be used individually or in combination.
[0111] Training domain At training time points The corresponding static features of the nodes can be denoted as: Node static features are used to characterize the attribute and association states of the training domain name and its associated nodes at the corresponding training time points.
[0112] After completing the static features of nodes corresponding to each training time point, the model training device 400 can arrange the static features of nodes corresponding to the same training domain name at different training time points in chronological order of training time to obtain the node time series corresponding to the training domain name:
[0113] in, Indicates the training domain name The corresponding node time series, This indicates the number of training time points. Indicates the training domain name At training time points The corresponding node static features.
[0114] The model training device 400 can be trained according to Figure 5 The processing method shown applies to node time series. Temporal processing is performed to obtain the fusion features used to train the initial teacher model.
[0115] See Figure 5 , Figure 5This demonstrates the processing steps of a temporal convolutional network and a long short-term memory network. Figure 5 The upper part illustrates the formation process of the node time series corresponding to the training domain name. After feature mapping, the domain name nodes and associated node features corresponding to each training time point are arranged according to the node identifier and training time order to form the node time series corresponding to the training domain name. Figure 5 The time arrangement in the table is used to represent the temporal relationship between different training time points, and does not mean that the node features corresponding to the previous training time point are used to generate the node features corresponding to the next training time point.
[0116] In one possible implementation, the temporal neural network used to process the node time series includes a temporal convolutional network (TCN) and a long short-term memory (LSTM) network. The model training device 400 can process the node time series... The input to the TCN (Tracking Network Convolutional Network) is processed by dilated causal convolution through multiple stacked temporal convolutional layers to obtain time-series output features corresponding to different time positions. Specifically, the first temporal convolutional layer receives the node's time series, and subsequent temporal convolutional layers receive the features output from the previous temporal convolutional layer. The TCN can extract local change features formed between adjacent or close training time points, such as local temporal features corresponding to IP address switching, certificate replacement, domain name status changes, or changes in resolvable status.
[0117] For the The temporal convolutional layer is at the _th ... The output feature corresponding to each time position can be represented as:
[0118] in, Indicates the first The temporal convolutional layer at the ... Output features corresponding to each time position This represents the output feature of the convolutional layer at the corresponding historical time position in the previous time step. Indicates the first The convolutional layer at each time point is located at the convolutional kernel position. The corresponding convolution parameters, Indicates the kernel size. Indicates the first The dilation coefficient corresponding to each convolutional layer at each time step.
[0119] For the first convolutional layer, the training domain name At training time points The corresponding static features of the nodes can be used as input features for the corresponding time positions, that is:
[0120] in, Indicates the training domain name At training time points The corresponding node static features.
[0121] Dilated convolution can expand the receptive range of temporal convolutional layers without continuously increasing the kernel size, enabling different temporal convolutional layers to extract changing features across different time spans. Causal convolution allows training time points to be... The corresponding output features are based only on the training time points. The static features of nodes corresponding to previous training time points are determined, thereby maintaining the temporal direction of the node time series.
[0122] After completing the node time series processing, TCN can output time series output features corresponding to multiple time positions. The model training device 400 can input the time series output features into the LSTM in chronological order, and the LSTM can associate feature information over a longer time range to obtain the temporal dynamic features corresponding to the training domain.
[0123] In one possible implementation, TCN is in the... The time series output features of each time position are: The hidden state of the LSTM at the previous time position is The cell state corresponding to the previous time position was The LSTM processing procedure can be represented as follows:
[0124]
[0125]
[0126]
[0127]
[0128]
[0129] in, Indicates the output of the forget gate. Indicates the input gate output. Indicates the state of candidate cells. Indicates the current cell state. Indicates the output of the output gate. Indicates the current hidden state. This indicates that corresponding elements are multiplied.
[0130] The forget gate controls the degree to which historical information is retained, the input gate controls the degree to which new information corresponding to the current time position is written, the cell state can fuse historical information and current information, and the output gate controls the output content of the current hidden state.
[0131] As an implementation example, the model training device 400 can use the hidden state of the LSTM at the last training time point as the temporal dynamic feature corresponding to the training domain name. Using the training domain name... For example, the time-series dynamic features can be represented as:
[0132] in, This represents the hidden state of the LSTM at the last training time point. Indicates the training domain name The corresponding temporal dynamic features. In another possible implementation, the model training device 400 can also perform pooling, weighted aggregation, or concatenation on the hidden states corresponding to multiple training time points to obtain temporal dynamic features.
[0133] The model training device 400 can directly use temporal dynamic features as the fused features corresponding to the training domain, such as... Figure 5 As shown. In Figure 5 In another implementation not shown, the model training device 400 can also fuse the temporal dynamic features with the summaries of the node static features corresponding to multiple training time points, so that the resulting fused features simultaneously characterize the state information of the training domain name and its associated nodes and the change pattern of the aforementioned state over time.
[0134] return Figure 3 In step S304, the aforementioned fusion features are formed based on the effective node associations at each training time point in the domain name knowledge graph. These features can carry association information between the training domain name and IP address nodes, autonomous system nodes, and certificate nodes, as well as the temporal information of these associations as they develop over training time. The model training device 400 can input the fusion features corresponding to different training domain names into the initial teacher model, which then outputs a prediction result indicating whether the corresponding training domain name is a normal or problematic domain name.
[0135] The training domain name samples used to train the initial teacher model can simultaneously include normal domain name samples and bad domain name samples. The model training device 400 can determine the teacher model training loss based on the difference between the prediction results output by the initial teacher model and the corresponding labels, and adjust the model parameters of the initial teacher model based on the teacher model training loss, so that the initial teacher model learns the features of normal domain names, the features of bad domain names, and the feature differences between the two types of domain names.
[0136] The model training device 400 can repeatedly execute the processes of fused feature input, prediction result output, teacher model training loss calculation, and model parameter adjustment. When the preset number of training rounds is reached, the teacher model training loss meets the preset conditions, or the recognition performance corresponding to the validation subset meets the preset termination conditions, the model training device 400 can end the parameter adjustment and obtain the teacher model.
[0137] S305: The model training device 400 uses the teacher model to perform knowledge distillation on the initial student model to obtain the student model.
[0138] After the teacher model is trained, the model training device 400 can enter the knowledge distillation stage. In the knowledge distillation stage, the model parameters of the teacher model can remain unchanged, and the model training device 400 can adjust the model parameters of the initial student model based on the output results of the teacher model, the output results of the initial student model, and the labels corresponding to the training domain name samples.
[0139] In practical applications, the training domain name samples used for knowledge distillation training can include both normal domain name samples and bad domain name samples. The model training device 400 can input the fused features corresponding to the normal domain name samples and bad domain name samples into the teacher model and the initial student model, respectively, so that the initial student model learns the category features corresponding to the two types of domain names and the category differences between the two types of domain names.
[0140] The initial student model has fewer model parameters than the teacher model. For example, the initial student model can have fewer network layers, fewer hidden units, fewer network channels, or smaller feature dimensions, thereby reducing the computational and storage resources required for model inference.
[0141] In one possible implementation, the teacher model may include a teacher-side TCN-LSTM module, a Transformer multi-head self-attention encoding layer, and a teacher classification layer. The teacher-side TCN-LSTM module can further encode the fused features of the input teacher model to obtain teacher temporal features; the Transformer multi-head self-attention encoding layer can perform attention encoding on the teacher temporal features to obtain teacher attention-encoded features; and the teacher classification layer can output classification scores based on the teacher attention-encoded features, indicating whether the training domain name belongs to the normal domain name category or the bad domain name category.
[0142] Accordingly, the initial student model can include a student-side TCN-LSTM module, a shallow convolutional mapping layer, and a student classification layer. The student-side TCN-LSTM module can further encode the fused features input to the initial student model to obtain student temporal features; the shallow convolutional mapping layer can perform convolutional mapping on the student temporal features to obtain student mapping features; the student classification layer can output classification scores based on the student mapping features, indicating whether the training domain name belongs to the normal domain name category or the bad domain name category.
[0143] The teacher-side TCN-LSTM module and the student-side TCN-LSTM module can use the same or different network structures. Specifically, the student-side TCN-LSTM module can have fewer network layers, fewer convolutional channels, fewer hidden units, or fewer feature dimensions than the teacher-side TCN-LSTM module. The initial student model may not include a Transformer multi-head self-attention encoding layer, but instead uses shallow convolutional mapping layers with fewer parameters to process the student's temporal features, thereby reducing the number of model parameters and inference computation in the student model. For example, the shallow convolutional mapping layer can include one or more convolutional layers.
[0144] It should be noted that, Figure 5 The TCN and LSTM shown are used to form fusion features corresponding to training domain names based on node time series. The aforementioned teacher-side TCN-LSTM module and student-side TCN-LSTM module are optional further feature encoding structures within the teacher model and the initial student model, and they are at different processing stages. In this optional implementation, the fusion features input to the teacher model and the initial student model can include multiple feature components organized in training time order, for further feature encoding by the teacher-side TCN-LSTM module and the student-side TCN-LSTM module.
[0145] See Figure 6 , Figure 6 This illustrates a knowledge distillation process for the output layer. The model training device 400 can input the fused features corresponding to the same training domain name obtained in step S304 into the teacher model and the initial student model, respectively. The teacher model and the initial student model then output classification scores indicating whether the training domain name belongs to each candidate category. These classification scores reflect the predictive tendency of the corresponding model for the training domain name belonging to each candidate category.
[0146] Teacher model for categories The output classification score can be denoted as The initial student model targets categories. The output classification score can be denoted as During the knowledge distillation training process, the model training device 400 can use the high-temperature Softmax algorithm to process the classification scores output by the teacher model and the initial student model according to the distillation temperature coefficient, and obtain the softened class probabilities corresponding to the teacher model and the initial student model respectively.
[0147] Teacher model for categories The output softening class probability can be expressed as:
[0148] The initial student model is categorized. The output softening class probability can be expressed as:
[0149] in, This indicates that the distillation temperature coefficient is At that time, the training domain name output by the teacher model Category The probability of the softening category; This indicates that the distillation temperature coefficient is At that time, the initial student model outputs the training domain name Category The probability of the softening category; This indicates that the teacher model is for the training domain name. Category Output classification score; This indicates that the initial student model is for the training domain. Category Output classification score; Indicates the distillation temperature coefficient; Indicates the number of candidate categories; This represents the category index used to iterate through each candidate category in the Softmax denominator; This indicates the category index for which the current calculation of the softening category probability is performed.
[0150] In this embodiment of the application, the candidate categories include normal domain name categories and bad domain name categories, therefore , and Both can be 1 or 2. For example, category index 1 can correspond to the normal domain category, and category index 2 can correspond to the bad domain category.
[0151] Compared to the case where the distillation temperature coefficient is 1, when the distillation temperature coefficient... When the value is greater than 1, the class probability distributions output by the teacher model and the initial student model are smoother. The softened class probability output by the teacher model can serve as a soft label, indicating both the predicted class determined by the teacher model and reflecting the teacher model's relative predictive tendency towards normal and malicious domain name categories. The model training device 400 can make the softened class probability output by the initial student model close to the softened class probability output by the teacher model, so that the initial student model learns the class judgment information formed by the teacher model.
[0152] The model training device 400 can determine the distillation loss based on the difference between the softened class probabilities output by the teacher model and the initial student model. Specifically, the distillation loss can be calculated using Kullback-Leibler divergence, also known as relative entropy. This applies to models including... For a training batch of training domain name samples, the KL divergence distillation loss can be expressed as:
[0153] in, This indicates the KL divergence distillation loss. This indicates the number of training domain name samples included in the training batch. This represents the index of the training domain sample. Indicates the first in the training batch One training domain name, Indicates the number of candidate categories. Indicates the index of the candidate category.
[0154] KL divergence distillation loss is used to characterize the difference between the softened class probability distributions output by the teacher model and the initial student model. By reducing the KL divergence distillation loss, the model training device 400 can make the softened class probability distribution output by the initial student model closer to that output by the teacher model, thereby enabling the initial student model to learn the class judgment information formed by the teacher model based on the graph association information and temporal change information of the training domain names.
[0155] In addition to the distillation loss, the model training device 400 can also determine the classification loss based on the difference between the class prediction probabilities output by the initial student model and the true labels corresponding to the training domain name samples. For example, the model training device 400 can set the temperature to 1 and convert the classification scores output by the initial student model into class prediction probabilities using the standard Softmax algorithm:
[0156] in, This indicates the training domain name output by the initial student model when the temperature is 1. Category The predicted probability of the category.
[0157] As an example implementation, the classification loss can be the cross-entropy classification loss. The cross-entropy classification loss can be expressed as:
[0158] in, Represents the cross-entropy classification loss. Indicates the first training domains Does it belong to a category? The label value. When training domains Category hour, It can be set to 1; when the training domain name Not belonging to category hour, It can be 0.
[0159] Cross-entropy classification loss measures the difference between the class prediction probability output by the initial student model and the true label corresponding to the training domain sample. Reducing the cross-entropy classification loss can make the class prediction result output by the initial student model closer to the true class corresponding to the training domain sample.
[0160] The model training device 400 can determine the joint training loss corresponding to the initial student model based on the distillation loss and the classification loss. When the distillation loss uses KL divergence distillation loss and the classification loss uses cross-entropy classification loss, the joint training loss can be expressed as:
[0161] in, Indicates joint training losses, This represents the weight corresponding to the KL divergence distillation loss. The weights represent the weights corresponding to the cross-entropy classification loss. and All are non-negative numbers. For example, and The sum can be 1. Factors It can be used to compensate for the effect of the distillation temperature coefficient on the magnitude of the distillation loss gradient, so that the distillation loss and classification loss maintain an appropriate order of magnitude during joint training.
[0162] The model training device 400, by jointly using KL divergence distillation loss and cross-entropy classification loss, enables the initial student model to simultaneously learn the class probability relationship output by the teacher model and the true class corresponding to the training domain samples. The model training device 400 can adjust the model parameters of the initial student model based on the joint training loss and repeatedly execute the processes of fusing feature input, teacher model inference, initial student model inference, joint training loss calculation, and initial student model parameter adjustment. When a preset number of distillation training rounds are reached, the joint training loss meets preset conditions, or the recognition performance corresponding to the validation subset meets preset termination conditions, the model training device 400 can terminate parameter adjustment, resulting in the student model.
[0163] See Figure 7 , Figure 7 This demonstrates a hierarchical knowledge distillation training process. Specifically, the model training device 400 can input the fused features corresponding to the same training domain name into the teacher model and the initial student model respectively. Figure 6 Based on the knowledge distillation of the output layer, the intermediate features of the teacher and the student that need to be aligned are determined according to the preset correspondence between different network layers in the teacher model and the initial student model, and the intermediate feature distillation loss is determined based on the difference between the two. When the feature dimensions of the teacher's intermediate features and the student's intermediate features are different, the model training device 400 can perform feature mapping on at least one of them to make them have the same feature dimensions.
[0164] In one possible implementation, when the teacher model includes a teacher-side TCN-LSTM module and a Transformer multi-head self-attention encoding layer, and the initial student model includes a student-side TCN-LSTM module and a shallow convolutional mapping layer... Figure 7 The hierarchical knowledge distillation training process shown can include TCN-LSTM feature distillation, attention-encoded feature distillation, and output layer soft label distillation.
[0165] During the TCN-LSTM feature distillation process, the model training device 400 can acquire the teacher's temporal features output by the teacher-side TCN-LSTM module and the student's temporal features output by the student-side TCN-LSTM module, and determine the first intermediate feature distillation loss based on the difference between the teacher's temporal features and the student's temporal features. The first intermediate feature distillation loss can be used to enable the student-side TCN-LSTM module to learn the temporal feature representation information formed by the teacher-side TCN-LSTM module.
[0166] During the attention-encoding feature distillation process, the model training device 400 can acquire the teacher attention-encoding features output by the Transformer multi-head self-attention encoding layer and the student mapping features output by the shallow convolutional mapping layer, and determine the second intermediate feature distillation loss based on the difference between the teacher attention-encoding features and the student mapping features. The second intermediate feature distillation loss can be used to enable the shallow convolutional mapping layer to learn the feature association representation information formed by the Transformer multi-head self-attention encoding layer.
[0167] During the output layer soft label distillation process, the model training device 400 can employ... Figure 6 The high-temperature Softmax processing method described above uses the softening category probability output by the teacher model as a soft label, and determines the output layer distillation loss based on the difference between the softening category probabilities output by the teacher model and the initial student model. The output layer distillation loss can be the aforementioned KL divergence distillation loss.
[0168] The model training device 400 can determine the joint training loss corresponding to hierarchical knowledge distillation based on the first intermediate feature distillation loss, the second intermediate feature distillation loss, the output layer distillation loss, and the classification loss, and adjust the model parameters of the initial student model based on the joint training loss.
[0169] The aforementioned hierarchical knowledge distillation is one possible implementation. In other implementations, the model training device 400 may also omit the first intermediate feature distillation loss and the second intermediate feature distillation loss, and instead adjust the model parameters of the initial student model solely based on the aforementioned output layer distillation loss and classification loss.
[0170] pass Figures 3 to 7 After obtaining the student model through the processing shown, the model training device 400 can also perform incremental distillation training on the student model based on the newly added training domain name samples to obtain an updated student model.
[0171] In one possible implementation, the newly added training domain name samples may include newly discovered malicious domain name samples, newly confirmed normal domain name samples, and domain name resolution data and tags corresponding to the newly added domain names at multiple time points. The newly added training domain name samples may also include domain name resolution data and corresponding tags generated by domain names already included in the training at the newly added time points. For example, when the resolution address, associated certificate, autonomous system, registration status, or other attributes of existing training domain names change, the model training device 400 can generate new training domain name samples based on the changed domain name resolution data and corresponding tags.
[0172] The model training device 400 can perform attribute processing, domain name knowledge graph construction, node feature mapping, node time series generation, and time series feature fusion on the newly added training domain name samples according to the data processing methods adopted in the aforementioned steps S301 to S304, so as to obtain the fused features corresponding to the newly added training domain name samples.
[0173] For example, the model training device 400 can use the current student model as the student model to be updated, and the teacher model obtained in step S304 as the teacher model in the incremental distillation training stage, and input the fusion features corresponding to the newly added training domain name samples into the teacher model and the student model to be updated, respectively. The model training device 400 can determine the incremental training loss based on the output results of the two and the labels corresponding to the newly added training domain name samples.
[0174] As an example implementation, incremental training loss can include distillation loss and classification loss determined for newly added training domain name samples. Incremental training loss can be expressed as:
[0175] in, This represents the incremental training loss; This represents the KL divergence distillation loss calculated for newly added training domain name samples; This represents the cross-entropy classification loss calculated for newly added training domain name samples; Indicates the distillation temperature coefficient; Indicates the weight of distillation loss; This represents the classification loss weight. and The calculation method can be referred to Figure 6 The calculation method.
[0176] In practical applications, the distillation temperature coefficient, distillation loss weight, and classification loss weight used in incremental distillation training can be the same as the corresponding parameters used in the aforementioned knowledge distillation training. The student models before and after the update can use the same model structure, and incremental distillation training is mainly used to update the model parameters of the student models. In other implementations, the model training device 400 can also adjust at least one of the above training parameters according to the number of newly added training domain name samples, category distribution, or feature changes.
[0177] In the case of hierarchical knowledge distillation in the aforementioned knowledge distillation training, incremental distillation training can also utilize the correspondence and intermediate feature alignment between the intermediate network layers of the teacher model and the student model. The model training device 400 can determine the intermediate feature distillation loss based on the difference between the intermediate features of the teacher model and the intermediate features of the student model to be updated, and then correlate the intermediate feature distillation loss with... and Together they are used to adjust the model parameters of the student model that is to be updated.
[0178] As another implementation example, the model training device 400 can also combine newly added training domain name samples with at least a portion of historical training domain name samples to form an incremental training dataset, and adjust the parameters of the student model to be updated based on the incremental training dataset. Historical training domain name samples can be used to maintain the learning results of the student model on the features and category judgment information of existing training domain names.
[0179] Through incremental distillation training, the model training device 400 can learn the normal domain characteristics, bad domain characteristics, and category judgment information between the two types of domains contained in newly added training domain samples based on the current student model, without having to re-execute based on all historical training domain samples each time. Figures 3 to 7 The complete model training process is shown, thereby reducing the training overhead required for student model updates.
[0180] In summary, through Figures 3 to 7 The processing steps shown can produce a student model with fewer parameters than the teacher model. The student model can learn the correspondence between fused features and true categories through classification loss, and learn the category judgment information formed by the teacher model through distillation loss.
[0181] After obtaining the student model or an updated student model, the model training device 400 can deploy it to the domain name recognition device 500, and the processing flow can then return. Figure 2 Execute step S203. The domain name recognition device 500 can use the student model to output the predicted probability or classification result that the domain name to be identified is a bad domain name, without calling the teacher model.
[0182] The above combination Figures 1 to 7 This section explains the malicious domain name identification system, domain name identification methods, teacher model training, and knowledge distillation process. The following section combines... Figure 8 Explain the functional structure of the malicious domain name identification device.
[0183] See Figure 8 , Figure 8 A schematic diagram of a malicious domain name identification device 800 is shown. In one possible embodiment, the malicious domain name identification device 800 may include an acquisition module 801, a fusion module 802, and an inference module 803.
[0184] The acquisition module 801 can be used to acquire the domain name to be identified and the domain name resolution data corresponding to the domain name at multiple time points.
[0185] The fusion module 802 can be used to perform feature fusion on the domain name resolution data corresponding to the domain name to be identified at multiple time points and the time sequence relationship of multiple time points using a temporal neural network, so as to obtain fused features.
[0186] The inference module 803 can be used to infer the fused features using the student model to obtain the recognition result. The recognition result is used to indicate whether the domain name to be identified is a malicious domain name. The student model is obtained by training the teacher model through knowledge distillation based on training samples. The training samples include sample fusion features and training labels. The sample fusion features are obtained by fusing the domain name resolution data corresponding to the sample domain name at multiple training time points and the temporal order relationship of multiple training time points. The training labels are used to indicate whether the sample domain name is a malicious domain name. The number of model parameters in the student model is less than that in the teacher model.
[0187] In one possible implementation, the fusion module 802 can use a temporal neural network to extract features from the domain name resolution data corresponding to the domain name to be identified at multiple time points to obtain a first feature; use the temporal neural network to extract features from the temporal order relationship at multiple time points to obtain a second feature; and fuse the first feature and the second feature to obtain a fused feature.
[0188] In one possible embodiment, the bad domain name identification device 800 may further include a map construction module 804, a feature determination module 805, a teacher training module 806, and a distillation training module 807.
[0189] Before using the student model to infer the fused features, the acquisition module 801 can also be used to acquire multiple labeled training domain name samples. These training domain name samples include the training domain name and the domain name resolution data corresponding to the training domain name at multiple training time points. The labels are used to characterize whether the corresponding training domain name is a normal domain name or a malicious domain name.
[0190] The knowledge graph construction module 804 can be used to construct a domain knowledge graph based on multiple labeled training domain name samples. The domain knowledge graph includes multiple domain name nodes and associated nodes of the domain name nodes. The domain name nodes are used to indicate training domain names, and the associated nodes include at least one of the following: Internet Protocol (IP) address nodes, Autonomous System (AS) nodes, domain name registrant nodes, geographic location nodes, and certificate nodes.
[0191] The feature determination module 805 can be used to determine the node features of multiple nodes in the domain name knowledge graph at different training time points based on the domain name resolution data corresponding to multiple training domain names at multiple training time points.
[0192] The teacher training module 806 can be used to train the initial teacher model by utilizing the node features corresponding to multiple nodes in the domain knowledge graph at different training time points and the labels corresponding to multiple training domain samples, thereby obtaining the teacher model.
[0193] The distillation training module 807 can be used to perform knowledge distillation training on the initial student model based on the teacher model to obtain the student model. The number of model parameters in the initial student model is less than the number of model parameters in the initial teacher model.
[0194] In one possible implementation, the associated nodes may include IP address nodes, autonomous system nodes, and certificate nodes, and the domain name knowledge graph can be an undirected weighted graph. The domain name knowledge graph may include a first edge connecting domain name nodes and IP address nodes, a second edge connecting domain name nodes and certificate nodes, and a third edge connecting IP address nodes and autonomous system nodes. The weight of the first edge is greater than the weight of the second edge, and the weight of the second edge is greater than the weight of the third edge.
[0195] In one possible implementation, the teacher training module 806 can map the node features corresponding to different types of nodes in the domain knowledge graph at different training time points to the same feature space, thereby obtaining node feature representations corresponding to different types of nodes.
[0196] The teacher training module 806 can determine the node time series corresponding to different training domains by arranging the node feature representations of different types of nodes corresponding to the same training domain at different training time points in chronological order based on the node identifiers and training time points corresponding to the node feature representations.
[0197] The teacher training module 806 can use the domain name knowledge graph, the node time series corresponding to different training domain names, and the labels corresponding to the training domain name samples to train the initial teacher model and obtain the teacher model.
[0198] In one possible implementation, the temporal neural network can include a temporal convolutional network and a long short-term memory network. The teacher training module 806 can, based on the temporal convolutional network, perform convolutional processing on the time series of nodes corresponding to different training domains to obtain time series output features corresponding to different training domains; based on the long short-term memory network, it can perform temporal fusion on the time series output features corresponding to different training domains to obtain fused features corresponding to different training domains; and using the domain knowledge graph, the fused features corresponding to different training domains, and the labels corresponding to the training domain samples, it can train the initial teacher model to obtain the teacher model.
[0199] In one possible implementation, the temporal convolutional network may include multiple stacked temporal convolutional layers. The teacher training module 806 can perform dilated causal convolution processing on the time series of nodes corresponding to different training domains based on the multiple stacked temporal convolutional layers to obtain the time series output features corresponding to different training domains.
[0200] In one possible implementation, the distillation training module 807 can use the teacher model and the initial student model to infer the fusion features corresponding to different training domain names, respectively, to obtain the output results of the teacher model and the initial student model; compare the differences between the output results of the teacher model and the initial student model to obtain the distillation loss; determine the classification loss based on the output results of the initial student model and the labels corresponding to multiple training domain name samples; and adjust the model parameters of the initial student model based on the distillation loss and the classification loss to obtain the student model.
[0201] In one possible embodiment, the bad domain name identification device 800 may further include an incremental distillation module 808.
[0202] The acquisition module 801 can also be used to acquire new training domain name samples before using the student model to infer the fused features.
[0203] The incremental distillation module 808 can be used to perform incremental distillation on the student model using newly added training domain name samples to obtain an updated student model. The hyperparameters of the updated student model are the same as those of the original student model, including the distillation temperature coefficient, distillation loss weights, and classification loss weights.
[0204] The inference module 803 can be used to infer the fused features using the updated student model to obtain the recognition result.
[0205] In one possible implementation, the identification result may include a predicted probability that the domain name to be identified is a bad domain name, or a classification result used to characterize whether the domain name to be identified is a normal domain name or a bad domain name.
[0206] It should be noted that, Figure 8 The acquisition module 801, fusion module 802, reasoning module 803, graph construction module 804, feature determination module 805, teacher training module 806, distillation training module 807, and incremental distillation module 808 shown are mainly used to illustrate the corresponding data processing functions, and do not limit each module to be implemented by a separate hardware structure. The graph construction module 804, feature determination module 805, teacher training module 806, distillation training module 807, and incremental distillation module 808 are optional functional modules, and do not mean that the malicious domain name identification device 800 must include all of the above functional modules at the same time.
[0207] In actual implementation, the above functional modules can be implemented by independent modules, or at least two functional modules can be integrated. For example, the map construction module 804, the feature determination module 805, and the fusion module 802 can be integrated into a feature processing module; the teacher training module 806, the distillation training module 807, and the incremental distillation module 808 can be integrated into a model training and update module.
[0208] In practical applications, the map construction module 804, feature determination module 805, teacher training module 806, distillation training module 807, and incremental distillation module 808 can be set in... Figure 1 In the model training device 400 shown, the acquisition module 801, fusion module 802, and inference module 803 can be located in the domain name recognition device 500. The model training device 400 can send the trained or updated student model to the domain name recognition device 500, so that the domain name recognition device 500 can use the student model to recognize the domain name to be recognized.
[0209] Based on the above device embodiments, to illustrate the hardware implementation that the aforementioned malicious domain name identification device or method can rely on, please refer to... Figure 9 , Figure 9 A schematic diagram of the hardware structure of a computing device 900 is shown. (For example...) Figure 9 As shown, the computing device 900 may include a processor 901, a memory 902, a communication interface 903, and a bus 904. The processor 901, the memory 902, and the communication interface 903 can communicate via the bus 904.
[0210] The processor 901 may be a central processing unit, a graphics processing unit, a neural network processor, a microprocessor, an application-specific integrated circuit, a field-programmable gate array, or other devices with data processing capabilities. The processor 901 may execute computer programs stored in the memory 902 to implement the malicious domain name identification method in any of the method embodiments of this application.
[0211] The memory 902 can be used to store data used or generated during the computer program and the process of identifying malicious domain names. For example, the memory 902 can store at least one of the following: training domain name samples, domain names to be identified, domain name resolution data corresponding to multiple time points, labels corresponding to training domain name samples, domain name knowledge graph, node features, node time series, fusion features, teacher model, student model, model parameters, model training loss, or domain name identification results.
[0212] The memory 902 may include volatile memory and / or non-volatile memory. For example, the memory 902 may include random access memory, read-only memory, flash memory, solid-state drive, disk storage, or other storage media capable of storing computer programs and data.
[0213] The communication interface 903 can be used to interact with the domain name data source 100, the sample database 200, or other computing devices. For example, the communication interface 903 can be used to receive the domain name to be identified and its corresponding domain name resolution data at multiple time points, receive training domain name samples, teacher model parameters, or student model parameters, and send the domain name identification result or the updated student model parameters.
[0214] Bus 904 may include an address bus, a data bus, a control bus, or other connection structures that enable data exchange between the various components within the computing device 900. Figure 9 The connection between the processor 901, memory 902 and communication interface 903 is shown as a single bus, but this does not mean that the computing device 900 includes only one bus or uses only one type of bus structure.
[0215] The memory 902 may store executable code, which the processor 901 may execute to perform the method executed by the corresponding device in any of the foregoing method embodiments. For example, the processor 901 may perform at least one of the following processes: data acquisition, feature fusion, domain name knowledge graph construction, node feature determination, teacher model training, knowledge distillation training, incremental distillation training, and domain name recognition.
[0216] It should be understood that the computing device 900 in the embodiments of this application can correspond to Figure 1 Any one of the feature fusion device 300, model training device 400, or domain name recognition device 500 shown can also correspond to Figure 8 The malicious domain name identification device 800 shown is shown.
[0217] In one possible implementation, the feature fusion device 300, the model training device 400, and the domain name recognition device 500 can be implemented by different computing devices 900. For example, the first computing device can be used to generate fused features corresponding to the training domain name and train the teacher model and the student model, while the second computing device can load the trained student model and use the student model to recognize the domain name to be recognized.
[0218] In another possible implementation, at least two of the feature fusion device 300, model training device 400, and domain name recognition device 500 can be integrated into the same computing device 900. For example, the computing device 900 can complete the entire process of fusion feature generation, student model training, and domain name inference.
[0219] The operations and other operations and / or functions implemented by the computing device 900 are used to implement the corresponding processes in the foregoing method embodiments. The corresponding processing procedures and their technical effects can be found in the relevant descriptions in the foregoing method embodiments; to avoid repetition, they will not be repeated here.
[0220] This application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it can implement the malicious domain name identification method in any embodiment of this application. The computer-readable storage medium may be a read-only memory, random access memory, disk, optical disk, flash memory, solid-state drive, or other medium capable of storing program code.
[0221] This application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it can implement the malicious domain name identification method in any embodiment of this application.
[0222] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, modules, and equipment described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here in the embodiments of this application.
[0223] In the several embodiments provided in this application, it should be understood that the disclosed methods, apparatus, and devices can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative, and the division of modules is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.
[0224] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for identifying malicious domain names, characterized in that, The method includes: Obtain the domain name to be identified and the domain name resolution data corresponding to the domain name at multiple time points; Using a temporal neural network, feature fusion is performed on the domain name resolution data corresponding to the domain name to be identified at multiple time points and the temporal order relationship of the multiple time points to obtain fused features; The student model is used to infer the fused features to obtain an identification result. The identification result is used to indicate whether the domain name to be identified is a bad domain name. The student model is obtained by training the teacher model through knowledge distillation based on training samples. The training samples include sample fusion features and training labels. The sample fusion features are obtained by fusing the domain name resolution data corresponding to the sample domain name at multiple training time points and the temporal order relationship of the multiple training time points. The training labels are used to indicate whether the sample domain name is a bad domain name. The number of parameters in the student model is less than the number of parameters in the teacher model.
2. The method according to claim 1, characterized in that, The method utilizes a temporal neural network to perform feature fusion on the domain name resolution data corresponding to the domain name to be identified at multiple time points and the temporal order relationship of the multiple time points, to obtain fused features, including: Using the temporal neural network, feature extraction is performed on the domain name resolution data corresponding to the domain name to be identified at the multiple time points to obtain the first feature; Using the aforementioned temporal neural network, feature extraction is performed on the temporal sequence relationship of the multiple time points to obtain a second feature; The first feature and the second feature are fused together to obtain the fused feature.
3. The method according to claim 1, characterized in that, Before using the student model to infer the fused features, the method further includes: Obtain multiple labeled training domain name samples. The training domain name samples include training domain names and domain name resolution data corresponding to the training domain names at multiple training time points. The labels are used to characterize whether the corresponding training domain name is a normal domain name or an undesirable domain name. Based on the multiple labeled training domain name samples, a domain name knowledge graph is constructed. The domain name knowledge graph includes multiple domain name nodes and associated nodes of the domain name nodes. The domain name nodes are used to indicate training domain names. The associated nodes include at least one of Internet Protocol IP address nodes, autonomous system nodes, domain name registrant nodes, geographic location nodes, and certificate nodes. Based on the domain name resolution data of the multiple training domain names at the multiple training time points, determine the node features of multiple nodes in the domain name knowledge graph at different training time points. The initial teacher model is trained using the node features corresponding to multiple nodes in the domain knowledge graph at different training time points, as well as the labels corresponding to the multiple training domain samples, to obtain the teacher model. Based on the teacher model, the initial student model is trained by knowledge distillation to obtain the student model, wherein the number of parameters of the initial student model is less than the number of parameters of the initial teacher model.
4. The method according to claim 3, characterized in that, The associated nodes include the IP address node, the autonomous system node, and the certificate node. The domain name knowledge graph is an undirected weighted graph. The domain name knowledge graph includes a first edge connecting the domain name node and the IP address node, a second edge connecting the domain name node and the certificate node, and a third edge connecting the IP address node and the autonomous system node. The weight of the first edge is greater than the weight of the second edge, and the weight of the second edge is greater than the weight of the third edge.
5. The method according to claim 3, characterized in that, The method of training the initial teacher model by utilizing the node features corresponding to multiple nodes in the domain name knowledge graph at different training time points, and the labels corresponding to the multiple training domain name samples, to obtain the teacher model includes: The node features corresponding to different types of nodes in the domain knowledge graph at different training time points are mapped to the same feature space to obtain the node feature representations corresponding to the different types of nodes. Based on the node identifier and training time point corresponding to the node feature representation, the node feature representations corresponding to different types of nodes for the same training domain name at different training time points are arranged in chronological order to determine the node time sequence corresponding to the different training domain names. The initial teacher model is trained using the domain knowledge graph, the node time series corresponding to the different training domains, and the labels corresponding to the training domain samples to obtain the teacher model.
6. The method according to claim 5, characterized in that, The temporal neural network includes a temporal convolutional network and a long short-term memory network; The process of training the initial teacher model using the domain name knowledge graph, the node time series corresponding to the different training domain names, and the labels corresponding to the training domain name samples to obtain the teacher model includes: Based on the temporal convolutional network, the node time series corresponding to the different training domain names are convolved to obtain the time series output features corresponding to the different training domain names; Based on the Long Short-Term Memory Network, the time-series output features corresponding to the different training domain names are fused in time series to obtain the fused features corresponding to the different training domain names; The initial teacher model is trained using the domain knowledge graph, the fusion features corresponding to the different training domains, and the labels corresponding to the training domain samples to obtain the teacher model.
7. The method according to claim 6, characterized in that, The step of performing convolution processing on the time series of nodes corresponding to different training domains based on the temporal convolutional network to obtain the time series output features corresponding to the different training domains includes: Dilated causal convolution is performed on the time series of nodes corresponding to different training domains based on multi-layered stacked temporal convolutional layers to obtain the time series output features corresponding to the different training domains.
8. The method according to claim 6, characterized in that, The step of training the initial student model with knowledge distillation based on the teacher model to obtain the student model includes: The teacher model and the initial student model are used to infer the fusion features corresponding to the different training domain names, respectively, to obtain the output results of the teacher model and the output results of the initial student model. The difference between the output of the teacher model and the output of the initial student model is compared to obtain the distillation loss value; Based on the output of the initial student model and the labels corresponding to the multiple training domain name samples, a classification loss value is determined. The classification loss value is used to characterize the degree of difference between the output of the initial student model and the labels. Based on the distillation loss value and the classification loss value, the model parameters of the initial student model are adjusted to obtain the student model.
9. The method according to claim 1, characterized in that, Before using the student model to infer the fused features, the method further includes: Obtain newly added training domain name samples; The student model is incrementally distilled using the newly added training domain name samples to obtain an updated student model. The hyperparameters of the updated student model are the same as those of the original student model, and the hyperparameters include distillation temperature coefficient, distillation loss weight, and classification loss weight. The process of using a student model to infer the fused features and obtain the recognition result includes: The updated student model is used to infer the fused features to obtain the recognition result.
10. The method according to claim 1, characterized in that, The identification results include the predicted probability that the domain name to be identified is a bad domain name, or the classification result used to characterize whether the domain name to be identified is a normal domain name or a bad domain name.
11. A device for identifying malicious domain names, characterized in that, The device includes: The acquisition module is used to acquire the domain name to be identified and the domain name resolution data corresponding to the domain name to be identified at multiple time points; The fusion module is used to use a temporal neural network to perform feature fusion on the domain name resolution data corresponding to the domain name to be identified at multiple time points and the temporal order relationship of the multiple time points to obtain fused features. The inference module is used to infer the fused features using the student model to obtain a recognition result. The recognition result is used to indicate whether the domain name to be identified is a bad domain name. The student model is obtained by training the teacher model through knowledge distillation based on training samples. The training samples include sample fusion features and training labels. The sample fusion features are obtained by fusing the domain name resolution data corresponding to the sample domain name at multiple training time points and the temporal order relationship of the multiple training time points. The training labels are used to indicate whether the sample domain name is a bad domain name. The number of parameters in the student model is less than the number of parameters in the teacher model.
12. A computing device comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.