Feature construction method for encrypted domain name service discovery and encrypted domain name service discovery method
By constructing an encrypted domain name service discovery method that includes statistical features, sequence features, and fluctuation point sequence features, and combining it with a deep learning model, the problem of insufficient accuracy in the existing technology for discovering encrypted domain name services is solved, and accurate identification of encrypted domain name service behaviors and timely discovery of malicious activities are achieved.
Patent Information
- Application Number
- CN202411831648.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-12
AI Technical Summary
Existing technologies lack in-depth detection of sequence characteristics in encrypted domain name service traffic data, resulting in insufficient accuracy and effectiveness in the discovery of encrypted domain name services. In particular, the neglect of data volatility affects the identification of potential patterns and cyclical changes.
By acquiring and classifying data streams in real time, extracting the field features and sequence features of data packets, performing volatility analysis, constructing feature splicing including statistical features, sequence features, and fluctuation point sequence features, and combining deep learning models for encrypted domain name service discovery.
It improves the accuracy of encrypted domain name service discovery and the ability to identify malicious activities, reduces network security risks, and can more accurately capture behavioral relationship patterns and abnormal events.
Smart Images

Figure CN119892671B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network traffic feature analysis, and in particular to a feature construction method for encrypted domain name service discovery, encrypted domain name service discovery model training, and an encrypted domain name service discovery method. Background Art
[0002] Encrypted domain name services (DNS) are used to identify and distinguish DNS services that utilize encryption technology within network traffic. Discovering encrypted DNS services is crucial for maintaining network security and user privacy. Given that attackers can potentially exploit encrypted DNS services to conceal their malicious activities, discovering and analyzing encrypted DNS services enables network administrators to effectively identify and manage encrypted traffic, thereby helping to identify potential malicious activity and data breaches. During the execution of encrypted DNS services, corresponding traffic data is generated. Service characteristics refer to the specific attributes within this traffic data that describe the type of encrypted DNS service and its operational processes, as well as unique information that can be used to identify the service. When used for service discovery, these characteristics primarily represent the specific behaviors generated during the service process to distinguish different service types, focusing on describing the service's operational patterns and usage habits—that is, how the service performs its functions within the network. Service characteristics essentially extract and represent specific features that characterize service behavior during the service process. Service analysis typically relies on features such as packet length and packet timing in IP communication between the server and client, reflecting the service's operational characteristics and usage patterns.
[0003] Current technologies typically rely primarily on the statistical characteristics of the aforementioned service features to detect traffic packets, thereby achieving service discovery. However, service features not only contain statistical characteristics, such as numerical characteristics such as packet length in traffic packets, but also sequence characteristics that characterize the periodic nature of the service. Sequence characteristics are typically a periodic measurement variable, referring to the features extracted from multiple traffic packets generated by an encrypted domain name service from the start of service to the end of service. This temporal characteristic can reflect the activity cycle and continuity of the service over a period of time. For example, the arrival time of a packet can reflect the activity level of the service during a specific time period.
[0004] Current technologies lack the ability to detect and fully exploit sequence features, particularly data volatility. This can lead to overlooking underlying patterns and cyclical changes in traffic data generated by encrypted domain name services, impacting the accuracy and effectiveness of service discovery. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a feature construction method for encrypted domain name service discovery, encrypted domain name service discovery model training, and encrypted domain name service discovery method to eliminate or improve one or more defects in the prior art.
[0006] One aspect of the present invention provides a feature construction method for encrypted domain name service discovery, the method comprising the following steps:
[0007] Acquire multiple data streams generated by running different domain name services corresponding to different target IP addresses in real time, classify the multiple data streams according to the target IP addresses, and obtain at least one data stream for each target IP address, wherein the data stream is a data packet sequence composed of multiple data packets;
[0008] Marking each data flow of each target IP address with a corresponding label based on the correspondence between the target IP address and the running domain name service, wherein the label includes an encrypted domain name service and a non-encrypted domain name service;
[0009] Extracting field features of each data packet from each data flow of the label corresponding to each target IP address's mark, and obtaining statistical features and sequence features of each data flow of the label corresponding to each target IP address's mark based on the field features;
[0010] performing a volatility analysis on the sequence characteristics to obtain a set of fluctuation point positions of the sequence characteristics, obtaining a horizontal volatility degree estimate and a vertical volatility degree estimate for each fluctuation point based on the set of fluctuation point positions, and obtaining a fluctuation point sequence characteristic based on the horizontal volatility degree estimates and the vertical volatility degree estimates for all fluctuation points;
[0011] The statistical features, sequence features and the fluctuation point sequence features of each data stream corresponding to the tag of each target IP address are spliced together to obtain multiple features for encrypted domain name service discovery.
[0012] In some embodiments of the present invention, the field characteristics include packet length, packet direction, packet arrival time, and packet response time; obtaining statistical characteristics and sequence characteristics of each data flow corresponding to the tag of each target IP address based on the field characteristics includes:
[0013] Based on the packet length, packet arrival time, and packet response time of all data packets in each data flow corresponding to the label of each target IP address, statistical features corresponding to the packet length, packet arrival time, and packet response time are obtained respectively, and the statistical features corresponding to the packet length, packet arrival time, and packet response time are spliced to obtain the statistical features of each data flow corresponding to the label of each target IP address;
[0014] The packet length, packet direction and packet arrival time of each packet in each data flow corresponding to the label based on the label corresponding to each target IP address are obtained to obtain corresponding triplet features. The triplet features of all packets in each data flow are used to form corresponding sequence features.
[0015] In some embodiments of the present application, volatility analysis is performed on the sequence features to obtain a set of volatility point positions of the sequence features, including:
[0016] Wavelet transform is performed on the sequence features.
[0017] The triplet features of every three consecutive packets in the sequence features after wavelet transform are uniformly sampled, and all packets belonging to the maximum value points and minimum value points in the sequence features after wavelet transform are determined based on the triplet features of every three consecutive packets.
[0018] The attribute of the packets belonging to the maximum value points is recorded as a first attribute value, and the attribute of the packets belonging to the minimum value points is recorded as a second attribute value. Every two consecutive packets in the sequence features after wavelet transform, whose attributes are the first attribute value, the second attribute value or the second attribute value, the first attribute value, are volatility points, and the positions of all volatility points in the corresponding packet sequence of the corresponding target IP address are recorded.
[0019] All volatility points are arranged according to the order of the positions of the volatility points in the corresponding packet sequence of the corresponding target IP address, and a set of volatility point positions is formed based on the positions of all arranged volatility points in the corresponding packet sequence of the corresponding target IP address.
[0020] In some embodiments of the present application, the transverse volatility degree estimate and the longitudinal volatility degree estimate of each volatility point are obtained based on the set of volatility point positions, including:
[0021] All packets as volatility points in the corresponding packet sequence of the corresponding target IP address are obtained based on the set of volatility point positions, and the transverse volatility degree estimate of each volatility point is obtained based on the triplet features of each volatility point and the triplet features of the packets before and after each volatility point in the corresponding packet sequence of the corresponding target IP address.
[0022] The longitudinal volatility degree estimate of each volatility point is obtained based on the triplet features of each volatility point and the average of the triplet features of all packets in the corresponding packet sequence of the corresponding target IP address.
[0023] In some embodiments of the present application, the sequence feature of the volatility points is obtained based on the transverse volatility degree estimates and the longitudinal volatility degree estimates of all volatility points according to the following formula:
[0024]
[0025]
[0026] Among them, B V Represents the characteristics of the fluctuation point sequence, sigmoid() represents the normalization function, HBD i It represents the horizontal volatility estimation of the ith fluctuation point, ZBD i It represents the vertical volatility estimation of the ith fluctuation point, y i represents the triplet feature of the i-th fluctuation point, bi represents the position of the fluctuation point obtained based on the i-th fluctuation point position in the fluctuation point position set in the corresponding data packet sequence of the corresponding target IP address, and n represents the number of fluctuation points.
[0027] In some embodiments of the present invention, the statistical features include variance, standard deviation, mean, median, mode, the inclination of the packet length, packet arrival time or packet response time relative to the median, the skewness of the packet length, packet arrival time or packet response time relative to the mode, and the coefficient of variation of the packet length, packet arrival time or packet response time.
[0028] Another aspect of the present invention provides a method for training an encrypted domain name service discovery model, the method comprising the following steps:
[0029] The multiple features for encrypted domain name service discovery constructed by the aforementioned feature construction method for encrypted domain name service discovery and the corresponding labels are input into a preset encrypted domain name service discovery model to train the encrypted domain name service discovery model, wherein the labels include encrypted domain name services and non-encrypted domain name services.
[0030] Another aspect of the present invention provides a method for discovering an encrypted domain name service, the method comprising the following steps:
[0031] Acquire multiple data streams generated by running different domain name services corresponding to different target IP addresses in real time, classify the multiple data streams according to the target IP addresses, and obtain at least one data stream for each target IP address, wherein the data stream is a data packet sequence composed of multiple data packets;
[0032] Extracting field features of each data packet from each data flow of each target IP address, and obtaining statistical features and sequence features of each data flow of each target IP address based on the field features;
[0033] performing a volatility analysis on the sequence characteristics to obtain a set of fluctuation point positions of the sequence characteristics, obtaining a horizontal volatility degree estimate and a vertical volatility degree estimate for each fluctuation point based on the set of fluctuation point positions, and obtaining a fluctuation point sequence characteristic based on the horizontal volatility degree estimates and the vertical volatility degree estimates for all fluctuation points;
[0034] The statistical features, sequence features, and fluctuation point sequence features of each data flow of each target IP address are spliced together to obtain multiple features for encrypted domain name service discovery;
[0035] The multiple features used for encrypted domain name service discovery are input into the encrypted domain name service discovery model trained by the aforementioned encrypted domain name service discovery model training method, so that the encrypted domain name service discovery model outputs a result of whether it is an encrypted domain name service.
[0036] In some embodiments of the present invention, the field characteristics include packet length, packet direction, packet arrival time, and packet response time; obtaining statistical characteristics and sequence characteristics of each data flow of each target IP address based on the field characteristics includes:
[0037] Based on the packet length, packet arrival time, and packet response time of all data packets in each data flow of each target IP address, statistical features corresponding to the packet length, packet arrival time, and packet response time are obtained respectively, and the statistical features corresponding to the packet length, packet arrival time, and packet response time are spliced together to obtain the statistical features of each data flow of each target IP address;
[0038] The corresponding triplet features are obtained based on the packet length, packet direction and packet arrival time of each data packet in each data flow of each target IP address, and the corresponding sequence features are formed by the triplet features of all data packets in each data flow.
[0039] Another aspect of the present invention provides an electronic device, which includes: a processor, a memory, and computer instructions stored in the memory, wherein the processor is used to execute the computer instructions. When the computer instructions are executed, the device implements the aforementioned feature construction method for encrypted domain name service discovery, encrypted domain name service discovery model training method, or encrypted domain name service discovery method steps.
[0040] Another aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aforementioned steps of the feature construction method for encrypted domain name service discovery, the encrypted domain name service discovery model training method, or the encrypted domain name service discovery method.
[0041] Another aspect of the present invention provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the aforementioned feature construction method for encrypted domain name service discovery, encrypted domain name service discovery model training method, or encrypted domain name service discovery method.
[0042] The feature construction method for encrypted domain name service discovery, encrypted domain name service discovery model training and encrypted domain name service discovery method of the present invention, wherein the feature construction method introduces volatility analysis of sequence features based on the numerical dimensions of statistical features and sequence features, can effectively capture the behavioral relationship patterns in the traffic data generated by the encrypted domain name service, and is of great significance for characterizing the encrypted domain name service behavior and further encrypted domain name service discovery tasks.
[0043] Additional advantages, objects, and features of the present invention will be set forth in part in the following description and will become apparent to those skilled in the art upon examination of the following or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained by the structures particularly pointed out in the description and drawings.
[0044] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other purposes that can be achieved by the present invention will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention.
[0046] Figure 1 1 is a flow chart of a feature construction method for encrypted domain name service discovery according to an embodiment of the present invention;
[0047] Figure 2 A flowchart of a method for training an encrypted domain name service discovery model according to an embodiment of the present invention is shown;
[0048] Figure 3 The figure is a flow chart of an encrypted domain name service discovery method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0050] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, the accompanying drawings only show structures and / or processing steps closely related to the solutions according to the present invention, while other details that are not closely related to the present invention are omitted.
[0051] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.
[0052] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0053] Given that existing encrypted domain name service discovery methods lack analysis and exploration of the volatility of sequence characteristics, and that some methods rely on sequence characteristics that are only reflected in the data dimension, which is too simplistic and cannot deeply analyze the behavior patterns of domain name services from the perspective of the volatility of the time series, the present invention takes into account the differences in specific events and activity patterns during the service resolution process between encrypted domain name services and non-encrypted domain name services, and the different special patterns of fluctuations in the time series of traffic data generated by encrypted domain name services compared to non-encrypted domain name services. For example, when some services use encrypted domain name services for malicious activities, the volatility of service traffic data is usually greater than that of encrypted domain name services under normal circumstances. At the same time, the startup, interruption, and termination of encrypted domain name services often lead to obvious feature changes in the corresponding generated traffic data. This change in service characteristics can be reflected and explored through the volatility of sequence characteristics.
[0054] Therefore, the embodiments of the present invention provide a feature construction method for discovering encrypted domain name services, an encrypted domain name service discovery model training, and an encrypted domain name service discovery method that combine sequence fluctuation characteristics. By introducing sequence feature volatility analysis based on the numerical dimensions of statistical features and sequence features, the method can effectively capture behavioral relationship patterns in the traffic data generated by encrypted domain name services, which is of great significance for characterizing encrypted domain name service behavior and further tasks of discovering encrypted domain name services. The method can also play a positive role in timely identifying abnormal behavior and sequence feature traits generated by emergencies in the traffic data generated by encrypted domain name services. When malicious encrypted domain name services appear and cause traffic fluctuations, the method can promote the discovery and identification of malicious encrypted domain name services, greatly reducing network security risks.
[0055] Figure 1 FIG. 1 is a flow chart of a feature construction method for discovering encrypted domain name services in one embodiment of the present invention. Figure 1 As shown, the feature construction method includes the following steps:
[0056] Step S110, real-time acquisition of multiple data streams generated when different domain name services corresponding to different target IP addresses are running, and the multiple data streams are classified according to the target IP addresses to obtain at least one data stream for each target IP address, wherein the data stream is a data packet sequence composed of multiple data packets.
[0057] In this step, the domain name services can be different services provided by multiple servers with different IP addresses accessing the same domain name simultaneously, including encrypted domain name services and non-encrypted domain name services. In the specific application scenario of this method, a router with a network monitoring function (Netflow) can be deployed in the gateway server of the network. After manually testing the network access function, the Netflow router collects in real time the multiple traffic data streams (IP data streams) generated when multiple servers with different IP addresses run different domain name services (including encrypted domain name services and non-encrypted domain name services). The raw data streams collected by the Netflow router are then sent to the network behavior monitoring (NBOS) resident system for subsequent analysis.
[0058] The NBOS on-site system organizes and records the stored raw data streams, performing preprocessing on them. This preprocessing primarily converts the raw data streams into standardized data stream records. These standardized data stream records are parsed to obtain the server IP address of the corresponding service. This server IP address is defined as the target IP address, which serves as the identifier for the service activity. Each data stream record is then categorized by the target IP address, and all data stream records for each target IP address are statistically analyzed. All preprocessed data streams are then sent to an activity repository for storing key data and features. Each data stream record is considered a domain name service activity.
[0059] Step S120 , marking each data flow of each target IP address with a corresponding label based on the correspondence between the target IP address and the running domain name service, wherein the label includes an encrypted domain name service and a non-encrypted domain name service.
[0060] In this step, the correspondence between the target IP address and the running domain name service is preset. In a specific application scenario, each pre-processed data stream stored in the active library is labeled by the NBOS resident system. Specifically, if the server of a certain target IP address provides or runs an encrypted domain name service, the data stream generated when the server of the target IP address runs the service is marked with the label of the encrypted domain name service; if the server of a certain target IP address provides or runs a non-encrypted domain name service, the data stream generated when the server of the target IP address runs the service is marked with the label of the non-encrypted domain name service.
[0061] Step S130, extracting the field features of each data packet from each data flow corresponding to the label of each target IP address, and obtaining the statistical features and sequence features of each data flow corresponding to the label of each target IP address based on the field features.
[0062] In this step, in a specific application scenario, the field features used to depict the two types of service features of statistical features and sequence features are extracted from each preprocessed data stream stored in the active library by the NBOS premises system. Through this step, it can be explicitly known that the label to which the field features of each data packet belong is the encrypted domain name service or the non-encrypted domain name service, that is, it is explicitly known which field features among them belong to the field features of the encrypted domain name service and which field features belong to the field features of the non-encrypted domain name service.
[0063] In an embodiment of the application, the field features include packet length, packet direction, packet arrival time and packet response time.
[0064] For a single domain name service activity of a server of a target IP address, the direction of the data packet flowing into the target IP address is set as the entry direction of the domain name service, the direction of the data packet flowing out of the target IP address is set as the exit direction of the domain name service, and the mapping from the entry direction to the exit direction is the packet direction of each data packet in the data stream. The packet arrival time refers to the time when the data packet arrives at the target IP address. The packet response time refers to the time difference between the sending and the receiving of the reply, that is, the time difference obtained by subtracting the time when the sender sends the data packet from the time when the sender receives the response packet. This time difference can reflect the response performance of the network and the delay characteristics of the service. The field features can further include source IP address, destination IP address, source port and destination port for determining session flow features.
[0065] In an embodiment, the statistical features and sequence features of each data stream corresponding to the label of each target IP address based on the field features in step S130 specifically include the following steps:
[0066] In step S133, the statistical features corresponding to the packet length, the packet arrival time and the packet response time are respectively obtained based on the packet length, the packet arrival time and the packet response time of all data packets in each data stream corresponding to the label of each target IP address, and the statistical features corresponding to the packet length, the packet arrival time and the packet response time are spliced to obtain the statistical features of each data stream corresponding to the label of each target IP address.
[0067] In this step, the statistical features include variance, standard deviation, mean value, median, mode, skewness of the packet length, packet arrival time or packet response time relative to the median, kurtosis of the packet length, packet arrival time or packet response time relative to the mode, and coefficient of variation of the packet length, packet arrival time or packet response time.
[0068] For multiple data packets in a data stream generated when a server with a certain target IP address runs a domain name service activity, the statistical characteristics of the three variables corresponding to the data stream of the domain name service V, which include all data packets, are obtained as Tv = [S1, S2, ..., S8, A1, A2, ..., A8, R1, R2, ..., R8], where the three variables include packet length S, packet arrival time A, and packet response time R. S1, S2, ..., S8 represent the variance, standard deviation, mean, median, mode, and packet length relative to the response time of the packet, respectively. The median slope, packet length skewness relative to the mode, and packet length coefficient of variation are represented. A1, A2, …, A8 represent the variance, standard deviation, mean, median, mode, packet arrival time slope relative to the median, packet arrival time skewness relative to the mode, and packet arrival time coefficient of variation, respectively. R1, R2, …, R8 represent the variance, standard deviation, mean, median, mode, packet response time slope relative to the median, packet response time skewness relative to the mode, and packet response time coefficient of variation, respectively. The mode represents the degree of variation, distribution, average size, median, and most frequently occurring value of a variable. The slope and skewness relative to the median and mode can be used to calculate the degree of asymmetry in the variable's distribution. The coefficient of variation can be used to calculate the ratio of the standard deviation to the mean.
[0069] Because the statistical characteristics of the above three variables are crucial for detecting and identifying encrypted domain name services, their calculations, by numerically defining traffic data characteristics / domain name service characteristics from the perspective of packet statistics, can reflect, quantify, and describe the behavioral patterns of domain name service traffic data from a numerical statistical perspective. Different types of statistical calculations can reflect different data characteristics, so the above eight statistical calculations are performed for each variable.
[0070] Step S134, based on the packet length, packet direction and packet arrival time of each data packet in each data flow corresponding to the label of each target IP address, a corresponding triplet feature is obtained, and the corresponding sequence feature is formed by the triplet features of all data packets in each data flow.
[0071] This step describes and reveals the behavior pattern information of the encrypted domain name service traffic data within the periodic time period by selecting the packet length S, packet direction D and packet arrival time A of the traffic data packet in the domain name service V. Assume that the traffic data stream generated in a domain name service contains N data packet samples, and the sequence feature of the data stream is Hv = [(Sx1, Dx1, Ax1), (Sx2, Dx2, Ax2), ..., (Sx N ,Dx N ,AxN )], where the triplet feature of each data packet is y i =(Sx i ,Dx i ,Ax i ), x i Represents the i-th data packet, i = 1, 2,…, N.
[0072] Step S140, performing volatility analysis on the sequence characteristics to obtain a set of fluctuation point positions of the sequence characteristics, obtaining a horizontal volatility degree estimate and a vertical volatility degree estimate for each fluctuation point based on the fluctuation point position set, and obtaining a fluctuation point sequence characteristic based on the horizontal volatility degree estimate and the vertical volatility degree estimate of all fluctuation points.
[0073] This step extracts the fluctuation point sequence features from the sequence features to assess the overall trend of the sequence features. Sequence features are a key dimension for detecting and identifying encrypted domain name services. However, simple sequence features cannot measure fluctuations in traffic data streams and further identify anomalies and fluctuation points. For example, if an encrypted domain name service experiences a sudden increase in traffic at a certain point in time, simple sequence features cannot capture this abnormal behavior or reflect this pattern, making it difficult to accurately identify the encrypted domain name service. Therefore, the present invention introduces a volatility measurement indicator for sequence features, namely the fluctuation point sequence features extracted in this step.
[0074] In one embodiment of the present invention, step S140 performs a volatility analysis on the sequence feature to obtain a set of fluctuation point positions of the sequence feature, which specifically includes the following steps:
[0075] Step S141: performing wavelet transform on the sequence features.
[0076] Step S142, uniformly sample the triplet features of every three consecutive data packets in the sequence features after wavelet transformation, and determine all data packets belonging to the maximum and minimum points in the sequence features after wavelet transformation based on the triplet features of every three consecutive data packets. For example, the three consecutive data packets are x i-1 、x i and x i+1 , and the triplet features of these three data packets have the following relationship: (y i -y i-1 )(y i+1 -y i )﹤0, then the data packet x in the sequence feature after wavelet transformation i is an extreme point. Further, if x i ﹥x i-1 , then the data packet x i is a maximum point, if x i ﹤xi-1 then the data packet x i is a minimum point.
[0077] Step S143, the attribute of the data packet belonging to the maximum point is recorded as the first attribute value, the attribute of the data packet belonging to the minimum point is recorded as the second attribute value, and the attribute of each two continuous data packets in the sequence features after the wavelet transform is the first attribute value, the second attribute value or the second attribute value, the first attribute value, and all the data packets are the fluctuation points, and the positions of all the fluctuation points in the corresponding data packet sequence of the corresponding target IP address are recorded. For example, the attribute of the data packet belonging to the maximum point is marked as 1, and the attribute of the data packet belonging to the minimum point is marked as -1, if the attributes of the two continuous data packets are "1, -1" or "-1, 1", the two data packets are both called the fluctuation points, all the fluctuation points are screened out in this way, and the positions of all the fluctuation points in the corresponding data packet sequence of the corresponding target IP address are recorded as b1, b2, …, bn. The position is determined by the arrival time of the data packet corresponding to the fluctuation point at the target IP address.
[0078] Step S144, all the fluctuation points are arranged according to the order of the positions of the fluctuation points in the corresponding data packet sequence of the corresponding target IP address, and a fluctuation point position set is formed based on the positions of all the arranged fluctuation points in the corresponding data packet sequence of the corresponding target IP address. For example, the fluctuation point position set can be [b1, b2, …, bn].
[0079] In an embodiment of the present application, the transverse fluctuation degree estimation and the longitudinal fluctuation degree estimation of each fluctuation point are obtained based on the fluctuation point position set in step S140, and the specific steps include the following steps:
[0080] Step S145, all the data packets as the fluctuation points in the corresponding data packet sequence of the corresponding target IP address are obtained based on the fluctuation point position set, and the transverse fluctuation degree estimation of each fluctuation point is obtained based on the triplet feature of each fluctuation point and the triplet features of the data packets before and after each fluctuation point in the corresponding data packet sequence of the corresponding target IP address.
[0081] The transverse fluctuation degree estimation of each fluctuation point is obtained according to the following formula:
[0082]
[0083] wherein, HBD i represents the transverse fluctuation degree estimation of the i th fluctuation point, y i represents the triplet feature of the i th fluctuation point, y i+1 represents the triplet feature of the data packet in the position after the i th fluctuation point in the corresponding data packet sequence of the corresponding target IP address, y i-1The triplet feature representing the data packet located before the i-th fluctuation point in the data packet sequence corresponding to the target IP address.
[0084] Step S146 , obtaining an estimate of the longitudinal volatility of each fluctuation point based on the triplet feature of each fluctuation point and the average value of the triplet features of all data packets in the data packet sequence corresponding to the target IP address.
[0085] Specifically, the vertical volatility of each fluctuation point is estimated according to the following formula:
[0086]
[0087] Among them, ZBD i It represents the estimation of the vertical volatility of the ith fluctuation point, It represents the average value of the triplet features of all packets in the packet sequence corresponding to the target IP address where the fluctuation point is located, and n represents the number of fluctuation points.
[0088] In one embodiment of the present invention, in step S140, the fluctuation point sequence characteristics are obtained based on the horizontal volatility degree estimation and the vertical volatility degree estimation of all fluctuation points according to the following formula:
[0089]
[0090]
[0091] Among them, B V Represents the characteristics of the fluctuation point sequence, sigmoid() represents the normalization function, HBD i It represents the horizontal volatility estimation of the ith fluctuation point, ZBD i It represents the vertical volatility estimation of the ith fluctuation point, y i represents the triplet feature of the i-th fluctuation point, bi represents the position of the fluctuation point obtained based on the i-th fluctuation point position in the fluctuation point position set in the corresponding data packet sequence of the corresponding target IP address, and n represents the number of fluctuation points.
[0092] Through the above steps, the wavelet-transformed sequence features are defined to obtain individual fluctuation points. By defining indicators related to the fluctuation point and the triplet features of the packets preceding and following it in the original data packet sequence, as well as the degree of numerical variation between the fluctuation point and the average triplet features of the original data packet sequence as a whole, the above formulas for measuring the degree of fluctuation of data points in the sequence features are determined. The fluctuation point sequence contained in the fluctuation point sequence features obtained using this formula helps improve the accuracy of discovering and identifying encrypted domain name services.
[0093] Step S150 , the statistical features, sequence features and the fluctuation point sequence features of each data stream corresponding to the tag of each target IP address are spliced together to obtain a plurality of features for encrypted domain name service discovery.
[0094] In this step, specifically, by splicing the statistical features, sequence features, and fluctuation point sequence features of the data stream, which include statistical features corresponding to packet length, packet arrival time, and packet response time, the feature Zv = (Tv, Hv, Bv) finally constructed for encrypted domain name service discovery can be obtained. For the statistical features, sequence features, and fluctuation point sequence features obtained in this method, it can be clearly known whether the labels corresponding to these features are encrypted domain name services or non-encrypted domain name services. The multiple features finally obtained for encrypted domain name service discovery are multiple spliced features constructed from multiple data streams with labels corresponding to different target IP addresses (the statistical features, sequence features, and fluctuation point sequence features of each data stream are spliced).
[0095] Figure 2 FIG. 1 is a flow chart of a method for training an encrypted domain name service discovery model according to an embodiment of the present invention. Figure 2 As shown, an embodiment of the present invention also provides an encrypted domain name service discovery model training method, which includes the following steps:
[0096] In step S210, a plurality of features for encrypted domain name service discovery and corresponding labels constructed by the feature construction method for encrypted domain name service discovery of each of the aforementioned embodiments are input into a preset encrypted domain name service discovery model to train the encrypted domain name service discovery model, wherein the labels include encrypted domain name services and non-encrypted domain name services.
[0097] In one embodiment of the present invention, the preset encrypted domain name service discovery model is a Text Convolutional Neural Network (TextCNN) classifier. TextCNN is a general deep learning model that can extract temporal information from sequence features and is also effective in capturing discrete features. In other embodiments, the preset encrypted domain name service discovery model can also be other deep learning models.
[0098] Figure 3 FIG. 1 is a flow chart of an encrypted domain name service discovery method according to an embodiment of the present invention. Figure 3 As shown, an embodiment of the present invention further provides an encrypted domain name service discovery method, which includes the following steps:
[0099] Step S310, acquiring in real time multiple data streams generated by the operation of different domain name services corresponding to different target IP addresses, classifying the multiple data streams according to the target IP addresses, and obtaining at least one data stream for each target IP address, wherein the data stream is a data packet sequence consisting of multiple data packets;
[0100] Step S320, extracting field features of each data packet from each data flow of each target IP address, and obtaining statistical features and sequence features of each data flow of each target IP address based on the field features;
[0101] Step S330: performing a volatility analysis on the sequence feature to obtain a set of fluctuation point positions of the sequence feature, obtaining a horizontal volatility degree estimate and a vertical volatility degree estimate for each fluctuation point based on the set of fluctuation point positions, and obtaining a fluctuation point sequence feature based on the horizontal volatility degree estimates and the vertical volatility degree estimates for all fluctuation points;
[0102] Step S340: combining the statistical features, sequence features, and the fluctuation point sequence features of each data flow of each target IP address to obtain a plurality of features for encrypted domain name service discovery;
[0103] Step S350: Input the multiple features used for encrypted domain name service discovery into the encrypted domain name service discovery model trained by the encrypted domain name service discovery model training method in the aforementioned embodiments, so that the encrypted domain name service discovery model outputs a result of whether it is an encrypted domain name service.
[0104] In one embodiment of the present invention, the field characteristics include packet length, packet direction, packet arrival time, and packet response time; and obtaining statistical characteristics and sequence characteristics of each data flow of each target IP address based on the field characteristics in step S320 includes the following steps:
[0105] Step S323: Based on the packet lengths, packet arrival times, and packet response times of all data packets in each data flow of each target IP address, statistical features corresponding to the packet lengths, packet arrival times, and packet response times are obtained respectively, and the statistical features corresponding to the packet lengths, packet arrival times, and packet response times are concatenated to obtain statistical features of each data flow of each target IP address;
[0106] Step S324, based on the packet length, packet direction and packet arrival time of each data packet in each data flow of each target IP address, a corresponding triplet feature is obtained, and a corresponding sequence feature is formed by the triplet features of all data packets in each data flow.
[0107] In the embodiment of the present invention, a feature construction method, an encrypted domain name service discovery model training method, and an encrypted domain name service discovery method are provided. The features constructed by the feature construction method for encrypted domain name service discovery are based on the statistical features of the data stream, combined with the sequence features of the data stream, and introduce fluctuation point sequence features. By connecting or splicing the statistical features, sequence features, and fluctuation point sequence features to jointly construct features for encrypted domain name service discovery, dynamic insights into domain name service behavior can be increased, and encrypted domain name service behavior can be more accurately characterized from multiple perspectives, thereby improving the accuracy of encrypted domain name service discovery. The fluctuation point sequence feature is specifically obtained by defining and identifying fluctuation points in the sequence feature, and defining the degree of fluctuation of data packets (fluctuation points) in the data stream based on the horizontal and vertical fluctuations of the fluctuation points around the data stream. It can reflect the overall change trend and key point trend in the service traffic, and quantify the volatility of the sequence feature. This can help identify the periodic changes and trends in the traffic data generated by the domain name service, as well as the special behavioral patterns contained in the domain name service data, ultimately achieving accurate discovery of encrypted domain name services.
[0108] Corresponding to the above method, the present invention also provides an electronic device, which includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the aforementioned feature construction method for encrypted domain name service discovery, encrypted domain name service discovery model training method or encrypted domain name service discovery method steps.
[0109] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the aforementioned feature construction method for encrypted domain name service discovery, the encrypted domain name service discovery model training method, or the steps of the encrypted domain name service discovery method. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.
[0110] An embodiment of the present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the aforementioned feature construction method for encrypted domain name service discovery, encrypted domain name service discovery model training method, or encrypted domain name service discovery method.
[0111] It should be understood by those skilled in the art that the various exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is specifically performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.
[0112] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.
[0113] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.
[0114] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations to the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A feature construction method for encrypted domain name service discovery, characterized in that: The method comprises: Acquire multiple data streams generated by running different domain name services corresponding to different target IP addresses in real time, classify the multiple data streams according to the target IP addresses, and obtain at least one data stream for each target IP address, wherein the data stream is a data packet sequence composed of multiple data packets; Marking each data flow of each target IP address with a corresponding label based on the correspondence between the target IP address and the running domain name service, wherein the label includes an encrypted domain name service and a non-encrypted domain name service; Extracting field features of each data packet from each data flow of the label corresponding to each target IP address's mark, and obtaining statistical features and sequence features of each data flow of the label corresponding to each target IP address's mark based on the field features; performing a volatility analysis on the sequence characteristics to obtain a set of fluctuation point positions of the sequence characteristics, obtaining a horizontal volatility degree estimate and a vertical volatility degree estimate for each fluctuation point based on the set of fluctuation point positions, and obtaining a fluctuation point sequence characteristic based on the horizontal volatility degree estimates and the vertical volatility degree estimates for all fluctuation points; The statistical features, sequence features and the fluctuation point sequence features of each data stream corresponding to the tag of each target IP address are spliced together to obtain multiple features for encrypted domain name service discovery.
2. The method according to claim 1, characterized in that The field characteristics include packet length, packet direction, packet arrival time, and packet response time. Based on the field characteristics, the statistical characteristics and sequence characteristics of each data flow corresponding to the tag of each target IP address are obtained, including: Based on the packet length, packet arrival time, and packet response time of all data packets in each data flow corresponding to the label of each target IP address, statistical features corresponding to the packet length, packet arrival time, and packet response time are obtained respectively, and the statistical features corresponding to the packet length, packet arrival time, and packet response time are spliced to obtain the statistical features of each data flow corresponding to the label of each target IP address; The corresponding triplet feature is obtained based on the packet length, packet direction and packet arrival time of each data packet in each data flow corresponding to the label of each target IP address, and the corresponding sequence feature is formed by the triplet feature of all data packets in each data flow; Performing a volatility analysis on the sequence feature to obtain a set of fluctuation point positions of the sequence feature, including: Performing wavelet transform on the sequence features; Uniformly sampling the triplet features of every three consecutive data packets in the sequence features after wavelet transformation, and determining all data packets belonging to maximum value points and minimum value points in the sequence features after wavelet transformation based on the triplet features of every three consecutive data packets; The attribute of the data packet belonging to the maximum value point is recorded as the first attribute value, and the attribute of the data packet belonging to the minimum value point is recorded as the second attribute value. In the sequence features after wavelet transformation, every two consecutive data packets with the first attribute value, the second attribute value, or the second attribute value, the first attribute value are all fluctuation points, and the positions of all fluctuation points in the corresponding data packet sequence of the target IP address are recorded; All fluctuation points are arranged according to the order of their positions in the corresponding data packet sequence of the corresponding target IP address, and a fluctuation point position set is formed based on the positions of all arranged fluctuation points in the corresponding data packet sequence of the corresponding target IP address.
3. The method according to claim 2, characterized in that Obtaining a horizontal volatility degree estimate and a vertical volatility degree estimate for each fluctuation point based on the fluctuation point position set, including: Based on the set of fluctuation point positions, all data packets in the data packet sequence corresponding to the target IP address that serve as fluctuation points are obtained, and based on the triplet features of each fluctuation point and the triplet features of the data packets before and after each fluctuation point in the data packet sequence corresponding to the target IP address, a horizontal volatility estimate of each fluctuation point is obtained; The longitudinal volatility degree estimation of each fluctuation point is obtained based on the triplet feature of each fluctuation point and the average value of the triplet features of all data packets in the corresponding data packet sequence of the target IP address.
4. The method according to claim 3, characterized in that Based on the horizontal volatility estimation and vertical volatility estimation of all fluctuation points, the fluctuation point sequence characteristics are obtained according to the following formula: Among them, B V Represents the characteristics of the fluctuation point sequence, sigmoid() represents the normalization function, HBD i It represents the horizontal volatility estimation of the ith fluctuation point, ZBD i It represents the vertical volatility estimation of the ith fluctuation point, y i represents the triple feature of the i-th fluctuation point, bi represents the position of the fluctuation point obtained based on the i-th fluctuation point position in the fluctuation point position set in the corresponding data packet sequence of the corresponding target IP address, and n represents the number of fluctuation points.
5. The method according to claim 2, characterized in that The statistical features include variance, standard deviation, mean, median, mode, the inclination of the packet length, packet arrival time or packet response time relative to the median, the skewness of the packet length, packet arrival time or packet response time relative to the mode, and the coefficient of variation of the packet length, packet arrival time or packet response time.
6. A method for discovering encrypted domain name services, characterized in that: The method comprises: Acquire multiple data streams generated by running different domain name services corresponding to different target IP addresses in real time, classify the multiple data streams according to the target IP addresses, and obtain at least one data stream for each target IP address, wherein the data stream is a data packet sequence composed of multiple data packets; Extracting field features of each data packet from each data flow of each target IP address, and obtaining statistical features and sequence features of each data flow of each target IP address based on the field features; performing a volatility analysis on the sequence characteristics to obtain a set of fluctuation point positions of the sequence characteristics, obtaining a horizontal volatility degree estimate and a vertical volatility degree estimate for each fluctuation point based on the set of fluctuation point positions, and obtaining a fluctuation point sequence characteristic based on the horizontal volatility degree estimates and the vertical volatility degree estimates for all fluctuation points; The statistical features, sequence features, and fluctuation point sequence features of each data flow of each target IP address are spliced together to obtain multiple features for encrypted domain name service discovery; The multiple features for encrypted domain name service discovery are input into an encrypted domain name service discovery model trained by an encrypted domain name service discovery model training method, so that the encrypted domain name service discovery model outputs a result of whether it is an encrypted domain name service; the encrypted domain name service discovery model training method includes: inputting multiple features for encrypted domain name service discovery constructed by the method according to any one of claims 1 to 5 and corresponding labels into a preset encrypted domain name service discovery model to train the encrypted domain name service discovery model, wherein the labels include encrypted domain name services and non-encrypted domain name services.
7. The method according to claim 6, characterized in that The field characteristics include packet length, packet direction, packet arrival time, and packet response time; based on the field characteristics, statistical characteristics and sequence characteristics of each data flow of each target IP address are obtained, including: Based on the packet length, packet arrival time, and packet response time of all data packets in each data flow of each target IP address, statistical features corresponding to the packet length, packet arrival time, and packet response time are obtained respectively, and the statistical features corresponding to the packet length, packet arrival time, and packet response time are spliced together to obtain the statistical features of each data flow of each target IP address; The corresponding triplet features are obtained based on the packet length, packet direction and packet arrival time of each data packet in each data flow of each target IP address, and the corresponding sequence features are formed by the triplet features of all data packets in each data flow.
8. An electronic device comprising a processor, a memory, and computer instructions stored in the memory, characterized in that: The processor is configured to execute the computer instructions, and when the computer instructions are executed, the device implements the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method for classifying encrypted traffic and server, and computer readable storage medium
CN108768986A
Encrypted traffic classification system based on data packets
CN114866486A