A method, electronic device and storage medium for extracting traffic metadata based on damping increment statistics and Selenium
The method uses Selenium and Scapy with damping incremental statistics to extract DNS traffic metadata, addressing the challenge of distinguishing normal and tunnel traffic by simulating user behavior and efficiently capturing relevant features.
Patent Information
- Application Number
- CN202411234272.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-09-04
AI Technical Summary
The prior art is difficult to extract the metadata characteristics of DNS tunnel traffic accurately, comprehensively and efficiently, resulting in insufficient real-time and low-throughput tunnel detection performance.
The traffic metadata extraction method based on damping increment statistics and Selenium is adopted to generate original traffic through Selenium, and Scapy is used for acquisition and deep packet detection. The packet-level metadata is extracted in combination with deep packet detection technology, and session metadata is generated through damping increment statistics.
Simulate user network behavior, generate DNS raw traffic close to real traffic, and comprehensively extract packet-level and session-level features, improving detection accuracy and efficiency.
Smart Images

Figure CN119071197B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of traffic metadata extraction, and particularly relates to a method for extracting traffic metadata based on damping increment statistics and Selenium. Background Art
[0002] The Domain Name System (DNS), as a hierarchical distributed database system, is one of the most critical infrastructures of the Internet and an indispensable cornerstone for the vast majority of network activities. The original design intention of the DNS protocol is to convert complex IP addresses into domain names that are easy to remember and use, rather than for data transmission. Therefore, firewalls usually allow DNS services to pass through port 53 in the default configuration, which enables DNS traffic to spread freely in the network. For the above reasons, the DNS protocol has become an ideal choice for tunneling technology.
[0003] Due to their different uses, there are relatively significant differences in the characteristics between normal DNS traffic and DNS tunnel traffic. Normal DNS traffic is mainly used for domain name resolution, and the data exchange between the two communication parties is usually intermittent, non-persistent connection, and the packet size is relatively fixed and short. DNS tunnel traffic is mainly used for data transmission, and the two communication parties show characteristics such as continuity, high-frequency queries and responses, and larger packet lengths compared with normal traffic. Considering the requirements for real-time performance and low-throughput tunnel detection performance, the selected features need to cover load features (i.e., packet-level features) and traffic features (i.e., session-level features). Therefore, accurate, comprehensive, and efficient feature extraction is crucial for data analysis technology. Summary of the Invention
[0004] The problem to be solved by the present invention is to generate original traffic and accurately, comprehensively, and efficiently extract traffic metadata features, and a method for extracting traffic metadata based on damping increment statistics and Selenium is proposed.
[0005] To achieve the above object, the present invention is realized through the following technical solutions:
[0006] A method for extracting traffic metadata based on damping increment statistics and Selenium includes the following steps:
[0007] S1. Generate original traffic based on Selenium;
[0008] S2. Collect the original traffic generated in step S1 based on Scapy to obtain an original traffic file;
[0009] S3. First preprocess the original traffic file obtained in step S2, and then extract packet-level metadata based on deep packet inspection to obtain packet-level metadata;
[0010] S4. Preprocess the packet-level metadata obtained in step S3, and then generate session metadata based on metadata aggregation and damping increment statistics.
[0011] Furthermore, the specific implementation method of step S1 includes the following steps:
[0012] S1.1. Parameter setting: Set program-related parameters, including the number of search pages, the waiting time after recommended jump, the waiting time after page flipping, the web browsing time, the sleep time before clicking, and the number of keywords extracted from each thesaurus.
[0013] S1.2. Driver selection: Select a browser driver, such as ChromeDrive.
[0014] S1.3. Search engine selection: Select a search engine as the basis for subsequent searches, such as Baidu, Sohu, Bing.
[0015] S1.4. Thesaurus selection for keywords: Select a keyword thesaurus as the seed for the search.
[0016] S1.5. Determine whether there is an archived file: Adopt an interrupted continuation mechanism. By judging whether there is a residual archived file, determine whether the program has crashed, and then decide whether this operation is a new round of crawling or a continuation of the previous round of crawling; The archived files are divided into a total keyword set and a completed keyword set; The keyword archive records all keywords used in each round of crawling, and the completed keyword archive records the keywords that have been used.
[0017] S1.6. Random extraction of keywords: If there is no residual archived file, consider this operation as a new round of crawling. Randomly extract several pieces of data from the keyword thesaurus files in each field.
[0018] S1.7. Keyword archiving: Store the keywords extracted in this round of operation into the total keyword set.
[0019] S1.8. Keyword reading from archive: If there is a residual archived file, consider this operation as a continuation of the previous round of crawling. By reading the total keyword set and the completed keyword set and taking their difference set, obtain the keywords that have not been used.
[0020] S1.9. Traversal search for keywords: Conduct a traversal search for the keywords used in this round. Consider this keyword as the seed keyword, and store it in the completed keyword set after the keyword search.
[0021] S1.10. Pre-avoidance of anti-spider measures: Pre-avoid possible anti-spider detection measures by disabling the automation control features of the Blink rendering engine, setting the User-Agent option, modifying the experimental options excludeSwitches and useAutomationExtension, and using JavaScript to overwrite the navigator.webdriver property;
[0022] S1.11. Determine whether the spider is countered: Determine whether the anti-spider measures of the search engine are effective by detecting the web page structure obtained by the search;
[0023] S1.12. Introduce a recommended jump strategy. When the current keyword can no longer produce search results, jump through the relevant recommendations of the Baidu search engine to ensure that the number of browsing topics in each field is roughly the same;
[0024] S1.13. Traverse search results: After obtaining the search results, in order to simulate user behavior and avoid anti-spider measures, the Selenium spider will randomly stay on the page for a period of time;
[0025] S1.14. Search page jump: Select sequential page jump or random page jump through the parameters set by the user;
[0026] S1.15. Delete the archive: After all the keywords used in this round of operation are traversed, delete the archive file that records the total keyword set and the completed keyword set to obtain the original traffic.
[0027] Furthermore, the specific implementation method of step S2 includes the following steps:
[0028] S2.1. Listening protocol selection: Select the network protocol type to be collected, including one of DNS and HTTP;
[0029] S2.2. BPF filtering condition setting: Set BPF filtering conditions to quickly filter underlying protocols such as the transport layer, network layer, data link layer, and physical layer;
[0030] S2.3. Packet sniffing: Use Scapy to listen to the network traffic of the specified network adapter;
[0031] S2.4. BPF quick filtering: According to the BPF filtering conditions, quickly filter the network traffic;
[0032] S2.5. Deep filtering: Filter the network traffic that meets the requirements of the underlying protocol at the application layer level;
[0033] S2.6. Traffic preservation: Save the traffic that meets the conditions in an additive manner to obtain the original traffic file.
[0034] Furthermore, the specific implementation method of step S3 includes the following steps:
[0035] S3.1. Reading the original traffic file: All original traffic files are stored in the same specified folder path. By specifying the folder path, extract the file names in the folder and save them in the form of a list to obtain the file name list.
[0036] S3.2. File traversal: Use the file names and paths provided by the file name list to read the traffic data stored in the files one by one through an iterator.
[0037] S3.3. Process the traffic data of a single file obtained in step S3.2 as follows:
[0038] S3.3.1. Database initialization: Create a corresponding data packet metadata table in the specified database pcap_data.db through the file name and path.
[0039] S3.3.2. Data packet filtering: Set the filtering conditions as DNS packets with the transport layer being UDP or TCP, the network layer being IPv4 or IPv6, and the packet types being A, AAAA, CNAME, MX, TXT, HTTPS types. Filter the data in the traffic file read in step S3.2 to obtain the filtered DNS data packets.
[0040] S3.3.3. Aggregation key extraction: Extract the address information and protocol of the filtered DNS data packets to form a five-tuple aggregation key, including source IP, source port, protocol, destination port, and destination IP.
[0041] Normalize the tuple order according to the port number size, place the address with the smaller port number in the front for subsequent aggregation calculation. The existence of the protocol attribute in the aggregation key is to leave room for the future expansion of the detection scope.
[0042] S3.3.4. Data packet-level metadata extraction: Extract and calculate the data packet-level metadata from the filtered DNS data packets, including domain name entropy, timestamp, transport layer payload length, number of domain name labels, and ratio of non-lowercase characters in the domain name.
[0043] S3.4. Database update: Insert the aggregation key and packet-level metadata of a single data packet into the table. Set to submit an update to the database every 10,000 inserted data, and use this strategy as a heartbeat mechanism to observe whether the program has a memory explosion phenomenon.
[0044] Further, the specific implementation method of step S4 includes the following steps:
[0045] S4.1. Set extraction-related parameters, including session timeout length Timeout_Num, time window sliding step Time_Step, time window length Time_Window, maximum number of empty windows Max_Null_Num, and attenuation factor Weak_Factor;
[0046] S4.2. Read the packet-level metadata table obtained in step S3.4, read the names of all user-created tables in the specified database, and store them in a list to obtain a table name list;
[0047] S4.3. Table traversal: Traverse the packet metadata table through the packet metadata table names provided by the table name list obtained in step S4.2;
[0048] S4.4. Process a single packet metadata table, including the following steps:
[0049] S4.4.1. Database initialization: Create a corresponding session metadata table and session time table in the specified database dns_tunnel.db through the metadata table names provided by the table name list. The session time table records the start time and end time of each session, and uses a five-tuple aggregation key as the primary key; the session metadata table records the session metadata after damped increment calculation;
[0050] S4.4.2. Configure the time window: Retrieve the timestamps of the data in the packet metadata table, with its minimum value as the start time first_time and the maximum value as the end time last_time; Based on the time window length Time_Window, the value range of the time window ranges from the starting window [0, Time_Window] to the ending window [last_time, last_time + Time_Window] through window sliding / jumping;
[0051] In the case of normal window sliding, the relevant calculation formulas for the time window are shown in Formulas 1, 2, and 3:
[0052] window_right0 = first_time + Time_Window (1)
[0053] window_right i = window_right i-1 + Time_Step (2)
[0054] window_left i= window_right i - Time_Window(3)
[0055] where window_right0 is the right boundary of the initial window, and window_right i is the right boundary of the i-th time window, window_left i is the left boundary of the i-th time window, window_right i-1 is the right boundary of the (i - 1)-th time window, Time_Window is the time window size, Time_Step is the time window sliding step, and first_time is the start time;
[0056] S4.4.3. Extract packet-level metadata: Retrieve the packet data metadata table, and extract the metadata with timestamps within the time window interval configured in step S4.4.2;
[0057] S4.4.4. Metadata aggregation: Consider the set of data packets with the same aggregation key as a session, and perform aggregation calculations on the packet metadata under the same key using the five-tuple aggregation key;
[0058] During the calculation process, the packet timestamps are separated from other metadata, and an integer 1 representing the number of individual packets is inserted into the metadata header. After the aggregation calculation is completed, the session metadata dictionary and session time dictionary for the current window are obtained, corresponding to the session metadata table and session time table created in step S4.4.1;
[0059] S4.4.5. Damping increment calculation: Dampen and decay the historical session metadata in the database, and perform increment calculations on the session metadata under the current window; Introduce the concept of session timeout length Timeout_Num, that is, if the time interval between a new data packet and the previous data packet between two communication parties is greater than the timeout length Timeout_Num, then this communication is considered a new session; The session metadata replaces the update in an insert manner, changing the session metadata table, and the damping increment d λ (t) is calculated according to formula 4 as follows:
[0060] d λ (t)= 2 -λt (4)
[0061] where λ is the attenuation factor and t is the interval time;
[0062] The calculation formula for the interval time t is as shown in formula 5:
[0063] t = cur_time - session_last_time (5)
[0064] Among them, session_last_time is the last update time of the session in this historical data, and cur_time is the time of the current time window;
[0065] S4.4.5. Database update: Update the metadata calculated in step S4.4.4 in the database;
[0066] S4.4.6. Window sliding or jumping: The sliding step of the time window is Time_Step. That is, when there is still no traffic in the window after multiple window slides, retrieve the time stamp in the packet metadata table that is the closest to the current window and greater than the current window, and assign it to the next window to achieve window jumping.
[0067] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method for extracting traffic metadata based on damping increment statistics and Selenium are implemented.
[0068] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the method for extracting traffic metadata based on damping increment statistics and Selenium is implemented.
[0069] Advantages of the present invention:
[0070] For the method for extracting traffic metadata based on damping increment statistics and Selenium of the present invention, Selenium crawlers are used to generate DNS raw traffic, and the Scapy traffic capture module is used to automatically collect the traffic. Subsequently, deep packet inspection technology is used to extract the load characteristics of the raw traffic, and data aggregation technology is used to aggregate some characteristics into traffic characteristics. Finally, an improved damping increment statistical algorithm is used to form session metadata (Session Metadata).
[0071] For the method for extracting traffic metadata based on damping increment statistics and Selenium of the present invention, the behavior of users in a real network environment is better simulated during the traffic generation process, and the collected traffic is closer to real traffic. In addition, during the metadata extraction process, the feature metadata at the packet level and session level are more comprehensively considered, and the designed algorithm can efficiently complete the extraction and calculation of feature metadata. Description of the drawings
[0072] Figure 1 It is a flowchart of the method for extracting traffic metadata based on damping increment statistics and Selenium of the present invention;
[0073] Figure 2 This is the working mechanism diagram of the Selenium crawler of the present invention;
[0074] Figure 3 This is the flowchart of the network traffic capture program based on Scapy of the present invention;
[0075] Figure 4 This is the diagram of packet metadata extraction based on deep packet inspection of the present invention;
[0076] Figure 5 This is the diagram of session metadata generation based on damping increment statistics and data aggregation of the present invention. Detailed implementation manners
[0077] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners. It should be understood that the specific implementation manners described herein are only used to explain the present invention and are not used to limit the present invention, that is, the described specific implementation manners are only a part of the implementation manners of the present invention, rather than all of the specific implementation manners. The components of the specific implementation manners of the present invention usually described and shown in the accompanying drawings here can be arranged and designed in various different configurations, and the present invention can also have other implementation manners.
[0078] Therefore, the detailed description of the specific implementation manners of the present invention provided in the accompanying drawings below is not intended to limit the scope of the claimed present invention, but only represents the selected specific implementation manners of the present invention. All other specific implementation manners obtained by those skilled in the art based on the specific implementation manners of the present invention without creative efforts fall within the scope of protection of the present invention.
[0079] To further understand the content, features and effects of the present invention, the following specific implementation manners are exemplified and combined with the attached Figure 1 - Attached Figure 5 The details are as follows: Specific implementation manner 1:
[0081] A method for extracting traffic metadata based on damping increment statistics and Selenium, comprising the following steps:
[0082] S1. Generate original traffic based on Selenium;
[0083] Further, the specific implementation method of step S1 includes the following steps:
[0084] S1.1. Parameter setting: Set the parameters related to the program, including the number of search pages, the waiting time after recommended jump, the waiting time after page flipping, the web page browsing time, the sleep time before clicking, and the number of keywords extracted from each thesaurus;
[0085] S1.2. Driver Selection: Select a browser driver, such as ChromeDriver;
[0086] S1.3. Search Engine Selection: Select a search engine as the basis for subsequent searches, such as Baidu, Sohu, Bing;
[0087] S1.4. Keyword Thesaurus Selection: Select a keyword thesaurus as the seed for the search;
[0088] S1.5. Determine if there is an archived file: Adopt an interrupted line continuation mechanism. By judging whether there is a residual archived file, determine whether the program has crashed, and then decide whether this operation is a new round of crawling or a continuation of the previous round of crawling; The archived files are divided into the total keyword set and the completed keyword set; The keyword archive records all the keywords used in each round of crawling, and the completed keyword archive records the keywords that have been used;
[0089] S1.6. Random Extraction of Keywords: If there is no residual archived file, consider this operation as a new round of crawling. Randomly extract several pieces of data from the keyword thesaurus files in each field;
[0090] S1.7. Keyword Archiving: Store the keywords extracted in this round of operation into the total keyword set;
[0091] S1.8. Keyword Reading from Archive: If there is a residual archived file, consider this operation as a continuation of the previous round of crawling. By reading the total keyword set and the completed keyword set and taking their difference set, obtain the keywords that have not been used;
[0092] S1.9. Traversal Search of Keywords: Conduct a traversal search on the keywords used in this round. Consider this keyword as the seed keyword, and store it in the completed keyword set after the keyword search;
[0093] S1.10. Pre - avoidance of Anti - Crawler Measures: Pre - avoid possible crawler detection measures by disabling the automation control features of the Blink rendering engine, setting the User - Agent option, modifying the experimental options excludeSwitches, useAutomationExtension, and using JavaScript to override the navigator.webdriver property;
[0094] S1.11. Determine if the Crawler is Countered: Determine whether the anti - crawler measures of the search engine are effective by detecting the web page structure obtained from the search;
[0095] S1.12. Introduce a recommended jump strategy. When the current keyword cannot continue to produce search results, jump through the relevant recommendations of the Baidu search engine to ensure that the number of browsed topics in each field is roughly the same;
[0096] S1.13. Search result traversal: After obtaining the search results, in order to simulate user behavior and avoid anti-crawler measures, the Selenium crawler will randomly stay on the page for a period of time;
[0097] S1.14. Search page jump: Select sequential page jump or random page jump through the parameters set by the user;
[0098] Furthermore, select sequential page jump or random page jump. Among them, sequential page jump is more in line with the user's usage habits. Users often only focus on the web pages with higher rankings when using search engines, that is, sequential page browsing. While random page jump can expand the web page coverage of traffic data collection and solve the problem of the low ranking of web pages with less traffic in search results, making the traffic more comprehensive;
[0099] S1.15. Delete archive: After all the keywords used in this round of operation are traversed, delete the archive file that records the total keyword set and the completed keyword set to obtain the original traffic;
[0100] Furthermore, the main goal of the Selenium crawler is to collect as much normal network traffic as possible in an environment that simulates the user's daily network behavior. In order to increase the diversity of the collected network traffic, the THUOCL open Chinese word library of Tsinghua University is used in the selection of keywords. The relevant information of THUOCL is shown in Table 1.
[0101] Table 1 THUOCL Word Library
[0102]
[0103] S2. Use Scapy to collect the original traffic generated in step S1 to obtain the original traffic file;
[0104] Furthermore, the specific implementation method of step S2 includes the following steps:
[0105] S2.1. Listening protocol selection: Select the type of network protocol to be collected, including one of DNS and HTTP;
[0106] S2.2. BPF filtering condition setting: Set BPF filtering conditions to quickly filter the underlying protocols such as the transport layer, network layer, data link layer, and physical layer;
[0107] S2.3. Packet sniffing: Use Scapy to listen to the network traffic of the specified network adapter;
[0108] S2.4. BPF quick filtering: According to the BPF filtering conditions, quickly filter the network traffic;
[0109] S2.5. Deep filtering: Filter network traffic that meets the requirements of the underlying protocol at the application layer;
[0110] S2.6. Traffic preservation: Preserve eligible traffic in an additive manner to obtain the original traffic file;
[0111] S3. First preprocess the original traffic file obtained in step S2, and then extract packet-level metadata based on deep packet inspection to obtain packet-level metadata;
[0112] The typical characteristics of DNS tunnels can be divided into two categories: payload characteristics and traffic characteristics. Payload characteristics extract features from a more microscopic and specific perspective by analyzing the load information of a single traffic packet or multiple traffic packets. Traffic characteristics extract features from a more macroscopic perspective by analyzing the traffic information of the overall traffic over a period of time.
[0113] In view of the requirements for real-time performance and the detection performance of low-throughput tunnels, load characteristics and traffic characteristics are comprehensively selected based on technologies such as deep packet inspection, data aggregation, and damping increment statistics. The selected characteristics are shown in Table 2.
[0114] Table 2 Related characteristics
[0115]
[0116] In view of the characteristics of low-throughput DNS tunnels, the traffic characteristics packet_num and dur_time are selected as analysis indicators. On the premise that the length of the leaked data is fixed, in order to avoid detection, low-throughput DNS tunnels usually take measures such as reducing the length of the leaked data per time, increasing the total number of packets, and extending the total leakage time. The increase in the total number of packets will be abnormal in the traffic characteristic packet_num, and the extension of the leakage time will make the traffic characteristic dur_time show abnormality.
[0117] DNS tunnel packets can be divided into tunnel query packets for transmitting leaked data and tunnel response packets for transmitting control commands. The above two types of packets are different from normal DNS query and response packets in function and also have corresponding differences in structural content. Therefore, in view of the differences in the transport layer protocol load level of DNS tunnel traffic, the load characteristic payload_size is selected.
[0118] As the core content of DNS messages, domain name-related features play an important role in the detection of DNS tunnels. According to the principle of DNS tunnels, the leaked data is mainly carried by the domain name fields of DNS tunnel messages. Considering the characteristics that the leaked data in DNS tunnels is usually encoded and large in volume, the load features domain_entropy, domain_lable, and domain_NL_rate are selected.
[0119] Character statistics were performed on 248,726 benign domain names from the Alexa domain ranking and 105,083 tunnel domain names generated by DNS tunnel tools. The statistical results show that lowercase characters in benign domain names account for approximately 97.9% of the total proportion, numbers account for approximately 1%, and the total proportion of underscores and hyphens is approximately 1.07%. In contrast, in tunnel domain names, the proportion of lowercase letters is approximately 28.97%, numbers are approximately 16%, uppercase letters are approximately 54%, and the proportion of underscores and hyphens is only about 0.0038%. In addition, special characters (excluding underscores and hyphens) account for approximately 1.15% of the total proportion. From these data, it can be seen that compared with benign domain names, the characteristics of tunnel domain names show some obvious changes. In tunnel domain names, the proportion of lowercase letters, underscores, and hyphens has decreased significantly, while the proportion of numbers has increased significantly. In addition, uppercase letters and special characters other than underscores and hyphens have also appeared in tunnel domain names. Therefore, the non-lowercase character ratio domain_NL_rate in the domain name is selected as one of the load features.
[0120] In the research stage, the extraction of session metadata based on deep packet inspection and damped incremental statistics is divided into two stages: the extraction of packet metadata based on deep packet inspection and the generation of session metadata based on damped incremental statistics and data aggregation. The purpose of the step-by-step processing is to use the database to process large traffic files, thereby reducing the requirements for the hardware performance of the training device and improving the efficiency of data retrieval in subsequent processes. Since the SQLite database has advantages such as being open-source, lightweight, easy to deploy, and having good cross-platform performance, it is adopted in the research stage.
[0121] Furthermore, the specific implementation method of step S3 includes the following steps:
[0122] S3.1. Reading of original traffic files: All original traffic files are stored in the same specified folder path. By specifying the folder path, the file names in the folder are extracted and saved in the form of a list to obtain a file name list;
[0123] S3.2. File traversal: Using the file names and paths provided by the file name list, the traffic data stored in the files is read one by one through an iterator;
[0124] S3.3. Process the traffic data of the single file obtained in step S3.2 as follows:
[0125] S3.3.1. Database initialization: Create a corresponding packet metadata table in the specified database pcap_data.db through the file name and path, as shown in Table 3:
[0126] Table 3 Packet Metadata Table
[0127]
[0128]
[0129] S3.3.2. Packet filtering: Set the filtering conditions as DNS packets with the transport layer being UDP or TCP, the network layer being IPv4 or IPv6, and the packet types being A, AAAA, CNAME, MX, TXT, HTTPS types. Filter the data in the traffic file read in step S3.2 to obtain the filtered DNS packets;
[0130] S3.3.3. Aggregation key extraction: Extract the address information and protocol of the filtered DNS packets to form a five-tuple aggregation key, including source IP, source port, protocol, destination port, and destination IP;
[0131] Normalize the tuple order according to the port number size, place the address with the smaller port number in the front for subsequent aggregation calculation. The existence of the protocol attribute in the aggregation key is to leave room for future expansion of the detection scope;
[0132] S3.3.4. Packet-level metadata extraction: Extract and calculate the packet-level metadata from the filtered DNS packets, including domain name entropy, timestamp, transport layer payload length, number of domain name labels, and ratio of non-lowercase characters in the domain name;
[0133] S3.4. Database update: Insert the aggregation key and packet-level metadata of a single packet into the table. Set to submit an update to the database every 10,000 inserted data, and use this strategy as a heartbeat mechanism to observe whether there is a memory explosion phenomenon in the program.
[0134] S4. For the packet-level metadata obtained in step S3, first perform preprocessing, and then generate session metadata based on metadata aggregation and damped incremental statistics;
[0135] Furthermore, the specific implementation method of step S4 includes the following steps:
[0136] S4.1. Set extraction-related parameters, including session timeout length Timeout_Num, time window sliding step Time_Step, time window length Time_Window, maximum number of consecutive empty windows Max_Null_Num, and attenuation factor Weak_Factor;
[0137] S4.2. Read the packet-level metadata table obtained in step S3.4, read the names of all user-created tables in the specified database, and store them in a list to obtain a table name list;
[0138] S4.3. Table traversal: Traverse the packet metadata table using the packet metadata table names provided by the table name list obtained in step S4.2;
[0139] S4.4. Process a single packet metadata table, including the following steps:
[0140] S4.4.1. Database initialization: Create corresponding session metadata tables and session time tables in the specified database dns_tunnel.db using the metadata table names provided by the table name list. The session time table records the start time and end time of each session, and uses a five-tuple aggregation key as the primary key; the session metadata table records the session metadata after damped increment calculation;
[0141] Furthermore, use the five-tuple aggregation key (source IP, source port, protocol, destination port, destination IP) as the primary key, and each record corresponding to a session is unique. The session metadata table records the session metadata after damped increment calculation. Separating the two tables has many advantages: Considering the time cost, separating the two tables is conducive to quickly retrieving session time information in the time table. Considering the space cost, since the start times of the same sessions are the same, separating the two tables is conducive to saving space. The session time table is shown in Table 4, and the session metadata table is shown as follows:
[0142] Table 4 Session Time Table
[0143]
[0144] Table 5 Session Metadata Table
[0145]
[0146] S4.4.2. Configure the time window: Retrieve the timestamps of the data in the data packet metadata table, with the minimum value as the start time first_time and the maximum value as the end time last_time; Based on the time window length Time_Window, the value range of the time window is from the initial window [0, Time_Window] to the final window [last_time, last_time + Time_Window] through window sliding / jumping;
[0147] In the case of normal window sliding, the relevant calculation formulas for the time window are shown in Formulas 1, 2, and 3:
[0148] window_right0 = first_time + Time_Window (1)
[0149] window_right i = window_right i-1 + Time_Step (2)
[0150] window_left i = window_right i - Time_Window (3)
[0151] Among them, window_right0 is the right boundary of the initial window, window_right i is the right boundary of the i-th time window, window_left i is the left boundary of the i-th time window, window_right i-1 is the right boundary of the (i - 1)-th time window, Time_Window is the time window size, Time_Step is the time window sliding step, and first_time is the start time;
[0152] S4.4.3. Extract packet-level metadata: Retrieve the data packet metadata table and extract the metadata with timestamps within the time window interval configured in step S4.4.2;
[0153] S4.4.4. Metadata aggregation: Consider the data packet sets with the same aggregation key as a session, and use the five-tuple aggregation key to perform aggregation calculations on the data packet metadata under the same key;
[0154] During the calculation process, the data packet timestamp is separated from other metadata, and the integer 1 representing the quantity of a single data packet is inserted into the metadata header. After the aggregation calculation is completed, the session metadata dictionary and session time dictionary of the current window are obtained, corresponding to the session metadata table and session time table created in step S4.4.1; the metadata aggregation is as follows:
[0155] Algorithm metadata aggregation calculation Metadata_Aggregation
[0156] Input: Packet set Packets
[0157] Output: Session metadata dictionary Sessions_Metadata, session time dictionary Sessions_Time
[0158] 1: Data aggregation stage:
[0159] 2: Create dictionary Packets_Metadata
[0160] 3: for packet in Packets do
[0161] 4: Extract the aggregation key packet_key of the current data packet
[0162] 5: Extract the metadata packet_metadata of the current data packet
[0163] 6: Put the data packets with the same aggregation key into the set under the same key in the dictionary Packest_Metadata
[0164] 7: end for
[0165] 8: Aggregation calculation stage:
[0166] 9: Create session metadata dictionary Sessions_Metadata
[0167] 10: Create session time dictionary Sessions_Time
[0168] 11: for packets_metadata in Packets_Metadata do
[0169] 12: for packet_metadata in packets_metadata do
[0170] 13: Extract the current data packet aggregation key aggregation_key
[0171] 14: Extract the current data packet timestamp time
[0172] 15: Extract the metadata of the current data packet, delete the timestamp, and add 1 representing the quantity at the beginning.
[0173] 16: If the current aggregation key does not exist in the session time dictionary Sessions_Time
[0174] 17: Create an entry in the session time data dictionary with the current aggregation key as the key and (time, time) as the value.
[0175] 18: Create an entry in the session metadata dictionary with the current aggregation key as the key and the processed metadata as the value.
[0176] 19: Otherwise
[0177] 20: Extract the historical metadata tuple session_data of this key from the session metadata dictionary.
[0178] 21: Modify the session time dictionary, and change the last element of the tuple representing the latest time under this key to time.
[0179] 22: Modify the session metadata dictionary, add session_data and metadata bit by bit and modify the corresponding values.
[0180] 23: end for
[0181] 24: end for.
[0182] S4.4.5. Damping Increment Calculation: Dampen the historical session metadata in the database and perform an incremental calculation on the session metadata under the current window; introduce the concept of session timeout length Timeout_Num, that is, if the time interval between the new data packet and the previous data packet between the two communication parties is greater than the timeout length Timeout_Num, then this communication is considered a new session; the session metadata replaces the update in an inserted manner, change the session metadata table, and the damping increment d λ (t) is calculated as shown in Equation 4:
[0183] d λ (t) = 2 -λt (4)
[0184] Among them, λ is the attenuation factor and t is the interval time;
[0185] The calculation formula for the interval time t is shown in Equation 5:
[0186] t = cur_time - session_last_time(5)
[0187] Among them, session_last_time is the last update time of the session in this historical data, and cur_time is the time of the current time window;
[0188] The steps for calculating the damping increment are as follows:
[0189] Algorithm Damping Decay and Increment Calculation Decay_and_Caculation
[0190] Input: Session metadata dictionary Sessions_Metadata, session time dictionary Sessions_Time, database Database, session timeout length Timeout_Num
[0191] Output: Database Database
[0192] 1: for key in Sessions_Time.Keys do
[0193] 2: Extract the session start time session_start_time of this session recorded in the database session time table
[0194] 3: Extract the session latest update time session_update_time of this session recorded in the database session time table
[0195] 4: Extract the start time cur_start_time of this session in the session time dictionary
[0196] 5: Extract the latest update time cur_end_time of this session in the session time dictionary
[0197] 6: If this key does not exist in the database
[0198] 7: The session duration dur_time is the time difference cur_end_time - cur_start_time of the session in the dictionary
[0199] 8: The inserted data insert_data is the union of this session metadata metadata and the session duration dur_time
[0200] 9: Insert the union of the aggregation key key and the inserted data insert_data into the database session time table
[0201] 10: Otherwise
[0202] 11: The time interval interval is the difference between cur_start_time and session_update_time
[0203] 12: If the time interval is greater than the session timeout length, it is regarded as a new session.
[0204] 13: The session duration dur_time is the time difference cur_end_time - cur_start_time of the session in the dictionary.
[0205] 14: The data to be inserted insert_data is the union of the session metadata metadata and the duration dur_time.
[0206] 15: Update the database time table, update the start and latest times, with the values being cur_start_time and cur_end_time.
[0207] 16: Otherwise
[0208] 17: The session duration dur_time is the difference between cur_end_time and session_start_time.
[0209] 18: The data to be calculated cal_data is the union of the session metadata metadata and the duration dur_time.
[0210] 19: Extract the historical latest data history_data from the database metadata table.
[0211] 20: Calculate the data to be inserted insert_data according to the following formula:
[0212] insert_data ← cal_data + history_data · 2 -λ′interval
[0213] 21: Update the database session time table, only update the latest value of the session, with the value being cur_end_time.
[0214] 24: Insert into the database session metadata table, update the value to be the union of the aggregation key key and the data to be inserted insert_data.
[0215] 25: end for.
[0216] S4.4.5. Database update: Update the metadata calculated in step S4.4.4 in the database;
[0217] S4.4.6. Window Sliding or Jumping: The sliding step of the time window is Time_Step. That is, when there is still no traffic in the window after multiple window slides, retrieve the time stamp in the data packet metadata table that is the closest to the current window and greater than the current window, and assign it to the next window to achieve window jumping.
[0218] Embodiment 2:
[0219] An electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of a method for extracting traffic metadata based on damping increment statistics and Selenium described in Embodiment 1 are implemented.
[0220] The computer device of the present invention may be a device including a processor and a memory, such as a single-chip microcomputer including a central processing unit. And, when the processor is used to execute the computer program stored in the memory, the steps of the above-mentioned recommendation method for modifying relationship-driven recommendation data based on CREO software are implemented.
[0221] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0222] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0223] Embodiment 3:
[0224] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, it implements a method for extracting traffic metadata based on damping increment statistics and Selenium described in Embodiment 1.
[0225] The computer-readable storage medium of the present invention can be any form of storage medium readable by the processor of a computer device, including but not limited to non-volatile memory, volatile memory, ferroelectric memory, etc. A computer program is stored on the computer-readable storage medium. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of the above-mentioned modeling method for modifying relationship-driven modeling data based on CREO software can be implemented.
[0226] The computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0227] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0228] Although the present application has been described above with reference to specific embodiments, various improvements can be made thereto and components thereof can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the features in the specific embodiments disclosed in the present application can be combined with each other in any way, and the exhaustive description of these combinations is not given in this specification only for the consideration of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A method for extracting traffic metadata based on damping increment statistics and Selenium, characterized in that It includes the following steps: S1. Generate original traffic based on Selenium; S2. Collect the original traffic generated in step S1 based on Scapy to obtain an original traffic file; S3. First preprocess the original traffic file obtained in step S2, and then extract packet-level metadata based on deep packet inspection to obtain packet-level metadata; The metadata of the packet includes domain name entropy, timestamp, transport layer payload length, number of domain name labels, and ratio of non-lowercase characters in the domain name; Aggregation key extraction: Extract the address information and protocol of the filtered DNS packets to form a five-tuple aggregation key, including source IP, source port, protocol, destination port, and destination IP; S4. First preprocess the packet-level metadata obtained in step S3, and then generate session metadata based on metadata aggregation and damping increment calculation; Among them, metadata aggregation: Consider the set of packets with the same aggregation key as a session, and perform aggregation calculation on the packet metadata under the same key using the five-tuple aggregation key; During the calculation process, the packet timestamp is separated from other metadata, and an integer 1 representing the number of individual packets is inserted into the metadata header. After the aggregation calculation is completed, a session metadata dictionary and a session time dictionary for the current window are obtained, corresponding to the created session metadata table and session time table; Among them, damping increment calculation: Dampen and decay the historical session metadata in the database, and perform increment calculation on the session metadata under the current window; Introduce the concept of session timeout length Timeout_Num, that is, if the time interval between a new packet and the previous packet between two communication parties is greater than the timeout length Timeout_Num, then this communication is considered a new session; The session metadata replaces the update in an insert manner, and changes are made to the session metadata table; Calculate the data to be inserted insert_data according to the following formula: insert_data←cal_data+history_data·2 -λ·interval Among them, the data to be calculated cal_data is the union of session metadata metadata and session duration dur_time; history_data is the latest historical data in the database metadata table; the session duration dur_time is the difference between cur_end_time and session_start_time; cur_end_time is the latest update time of this session in the session time dictionary, session_start_time is the session start time of the session recorded in the database session time table, metadata is the session metadata in the session metadata dictionary Sessions_Metadata; the time interval interval is the difference between cur_start_time and session_update_time; cur_start_time is the start time of this session in the session time dictionary; session_update_time is the latest update time of the session recorded in the database session time table.
2. The method for extracting traffic metadata based on damping increment statistics and Selenium according to claim 1, wherein The specific implementation method of step S1 includes the following steps: S1.
1. Parameter setting: Set the parameters related to the program, including the number of search pages, the waiting time after recommended jump, the waiting time after page turning, the web page browsing time, the sleep time before clicking, and the number of keywords extracted from each thesaurus; S1.
2. Driver selection: Select the browser driver; S1.
3. Search engine selection: Select a search engine as the basis for subsequent searches; S1.
4. Keyword thesaurus selection: Select a keyword thesaurus as the seed for the search; S1.
5. Determine whether there is an archived file: Adopt an interruption and continuation mechanism. By judging whether there is a residual archived file, determine whether the program has crashed, and then decide whether this operation is a new round of crawling or a continuation of the previous round of crawling; The archived files are divided into the total keyword set and the completed keyword set; The keyword archive records all the keywords used in each round of crawling, and the completed keyword archive records the keywords that have been used; S1.
6. Random extraction of keywords: If there is no residual archived file, regard this operation as a new round of crawling, and randomly extract several pieces of data from the keyword thesaurus files in each field; S1.
7. Keyword archiving: Store the keywords extracted in this round of operation into the total keyword set; S1.
8. Keyword reading from archive: If there is a residual archived file, regard this operation as a continuation of the previous round of crawling. By reading the total keyword set and the completed keyword set and taking their difference set, obtain the keywords that have not been used; S1.
9. Traversal search of keywords: Conduct a traversal search on the keywords used in this round. This keyword is regarded as the seed keyword. After the keyword search, store it into the completed keyword set; S1.
10. Pre-avoidance of anti-crawler measures: Pre-avoid possible crawler detection measures by disabling the automation control features of the Blink rendering engine, setting the User-Agent option, modifying the experimental options excludeSwitches, useAutomationExtension, and using JavaScript to overwrite the navigator.webdriver property; S1.
11. Determine whether the crawler is countered: Determine whether the anti-crawler measures of the search engine take effect by detecting the web page structure obtained from the search; S1.
12. Introduce a recommended jump strategy. When the current keyword can no longer produce search results, jump through the relevant recommendations of the search engine to ensure that the number of browsing topics in each field is roughly the same; S1.
13. Traversal of search results: After obtaining the search results, in order to simulate user behavior and avoid anti-crawler measures, the Selenium crawler will randomly stay on the page for a period of time; S1.
14. Search page jump: Select sequential page jump or random page jump according to the parameters set by the user; S1.
15. Delete archive: When all the keywords used in this round of operation have been traversed, delete the archived files that record the total keyword set and the completed keyword set to obtain the original traffic.
3. The method for extracting traffic metadata based on damping increment statistics and Selenium according to claim 2, wherein The specific implementation method of step S2 includes the following steps: S2.
1. Listening protocol selection: Select the type of network protocol to be collected, including one of DNS and HTTP; S2.
2. BPF Filter Condition Setting: Set the BPF filter condition to quickly filter underlying protocols such as the transport layer, network layer, data link layer, and physical layer; S2.
3. Packet Sniffing: Use Scapy to monitor the network traffic of a specified network adapter; S2.
4. BPF Quick Filtering: According to the BPF filter condition, quickly filter the network traffic; S2.
5. Deep Filtering: Filter the network traffic that meets the requirements of the underlying protocol at the application layer; S2.
6. Traffic Saving: Save the traffic that meets the conditions in an additive manner to obtain the original traffic file.
4. A method for extracting traffic metadata based on damping increment statistics and Selenium according to claim 3, characterized in that The specific implementation method of step S3 includes the following steps: S3.
1. Reading the Original Traffic File: All original traffic files are stored in the same specified folder path. By specifying the folder path, extract the file names in the folder and save them in the form of a list to obtain the file name list; S3.
2. File Traversal: Use the file names and paths provided by the file name list to read the traffic data stored in the files one by one through an iterator; S3.
3. Perform the following processing on the traffic data of a single file obtained in step S3.2: S3.3.
1. Database Initialization: Create a corresponding packet metadata table in the specified database pcap_data.db through the file name and path; S3.3.
2. Packet Filtering: Set the filtering condition to DNS packets with UDP or TCP at the transport layer, IPv4 or IPv6 at the network layer, and the packet types being A, AAAA, CNAME, MX, TXT, or HTTPS type. Filter the data in the traffic file read in step S3.2 to obtain the filtered DNS packets; S3.3.
3. Aggregation Key Extraction: Extract the address information and protocol of the filtered DNS packets to form a five-tuple aggregation key, including source IP, source port, protocol, destination port, and destination IP; Normalize the tuple order according to the port number size, place the address with the smaller port number at the front for subsequent aggregation calculation. The existence of the protocol attribute in the aggregation key is to leave room for future expansion of the detection scope; S3.3.
4. Packet-Level Metadata Extraction: Extract and calculate the packet-level metadata from the filtered DNS packets, including domain name entropy, timestamp, transport layer payload length, number of domain name labels, and ratio of non-lowercase characters in the domain name; S3.
4. Database Update: Insert the aggregation key and packet-level metadata of a single packet into the table. Set to submit an update to the database every 10,000 inserted data, and use this strategy as a heartbeat mechanism to observe whether there is a memory explosion phenomenon in the program.
5. A method for extracting traffic metadata based on damping increment statistics and Selenium according to claim 4, characterized in that The specific implementation method of step S4 also includes the following steps: S4.
1. Set the relevant extraction parameters, including session timeout length Timeout_Num, time window sliding step Time_Step, time window length Time_Window, maximum number of empty windows Max_Null_Num, and attenuation factor Weak_Factor; S4.
2. Read the packet-level metadata table obtained in step S3.4, read the names of all tables created by the user in the specified database, and store them in a list to obtain a table name list; S4.
3. Table traversal: Traverse the packet-level metadata table through the packet-level metadata table names provided by the table name list obtained in step S4.2; S4.
4. Process the individual packet-level metadata tables, including the following steps: Database initialization: Create corresponding session metadata tables and session time tables in the specified database dns_tunnel.db through the metadata table names provided by the table name list. The session time table records the start time and end time of each session, and uses a five-tuple aggregation key as the primary key; The session metadata table records the session metadata after damped increment calculation; Configure the time window: Retrieve the timestamps of the data in the packet-level metadata table, with its minimum value as the start time first_time and the maximum value as the end time last_time; Based on the time window length Time_Window, the time window value range slides or jumps from the start window [0, Time_Window] to the end window [last_time, last_time + Time_Window]; In the case of normal window sliding, the relevant calculation formulas for the time window are shown in Formulas 1, 2, and 3: window_right0 = first_time + Time_Window (1) window_right i = window_right i-1 + Time_Step (2) window_left i = window_right i - Time_Window (3) Among them, window_right0 is the right boundary of the initial window, and window_right i is the right boundary of the i-th time window, and window_left i is the left boundary of the i-th time window, and window_right i-1 is the right boundary of the (i - 1)-th time window, Time_Window is the length of the time window, Time_Step is the sliding step of the time window, and first_time is the starting time; Extract packet-level metadata: Retrieve the packet-level metadata table, and extract the metadata with timestamps in the configured time window interval; The damping increment calculation also includes the damping increment d λ (t), and its calculation formula is as shown in Formula 4: d λ (t) = 2 -λt (4) Among them, λ is the attenuation factor and t is the interval time; The calculation formula for the interval time t is shown in Formula 5: t = cur_time - session_update_time (5) Among them, cur_time is the current time window time; Database update: Update the calculated metadata in the database; Window sliding or jumping: The sliding step of the time window is Time_Step. When there is still no traffic in the window after multiple window slides, retrieve the timestamp in the packet-level metadata table that is the closest to the current window and greater than the current window, and assign it to the next window to achieve window jumping.
6. An electronic device, characterized in that, It includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the steps of a method for extracting traffic metadata based on damped increment statistics and Selenium according to any one of claims 1-5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements a method for extracting traffic metadata based on damped increment statistics and Selenium according to any one of claims 1-5.
Citation Information
Patent Citations
Method for extracting, analyzing and searching network flow and content
CN103281213A
Network traffic retrieval method and device
CN113590910A