Domain name system (DNS) traffic log lossless compression method and device based on field sorting

Through segmented reading, field encoding and multi-field sorting, DNS traffic logs are processed, and the problems of limited compression effect and lack of field optimization in the existing technology are solved, and efficient data redundancy reduction and compression ratio improvement are achieved.

CN120201093APending Publication Date: 2025-06-24JIANGSU FUTURE NETWORKS INNOVATION
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510373734.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When compressing large-scale DNS traffic logs, the compression effect is limited and the field optimization and sorting strategies are lacking, resulting in poor redundancy elimination effect and low compression ratio.

Method used

Read the DNS traffic log in segments, extract log information line by line, find and replace the number of domain names and DNS addresses, sort based on multiple fields, convert it into hexadecimal and write it in groups, and convert the time and date into timestamps or differences to form the processed DNS traffic log, and use redundancy to eliminate the compression.

Benefits of technology

Effectively reduce data redundancy, reduce the use of repeated strings, compress time and field space, optimize storage efficiency, improve compression ratio, ensure compact data storage under low entropy characteristics, significantly reduce file size, and facilitate efficient storage and transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201093A_ABST
    Figure CN120201093A_ABST
Patent Text Reader

Abstract

The invention relates to a domain name system (DNS) traffic log lossless compression method and device based on field sorting, and the method comprises the steps: responding to a DNS traffic log compression request, reading data in a segmented manner from a DNS traffic log according to a preset line number, and extracting log information line by line; the codes are replaced by searching the numbers of the domain names and the DNS addresses, and the processed log lines are stored in a memory; sequencing each log segment based on a plurality of fields according to a predetermined rule, converting the sequenced log line number into a 36-system and grouping and writing, converting the time and date into a timestamp or a difference value, and adding the timestamp or the difference value to an index to form a processed DNS traffic log; and performing redundancy elimination class compression on the processed DNS traffic log to generate a compressed file. According to the method, data redundancy is effectively reduced through segmented reading, field coding and multi-field sorting, and meanwhile, the compression ratio is further improved by combining domain name reversal sorting and gzip compression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data compression, and in particular to a lossless compression method and device for DNS traffic logs based on field sorting. Background Art

[0002] The gzip algorithm is the mainstream method for compressing DNS traffic logs, and it has a significant compression ratio advantage when dealing with small-scale and highly redundant text data. However, for large-volume DNS traffic logs with large size and uneven redundancy distribution, the compression effect of gzip is significantly limited, and it cannot fully save storage space, especially when fields such as domain names, DNS addresses, and time information occupy a large number of bytes. In addition, the prior art has not been optimized for the characteristics of DNS traffic logs, such as the repeatability of domain names and DNS addresses, and the verbosity of time information, resulting in low compression efficiency.

[0003] In terms of domain name compression, the existing methods have not effectively utilized the characteristics of high-frequency domain names and have not significantly reduced the storage requirements of dozens to hundreds of bytes. DNS address compression also lacks a systematic coding strategy and fails to streamline the data volume. In timestamp processing, traditional methods directly store the complete time and date string, which occupies at least 14 bytes, and do not optimize the space occupancy through methods such as differences. For large-scale logs, the existing compression technologies have not improved data similarity through field sorting, limiting the effect of redundancy elimination, and the potential of algorithms such as gzip has not been fully exploited.

[0004] In addition, when compressing large-scale DNS traffic logs, the prior art lacks targeted segmentation and sorting strategies, resulting in scattered similar log lines and a low compression ratio. At the same time, there is a lack of an efficient sequential recovery mechanism during the decompression process, increasing the complexity of data reconstruction. These defects jointly restrict the efficient utilization of storage space and the improvement of compression performance. Summary of the Invention

[0005] The purpose of the present invention is to provide a lossless compression method and device for DNS traffic logs based on field sorting to solve the problems of limited compression effect and lack of field optimization and sorting strategies in the prior art.

[0006] To achieve one of the above-mentioned invention purposes, an embodiment of the present invention provides a lossless compression method for DNS traffic logs based on field sorting, which is characterized in that it includes: In response to a DNS traffic log compression request, read data in segments by a preset number of lines from the DNS traffic log, and extract log information line by line; Perform replacement encoding by looking up the numbers of domain names and DNS addresses, and save the processed log lines to memory; Sort each log segment based on multiple fields according to a predetermined rule, convert the sorted log line numbers to base 36 and group them for writing. At the same time, convert the time and date to a timestamp or a difference and append it to the index to form the processed DNS traffic log; Use redundancy elimination type compression on the processed DNS traffic log to generate a compressed file.

[0007] As a further improvement of an embodiment of the present invention, the method further includes obtaining the number of requests for the requested domain name and the number of requests for the DNS address in the historical DNS traffic log, and respectively establishing a domain name coding table and a DNS address coding table, specifically including, The domain name coding table assigns a unique code to the common domain names in descending order of the occurrence frequency of the requested domain names in the historical DNS traffic log, and adds an identifier to distinguish the code from the unencoded domain names; The DNS address coding table sorts all the DNS addresses in the historical DNS traffic log in descending order of the occurrence frequency, assigns a unique code starting from 0 to each DNS address, and replaces the original DNS address with this code.

[0008] As a further improvement of an embodiment of the present invention, the method further includes that the replacement coding by looking up the numbers of the domain name and the DNS address includes, Read the log segments of the preset number of lines in the DNS traffic log compression request; Extract the domain name, time and date, record type and DNS address in the log line by line, look up the domain name number and DNS address number in the domain name coding table and the DNS address coding table for replacement, and save it to the memory after the replacement is completed; If the corresponding domain name number and DNS address code can be found, use the identifier and the number to replace the domain name, and use the number to replace the DNS address; if not, keep the original domain name or DNS address unchanged.

[0009] As a further improvement of an embodiment of the present invention, the method further includes that the "sorting each log segment based on multiple fields according to a predetermined rule" includes, Divide the DNS traffic log of the preset number of lines into multiple log segments. After encoding and replacing each log segment, sort them in turn according to the reversed order of the domain name string, request type, DNS address and the corresponding log segment line number.

[0010] As a further improvement of an embodiment of the present invention, the method further includes that the "converting the sorted log line numbers to base 36 and grouping them for writing, and converting the time and date to a timestamp or a difference and appending it to the index" includes, Convert the line numbers of the sorted log segments corresponding to the original log segments into 36 - character strings, and write the flag indicating the end of the index into the corresponding lines; Convert all the time and dates of each sorted log segment into timestamps, specifically including, Determine whether the current line is the first line other than the index in the processed DNS traffic log; If so, replace the time and date with the timestamp, and write the sorted log after the index; If not, calculate the difference between the timestamp of this line and the timestamp of the first line, replace the time and date with the difference, and write the sorted log after the index.

[0011] As a further improvement of an embodiment of the present invention, the method further includes that the redundancy elimination of the sorted log segments includes, After the timestamp conversion is completed, continue to read the next preset number of lines of log segments until the DNS traffic log is read to the end; if the number of lines of DNS traffic log read at one time is less than the preset number of lines, it is determined that the reading is completed; After the DNS traffic log is processed, apply redundancy elimination - type compression to the processed DNS traffic log to generate a compressed file.

[0012] As a further improvement of an embodiment of the present invention, the method further includes that during decompression, use the index to restore the original log order and reconstruct the original field content of the DNS traffic log through reverse operations, including, When decompressing the compressed DNS traffic log file, convert the 36 - character line numbers in the index back to decimal to restore the line number order in the original log; Add the timestamp difference back to the timestamp of the first line to restore the original time and date string, and replace the encoding with the original domain name and DNS address according to the domain name coding table and the DNS address coding table to ensure the accurate reconstruction of the original content of the DNS traffic log after decompression.

[0013] To achieve one of the above - mentioned invention purposes, an embodiment of the present invention also provides a lossless compression device for DNS traffic logs based on field sorting, which is characterized in that it includes a reading module, a replacement module, a conversion module, and a compression module; The reading module is used to respond to the DNS traffic log compression request, read data in segments by a preset number of lines from the DNS traffic log, and extract the domain name, time and date, record type, and DNS address line by line; The replacement module is used to perform replacement encoding by looking up the numbers of the domain name and DNS address, and save the processed log lines to the memory; The conversion module is used to sort each log segment in the order of reverse domain name, record type, DNS address, and line number, convert the sorted log line numbers into base-36 and group them for writing, and at the same time convert the time and date into timestamps or differences and append them to the index to form processed DNS traffic logs; The compression module is used to perform redundancy elimination-based compression on the processed DNS traffic logs to generate compressed files.

[0014] To achieve one of the above-mentioned invention purposes, an embodiment of the present invention further provides an electronic device, including a memory and a processor, characterized in that the memory stores a computer program that can run on the processor, and when the program is executed on the processor, the steps in the above-mentioned lossless compression method for DNS traffic logs based on field sorting are implemented.

[0015] To achieve one of the above-mentioned invention purposes, an embodiment of the present invention further provides a storage medium, the storage medium stores a computer program, characterized in that when the computer program is executed by a processor, the steps in the above-mentioned lossless compression method for DNS traffic logs based on field sorting are implemented.

[0016] Compared with the prior art, the present invention provides a lossless compression method and device for DNS traffic logs based on field sorting. By segment reading, field encoding, and multi-field sorting, data redundancy is effectively reduced. Among them, the number replacement of domain names and DNS addresses reduces the occupation of repeated strings, and the timestamp difference encoding compresses the time field space. The base-36 line number conversion optimizes the storage efficiency. Combining domain name reverse sorting and gzip compression further improves the compression ratio and ensures compact storage of data under low entropy characteristics. This method significantly reduces the file size while maintaining the integrity of the logs, facilitating efficient storage and transmission, and is particularly suitable for the management and analysis of large-scale DNS traffic logs. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is the overall flowchart of the lossless compression method for DNS traffic logs based on field sorting according to the present invention.

[0018] Figure 2 is the specific flowchart of an implementation scenario of the lossless compression method for DNS traffic logs based on field sorting according to the present invention.

[0019] Figure 3 is the schematic architecture diagram of the lossless compression device for DNS traffic logs based on field sorting according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The present invention will be described in detail below in conjunction with the specific embodiments shown in the accompanying drawings. However, these embodiments do not limit the present invention, and any structural, methodical, or functional transformations made by those of ordinary skill in the art based on these embodiments are included within the protection scope of the present invention.

[0021] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.

[0022] In the first embodiment of the present invention, the present invention provides a lossless compression method for DNS traffic logs based on field sorting, as Figure 1 shown, the method includes, S1: In response to a DNS traffic log compression request, read data in segments by a preset number of lines from the DNS traffic log, and extract log information line by line; S2: Perform replacement encoding by looking up the numbers of domain names and DNS addresses, and save the processed log lines to memory; S3: Sort each log segment based on multiple fields according to a predetermined rule, convert the sorted log line numbers to base 36 and group them for writing, and at the same time convert the time and date to a timestamp or a difference and append it to the index to form a processed DNS traffic log; S4: Apply redundancy elimination type compression to the processed DNS traffic log to generate a compressed file.

[0023] In a specific embodiment of the present invention, the request times of the requested domain names and the request times of the DNS addresses in the historical DNS traffic logs are obtained, and a domain name encoding table and a DNS address encoding table are established respectively. Specifically, The domain name encoding table assigns unique codes to common domain names in descending order of the occurrence frequencies of the requested domain names in the historical DNS traffic log, and adds identifiers to distinguish the codes from the unencoded domain names; The DNS address encoding table sorts the occurrence frequencies of all DNS addresses in the historical DNS traffic log from high to low, assigns a unique code starting from 0 to each DNS address, and replaces the original DNS address with this code.

[0024] Performing replacement encoding by looking up the numbers of domain names and DNS addresses, specifically, Read the log segment of the preset number of lines in the DNS traffic log compression request; Extract the domain name, date and time, record type, and DNS address from the log line by line, look up the domain name number and DNS address number in the domain name coding table and the DNS address coding table for replacement, and save it to the memory after the replacement is completed; If the corresponding domain name number and DNS address code can be found, replace the domain name with the identifier and the number, and replace the DNS address with the number; if not, keep the original domain name or DNS address unchanged.

[0025] In a specific implementation scenario of the present invention, sample the logs in the system to count the number of requests for the domain names, obtain the top n requested domain names ranked, and encode them as 0, 1, 2, 3,... in sequence. Then replace the top n domain names appearing in the logs with the encoding and add the identifier.

[0026] It should be noted that the value of n in the top n domain names needs to be determined according to the current scenario. If n is too large, more storage space is required to represent the encoding. If n is too small, some requested domain names will still occupy more byte space because the original strings are retained. In a specific implementation scenario of the present invention, the total number of resolutions of the top 100,000 domain names selected generally reaches more than 95% of the total number of resolutions of all domain names in the DNS traffic logs of most systems, and the encoding occupies at most 6 bytes, which can save a large amount of storage space.

[0027] Further, part of the code for generating the domain name field encoding is as follows: if domain in domain_top_num.keys(): domain_string = “:” + str(domain_top_num[domain]) else: domain_string = domain In a specific implementation scenario of the present invention, sample the logs in the system to count all the DNS addresses, then sort them according to the number of requests, and encode them as 0, 1, 2, 3,... in sequence. Then replace the DNS addresses appearing in the logs with the encoding.

[0028] It should be noted that in most cases, the number of DNS addresses is within 100,000, and the encoding occupies at most 5 bytes. Using the original strategy of string to record DNS addresses, it occupies about 10 bytes for IPv4 and about a dozen to two dozen bytes for IPv6. Since the DNS addresses in the requests are highly repetitive, a large amount of space can be saved after replacement. In a specific implementation scenario of the present invention, the total parsing times of the DNS address encoding in the logs of the same system can reach more than 99% of the total parsing times of all DNS addresses in the logs, and the number generally only needs at most 5 bytes, which can save a large amount of storage space.

[0029] Furthermore, the code for generating the encoding part of the DNS address field is as follows: if dns_ip in dns_ip_num: dns_string = str(dns_ip_num[dns_ip]) else: dns_string = dns_ip In a specific implementation manner of the present invention, each log segment is sorted based on multiple fields according to a predetermined rule. Specifically, The DNS traffic logs with a preset number of lines are divided into multiple log segments. After encoding and replacing each log segment, they are sorted in the order of the reversed domain name string, request type, DNS address, and the corresponding log segment line number in sequence.

[0030] In a specific implementation scenario of the present invention, as Figure 2 shown, first, the DNS traffic logs are divided into multiple log segments with a maximum of 30,000 lines. Each line of the log segment is sorted in sequence according to the three fields of the reversed domain name string, record type, and DNS address; then, the line numbers in the log segment before sorting are converted into 36 - bit characters as indexes according to the corresponding line numbers after sorting, and are divided into multiple lines with 100 characters per line and inserted into the front part of the segmented log; finally, a special string is added on the next line of the index to mark the end of the index.

[0031] For example, a log segment after splitting a log with a maximum of 30,000 lines. Assume that this is the last segment with only 10 lines in the log. The domain name, request type, and DNS address in the first line are respectively www.baidu.com, 1, 221.6.4.66; the domain name, request type, and DNS address in the second line are respectively www.baidu.com, 1, 221.6.4.67; the domain name, request type, and DNS address in the third line are respectively www.baidu.com, 5, 221.6.4.66; the domain name, request type, and DNS address in the fourth line are respectively map.baidu.com, 1, 221.6.4.67; the domain name, request type, and DNS address in the fifth line are respectively map.baidu.com, 5, 221.6.4.67; the domain name, request type, and DNS address in the sixth line are respectively www.sina.com, 5, 221.6.4.67; the domain name, request type, and DNS address in the seventh line are respectively www.sina.com, 1, 221.6.4.66; the domain name, request type, and DNS address in the eighth line are respectively www.sina.com, 1, 221.6.4.66; the domain name, request type, and DNS address in the ninth line are respectively mail.sina.com, 5, 221.6.4.66; the domain name, request type, and DNS address in the tenth line are respectively mail.sina.com, 1, 221.6.4.67.

[0032] It should be noted that the sorting is first in reverse order according to the domain name string (for example, www.baidu.com in reverse order is moc.udiab.www), then sorted according to the record type, then sorted according to the DNS address, and finally sorted according to the original line number of the log segment.

[0033] It should be noted that for the sorted log segments, the domain name, request type, and DNS address in the first line of this part are mail.sina.com 1 221.6.4.67, corresponding to the 10th line of the original log segment; the domain name, request type, and DNS address in the second line are mail.sina.com 5 221.6.4.66, corresponding to the 9th line of the original log segment; the domain name, request type, and DNS address in the third line are www.sina.com 1 221.6.4.66, corresponding to the 7th line of the original log segment; the domain name, request type, and DNS address in the fourth line are www.sina.com 1 221.6.4.66, corresponding to the 8th line of the original log segment; the domain name, request type, and DNS address in the fifth line are www.sina.com 5 221.6.4.67, corresponding to the 6th line of the original log segment; the domain name, request type, and DNS address in the sixth line are map.baidu.com 1 221.6.4.67, corresponding to the 4th line of the original log segment; the domain name, request type, and DNS address in the seventh line are map.baidu.com 5 221.6.4.67, corresponding to the 5th line of the original log segment; the domain name, request type, and DNS address in the eighth line are www.baidu.com 1 221.6.4.66, corresponding to the 1st line of the original log segment; the domain name, request type, and DNS address in the ninth line are www.baidu.com 1 221.6.4.67, corresponding to the 2nd line of the original log segment; the domain name, request type, and DNS address in the tenth line are www.baidu.com 5 221.6.4.66, corresponding to the 3rd line of the original log segment.

[0034] It should be noted that the generated index is preset to 100 per line, and the content is: a 9 7 8 6 4 5 1 2 3 ←END_OF_INDEX→ Insert it in front of the newly generated log segment.

[0035] Furthermore, the code for generating the cache sorting based on fields is as follows: idx1=''.join(reversed(domain)) idx2=query_type idx3=dns_ip idx4=part_line_idx idx_tuple=(idx1,idx2,idx3,idx4) cache_line_dict[idx_tuple]=new_line In a specific embodiment of the present invention, the sorted log line numbers are converted to base-36 and grouped for writing, and the time and date are converted to a timestamp or a difference value and appended to the index. Specifically, Convert the line numbers of the sorted log segments corresponding to the original log segments into base-36 characters, and at the same time write the flag indicating the end of the index into the corresponding lines; Convert all the time and dates of each sorted log segment into timestamps. Specifically, it includes Determine whether the current line is the first line other than the index in the processed DNS traffic log; If so, replace the time and date with the timestamp and write the sorted log after the index; If not, calculate the difference between the timestamp of this line and the timestamp of the first line, use the difference to replace the time and date, and write the sorted log after the index.

[0036] It should be noted that the string of time and date is converted into a timestamp. The timestamp of the first line of the compressed file is used as the standard timestamp, and the timestamp of each subsequent line is calculated for the difference from the standard timestamp.

[0037] For example: The time and date string of the first line of the log is 20241021094909, that is, 09:49:09 on October 21, 2024, then it is converted into the timestamp 1729475349 and replaces the time and date string of the first line. The time and date string of the second line of the log is 20241021094910, that is, 09:49:10 on October 21, 2024, which is converted into the timestamp 1729475350, subtract the timestamp 1729475349 of the first line, and get the difference 1, which replaces the time and date string of the second line... The time and date string of the nth line of the log is Sn, subtract the timestamp 1729475349 of the first line, and get the difference Dn, which replaces the time and date string of the nth line.

[0038] Furthermore, the code for generating the new time field is as follows: ts = time.strptime(daytime, "%Y%m%d%H%M%S") line_count[write_file_id] += 1 if line_count[write_file_id] == 1: first_ts = ts timestring = str(ts) else: delta_ts = ts - first_ts timestring = str(delta_ts) In a specific embodiment of the present invention, redundancy elimination is performed on the sorted log segments. Specifically, After the timestamp conversion is completed, continue to read the next preset number of lines of log segments until the DNS traffic log reading is completed; if the number of DNS traffic logs read at one time is less than the preset number of lines, it is determined that the reading is completed; When the DNS traffic log processing is completed, apply redundancy elimination type compression to the processed DNS traffic log to generate a compressed file.

[0039] It should be noted that the redundancy elimination type compression adopted by the present invention is gzip compression.

[0040] It should be noted that the request types of similar domain names and the DNS traffic log content of DNS addresses have a high similarity. After reordering, the concentration of similar strings is greatly improved. Therefore, when the present invention uses gzip compression, the string redundancy within a specific length is higher, and the compression effect is better. According to the number of load files, the gzip compression effect can be increased by up to about 70%.

[0041] In a specific embodiment of the present invention, during decompression, use the index to restore the original log order and reconstruct the original field content of the DNS traffic log through reverse operations. Specifically, When decompressing the compressed DNS traffic log file, convert the 36 - base line number in the index back to decimal to restore the line number order in the original log; Add the timestamp difference back to the timestamp of the first line to restore the original time - date string, and replace the encoding with the original domain name and DNS address according to the domain name coding table and DNS address coding table to ensure the accurate reconstruction of the original content of the DNS traffic log after decompression.

[0042] It should be noted that, first, when decompressing the compressed DNS traffic log file, the index data contained therein is read and parsed. According to the 36 - based line numbers recorded in the index, they are converted into corresponding decimal values one by one, so as to accurately restore the order of each log line in the original log file. Subsequently, for the restoration of the time field, the first timestamp in the generated file except the index is obtained as the reference value, and the timestamp differences stored in the subsequent lines are successively added to this reference timestamp to recalculate the complete timestamp of each line, and these timestamps are converted back to the original time - date string format to ensure the accurate restoration of time information. At the same time, for the reconstruction of the encoding field, based on the pre - established domain name encoding table and DNS address encoding table, the domain name numbers and DNS address numbers stored in each line of the log are respectively searched, and they are inversely mapped to the corresponding original domain name strings and DNS address values. During this process, if the original domain name or DNS address that has not been encoded is encountered, its original content is directly retained. In addition, after the above operations are completed, it is further verified whether the reconstructed log lines are consistent with the original field content, including checking whether the unencoded fields such as the record type remain complete, ensuring that the entire decompression process is lossless and accurate. Finally, through the above reverse operations, a DNS traffic log file that is exactly the same as the one before compression is generated, thus realizing the complete restoration from the compressed file to the original data.

[0043] In the second embodiment of the present invention, the present invention provides a DNS traffic log lossless compression device based on field sorting, as Figure 2 shown, including a reading module 1, a replacement module 2, a conversion module 3, and a compression module 4; The reading module 1 is used to respond to a DNS traffic log compression request, read data in segments from the DNS traffic log according to a preset number of lines, and extract the domain name, time - date, record type, and DNS address line by line; The replacement module 2 is used to perform replacement encoding by looking up the numbers of the domain name and DNS address, and save the processed log lines to the memory; The conversion module 3 is used to sort each log segment according to the reversed order of the domain name, record type, DNS address, and line number, convert the sorted log line numbers into 36 - based and write them in groups, and at the same time convert the time - date into a timestamp or a difference and append it to the index to form a processed DNS traffic log; The compression module 4 is used to perform redundancy - elimination - type compression on the processed DNS traffic log to generate a compressed file.

[0044] In the third embodiment of the present invention, the present invention provides an electronic device, including a memory and a processor, characterized in that a computer program that can run on the processor is stored in the memory, and the steps in the above-mentioned DNS traffic log lossless compression method based on field sorting are implemented when the program is executed on the processor.

[0045] In the fourth embodiment of the present invention, the present invention provides a storage medium, which stores a computer program, characterized in that the steps in the above-mentioned DNS traffic log lossless compression method based on field sorting are implemented when the computer program is executed by a processor.

[0046] In summary, the present invention provides a DNS traffic log lossless compression method and device based on field sorting. By segment reading, field encoding, and multi-field sorting, data redundancy is effectively reduced. Among them, the number replacement of domain names and DNS addresses reduces the occupation of repeated strings, the timestamp difference encoding compresses the time field space, and the 36-base line number conversion optimizes the storage efficiency. Combining the reverse sorting of domain names and gzip compression further improves the compression ratio and ensures compact storage of data under low entropy characteristics. This method significantly reduces the file size while maintaining the integrity of the log, facilitating efficient storage and transmission, and is particularly suitable for the management and analysis of large-scale DNS traffic logs.

[0047] It should be understood that although this specification is described according to embodiments, not each embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0048] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described modules can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0049] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0050] In addition, the functional modules in each embodiment of the present application can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0051] The integrated module implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium and include several instructions for causing a computer system (which may be a personal computer, a server, or a network system, etc.) or a processor to execute some steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0052] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of the present application.

Claims

1. A method for lossless compression of DNS traffic logs based on field sorting, characterized by: include, In response to the DNS traffic log compression request, data is read from the DNS traffic log in segments according to a preset number of lines, and log information is extracted line by line; By looking up the domain name and DNS address number, the replaced code is saved in the memory. Each log segment is sorted according to the predetermined rules based on multiple fields, the sorted log line numbers are converted to hexadecimal and written in groups, and the time date is converted to a timestamp or difference and appended to the index to form a processed DNS traffic log; The processed DNS traffic log is compressed using redundancy elimination to generate a compressed file.

2. The method for lossless compression of DNS traffic logs based on field sorting according to claim 1 is characterized in that: Also includes, Obtain the number of domain name requests and DNS address requests in the historical DNS traffic log, and establish a domain name encoding table and a DNS address encoding table respectively, including: The domain name encoding table counts the occurrence frequencies of requested domain names in the historical DNS traffic logs, sorts them from high to low, and assigns unique codes to commonly used domain names, and adds identifiers to distinguish between encoded and unencoded domain names; The DNS address encoding table counts the occurrence frequencies of all DNS addresses in the historical DNS traffic log, sorts them from high to low, assigns a unique code starting from 0 to each DNS address, and replaces the original DNS address with the code.

3. The method for lossless compression of DNS traffic logs based on field sorting according to claim 1, characterized in that: The replacement encoding by searching the domain name and the DNS address number includes: Reading a log segment of a preset number of lines in the DNS traffic log compression request; Extract the domain name, time date, record type and DNS address in the log line by line, find the domain name number and DNS address number from the domain name encoding table and the DNS address encoding table to replace them, and save them to the memory after the replacement is completed; If the corresponding domain name number and DNS address code can be found, the identifier and number are used to replace the domain name, and the number is used to replace the DNS address; if not found, the original domain name or DNS address remains unchanged.

4. The method for lossless compression of DNS traffic logs based on field sorting according to claim 3 is characterized in that: The "sorting each log segment based on multiple fields according to a predetermined rule" includes: Divide the DNS traffic log with a preset number of lines into multiple log segments. After encoding and replacing each log segment, sort them in the reversed order of the domain name string, request type, DNS address, and corresponding log segment line number.

5. The method for lossless compression of DNS traffic logs based on field sorting according to claim 4 is characterized in that: The "converting the sorted log line numbers into hexadecimal and writing them in groups, converting the time date into a timestamp or a difference and appending them to the index" includes: Convert the line number of the sorted log segment corresponding to the original log segment into hexadecimal characters, and write the end of index mark into the corresponding line; Convert all the time dates of each sorted log segment into timestamps, including: Determine whether the current line is the first line in the processed DNS traffic log excluding the index; If yes, replace the time date with the timestamp and write the sorted log after the index; If not, calculate the difference between the timestamp of this row and the timestamp of the first row, use the difference to replace the time date, and write the sorted log after the index.

6. The method for lossless compression of DNS traffic logs based on field sorting according to claim 5 is characterized in that: The redundancy elimination of the sorted log segments includes: After the timestamp conversion is completed, continue to read the next preset number of log segments until the reading of the DNS traffic log is completed; if the number of DNS traffic logs read at one time is less than the preset number of lines, it is determined that the reading is completed; When the DNS traffic log processing is completed, redundancy elimination type compression is applied to the processed DNS traffic log to generate a compressed file.

7. The method for lossless compression of DNS traffic logs based on field sorting according to claim 1, characterized in that: Also includes, During decompression, the original log order is restored using the index, and the original field content of the DNS traffic log is reconstructed through reverse operations. include, When decompressing the compressed DNS traffic log file, converting the hexadecimal line number in the index back to decimal to restore the line number sequence in the original log; The timestamp difference is added back to the timestamp of the first line to restore the original time and date string, and the encoding is replaced with the original domain name and DNS address according to the domain name encoding table and the DNS address encoding table to ensure that the original content of the DNS traffic log is accurately reconstructed after decompression.

8. A DNS traffic log lossless compression device based on field sorting, characterized by: It includes a reading module, a replacing module, a converting module and a compressing module; The reading module is used to respond to the DNS traffic log compression request, read data from the DNS traffic log in segments according to a preset number of lines, and extract the domain name, time and date, record type and DNS address line by line; The replacement module is used to replace the code by looking up the domain name and the DNS address number, and save the processed log line to the memory; The conversion module is used to sort each log segment according to the reversed order of the domain name, record type, DNS address and line number, convert the sorted log line number into hexadecimal and write it in groups, and convert the time date into a timestamp or difference and append it to the index to form a processed DNS traffic log; The compression module is used to perform redundancy elimination compression on the processed DNS traffic log to generate a compressed file.

9. An electronic device, comprising a memory and a processor, characterized in that: The memory stores a computer program that can be run on the processor, and when the program is executed on the processor, the steps in the field sorting-based DNS traffic log lossless compression method as described in any one of claims 1 to 7 are implemented.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps in the method for lossless compression of DNS traffic logs based on field sorting as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Log compression method and device, equipment, medium and product

    CN121530947A