Http traffic data deduplication method, device and storage medium
Patent Information
- Application Number
- CN202611224792.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-13
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]本申请的主要目的在于提供一种HTTP流量数据去重方法、设备和存储介质,旨在解决去重粒度较粗导致的HTTP流量数据内存占用大的技术问题
[0014]本申请提供了一种HTTP流量数据去重方法,通过解析HTTP流量数据得到请求体和响应体;使用杂凑函数分别计算请求体的杂凑值以及响应体的杂凑值;在布隆过滤器中分别查找请求体的杂凑值、响应体的杂凑值;若在布隆过滤器中查找到杂凑值,则在正文哈希表中查找杂凑值;若在正文哈希表中查找到杂凑值,则从正文哈希表中取出杂凑值对应的请求索引,用请求索引替代HTTP流量数据中的请求体或响应体,得到目标HTTP流量数据。通过布隆过滤器和正文哈希表的定位,将重复的请求体或响应体替换为轻量的请求索引,实现了对海量HTTP流量数据的高效去重存储,大幅减少了HTTP流量数据的内存占用。
Smart Images

Figure CN122845657A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, device and storage medium for deduplicating HTTP traffic data. Background Technology
[0002] Currently, in applications such as network security, performance monitoring, and user behavior analysis, systems need to continuously capture and analyze large amounts of HTTP traffic data, which contains a significant amount of duplicate content. Related technologies deduplicate individual requests, treating the entire HTTP request as a whole. However, due to this coarse-grained dedupplication, the deduplication process still results in the HTTP traffic data consuming a large amount of memory. Summary of the Invention
[0003] The main purpose of this application is to provide a method, device and storage medium for deduplicating HTTP traffic data, which aims to solve the technical problem of large memory consumption of HTTP traffic data caused by coarse deduplication granularity.
[0004] To achieve the above objectives, this application provides a method for deduplicating HTTP traffic data, which includes: Parse HTTP traffic data to obtain the request body and response body; Use hash functions to calculate the hash value of the request body and the hash value of the response body respectively; Search for the hash values of the request body and the response body in the Bloom filter respectively; If a hash value is found in the Bloom filter, then the hash value is searched in the text hash table. If a hash value is found in the text hash table, the request index corresponding to the hash value is retrieved from the text hash table, and the request index is used to replace the request body or response body in the HTTP traffic data to obtain the target HTTP traffic data.
[0005] In one embodiment, the hash values of the request body and the response body are calculated using hash functions, including: Retrieve the raw byte data of the request body or the raw byte data of the response body; The original byte data is padded to obtain the padded data. The original data length information is appended to the end of the padded data to obtain the preprocessed data. The length of the preprocessed data is an integer multiple of the preset block size. The preprocessed data is divided into multiple data blocks of a preset block size; For each data block, perform multiple rounds of iterative compression and store the intermediate state in a register. In each round of iteration, update the intermediate state of the register through logical operations, circular shifts, and modulo addition operations. After all data blocks have been iteratively compressed, the final state of each register is obtained. The final states of each register are then concatenated to obtain the hash value.
[0006] In one embodiment, before searching for the hash value of the request body and the hash value of the response body in the Bloom filter, the method further includes: The required bit array length is calculated based on the total number of request and response bodies and the preset false positive rate; Calculate the number of hash functions needed based on the length of the bit array and the total number of bits; Construct a bit array of length equal to the length of the bit array, and construct the hash functions corresponding to the number of hash functions.
[0007] In one embodiment, after searching for the hash values of the request body and the response body in the Bloom filter, the method further includes: If the hash value is not found in the Bloom filter, add the hash value to the Bloom filter and store the HTTP traffic data in a memory block.
[0008] In one embodiment, after looking up the hash value in the text hash table, the method further includes: If the hash value cannot be found in the text hash table, the hash value is stored in the text hash table as the key and the corresponding request index is stored as the value, and the HTTP traffic data is stored in the memory block.
[0009] In one embodiment, after obtaining the target HTTP traffic data, the method further includes: The target HTTP traffic data is stored in a memory block. When the size of the data in the memory block exceeds a preset threshold, the data in the memory block is compressed and the compressed data block is written to the disk.
[0010] In one embodiment, the HTTP traffic data deduplication method further includes: Assign a session identifier to each HTTP traffic data request-response pair; Create a cross-domain hash table. The key of the cross-domain hash table is the hash value of the request body or the hash value of the response body, and the value of the cross-domain hash table is a list containing the session identifier, type tag, and request index. When a hash value is found in the Bloom filter, the hash value is then looked up in the cross-domain hash table. If a hash value is found in the cross-domain hash table and the type tag is different from the current type, the corresponding request body or response body is replaced with the retrieved request index to obtain the target HTTP traffic data.
[0011] In one embodiment, after looking up the hash value in the cross-domain hash table, the method further includes: If a hash value is found in the cross-domain hash table but the type tag is the same as the current type, or if no hash value is found, then the hash value is searched in the text hash table. After deduplication of the request or response body is completed, the hash value of the request or response body, the session identifier, the type tag, and the request index are stored in a cross-domain hash table.
[0012] In addition, to achieve the above objectives, this application also provides an HTTP traffic data deduplication device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the above-described HTTP traffic data deduplication method.
[0013] In addition, to achieve the above objectives, this application also provides a storage medium, which is a computer-readable storage medium, storing a program that implements the HTTP traffic data deduplication method. The program that implements the HTTP traffic data deduplication method is executed by a processor to implement the steps of the above-mentioned HTTP traffic data deduplication method.
[0014] This application provides a method for deduplicating HTTP traffic data. It involves parsing HTTP traffic data to obtain the request body and response body; using hash functions to calculate the hash values of the request body and response body respectively; searching for the hash values of the request body and response body in a Bloom filter; if a hash value is found in the Bloom filter, searching for the hash value in the text hash table; if a hash value is found in the text hash table, retrieving the corresponding request index from the text hash table and replacing the request body or response body in the HTTP traffic data with the request index to obtain the target HTTP traffic data. By using the Bloom filter and text hash table to locate duplicate request bodies or response bodies and replacing them with lightweight request indexes, this method achieves efficient deduplication and storage of massive amounts of HTTP traffic data, significantly reducing the memory footprint of HTTP traffic data. Attached Figure Description
[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating an embodiment of the HTTP traffic data deduplication method of this application. Figure 2 A simplified flowchart illustrating an embodiment of the HTTP traffic data deduplication method of this application; Figure 3 This is a schematic diagram of the HTTP traffic data deduplication device in the embodiments of this application.
[0018] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0019] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not intended to limit this application.
[0020] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0021] Currently, in applications such as network security, performance monitoring, and user behavior analysis, systems need to continuously capture and analyze large amounts of HTTP traffic data, which contains a significant amount of duplicate content. Related technologies deduplicate individual requests, treating the entire HTTP request as a whole. However, due to this coarse-grained dedupplication, the deduplication process still results in the HTTP traffic data consuming a large amount of memory.
[0022] The main solution of this application is as follows: Obtain the request body and response body by parsing HTTP traffic data; calculate the hash value of the request body and the response body using hash functions; search for the hash value of the request body and the response body in a Bloom filter; if a hash value is found in the Bloom filter, search for the hash value in the text hash table; if a hash value is found in the text hash table, retrieve the request index corresponding to the hash value from the text hash table, and replace the request body or response body in the HTTP traffic data with the request index to obtain the target HTTP traffic data. By using the Bloom filter and text hash table to locate duplicate request bodies or response bodies, duplicate request bodies or response bodies are replaced with lightweight request indexes, achieving efficient deduplication and storage of massive HTTP traffic data, and significantly reducing the memory footprint of HTTP traffic data.
[0023] It should be noted that the execution subject in this embodiment can be an HTTP traffic data deduplication device, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an HTTP traffic data deduplication device capable of performing the above functions. This embodiment does not specifically limit it in this way. The following uses an HTTP traffic data deduplication device as the execution subject as an example to describe this embodiment and the following embodiments.
[0024] Based on this, Embodiment 1 of this application proposes a method for deduplicating HTTP traffic data. Please refer to... Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the HTTP traffic data deduplication method of this application. The HTTP traffic data deduplication method includes steps S10 to S50: Step S10: Parse the HTTP traffic data to obtain the request body and response body.
[0025] In this embodiment, HTTP (Hypertext Transfer Protocol) traffic data refers to the raw HTTP request and response messages obtained through network packet capture or proxies, which are the raw data sources to be processed and stored. The request header refers to the header fields in the HTTP request message, containing information such as the request method, URL (Uniform Resource Locator), Host, and User-Agent, used to describe the request's metadata. The request body refers to the actual data content carried in the HTTP request message, such as POST form data, JSON (JavaScript Object Notation) strings, uploaded files, etc., and is located after the request header. The response header refers to the header fields in the HTTP response message, containing information such as the status code, Content-Type, and Cache-Control, used to describe the response's metadata. The response body refers to the actual data content carried in the HTTP response message, such as HTML (Hypertext Markup Language) pages, JSON data, image binary streams, etc., and is located after the response header.
[0026] As an optional implementation, HTTP traffic data is parsed into request headers, request bodies, response headers, and response bodies according to the HTTP protocol standard.
[0027] Specifically, following the HTTP protocol standard, HTTP traffic data is scanned line by line using carriage return and line feed characters as line separators. First, the request or response line is read, then each header field is read sequentially until a blank line is encountered, completing the parsing of the request or response headers. All data after the blank line constitutes the raw byte content of the request or response body. By parsing header fields according to the HTTP protocol specification, various transport encodings and header information can be accurately identified, providing a structured data foundation for subsequent content-type-based differential compression or header field analysis.
[0028] As another optional implementation, HTTP traffic data can be quickly segmented based on a preset delimiter to obtain request headers, request bodies, response headers, and response bodies.
[0029] Specifically, the system directly locates the first blank line in the HTTP traffic data. The data before the blank line is treated as either a request or response header, and the data after the blank line is treated as either a request or response body. The distinction between request and response headers is made by checking if the data begins with an HTTP version number, such as "HTTP / 1.1". If it does, it's a response line; otherwise, it's a request line. By skipping fine-grained parsing of header fields and directly locating the blank line for rapid segmentation, the system significantly reduces parsing overhead.
[0030] Step S20: Use hash functions to calculate the hash value of the request body and the hash value of the response body respectively.
[0031] In this embodiment, a hash function refers to a one-way encryption function that maps data of arbitrary length to a fixed-length unique fingerprint, such as SHA1 (Secure Hash Algorithm 1), SHA-256 (Secure Hash Algorithm 256-bit), MD5 (Message Digest Algorithm 5), etc., used to verify data integrity and uniquely identify content. The hash value refers to the fixed-length binary output calculated by the secure hash algorithm, serving as a unique fingerprint of the request or response body content for subsequent duplicate detection.
[0032] As an optional implementation, the SHA1 secure hash algorithm is used to calculate the hash values of the request body and the response body, respectively.
[0033] Specifically, the original byte data of the request or response body is obtained and padded to ensure the padded data length satisfies L mod 512 = 448. Then, 64 bits of the original data length information are appended to obtain the preprocessed data. Next, the preprocessed data is divided into multiple 512-bit data blocks. For each data block, 80 rounds of iterative compression operations are performed. In each round, the states of the five registers are updated through logical operations, circular shifts, and modulo addition operations. After all data blocks have been processed, the final states of the five registers are concatenated into a 160-bit hash value. The SHA1 secure hash algorithm is fast and has a short hash value, making it suitable for quickly generating content fingerprints in high-throughput scenarios and reducing index storage overhead.
[0034] As an alternative implementation, the SHA-256 secure hash algorithm is used to calculate the hash values of the request body and the response body, respectively.
[0035] Specifically, the original byte data of the request or response body is obtained and padded to ensure that the padded data length satisfies L mod 512 = 448. Then, 64 bits of the original data length information are appended to obtain the preprocessed data. Next, the preprocessed data is divided into multiple 512-bit data blocks. For each data block, 64 rounds of iterative compression operations are performed. In each iteration, the state of eight 32-bit registers is updated through logical operations, circular right shift operations, and modulo addition operations. After all data blocks have been processed, the final states of the eight registers are concatenated into a 256-bit hash value. The SHA-256 secure hash algorithm has stronger collision resistance and higher security, making it suitable for scenarios with strict data integrity requirements, avoiding erroneous deduplication due to hash collisions.
[0036] Step S30: Search for the hash values of the request body and the response body in the Bloom filter.
[0037] In this embodiment, a Bloom filter refers to a space-efficient probabilistic data structure that uses a bit array and multiple hash functions to determine whether an element is likely to exist in the set, and is used to quickly screen whether a request body or response body has appeared before.
[0038] As one implementation, the hash values of the request body and the response body are queried in a Bloom filter.
[0039] Specifically, when processing current HTTP traffic data, the hash values of the request body and response body are queried separately using a Bloom filter. The request body hash value is used as input, and multiple bit array positions are calculated through multiple hash functions of the Bloom filter. It is then checked whether all these positions are 1; if all are 1, the hash value is found, indicating that the request body may have appeared; if any position is 0, the hash value is not found, and the request body definitely has not appeared. After querying the request body, the same steps are performed to query the response body hash value. These two queries are independent of each other, each yielding a result of finding or not finding the request body. By sharing a Bloom filter for the request and response bodies and querying them independently, the system maintains a low memory footprint while ensuring the independence of the duplicate detection for the two body types, providing a unified initial screening basis for subsequent deduplication.
[0040] Step S40: If a hash value is found in the Bloom filter, then search for the hash value in the text hash table.
[0041] In this embodiment, the text hash table refers to a precise key-value storage structure, with the hash value of the request body or response body as the key and the corresponding request index as the value, used to accurately locate duplicate content and perform deduplication and replacement after the Bloom filter hits.
[0042] As an alternative implementation, if a hash value is found in the Bloom filter, a lookup is performed in the body hash table using the hash value of the current request body or response body as the key.
[0043] Specifically, when the Bloom filter finds a hash value, it searches the text hash table using the hash value of the current request or response body as the key. The text hash table uses chaining to resolve hash collisions. Each bucket stores a linked list containing the hash value and its corresponding request index. The corresponding bucket is located by calculating the hash address of the hash value, and then the linked list within that bucket is traversed, comparing hash values one by one. By finding a hash value in the Bloom filter and searching the text hash table using that hash value as the key, hash collisions are effectively resolved using chaining. The hash address is calculated to quickly locate the bucket, and the linked list is traversed to compare hash values, thereby obtaining the corresponding request index to support subsequent deduplication and replacement operations.
[0044] As another alternative implementation, if a hash value is found in the Bloom filter, the hash value is queried in the hotspot cache; if the hash value is not found in the hotspot cache, the hash value is searched in the text hash table.
[0045] Specifically, a hotspot cache is introduced between the Bloom filter and the text hash table. This cache uses an LRU eviction policy and stores the hash values and their corresponding request indexes that have been successfully queried in the text hash table the N most recently. When the Bloom filter finds a hash value, it first queries the hotspot cache. If a match is found, the request index is retrieved directly without accessing the text hash table. If a match is not found, the text hash table is queried, and the successful result is written to the hotspot cache. By leveraging the principle of temporal locality, frequently repeated Body content is quickly retrieved, reducing the access pressure on the text hash table and improving the overall deduplication throughput.
[0046] Step S50: If a hash value is found in the text hash table, the request index corresponding to the hash value is retrieved from the text hash table, and the request index is used to replace the request body or response body in the HTTP traffic data to obtain the target HTTP traffic data.
[0047] In this embodiment, the request index is an identifier used to uniquely locate the data storage location of a certain HTTP traffic data. The corresponding request body or response body data can be found through the request index.
[0048] As one implementation, when a hash value is found in the text hash table, the request index in the text hash table is used to replace the request body or response body in the HTTP traffic data to obtain the target HTTP traffic data.
[0049] Specifically, when the hash value of the current request body or response body is found in the text hash table, the request index corresponding to that hash value is retrieved from the text hash table. This request index points to the location in memory of a previously stored body of the same type. In the current HTTP traffic data, the request body or response body is replaced with this request index to obtain the target HTTP traffic data. By replacing duplicate body content with a lightweight index, only the index is retained during storage, rather than the actual data, thereby reducing the storage space occupied by duplicate content.
[0050] This embodiment provides a method for deduplicating HTTP traffic data. First, the HTTP traffic data is parsed to obtain the request body and response body. Hash functions are then used to calculate the hash values of the request body and response body, respectively. The hash values of the request body and response body are then searched in a Bloom filter. If a hash value is found in the Bloom filter, it is then searched in a text hash table. If a hash value is found in the text hash table, the corresponding request index is retrieved from the text hash table, and this request index replaces the request body or response body in the HTTP traffic data to obtain the target HTTP traffic data. By using the Bloom filter and text hash table to locate duplicate request bodies or response bodies, duplicate request bodies or response bodies are replaced with lightweight request indexes, achieving efficient deduplication and storage of massive amounts of HTTP traffic data and significantly reducing the memory footprint of HTTP traffic data.
[0051] Based on Embodiment 1, in Embodiment 2 of this application, the content that is the same as or similar to that in Embodiment 1 can be referred to the above description, and will not be repeated hereafter. On this basis, hash functions are used to calculate the hash value of the request body and the hash value of the response body, including: Step S21: Obtain the raw byte data of the request body or the raw byte data of the response body.
[0052] Step S22: Padded data is performed on the original byte data to obtain padded data.
[0053] Step S23: Append the original data length information to the end of the padded data to obtain the preprocessed data. The length of the preprocessed data is an integer multiple of the preset block size.
[0054] In this embodiment, raw byte data refers to the binary content of the request body or response body directly parsed from HTTP traffic data without any processing or transformation. Specifically, the raw byte data of the request body refers to the raw binary data of the Body portion following the request header, extracted from the HTTP request message; the raw byte data of the response body refers to the raw binary data of the response body following the response header, extracted from the HTTP response message.
[0055] As one implementation method, the raw byte data of the request body or response body is extracted from the parsed HTTP traffic data as the object to be calculated. The raw byte data of the request body or response body is denoted as M={b1,b2,…,b n}, where b i Let represent the i-th byte of data in the Body, and n represent the total number of bytes in the Body. Since the SHA1 algorithm requires the input data length to be a multiple of 512 bits, and the length of the original byte data is uncertain, padding is necessary. First, a "1" bit is added to the end of the original byte data, then several "0" bits are added so that the length L of the padded data satisfies Lmod 512 = 448, which is 64 bits less than a multiple of 512. Then, 64 bits of the original data length information are appended to the end of the padded data, resulting in preprocessed data with a total length that is a multiple of 512 bits. Padding ensures the correctness and collision resistance of the algorithm, guaranteeing that original data of different lengths can be processed uniformly.
[0056] Step S24: Divide the preprocessed data into multiple data blocks of a preset block size.
[0057] Step S25: Perform multiple rounds of iterative compression on each data block and store the intermediate state in the register. In each round of iteration, update the intermediate state of the register through logical operations, circular shifts, and modulo addition operations.
[0058] Step S26: After all data blocks have been iteratively compressed, the final state of each register is obtained. The final states of each register are then concatenated to obtain the hash value.
[0059] In this embodiment, a data block refers to an independent processing unit formed by cutting preprocessed data into fixed-length segments. Each data block is sequentially fed into the iterative compression operation. A register refers to a set of fixed-length temporary storage units used to save and update intermediate calculation results during iteration. An intermediate state refers to the value stored in the register after each iteration, representing a partial hash result of the current data block at that stage, and serving as input for the next iteration. Logical operations refer to performing bitwise AND, OR, XOR, and NOT operations on the data in the register during each iteration, used for mixing and spreading the data. Circular shift refers to shifting the binary number in the register left or right by a specified number of bits, used to shuffle the data bits and enhance the spreading effect. Modulo addition refers to adding two values and then taking the modulo, used to accumulate register values within a limited range to prevent overflow.
[0060] In one implementation, the preprocessed data is divided into several 512-bit data blocks of a fixed size, resulting in multiple data blocks B={B1,B2,…,B...} of a preset block size. m}, where B1 refers to the first 512-bit data block, B2 refers to the second 512-bit data block, and B m This refers to the m-th 512-bit data block, where each data block contains 16 32-bit words. Each data block is independently fed into an iterative compression operation for processing. The SHA1 algorithm uses five 32-bit registers A, B, C, D, and E for iterative compression operations, where the registers are initialized to fixed preset constant values. For each 512-bit data block, 80 rounds of iterative compression operations are performed sequentially. In each iteration, the state of the five registers is updated through logical operations, circular shift operations, and modulo addition operations. After each data block is processed, the intermediate state value stored in the five registers is updated once and used as the initial state for processing the next data block. After all data blocks have completed 80 rounds of iterative compression operations, the final state value stored in the five registers is a 160-bit hash result. The five 32-bit register values are concatenated sequentially to obtain a hash value, with register A first and register E last, forming a 40-bit hexadecimal number, such as 2fd4e1c67a2d28fced849ee1bb76e7391b93eb12. Through multiple rounds of iterative computation, the data content and register states are mixed to generate highly dispersed intermediate results, ensuring that different inputs produce distinctly different hash values.
[0061] Specifically, each data block is 512 bits long, and a word is 32 bits long. This means a 512-bit data block can be divided into 16 32-bit words. In the 80 rounds of iterative compression in the SHA1 algorithm, the entire 512-bit data block is not used directly. Instead, 16 words corresponding to a database are used as seeds, and 80 32-bit words are generated through the message expansion algorithm. In each iteration, one word is used in a mixed operation with a register, and the hash value is obtained after all data blocks have been processed. By splitting the data block into 16 32-bit words and expanding them into 80 words for 80 iterations, the diffusion and obfuscation of the data are enhanced, ensuring that even a tiny change in input will result in a completely different hash value, thus improving the reliability of subsequent deduplication.
[0062] In addition, the same or similar functions can be achieved by using SHA-256, SHA-512, MD5, SM3 or other hash algorithms that can generate fixed-length content fingerprints.
[0063] In this embodiment, a hash value is calculated to provide a unique content identifier for the Bloom filter and the text hash table, supporting subsequent efficient HTTP traffic data deduplication and compression storage steps.
[0064] Based on any of the above embodiments of this application, Embodiment 3 of this application proposes a method for deduplicating HTTP traffic data, which can be referred to the above description and will not be repeated hereafter. In addition, before searching for the hash value of the request body and the hash value of the response body in the Bloom filter, the method further includes: Step S301: Calculate the required bit array length based on the total number of request and response bodies and the preset false positive rate.
[0065] In this embodiment, the Bloom filter has the characteristic of no false negatives, meaning it will not falsely classify existing elements as non-existent. However, there is a possibility of false positives, meaning it may falsely classify non-existent elements as present. Its false positive rate can be controlled by the bit array length *m* and the number of hash functions *k*. The preset false positive rate refers to the maximum false positive probability allowed by the Bloom filter, i.e., the upper limit of the probability of falsely classifying a non-existent element as present. The total number of request and response bodies refers to the sum of all request and response bodies expected to be inserted into the Bloom filter. The bit array refers to the fixed-length binary bit vector used by the underlying Bloom filter, with each bit initially set to 0, used to mark the existence status of elements through hash mapping.
[0066] As one implementation method, obtain the preset false positive rate p and the total number of request and response bodies n, and substitute them into the Bloom filter bit array length calculation formula m = -(n The required bit array length m is calculated by using ln(p) / (ln(2))². For example, when processing 100 million HTTP responses, i.e., the number of response bodies is 10... 8 n=10 8 The desired false positive rate is below 0.1%, i.e., p = 0.001, which yields m ≈ 1.44. 10 8 A bit array of approximately log2(1 / 0.001) ≈ 1.92 GB is required. By calculating the necessary bit array length, the bit array length m is minimized given the total number of request and response bodies n and the preset false positive rate p. The optimal bit array length is calculated based on the preset false positive rate and the total amount of data, minimizing memory usage while meeting the false positive rate requirement.
[0067] Step S302: Calculate the number of hash functions required based on the length of the bit array and the total number of bits.
[0068] In this embodiment, a hash function is a function that maps the hash value of input data, such as the request body or response body, to one or more position indices in a bit array, and is used to mark or query the existence of elements in a Bloom filter.
[0069] As one implementation method, after obtaining the bit array length m and the total number n of request and response bodies, substitute them into the formula for the optimal number of hash functions k=(m / n). ln(2) is used to calculate the theoretically lowest number of hash functions, k. For example, when n=10 8 m≈1.44 10 8 log2(1 / 0.001)≈1.44 10 9 At that time, the calculated number of hash functions, k ≈ 13. Too large a number of hash functions, k, will cause the bit array to fill up quickly, increasing the false positive rate; too small a number of hash functions will result in sparse hash positions, also leading to a high false positive rate. The goal is to calculate the optimal number of hash functions to make the actual false positive rate as close as possible to the preset false positive rate. This is achieved by calculating the optimal number of hash functions based on the bit array length and the total amount of data, thus bringing the actual false positive rate closer to the preset value.
[0070] Step S303: Construct a bit array with a length equal to the length of the bit array, and construct hash functions corresponding to the number of hash functions.
[0071] As one implementation, based on the calculated bit array length *m*, a contiguous bit array of length *m* is allocated in memory, and all bits are initialized to 0. When processing a single HTTP traffic data item, if the hash value of the request or response body is not found in the Bloom filter, an addition operation is performed. This involves calculating the hash value using *k* hash functions, obtaining *k* bit array position indices, and setting all *k* positions to 1. When processing subsequent HTTP traffic data, for the hash value of the current request or response body, a query operation is performed. This involves calculating the hash value using the same *k* hash functions, obtaining *k* position indices, and checking the values at these *k* positions in the bit array. If all are 1, it is determined that the Body may have appeared, requiring further querying of the Body Hash Table for confirmation. If any position is 0, it is determined that the Body has definitely not appeared, and its hash value is directly added to the Bloom filter. By constructing a specified bit array and selecting a corresponding number of hash functions, the Bloom filter is initialized, providing a data structure foundation for efficient element addition and querying.
[0072] In this embodiment, the optimal bit array length and the number of hash functions are calculated based on the total number of request and response bodies and the preset false positive rate. This constructs a space-efficient Bloom filter, which can significantly reduce memory overhead under a controllable false positive rate. For example, it can provide deduplication services for 100 million elements with less than 2GB of memory and a false positive rate of less than one in a thousand, thus saving memory and supporting efficient deduplication of massive HTTP traffic data.
[0073] Based on any of the above embodiments of this application, Embodiment 4 of this application proposes an HTTP traffic data deduplication method, which can be referred to the above description and will not be repeated hereafter. In addition, after searching for the hash values of the request body and the response body in the Bloom filter, the method further includes: Step S304: If no hash value is found in the Bloom filter, add the hash value to the Bloom filter and store the HTTP traffic data in the memory block.
[0074] As one implementation, when a query for the hash value of the current request or response body in the Bloom filter fails to find it, it indicates that the Body content has never appeared before. In this case, k hash functions of the Bloom filter are used to calculate the k position indices of the hash value, and all values at these k positions in the bit array are set to 1, completing the recording of the hash value in the Bloom filter so that subsequent identical Body content can be identified. No deduplication operation is performed on the current request or response body; its original data, along with the corresponding request headers, response headers, and other information, is stored in a memory block for subsequent compression. By adding the corresponding hash value to the Bloom filter, it ensures that the first appearance of the Body content is completely preserved, laying the foundation for the identification of subsequent duplicate data.
[0075] In this embodiment, by adding the hash value to the Bloom filter and storing the complete HTTP traffic data in the memory block when the query is not found, the first appearance of Body data is ensured to be completely preserved, and a comparison benchmark is provided for the repeated identification of the same content in the future, while avoiding data loss due to misjudgment.
[0076] Based on any of the above embodiments of this application, Embodiment 5 of this application proposes a method for deduplicating HTTP traffic data, which can be referred to the above description and will not be repeated hereafter. In addition, after searching for the hash value in the text hash table, the method further includes: Step S401: If the hash value cannot be found in the text hash table, store the hash value as the key and the corresponding request index as the value in the text hash table, and store the HTTP traffic data in the memory block.
[0077] As one implementation, when a Bloom filter query hits a hash value that is not yet recorded in the text hash table (i.e., the current request or response body appears for the second time), the hash value of the current request or response body is used as the key, and the request index of the current HTTP traffic data, such as the request sequence number, is used as the value. This key-value pair is then stored in the text hash table, establishing a mapping between the Body content and its first appearance. The current HTTP traffic data is stored in a memory block, and no deduplication or replacement operation is performed on the Body. By storing the hash value as the key and the current request index as the value in the text hash table, a precise index location is provided for subsequent appearances of the same Body, ensuring the accuracy and traceability of the deduplication operation.
[0078] In this embodiment, by marking and indexing the Body content that appears for the second time, the transition from probabilistic marking by the Bloom filter to precise indexing by the text hash table is completed, providing a directly referable location basis for subsequent repetitions.
[0079] Based on any of the above embodiments of this application, Embodiment Six of this application proposes a method for deduplicating HTTP traffic data, which can be referred to the above description and will not be repeated hereafter. In addition, after obtaining the target HTTP traffic data, the method further includes: Step S501: Store the target HTTP traffic data into a memory block. When the data size in the memory block exceeds a preset threshold, compress the data in the memory block and write the compressed data block to the disk.
[0080] As one implementation method, the total amount of HTTP traffic data temporarily stored in the current memory block is continuously monitored. When the cumulative data size exceeds a preset threshold, such as 8MB or 16MB, the reception of new HTTP data is paused and compression is performed. Specifically, a uniform compression algorithm is applied to all deduplicated HTTP data in the entire memory block, such as gzip (GNU Gzip), zstd (Zstandard), snappy, or lz4 (LZ4). Since different HTTP traffic data often contain a large number of duplicate strings in their request and response headers, compressing the data block can achieve a high compression ratio. After compression, the generated compressed data block is written to disk as a whole, and the memory block is cleared after successful writing to free up memory space for continued caching and processing of subsequent HTTP data. By further compressing the data volume after deduplication, disk usage can be significantly reduced.
[0081] In this embodiment, a high compression ratio is achieved by compressing the memory block as a whole and taking advantage of the large number of repeated strings in the header information after deduplication. This significantly reduces the number of disk writes and I / O overhead, and further saves disk storage space.
[0082] Based on any of the above embodiments of this application, Embodiment Seven of this application proposes a method for deduplicating HTTP traffic data, which can be referred to the above description and will not be repeated hereafter. In addition, the HTTP traffic data deduplication method further includes: Step S60: Assign a session identifier to each HTTP traffic data request-response pair.
[0083] In this embodiment, a request-response pair refers to a logical unit consisting of paired request and response data that appears in a complete HTTP interaction, including the request body and response body from the same HTTP traffic data. A session identifier is a unique identifier assigned to each request-response pair, used to associate and locate the corresponding request and response bodies within the session to which the pair belongs.
[0084] As one implementation method, when processing each HTTP traffic data piece, a globally unique session identifier is assigned to the corresponding request and response body pair. This session identifier can take the form of an incrementing sequence number or a distributed ID generated based on a timestamp and node ID, and is used to uniquely locate the HTTP traffic data piece in the subsequent cross-domain hash table and index structure. By assigning a unique session identifier to each HTTP request-response pair, a cross-request and response association anchor is established, providing an indexing foundation for subsequently locating the original Body data and implementing cross-type references.
[0085] Step S70: Establish a cross-domain hash table. The key of the cross-domain hash table is the hash value of the request body or the hash value of the response body, and the value of the cross-domain hash table is a list containing the session identifier, type tag, and request index.
[0086] In this embodiment, the cross-domain hash table refers to a unified hash table that spans the boundaries between request and response body types. It uses the hash value of the request or response body as the key and a list containing session identifiers, type flags, and request indexes as values. The type flag is a type attribute flag used to identify whether the hash value originates from the request or response body. During a cross-domain hash table query, it is used to determine whether the current Body and the historical Body have different types, thereby deciding whether to perform cross-domain substitution.
[0087] As one implementation, a cross-domain hash table is initialized, where the key of the hash table is the hash value of the request or response body, and the value is a list of structures containing the session identifier and type tag to which the hash value first appears. By establishing a cross-domain hash table, the type barrier between the request and response bodies is broken down, and a unified index structure supporting cross-matching is constructed.
[0088] Step S80: When a hash value is found in the Bloom filter, the hash value is searched in the cross-domain hash table.
[0089] As one implementation method, when processing the current request body or response body, its hash value is first queried in its corresponding Bloom filter. If the hash value is found in the Bloom filter, a lookup is performed in the cross-domain hash table using the hash value of the current Body as the key. By prioritizing the query in the cross-domain hash table after the initial Bloom filter hits, it is possible to quickly determine whether the current Body has cross-type duplicates.
[0090] Step S90: If a hash value is found in the cross-domain hash table and the type tag is different from the current type, the corresponding request body or response body is replaced with the found request index to obtain the target HTTP traffic data.
[0091] As one implementation, when the hash value of the current Body is successfully found in the cross-domain hash table, and the type marker in the query result is different from the type of the current Body (e.g., the current type is a request body while the type marker in the cross-domain hash table is a response body), it indicates that the hash value has appeared in another type. In this case, the session identifier is retrieved from the query result and mapped to its corresponding request index. This request index is then used to replace the request body or response body in the current HTTP traffic data. The resulting HTTP traffic data is the target HTTP traffic data, where duplicate Body content is replaced with a lightweight index, thus achieving cross-duplicate removal between the request body and response body. By replacing the current Body with the request index when a different type is found in the cross-domain hash table, cross-duplicate removal between the request body and response body is achieved.
[0092] In this embodiment, by assigning a session identifier to each request-response pair and establishing a cross-domain hash table, the cross-domain hash table is queried first after the Bloom filter initially detects a match, thereby achieving deduplication of cross-references between the request body and the response body, and further saving storage space for duplicate data of different types.
[0093] Based on any of the above embodiments of this application, Embodiment Eight of this application proposes a method for deduplicating HTTP traffic data, which can be referred to the above description and will not be repeated hereafter. In addition, after searching for the hash value in the cross-domain hash table, the method further includes: Step S81: If a hash value is found in the cross-domain hash table but the type tag is the same as the current type, or if no hash value is found, then search for the hash value in the text hash table.
[0094] As one implementation method, when a hash value for the current request body or response body is found in the cross-domain hash table but the type marker in the query result matches the type of the current Body, or when the hash value is not found in the cross-domain hash table, it indicates that the current Body does not have cross-type duplication and cannot be deduplicated using the cross-domain hash table. In this case, a precise search is performed in the text hash table using the hash value as the key. The text hash table stores historical indexes of the same type of Body. If the hash value is found in the text hash table, the corresponding request index is used to replace the current Body, achieving deduplication within the same type. If the hash value is not found, the hash value and the current request index are stored in the text hash table as the basis for subsequent deduplication within the same type. By falling back to the text hash table when the cross-domain hash table is not found or the type is the same, the seamless connection between cross-domain deduplication and deduplication within the same type is ensured, ensuring that all kinds of duplicate scenarios can be handled correctly.
[0095] Step S82: After completing the deduplication operation on the request body or response body, store the hash value of the request body or response body, the session identifier, the type tag, and the request index into a cross-domain hash table.
[0096] As one implementation method, after completing the deduplication operation on the request body or response body, the processing result of the Body is stored in a memory block, and then the hash value, session identifier, type tag, and request index of the current Body are recorded in a cross-domain hash table. By writing the hash value, session identifier, type tag, and request index of the current Body into the cross-domain hash table after processing, the index data required for cross-domain matching is continuously accumulated, providing a data foundation for subsequent cross-deduplication between request bodies and response bodies.
[0097] In this embodiment, when the cross-domain hash table is not found or the type is the same, the system falls back to the text hash table for querying, ensuring complete coverage of the cross-domain deduplication and same-type deduplication processes. At the same time, after each deduplication operation is completed, the current Body information is stored in the cross-domain hash table to continuously accumulate cross-type matching index data, so that the system has the ability to continuously enhance cross-type deduplication as the amount of data processed increases.
[0098] For example, to help understand the technical concept or principle of the HTTP traffic data deduplication method in the above embodiments, please refer to... Figure 2 , Figure 2 A simplified flowchart illustrating the HTTP traffic data deduplication method embodiment of this application is shown below: First, HTTP traffic data is received and parsed into four independent parts: request header, request body, response header, and response body. A hash function, such as SHA1, is applied to the request body and response body to calculate their respective hash values. These two hash values are then searched in a Bloom filter. If not found, the hash value is added to the Bloom filter, and the current HTTP traffic data is stored in a memory block. If found, the hash value is further searched in the text hash table. If not found in the text hash table, the hash value and the request index are stored in the text hash table, and the current HTTP traffic data is stored in a memory block. If the hash value is found in the text hash table, the corresponding request index is retrieved from the text hash table and used to replace the request body or response body in the current HTTP traffic data, obtaining the target HTTP traffic data, which is then stored in a memory block. When the accumulated data size in the memory block exceeds a preset threshold, the entire memory block is compressed, and the compressed data block is written to disk. This achieves efficient deduplication and compressed storage of duplicate bodies in massive HTTP traffic with minimal memory overhead.
[0099] This application provides an HTTP traffic data deduplication device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the HTTP traffic data deduplication method in the first embodiment described above.
[0100] The following is for reference. Figure 3 The diagram illustrates a structural schematic suitable for implementing an HTTP traffic data deduplication device according to embodiments of this application. The HTTP traffic data deduplication device in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablets, and in-vehicle terminals, as well as fixed terminals such as digital TVs and desktop computers. Figure 3 The HTTP traffic data deduplication device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0101] like Figure 3As shown, the HTTP traffic data deduplication device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1002 or a program loaded from storage device 1003 into random access memory (RAM) 1004. The random access memory 1004 also stores various programs and data required for the operation of the HTTP traffic data deduplication device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the HTTP traffic data deduplication device to communicate wirelessly or wiredly with other devices to exchange data. Although HTTP traffic data deduplication devices with various systems are shown in the figure, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.
[0102] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0103] The HTTP traffic data deduplication device provided in this application, employing the HTTP traffic data deduplication method in the above embodiments, can solve the technical problem of large memory consumption of HTTP traffic data caused by coarse deduplication granularity. Compared with the prior art, the beneficial effects of the HTTP traffic data deduplication device provided in this application are the same as those of the HTTP traffic data deduplication device provided in the above embodiments, and other technical features of this HTTP traffic data deduplication device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0104] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0105] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0106] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the HTTP traffic data deduplication method in the above embodiments.
[0107] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), or any suitable combination thereof.
[0108] The aforementioned computer-readable storage medium may be included in the HTTP traffic data deduplication device; or it may exist independently and not be assembled into the HTTP traffic data deduplication device.
[0109] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an HTTP traffic data deduplication device, cause the HTTP traffic data deduplication device to: parse the HTTP traffic data to obtain the request body and response body; use hash functions to calculate the hash value of the request body and the hash value of the response body respectively; search for the hash value of the request body and the hash value of the response body in a Bloom filter respectively; if a hash value is found in the Bloom filter, then search for the hash value in the text hash table; if a hash value is found in the text hash table, then retrieve the request index corresponding to the hash value from the text hash table, and replace the request body or response body in the HTTP traffic data with the request index to obtain the target HTTP traffic data.
[0110] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0112] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0113] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described HTTP traffic data deduplication method, which can solve the technical problem of large memory consumption of HTTP traffic data caused by coarse deduplication granularity. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the HTTP traffic data deduplication method provided in the above embodiments, and will not be repeated here.
[0114] This application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the HTTP traffic data deduplication method described above.
[0115] The computer program product provided in this application can solve the technical problem of large memory consumption of HTTP traffic data caused by coarse deduplication granularity. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the HTTP traffic data deduplication method provided in the above embodiments, and will not be repeated here.
[0116] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.
Claims
1. A method for deduplicating HTTP traffic data, characterized in that, The HTTP traffic data deduplication method includes: Parse HTTP traffic data to obtain the request body and response body; Calculate the hash value of the request body and the hash value of the response body using hash functions. Search the Bloom filter for the hash value of the request body and the hash value of the response body respectively; If the hash value is found in the Bloom filter, then the hash value is searched in the text hash table; If the hash value is found in the text hash table, the request index corresponding to the hash value is retrieved from the text hash table, and the request index is used to replace the request body or the response body in the HTTP traffic data to obtain the target HTTP traffic data.
2. The HTTP traffic data deduplication method as described in claim 1, characterized in that, The step of using hash functions to calculate the hash value of the request body and the hash value of the response body includes: Obtain the raw byte data of the request body or the raw byte data of the response body; The original byte data is padded to obtain padded data; The original data length information is appended to the end of the padded data to obtain preprocessed data. The length of the preprocessed data is an integer multiple of the preset block size. The preprocessed data is divided into multiple data blocks of the preset block size; For each data block, multiple rounds of iterative compression are performed and intermediate states are stored in registers. In each round of iteration, the intermediate states of the registers are updated through logical operations, circular shifts, and modulo addition operations. After all the data blocks have been iteratively compressed, the final state of each register is obtained. The final states of each register are then concatenated to obtain the hash value.
3. The HTTP traffic data deduplication method as described in claim 1, characterized in that, Before searching for the hash values of the request body and the response body in the Bloom filter, the method further includes: The required bit array length is calculated based on the total number of the request body and the response body, as well as the preset false positive rate; The required number of hash functions is calculated based on the length of the bit array and the total number. Construct a bit array with a length equal to the length of the bit array, and construct hash functions corresponding to the number of hash functions.
4. The HTTP traffic data deduplication method as described in claim 1, characterized in that, After searching for the hash values of the request body and the response body in the Bloom filter, the process further includes: If the hash value is not found in the Bloom filter, the hash value is added to the Bloom filter, and the HTTP traffic data is stored in a memory block.
5. The HTTP traffic data deduplication method as described in claim 1, characterized in that, After searching for the hash value in the text hash table, the process further includes: If the hash value cannot be found in the text hash table, the hash value is stored in the text hash table as the key and the request index corresponding to the hash value is stored as the value, and the HTTP traffic data is stored in the memory block.
6. The HTTP traffic data deduplication method as described in claim 1, characterized in that, After obtaining the target HTTP traffic data, the process also includes: The target HTTP traffic data is stored in a memory block. When the size of the data in the memory block exceeds a preset threshold, the data in the memory block is compressed and the compressed data block is written to the disk.
7. The HTTP traffic data deduplication method as described in claim 1, characterized in that, The HTTP traffic data deduplication method also includes: Assign a session identifier to each request-response pair of the HTTP traffic data; Establish a cross-domain hash table, wherein the key of the cross-domain hash table is the hash value of the request body or the hash value of the response body, and the value of the cross-domain hash table is a list containing the session identifier, type tag, and request index; When the hash value is found in the Bloom filter, the hash value is then searched in the cross-domain hash table. If the hash value is found in the cross-domain hash table and the type tag is different from the current type, then the corresponding request body or response body is replaced with the found request index to obtain the target HTTP traffic data.
8. The HTTP traffic data deduplication method as described in claim 7, characterized in that, After searching for the hash value in the cross-domain hash table, the process further includes: If the hash value is found in the cross-domain hash table but the type tag is the same as the current type, or if the hash value is not found, then the hash value is searched in the text hash table. After completing the deduplication operation on the request body or the response body, the hash value of the request body or the hash value of the response body, the session identifier, the type tag, and the request index are stored in the cross-domain hash table.
9. An HTTP traffic data deduplication device, characterized in that, The HTTP traffic data deduplication device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the HTTP traffic data deduplication method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the HTTP traffic data deduplication method as described in any one of claims 1 to 8.