Flow data set compression method and device based on point cloud centroid and storage medium

By constructing voxel division and centrome calculation methods, the traffic data set is compressed, which solves the problem of low training efficiency of large-scale data sets and achieves the effect of efficient training and real-time detection.

CN120236117APending Publication Date: 2025-07-01TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510133883.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

How to efficiently train intelligent attack traffic identification models on large-scale traffic datasets to reduce computing resource consumption and training latency, especially in high-bandwidth scenarios of backbone networks or enterprise gateways.

Method used

By constructing voxels, the high-dimensional space traffic data set is divided, the point cloud density is obtained and classified and processed, the center of the high-dimensional polygon is sampled, and the target compressed voxel is spliced ​​to form the target compressed voxel, and the compressed traffic data set is output.

Benefits of technology

It significantly improves training efficiency and improves training speed of 75.82 times, while maintaining detection accuracy and robustness, ensuring real-time processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236117A_ABST
    Figure CN120236117A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network security, in particular to a traffic data set compression method, device and equipment based on a point cloud centroid and a computer storage medium. The traffic data set compression method comprises the following steps: firstly, dividing a high-dimensional space traffic data set by using voxels; then, density measurement is carried out on voxels, and the voxels are classified according to the density; then, sampling point clouds in the voxels, and reducing the density; finally, centroids of points obtained through sampling are calculated, and a compressed data set is expressed; in the process, the high-dimensional spatial manifold is divided by adopting voxels, so that the real-time compression of the data set is ensured. Meanwhile, the extracted point cloud can keep the manifold of a high-dimensional space flow data set, and the accuracy of a model trained on a compressed data set is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular, to a method, device, equipment and computer storage medium for compressing a traffic data set based on the centroid of a point cloud. Background Art

[0002] In recent years, network security has gradually become an important part of national security. The statement "Without network security, there is no national security" fully demonstrates the important position of Internet security construction in national security construction. However, distributed denial-of-service attacks have been continuously threatening network security. These attacks flood a large amount of traffic to key Internet services, hindering legitimate users from using these services. How to quickly detect the flooded traffic has become an important issue in the field of network security research. For example, in high-bandwidth scenarios such as backbone networks or enterprise gateways, real-time detection and blocking of attack traffic can protect a large number of legitimate network users.

[0003] However, with the popularization of the Internet, the scale of traffic has increased rapidly. Therefore, the scale of the data set for training an intelligent attack traffic recognition system has also increased significantly. The training of artificial intelligence algorithms, especially deep learning algorithms, consumes a large amount of computing resources. This means that building an intelligent attack traffic recognition model on a large-scale traffic data set with a complex model will consume a large amount of time cost and computing power overhead. Therefore, the problem to be solved currently is whether it is possible to preprocess massive large-scale traffic and data sets, compress the scale of the data set, only save samples with rich information, and finally build an attack traffic recognition model efficiently and in real time on this compressed data set, reduce the computing overhead generated during training, and limit the delay of the deployed system. Summary of the Invention

[0004] Therefore, the technical problem to be solved by the present invention is how to compress a large-scale traffic data set.

[0005] To solve the above technical problem, the present invention provides a method for compressing a traffic data set, including:

[0006] Constructing voxels to divide a high-dimensional space traffic data set;

[0007] Obtaining the point cloud density in each voxel, and classifying and processing the voxels according to the point cloud density to obtain a plurality of reserved voxels and a plurality of target voxels;

[0008] Sampling the point cloud inside each target voxel multiple times to obtain a plurality of high-dimensional polygons;

[0009] Obtaining the centroid of each high-dimensional polygon in the high-dimensional traffic space, and splicing all the centroids to obtain a target compressed voxel;

[0010] Output the concatenated multiple retained voxels and the target compressed voxels as a target compressed flow dataset.

[0011] Preferably, the construction of voxels for partitioning the high-dimensional space flow dataset includes:

[0012] Based on the high-dimensional space flow dataset, perform logarithmic transformation on the flow-level feature vectors of N alarms and then perform normalization processing to obtain normalized feature vectors, where the length of the flow-level feature vectors is M;

[0013] Define a voxel as an M-dimensional cube with a fixed side length;

[0014] Partition each sample point in the normalized feature vectors into the space covered by the corresponding voxel.

[0015] Preferably, the space covered by the voxel is determined according to the Cartesian product.

[0016] Preferably, the normalization processing includes max-min normalization processing.

[0017] Preferably, obtaining the point cloud density in each voxel and classifying and processing the voxels according to the point cloud density to obtain multiple retained voxels and multiple target voxels includes:

[0018] Remove the voxels with fewer than the first preset number of sample points;

[0019] Define the voxels with no less than the first preset number and fewer than the second preset number of sample points as retained voxels, and retain the retained voxels;

[0020] Define the voxels with no less than the second preset number of sample points as target voxels.

[0021] Preferably, performing multiple samplings on the point cloud inside the target voxel to obtain multiple high-dimensional polygons includes:

[0022] Define a random sampling function;

[0023] Based on the random sampling function, perform T samplings on the point cloud inside the target voxel to obtain T sets each containing K points, and construct them into T high-dimensional polygons each containing K vertices.

[0024] Preferably, obtaining the centroid of each high-dimensional polygon in the high-dimensional flow space includes:

[0025] Define a centroid calculation function;

[0026] Based on the centroid calculation function, obtain the geometric center of all points in each high-dimensional polygon to obtain the centroid of each high-dimensional polygon in the high-dimensional flow space.

[0027] The present invention also provides a traffic data set compression device, comprising:

[0028] A traffic space voxel division module for dividing a high-dimensional space traffic data set by constructing voxels;

[0029] A voxel density screening module for obtaining the point cloud density in each voxel, classifying and processing the voxels according to the point cloud density to obtain a plurality of retained voxels and a plurality of target voxels;

[0030] A high-dimensional polyhedron sampling module for sampling the point cloud inside each target voxel multiple times to obtain a plurality of high-dimensional polygons;

[0031] A point cloud centroid calculation module for obtaining the centroid of each high-dimensional polygon in the high-dimensional traffic space and splicing all the centroids to obtain a target compressed voxel;

[0032] A data set scale compression module for outputting the plurality of retained voxels and the target compressed voxel after splicing as a target compressed traffic data set.

[0033] The present invention also provides a traffic data set compression device, comprising:

[0034] A memory for storing a computer program;

[0035] A processor for implementing the steps of the above-mentioned traffic data set compression method when executing the computer program.

[0036] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned traffic data set compression method are implemented.

[0037] The above technical solution of the present invention has the following advantages compared with the prior art:

[0038] In the traffic data set compression method of the present invention, first, a high-dimensional space traffic data set is divided by using voxels. Then, the density of the voxels is measured and classified according to the density. Subsequently, the point cloud inside the voxels is sampled to reduce the density. Finally, the centroid of the sampled points is calculated to represent the compressed data set; in this process, the high-dimensional space manifold is divided by using voxels, ensuring real-time compression of the data set. At the same time, the extracted point cloud can maintain the manifold of the high-dimensional space traffic data set, ensuring the accuracy of the model trained on the compressed data set. Description of the Drawings

[0039] To make the content of the present invention easier to be clearly understood, the following further elaborates on the present invention in detail according to specific embodiments of the present invention in combination with the accompanying drawings, where:

[0040] Figure 1 is the implementation flowchart of a method for compressing a traffic data set provided by the present invention;

[0041] Figure 2 is the system architecture diagram of the present invention. Specific Embodiments

[0042] The core of the present invention is to provide a method, device, equipment and computer storage medium for compressing a traffic data set, which effectively compresses a large-scale traffic data set.

[0043] To enable those skilled in the art to better understand the solution of the present invention, the following further elaborates on the present invention in detail in combination with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0044] Please refer to Figure 1 and Figure 2 , Figure 1 is the implementation flowchart of a method for compressing a traffic data set provided by the present invention, Figure 2 is the system architecture diagram of the present invention; the specific operation steps are as follows:

[0045] S101: Construct voxels to divide the high-dimensional space traffic data set;

[0046] S102: Obtain the point cloud density in each voxel, and classify and process the voxels according to the point cloud density to obtain a plurality of retained voxels and a plurality of target voxels;

[0047] S103: Sample the point cloud inside each target voxel multiple times to obtain a plurality of high-dimensional polygons;

[0048] S104: Obtain the centroid of each high-dimensional polygon in the high-dimensional traffic space, and splice all the centroids to obtain a target compressed voxel;

[0049] S105: Output the plurality of retained voxels and the target compressed voxel after splicing as a target compressed traffic data set.

[0050] Based on the above embodiments, this embodiment elaborates on step S101 in detail:

[0051] Based on the high-dimensional space traffic dataset, the flow-level feature vectors of N alarms are logarithmically transformed and then normalized to obtain normalized feature vectors, where the length of the flow-level feature vectors is M:

[0052] S11: Specifically, we use and M to represent the flow-level feature vector related to the i-th alarm and the length of this feature vector respectively. Methods using different numbers of flow-level features have different M values.

[0053] Let the matrix S represent the flow-level features of all alarms, where s i,j is defined as the j-th flow-level feature of the i-th sample.

[0054]

[0055] S12: To make the features numerically stable and facilitate further analysis, the present invention performs a logarithmic transformation to reduce the feature range and prevent arithmetic overflow during the point cloud analysis process. The present invention uses E = log2(1 + S) to represent the result. Then, by performing min-max normalization on each row of the matrix E, each element in E = [e ij (1 ≤ i ≤ N, 1 ≤ j ≤ M) is transformed to [0, 1]. Now, the present invention obtains the normalized feature vectors represented by N points represented.

[0056] Define a voxel as an M-dimensional cube with a fixed side length:

[0057] S13: The present invention defines a voxel as an M-dimensional small cube with a side length of ∈ in the normalized traffic feature space , where is the index of the voxel. Therefore, the space covered by this voxel can be represented by the Cartesian product:

[0058]

[0059] Divide each sample point in the normalized feature vector into the space covered by the corresponding voxel:

[0060] S14: We define a function to determine whether the point is located in the space covered by the voxel .

[0061] S15: We cluster the point set P into the voxels indexed by the complete set of index vectors . Then, we define a function This function outputs the set of indices of the points located within the voxel indexed by :

[0062]

[0063] So far, the present invention has divided the dense point cloud representing the traffic data set with voxels. By dividing the point cloud in this way, the present invention can process the complex manifolds formed by large-scale traffic data sets in the high-dimensional traffic feature space region by region.

[0064] Based on the above embodiments, this embodiment will elaborate on step S102 in detail:

[0065] Remove the voxels with the number of sample points less than the first preset number;

[0066] Define the voxels with the number of sample points not less than the first preset number and less than the second preset number as reserved voxels, and reserve the reserved voxels:

[0067] S21: For the voxels representing outliers, the present invention needs to exclude them. Specifically, the present invention excludes the voxels representing less than E points. The present invention defines the index of the reserved voxels as Satisfying for all Let U be the number of reserved voxels.

[0068] Define the voxels with the number of sample points not less than the second preset number as target voxels:

[0069] S22: On the other hand, we define the index of the voxels containing more than D>E points as Satisfying for all where V is the number of high-density voxels.

[0070] S23: Subsequently, the present invention will focus on processing these point clouds with high density, which are the key to reducing the scale of the data set. For the voxels with the number of points exceeding E but less than D The present invention retains all the points therein. Because these points are neither outliers to be removed nor dense points to be compressed, their density is moderate and basically no processing is required. We define the coordinates of these points as:

[0071]

[0072] S24: The features corresponding to these directly retained points are denoted as

[0073] Based on the above embodiments, this embodiment will elaborate on step S103 in detail:

[0074] Define a random sampling function:

[0075] S31: The present invention defines the function which randomly selects from a set containing all integer elements Sample a subset of size K, where n is the random number seed. This sampling function aims to make the sampling results as different as possible under different random number seeds.

[0076] Based on the random sampling function, sample the point cloud inside the target voxel T times to obtain T sets each containing K points, and construct them into T high-dimensional polygons each containing K vertices:

[0077] S32: For a single high-density voxel, the present invention samples it T times. Each sampling constructs a set containing K points, forming a high-dimensional polygon. The present invention defines a function representing this sampling process. This function will be a set of indices of T groups of sampled points.

[0078]

[0079]

[0080] The object processed in this embodiment is the voxel representing high-density points, and these voxels are sampled. Through sampling, we aim to reduce the processing overhead and avoid the excessive processing overhead caused by directly processing the high-density point cloud.

[0081] Based on the above embodiments, this embodiment details step S104:

[0082] Define the centroid calculation function;

[0083] S41: The present invention calculates the centroid of the polygon formed by the points in each group of samplings in the high-dimensional flow space, and uses these centroids to represent the points in all polygons. Therefore, for a voxel with points in it, we can effectively represent them with T centroids, thus significantly reducing the scale of the point cloud, that is, the scale of the flow data set.

[0084] S42: We define the centroid calculation function to calculate a set of matrix indices indicating the centroid of all points:

[0085]

[0086] Based on the centroid calculation function, obtain the geometric center of all points in each high-dimensional polygon, and get the centroid of each high-dimensional polygon in the high-dimensional flow space.

[0087] S43: Define the coordinates of the centroid obtained by the present invention for any voxel in the high-dimensional flow feature space:

[0088]

[0089]

[0090]

[0091] S44: Finally, the present invention splices the matrix formed by all the centroids to obtain a matrix representing the point cloud obtained after point cloud compression of the points corresponding to the high-density voxels. Here, the function Concatenate(·) performs a splicing operation on the first dimension:

[0092]

[0093] Based on the above embodiments, this embodiment details step S105:

[0094] S51: Finally, the output is the dataset after compressing the original dataset, which is

[0095] Artificial intelligence algorithms are widely used in attack traffic detection. They learn traffic patterns from traffic datasets and can detect hidden threat traffic, with significantly improved accuracy compared to traditional fixed-rule detection. However, with the rapid increase in the scale of Internet traffic, the scale of traffic datasets has also increased significantly, which means that training an attack traffic recognition model on a large-scale traffic dataset will consume a large amount of computing power, along with a huge deployment time delay.

[0096] The present invention proposes a method aimed at quickly training an attack traffic detection model on a large-scale traffic dataset using limited computing resources by compressing traffic data, enabling the system to be quickly deployed to the network. The specific data compression method is based on the geometric structure of the high-dimensional traffic feature space. The present invention maps all sample traffic in the dataset to point clouds in the high-dimensional traffic feature space, extends the concept of voxels to the high-dimensional traffic feature space, and represents all points (traffic samples) by calculating the centroids of the points in the voxels, thus significantly reducing the scale of the dataset. Experimental verification shows that this method can improve the training efficiency by 75.82 times for the current top 10 detection methods on 8 datasets, while hardly affecting the detection accuracy and robustness, and ensuring the real-time processing of compressed traffic data.

[0097] This embodiment of the present invention provides a traffic dataset compression device; the specific device may include:

[0098] A traffic space voxel division module for dividing the high-dimensional space traffic dataset by constructing voxels;

[0099] A voxel density screening module for obtaining the point cloud density in each voxel and classifying and processing the voxels according to the point cloud density to obtain a plurality of retained voxels and a plurality of target voxels;

[0100] A high-dimensional polyhedron sampling module, configured to perform multiple samplings on the point cloud inside each target voxel to obtain multiple high-dimensional polygons;

[0101] A point cloud centroid calculation module, configured to obtain the centroids of each high-dimensional polygon in the high-dimensional flow space, and splice all the centroids to obtain a target compressed voxel;

[0102] A data set scale compression module, configured to output the spliced multiple retained voxels and the target compressed voxel as a target compressed flow data set.

[0103] The present invention discloses a flow data set compression device based on the centroid of a point cloud. The device includes: a flow space voxel division module for dividing the high-dimensional flow feature space into blocks; voxel density screening for analyzing the voxel classification process of the point cloud density in the voxel; high-dimensional polyhedron sampling for sampling several polygons for each voxel to reduce the point cloud density; point cloud centroid calculation for calculating the centroid of each polygon and using the centroid to represent the samples in the compressed flow data set. This technology analyzes the point cloud corresponding to the flow data set, analyzes its topological relationship in the high-dimensional space by means of voxel analysis, and finally represents the manifold of the point cloud corresponding to the original data set by calculating the centroid. Since the number of centroids is much smaller than the scale of the original point cloud, using these centroids as flow feature samples can significantly reduce the data set scale. This technology efficiently analyzes the point cloud through voxels, with the advantages of high throughput, low latency, and high precision; at the same time, the centroid can maintain the geometric properties of the manifold in the high-dimensional flow space, ensuring the accuracy of training on the compressed data set.

[0104] The flow data set compression device in this embodiment is used to implement the foregoing flow data set compression method. Therefore, the specific implementation manners in the flow data set compression device can be seen in the embodiment part of the foregoing flow data set compression method. For example, the flow space voxel division module, the voxel density screening module, the high-dimensional polyhedron sampling module, the point cloud centroid calculation module, and the data set scale compression module are respectively used to implement steps S101, S102, S103, S104, and S105 in the foregoing flow data set compression method. Therefore, the specific implementation manners can refer to the descriptions of the corresponding various part embodiments and will not be elaborated here.

[0105] The specific embodiment of the present invention further provides a flow data set compression device, including: a memory for storing a computer program; a processor for implementing the steps of the foregoing flow data set compression method when executing the computer program.

[0106] A specific embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method for compressing a traffic data set are implemented.

[0107] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0108] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0109] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0110] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0111] Obviously, the above embodiments are merely examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A method for compressing a traffic data set, characterized in that: include: Construct voxels to partition high-dimensional spatial flow data sets; Acquiring a point cloud density in each voxel, and classifying and processing the voxels according to the point cloud density to obtain a plurality of retained voxels and a plurality of target voxels; The point cloud inside each target voxel is sampled multiple times to obtain multiple high-dimensional polygons; Obtain the centroid of each high-dimensional polygon in the high-dimensional flow space, and concatenate all the centroids to obtain the target compressed voxel; The spliced ​​plurality of retained voxels and the target compressed voxels are output as a target compressed flow data set.

2. The method for compressing a flow data set according to claim 1, characterized in that: The construction of voxels to divide the high-dimensional spatial flow data set includes: Based on the high-dimensional spatial traffic data set, the flow-level feature vectors of N alarms are logarithmically transformed and then normalized to obtain a normalized feature vector, wherein the length of the flow-level feature vector is M; Define voxel as an M-dimensional cube with fixed side length; Each sample point in the normalized feature vector is divided into a space covered by a corresponding voxel.

3. The method for compressing a flow data set according to claim 2, characterized in that: The space covered by the voxels is determined according to the Cartesian product.

4. The method for compressing a flow data set according to claim 2, characterized in that: The normalization process includes a maximum-minimum normalization process.

5. The method for compressing a flow data set according to claim 1, characterized in that: The step of obtaining the point cloud density in each voxel, and classifying and processing the voxels according to the point cloud density to obtain a plurality of retained voxels and a plurality of target voxels includes: Removing voxels with sample points less than a first preset number; defining voxels having sample points not less than a first preset number and less than a second preset number as reserved voxels, and retaining the reserved voxels; Voxels having sample points not less than a second preset number are defined as target voxels.

6. The method for compressing a flow data set according to claim 1, characterized in that: The point cloud inside the target voxel is sampled multiple times to obtain multiple high-dimensional polygons, including: Define a random sampling function; Based on the random sampling function, the point cloud inside the target voxel is sampled T times to obtain T sets containing K points, and then constructed into T high-dimensional polygons containing K vertices.

7. The method for compressing a flow data set according to claim 1, characterized in that: The obtaining of the centroid of each high-dimensional polygon in the high-dimensional flow space includes: Define the centroid calculation function; Based on the centroid calculation function, the geometric centers of all points in each high-dimensional polygon are obtained to obtain the centroid of each high-dimensional polygon in the high-dimensional flow space.

8. A flow data set compression device, characterized in that: include: The flow space voxel partitioning module is used to construct voxels to partition high-dimensional space flow data sets; A voxel density screening module, used for obtaining the point cloud density in each voxel, and classifying and processing the voxels according to the point cloud density to obtain a plurality of retained voxels and a plurality of target voxels; A high-dimensional polyhedron sampling module is used to perform multiple sampling on the point cloud inside each target voxel to obtain multiple high-dimensional polygons; Point cloud centroid calculation module, used to obtain the centroid of each high-dimensional polygon in the high-dimensional flow space, and splice all the centroids to obtain the target compressed voxel; The data set scale compression module is used to output the spliced ​​multiple retained voxels and the target compressed voxels as a target compressed flow data set.

9. A flow data set compression device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of a flow data set compression method as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of a traffic data set compression method as described in any one of claims 1 to 7 are implemented.