Network high-speed traffic acquisition method, system and device based on clouded environment
By building a two-level dynamic scaling mechanism and deep learning model in a cloud environment, combining consistent hashing and lightweight DPI technology, the problem of insufficient resource elasticity in a cloud environment is solved, efficient traffic acquisition and processing is achieved, resource utilization and processing efficiency are improved, and real-time analysis needs of high-speed networks are met.
Patent Information
- Application Number
- CN202510673823.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-23
AI Technical Summary
Traditional physical equipment-based traffic acquisition methods are difficult to meet the real-time, dynamic scalability and resource utilization requirements in high-speed network scenarios in a cloud-based environment, and there are problems such as insufficient resource elasticity, lack of traffic prediction mechanism, limited traffic distribution and processing efficiency, and low data aggregation and synchronization performance.
By building a two-level dynamic scaling mechanism, combining deep learning technology and dynamic filtering mechanism, using FCN prediction model for traffic prediction, combining consistent hashing algorithms and lightweight DPI technology, flexible allocation and efficient processing of resources are achieved, optimization algorithms are used to find the optimal learning rate solution, and design ring buffer management strategies and Kafka message queue synchronization mechanism.
It improves the resource utilization and response speed of cloud environments, reduces cost and prediction error rates, improves the flexibility and throughput of traffic processing, ensures the continuity and order of data, and meets the real-time analysis and long-term storage needs in high-speed network scenarios.
Smart Images

Figure CN120263758A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of network communication technology. Specifically, it particularly relates to a method, system, and device for network high-speed traffic collection based on a cloudified environment. Background Art
[0002] With the rapid development of cloud computing technology, the cloudified environment, with its characteristics such as elastic resource scheduling, virtualization isolation, and high availability, has gradually become the core infrastructure for network high-speed traffic collection and analysis. The cloudified environment can provide flexible support for large-scale network traffic processing by dynamically allocating computing, storage, and network resources. However, in high-speed network scenarios (such as 5G, Internet of Things, industrial Internet), the traffic scale grows exponentially, and traditional traffic collection methods based on physical devices face performance bottlenecks and are difficult to meet the requirements of real-time, dynamic scalability, and resource utilization. Therefore, how to achieve efficient and intelligent network high-speed traffic collection and processing based on the cloudified environment has become a key challenge in current research.
[0003] Chinese Patent CN117294533B, a method and system for network traffic collection based on a cloudified environment; in the present invention, a mirror virtual switch is created on a computing node, and the network traffic of a monitoring node is diverted to this switch through its patch port; subsequently, the traffic is forwarded to a deployed virtualized instance agent node, where the network traffic data is converted into a file format and analyzed in real time; the analyzed data is further forwarded to a data security monitoring platform to achieve real-time monitoring of events; at the same time, according to the monitoring policy and event changes, the flow table of the virtual switch is dynamically adjusted.
[0004] There are problems in the prior art such as insufficient elasticity of cloudified environment resources, lack of a traffic prediction mechanism, limited traffic distribution and processing efficiency, and low data aggregation and synchronization performance. Summary of the Invention
[0005] (I) Technical Problems to be Solved Aiming at the problems in the related art, the present invention provides a method for network high-speed traffic collection based on a cloudified environment. The present invention solves the problems of insufficient elasticity of cloudified environment resources, lack of a traffic prediction mechanism, limited traffic distribution and processing efficiency, and low data aggregation and synchronization performance through a two-level dynamic scaling mechanism, deep learning technology, and a dynamic filtering mechanism.
[0006] (II) Technical Solutions To solve the above technical problems, the present invention is implemented through the following technical solutions: S1. Construct a cloudified environment and set relevant parameters of the cloudified environment; S2. Mirror the real-time high-speed network traffic to the cloudified environment and perform feature extraction to obtain the real-time high-speed network traffic features; S3. Build an FCN prediction model, improve the FCN prediction model using historical high-speed network traffic feature data and an optimization algorithm to obtain an optimized FCN network high-speed traffic prediction model, and combine it with the real-time high-speed network traffic features to obtain the predicted high-speed network traffic; Determine whether to perform the first cloudified environment scaling based on the predicted high-speed network traffic to obtain the cloudified environment after the first scaling; S4. Collect the load index data of the cloudified environment after the first scaling to obtain the real-time load index and determine whether to perform the second cloudified environment scaling to obtain the cloudified environment after the second scaling; S5. Collect the high-speed network traffic packets of the cloudified environment after the second scaling and perform dynamic filtering processing to obtain a preprocessed high-speed network traffic data set; S6. Aggregate and synchronize the preprocessed high-speed network traffic data set to obtain a standard high-speed network traffic data stream.
[0007] Preferably, the S1 includes the following steps: S11. Create a virtualized resource pool in the cloud platform, build a collection node cluster based on Kubernetes; Configure the elastic scaling policy of the resource pool; The elastic scaling policy of the resource pool changes according to the size of the predicted high-speed network traffic and the change of the real-time load rate; Allocate a dedicated virtual network interface for each collection node in the resource pool and enable the SR-IOV technology to improve the network card performance; S12. Deploy the SDN controller to the cloud gateway, configure the high-speed network traffic mirroring rule, and mirror the target high-speed network traffic to the collection system; Set the mirror high-speed network traffic bandwidth limit to avoid overload; S13. Configure the initial collection policy in the control center, where the initial collection policy includes the high-speed network traffic distribution weight and resource threshold of each node. The high-speed network traffic distribution weight dynamically adjusts the weight according to the node latency, and the resource threshold limits the maximum and minimum numbers of the node cluster; S14. Obtain the cloudified environment through S11, S12, and S13; The above steps build an efficient and flexible cloudified environment by creating a virtualized resource pool in the cloud platform, building a collection node cluster based on Kubernetes, configuring the elastic scaling policy, allocating a dedicated virtual network interface and enabling the SR-IOV technology, deploying the SDN controller and configuring the high-speed network traffic mirroring rule, and configuring the initial collection policy in the control center.
[0008] Preferably, the S2 includes the following steps: S21. Mirror the real-time network high-speed traffic to the cloudified environment to obtain the real-time mirrored network high-speed traffic; S22. Extract the five-tuple feature dataset and the application layer content features of the real-time mirrored network high-speed traffic through lightweight DPI to obtain the real-time network high-speed traffic features; the five-tuple feature dataset includes the source IP address, destination IP address, protocol type, source port, and destination port, and the application layer content features reflect the priority level of the network high-speed traffic data; use the consistent hashing algorithm to map the real-time network high-speed traffic to the collection node cluster according to the hash value of the five-tuple feature dataset; The above steps realize an efficient and orderly real-time network high-speed traffic processing mechanism by mirroring the real-time network high-speed traffic to the cloudified environment, using lightweight DPI technology to extract key feature data, and using the consistent hashing algorithm for traffic mapping; effectively improve the processing efficiency of the network high-speed traffic, ensure the orderliness of data processing and the continuity of sessions, and lay a foundation for subsequent real-time analysis and long-term storage.
[0009] Preferably, the S3 includes the following steps: S31. Construct an FCN prediction model and set the initial learning rate q of the FCN prediction model; S32. Collect the feature data of the historical network high-speed traffic in the cloudified environment to obtain the historical network high-speed traffic feature data; S33. Set the maximum number of training iterations and the training accuracy threshold; input the historical network high-speed traffic feature data into the FCN prediction model for training to obtain the training accuracy; adjust the learning rate of the FCN prediction model according to the training accuracy; when the training accuracy ≥ the training accuracy threshold or reaches the maximum number of training iterations, obtain the FCN network high-speed traffic prediction model; S34. Use an optimization algorithm to find the learning rate of the FCN network high-speed traffic prediction model to obtain the optimal solution; use the optimal solution as the learning rate of the FCN network high-speed traffic prediction model to obtain the optimized FCN network high-speed traffic prediction model; S35. Input the real-time network high-speed traffic feature data into the optimized FCN network high-speed traffic prediction model to obtain the predicted network high-speed traffic; S36. Set the scaling rule for the predicted high-speed network traffic. According to the scaling rule for the predicted high-speed network traffic, the real-time high-speed network traffic, and the predicted high-speed network traffic, perform the first cloudification environment scaling or keep it unchanged to obtain the cloudification environment after the first scaling; the scaling rule for the predicted high-speed network traffic is to set the predicted high-speed network traffic expansion threshold d and the predicted high-speed network traffic reduction threshold f; when the ratio e of the predicted high-speed network traffic to the current high-speed network traffic is ≥ the predicted high-speed network traffic expansion threshold d, expand the collection node group in the cloudification environment to times of the current collection node group; when the ratio g of the predicted high-speed network traffic to the current high-speed network traffic is ≤ the predicted high-speed network traffic reduction threshold f, reduce the collection node group in the cloudification environment to the current collection node group times; The above steps realize the first cloudification environment scaling by constructing and optimizing the FCN prediction model, training and predicting in combination with historical and real-time high-speed network traffic feature data, and then combining the set scaling rule for the predicted high-speed network traffic; effectively improve the resource utilization rate, ensure that the cloudification environment can dynamically adjust resources according to real-time and predicted traffic, avoid resource waste and overload problems, and significantly improve the flexibility and efficiency of the system in dealing with high-speed network traffic.
[0010] Preferably, the step of using an optimization algorithm to find the learning rate of the FCN network high-speed traffic prediction model to obtain the optimal solution in S34 includes the following steps: S341. Construct a sunflower population and set the size of the sunflower population to n; S342. According to the learning rate of the FCN network high-speed traffic prediction model, randomly set the initial height of the sunflower population to obtain the initial height set of the sunflower population; S343. There are resource limitations in the height growth of sunflowers. Excessive height leads to excessive energy consumption or unstable structure. Set the height gain coefficient and height penalty coefficient, and define the fitness function; S344. Perform iterative operations on the initial height set of the sunflower population. In each round of iteration, update the height of each sunflower in the initial height set of the sunflower population according to the fitness function, and obtain the best sunflower individual and the global best sunflower in the sunflower population in each round of iteration; S345. Repeat S344. When the maximum optimization iteration number is reached, stop the iteration and use the global best sunflower as the optimal solution; The above steps use the sunflower optimization algorithm to find the optimal learning rate solution for the FCN network high-speed traffic prediction model; by constructing a sunflower population, setting the initial height, defining the fitness function, and performing iterative operations, the global best sunflower is finally determined as the optimal solution of the learning rate; significantly improving the accuracy and efficiency of the FCN prediction model, providing a more accurate prediction basis for the dynamic scaling of the cloudified environment, and enhancing the system resource utilization rate and response speed.
[0011] Preferably, the S4 includes the following steps: S41. Collect the real-time CPU utilization rate and real-time memory occupancy rate of the collection node group in the cloudified environment after the first scaling, and obtain the real-time load index; S42. Set the load index scaling rule, and perform the second cloudified environment scaling or keep it unchanged according to the load index scaling rule and the real-time load index, to obtain the cloudified environment after the second scaling; the load index scaling rule is to set the maximum CPU utilization threshold h1, the maximum memory occupancy threshold h2, the minimum CPU utilization threshold h3, and the minimum memory occupancy threshold h4; when the real-time CPU utilization rate k1 ≥ the maximum CPU utilization threshold h1 or the real-time memory occupancy rate k2 ≥ the maximum load memory occupancy threshold h2, expand the collection node group in the cloudified environment to times of the current collection node group; when the real-time CPU utilization rate k1 ≤ the minimum CPU utilization threshold h3 or the real-time load rate k2 ≤ the maximum load memory occupancy threshold h4, reduce the collection node group in the cloudified environment to times of the current collection node group; The above steps collect the real-time CPU utilization rate and memory occupancy rate of the collection node group in the cloudified environment after the first scaling as the load index, and set the scaling rule based on these indexes, realizing the secondary precise scaling of the cloudified environment; ensuring that the cloudified environment can efficiently and flexibly adapt to the actual work requirements, further optimizing the resource allocation, improving the system performance and stability, and reducing the operation cost at the same time.
[0012] Preferably, the S5 includes the following steps: S51. In the node cluster of the cloudified environment after the second scaling, batch obtain the network high-speed traffic data through the DPDK kernel bypass technology to obtain the network high-speed traffic packets; S52. Set the dynamic filtering rule engine to discard the ICMP protocol, block the blacklist IP, and match the attack feature library; perform three-level filtering on the network high-speed traffic packets in combination with the dynamic filtering rule engine to obtain the filtered network high-speed traffic packets; S53. Extract the metadata fields of the filtered network high-speed traffic packets, and add threat level tags to the filtered network high-speed traffic packets according to the metadata fields to generate a preprocessed network high-speed traffic data set; The above steps efficiently batch obtain network high-speed traffic data by using the DPDK kernel bypass technology, and perform multi-level deep filtering on the traffic packets through the set dynamic filtering rule engine, effectively discarding irrelevant protocol data, blocking malicious IPs, and identifying potential attack features; extracting the metadata fields of the filtered traffic packets and adding threat level tags to generate a preprocessed network high-speed traffic data set; significantly improving the data collection and processing efficiency and enhancing the network security protection ability.
[0013] Preferably, the S6 includes the following steps: S61. Set the ring buffer threshold; when the ring buffer of the cloud environment accumulates to the ring buffer threshold, discard the data with low priority in the preprocessed network high-speed traffic data set, compress and encrypt it into the Avro format to obtain a processed network high-speed traffic data set; S62. Write the processed network high-speed traffic data set into the Kafka message queue, and partition the data in the processed network high-speed traffic data set according to the source IP hash value to obtain a standard network high-speed traffic data stream; The above steps set the ring buffer threshold and implement an intelligent data management strategy. When the buffer data accumulates to the threshold, selectively discard the data with low priority in the preprocessed network high-speed traffic data set, compress and encrypt it into the Avro format, write the processed data set into the Kafka message queue, and partition it according to the source IP hash value to generate a standardized network high-speed traffic data stream; optimizing the data storage and transmission efficiency and ensuring the high availability and orderliness of the data.
[0014] A network high-speed traffic collection system based on a cloud environment for implementing the above network high-speed traffic collection method based on a cloud environment, including a cloud environment construction module, a traffic mirroring and feature processing module, an optimized FCN network high-speed traffic prediction model module, a first cloud environment scaling decision module, a second cloud environment scaling decision module, a dynamic filtering and threat marking module, and a data aggregation and synchronization module; The cloud environment construction module is used to create an elastic resource pool, configure a collection node cluster and network parameters to implement the deployment of the infrastructure in the cloud environment; The traffic mirroring and feature processing module is used to obtain real-time mirror network high-speed traffic and extract key features to obtain real-time network high-speed traffic features; The optimized FCN network high-speed traffic prediction model module improves the FCN network prediction model through historical network high-speed traffic characteristics and an optimization algorithm to obtain an optimized FCN network high-speed traffic prediction model; The first cloudification environment scaling decision module is used to obtain predicted network high-speed traffic according to the optimized FCN network high-speed traffic prediction model and real-time network high-speed traffic characteristics; judge whether to perform the first cloudification environment scaling according to the predicted network high-speed traffic and the predicted network high-speed traffic scaling rules, and obtain the cloudification environment after the first scaling; The second cloudification environment scaling decision module is used to judge whether to perform the second cloudification environment scaling according to the above load index scaling rules and real-time load indexes, and obtain the cloudification environment after the second scaling; The dynamic filtering and threat marking module is used to perform multi-level cleaning on network high-speed traffic packets and mark the threat levels; The data aggregation and synchronization module is used to efficiently aggregate multi-node data and realize the output of a standardized data stream with low latency and high throughput.
[0015] A network high-speed traffic collection device based on a cloudification environment stores a program, which when executed by a processor is used to implement the above-mentioned network high-speed traffic collection method based on a cloudification environment.
[0016] (III) Beneficial effects The present invention has the following beneficial effects: The present invention effectively solves the problem of insufficient resource elasticity in the traditional cloudification environment through a two-level dynamic scaling mechanism combining prediction-driven scaling and real-time load feedback scaling; the first traffic prediction based on the optimized FCN network high-speed traffic prediction model realizes resource pre-allocation and avoids processing delays caused by burst traffic; the second precise adjustment based on real-time CPU / memory indexes further optimizes resource utilization; compared with a single threshold trigger mechanism, it improves the resource supply response speed, reduces resource redundancy at the same time, and significantly reduces costs.
[0017] The present invention introduces an optimization algorithm to dynamically search for the optimal solution of the learning rate of the FCN network high-speed traffic prediction model to obtain an optimized FCN network high-speed traffic prediction model; breaks through the limitation that the traditional gradient descent method is prone to falling into local optima; realizes the bio-inspired adaptive adjustment of model parameters by defining a fitness function, improves the training convergence speed of the model, reduces the prediction error rate, and provides a reliable basis for resource dynamic scheduling.
[0018] The present invention ensures fixed processing of traffic in the same session by adopting the consistent hashing algorithm, avoiding out-of-order problems and improving session continuity; combined with three-level dynamic filtering rules, it realizes fine-grained traffic cleaning and reduces the amount of invalid data processing; by extracting application layer content features with lightweight DPI and dynamically adjusting the node distribution weight, it improves the load balancing degree of the node cluster and enhances the overall processing throughput.
[0019] The present invention designs a priority-aware circular buffer management strategy, which preferentially retains high-priority traffic when the buffer reaches the threshold, reducing the critical data loss rate; through Kafka message queue partitioning and synchronization according to the source IP hash value, while ensuring data orderliness, combined with Avro format compression and encryption technology, it reduces data synchronization latency, reduces storage space occupancy, and meets the real-time analysis and long-term storage requirements in high-speed network scenarios.
[0020] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the invention, the drawings required for describing the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the invention, and those of ordinary skill in the art can also obtain other drawings according to these drawings without creative efforts.
[0022] Figure 1 It is a schematic flow chart of a method for collecting high-speed network traffic based on a cloudified environment of the present invention; Figure 2 It is a schematic module diagram of a system for collecting high-speed network traffic based on a cloudified environment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The technical solutions in the embodiments of the invention will be clearly and completely described below with reference to the drawings in the embodiments of the invention. Obviously, the described embodiments are only some embodiments of the invention, not all of them. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the invention belong to the scope of protection of the invention.
[0024] In the description of the present invention, it should be understood that the terms "opening", "upper", "lower", "top", "middle", "inner", etc. indicating the orientation or position relationship are only for the convenience of describing the invention and simplifying the description, rather than indicating or implying that the components or elements referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0025] Embodiment 1 Please refer to Figure 1 , the present invention discloses a method for collecting high-speed network traffic in a cloudified environment, including the following steps: S1. Build a cloudified environment and set relevant parameters of the cloudified environment; The S1 includes the following steps: S11. Create a virtualized resource pool in the cloud platform, build a collection node cluster based on Kubernetes; configure the elastic scaling policy of the resource pool; the elastic scaling policy of the resource pool changes according to the size of the predicted high-speed network traffic and the change of the real-time load rate; Allocate a dedicated virtual network interface to each collection node of the resource pool, and enable the SR-IOV technology to improve the network card performance; S12. Deploy an SDN controller to the cloud gateway, configure the high-speed network traffic mirroring rule, and mirror the target high-speed network traffic to the collection system; set the mirror high-speed network traffic bandwidth limit to avoid overload; for example, only mirror 50% of the network high-speed traffic for sampling analysis; S13. Configure an initial collection policy in the control center, the initial collection policy includes the high-speed network traffic distribution weight and resource threshold of each node, the high-speed network traffic distribution weight dynamically adjusts the weight according to the node latency, and the resource threshold limits the maximum and minimum numbers of the node cluster; for example, the collection node cluster retains at least 2 collection nodes and can be expanded to a maximum of 50 collection nodes; CPU usage rate ≥ 70% triggers expansion, and filtering rules, define a basic rule template, such as discarding ICMP protocol and marking abnormal IP high-speed network traffic; S14. Obtain the cloudified environment through S11, S12, and S13; S2. Mirror the real-time high-speed network traffic to the cloudified environment and perform feature extraction to obtain real-time high-speed network traffic features; The S2 includes the following steps: S21. Mirror the real-time high-speed network traffic to the cloudified environment to obtain real-time mirrored high-speed network traffic; S22. Extract the five-tuple feature dataset and application layer content features of the real-time mirrored high-speed network traffic through lightweight DPI to obtain real-time high-speed network traffic features; the five-tuple feature dataset includes source IP address, destination IP address, protocol type, source port, and destination port, and the application layer content features reflect the priority level of the high-speed network traffic data; adopt the consistent hashing algorithm to map the real-time high-speed network traffic to the collection node cluster according to the hash value of the five-tuple feature dataset; S3. Construct an FCN prediction model, improve the FCN prediction model using historical network high-speed traffic feature data and an optimization algorithm to obtain an optimized FCN network high-speed traffic prediction model, and combine real-time network high-speed traffic features to obtain the predicted network high-speed traffic; determine whether to perform the first cloud environment scaling based on the predicted network high-speed traffic to obtain the cloud environment after the first scaling. The S3 includes the following steps: S31. Construct an FCN prediction model and set the initial learning rate q of the FCN prediction model. S32. Collect the feature data of the historical network high-speed traffic of the cloud environment to obtain historical network high-speed traffic feature data. S33. Set the maximum number of training iterations and the training accuracy threshold; input the historical network high-speed traffic feature data into the FCN prediction model for training to obtain the training accuracy; adjust the learning rate of the FCN prediction model according to the training accuracy; when the training accuracy ≥ the training accuracy threshold or reaches the maximum number of training iterations, obtain the FCN network high-speed traffic prediction model. S34. Use an optimization algorithm to find the learning rate of the FCN network high-speed traffic prediction model to obtain the optimal solution; use the optimal solution as the learning rate of the FCN network high-speed traffic prediction model to obtain an optimized FCN network high-speed traffic prediction model. The step of using an optimization algorithm to find the learning rate of the FCN network high-speed traffic prediction model to obtain the optimal solution in the S34 includes the following steps: S341. Construct a sunflower population, set the size of the sunflower population to n, then the sunflower population is represented as r = {r1, r2, …, r i , …, r n}, where r i represents the i-th sunflower in the sunflower population; set the maximum number of optimization iterations. S342. According to the learning rate of the FCN network high-speed traffic prediction model, randomly set the initial height of the sunflower population to obtain the sunflower population initial height set as m = {m1, m2, …, m i , …, m n}, where m i represents the height of the i-th sunflower in the sunflower population. S343. There are resource limitations in the height growth of sunflowers. Excessive height leads to excessive energy consumption or unstable structure. Set the height gain coefficient as k1 and the height penalty coefficient as k2, and define the fitness function. The fitness function formula is as follows, ; Among them, f(m) represents the fitness function, which is used to evaluate the fitness of the sunflower at the current height m, and m represents the current height of the sunflower; S344. Perform iterative operations on the initial height set of the sunflower population. During each round of iteration, update the heights of each sunflower in the initial height set of the sunflower population according to the fitness function, and obtain the best sunflower individual and the global best sunflower in the sunflower population during each round of iteration; S345. Repeat S344. When the maximum optimization iteration number is reached, stop the iteration and take the global best sunflower as the optimal solution; S35. Input the real-time network high-speed traffic feature data into the optimized FCN network high-speed traffic prediction model to obtain the predicted network high-speed traffic; S36. Set the prediction network high-speed traffic scaling rule; according to the prediction network high-speed traffic scaling rule, the real-time mirror network high-speed traffic, and the predicted mirror network high-speed traffic, perform the first expansion, reduction, or keep unchanged on the collection node group in the cloudified environment to obtain the cloudified environment after the first scaling; the prediction network high-speed traffic scaling rule is to set the prediction network high-speed traffic expansion threshold d and the prediction network high-speed traffic reduction threshold f; when the ratio e of the predicted network high-speed traffic to the current network high-speed traffic is ≥ the prediction network high-speed traffic expansion threshold d, expand the collection node group in the cloudified environment to times of the current collection node group; when the ratio g of the predicted network high-speed traffic to the current network high-speed traffic is ≤ the prediction network high-speed traffic reduction threshold f, reduce the collection node group in the cloudified environment to times of the current collection node group; for example, when the predicted mirror network high-speed traffic data is 140% of the real-time mirror network high-speed traffic data, expand the number of collection node groups in the cloudified environment to 1.2 times the number of the current collection node group; S4. Collect the load index data of the cloudified environment after the first scaling to obtain the real-time load index and determine whether to perform the second cloudified environment scaling to obtain the cloudified environment after the second scaling; The S4 includes the following steps: S41. Collect the real-time CPU utilization rate and real-time memory occupancy rate of the collection node group in the cloudified environment after the first scaling to obtain the real-time load index; S42. Set the scaling rule for load metrics. According to the scaling rule for load metrics and the real-time load metrics, perform the second scaling of the cloudified environment or keep it unchanged to obtain the cloudified environment after the second scaling; the scaling rule for load metrics is to set the maximum CPU utilization threshold h1, the maximum memory occupancy threshold h2, the minimum CPU utilization threshold h3, and the minimum memory occupancy threshold h4; when the real-time CPU utilization k1 ≥ the maximum CPU utilization threshold h1 or the real-time memory occupancy k2 ≥ the maximum load memory occupancy threshold h2, expand the collection node group in the cloudified environment to times of the current collection node group; when the real-time CPU utilization k1 ≤ the minimum CPU utilization threshold h3 or the real-time load rate k2 ≤ the maximum load memory occupancy threshold h4, reduce the collection node group in the cloudified environment to times of the current collection node group; if the maximum CPU utilization threshold h1 and the maximum load memory occupancy threshold h2 are both 75%, the real-time CPU utilization k1 is 85%, and the memory occupancy k2 is 70%, then the number of nodes is expanded to 1.133 times of the current number of nodes; S5. Collect the network high-speed traffic packets of the cloudified environment after the second scaling and perform dynamic filtering processing to obtain a preprocessed network high-speed traffic data set; S5 includes the following steps: S51. In the node cluster of the cloudified environment after the second scaling, batch obtain network high-speed traffic data through the DPDK kernel bypass technology to obtain network high-speed traffic packets; S52. Set the dynamic filtering rule engine to discard the ICMP protocol, block the blacklist IP, and match the attack feature library; perform three-level filtering on the network high-speed traffic packets in combination with the dynamic filtering rule engine to obtain the filtered network high-speed traffic packets; S53. Extract the metadata fields of the filtered network high-speed traffic packets, and add threat level labels to the filtered network high-speed traffic packets according to the metadata fields to generate a preprocessed network high-speed traffic data set; S6. Aggregate and synchronize the preprocessed network high-speed traffic data set to obtain a standard network high-speed traffic data stream; S6 includes the following steps: S61. Set the circular buffer threshold; when the circular buffer of the cloudified environment accumulates to the circular buffer threshold, discard the data with low priority in the preprocessed network high-speed traffic data set and compress and encrypt it into the Avro format to obtain the processed network high-speed traffic data set; S62. Write the processed network high-speed traffic dataset into the Kafka message queue, partition the data in the processed network high-speed traffic dataset according to the source IP hash value, and obtain the standard network high-speed traffic data stream.
[0026] Embodiment 2 Please refer to Figure 2 , a network high-speed traffic collection system based on a cloudified environment, which is used to implement the above-mentioned network high-speed traffic collection method based on a cloudified environment, including a cloudified environment construction module, a traffic mirroring and feature processing module, an optimized FCN network high-speed traffic prediction model module, a first cloudified environment scaling decision module, a second cloudified environment scaling decision module, a dynamic filtering and threat marking module, and a data aggregation and synchronization module; The cloudified environment construction module is used to create an elastic resource pool, configure the collection node cluster and network parameters, and implement the deployment of the infrastructure in the cloudified environment; The traffic mirroring and feature processing module is used to obtain the real-time mirror network high-speed traffic and extract key features to obtain the real-time network high-speed traffic features; The optimized FCN network high-speed traffic prediction model module improves the FCN network prediction model through historical network high-speed traffic features and optimization algorithms to obtain an optimized FCN network high-speed traffic prediction model; The first cloudified environment scaling decision module is used to obtain the predicted network high-speed traffic according to the optimized FCN network high-speed traffic prediction model and the real-time network high-speed traffic features; according to the predicted network high-speed traffic and the predicted network high-speed traffic scaling rules, determine whether to perform the first cloudified environment scaling to obtain the cloudified environment after the first scaling; The second cloudified environment scaling decision module is used to determine whether to perform the second cloudified environment scaling according to the above load index scaling rules and the real-time load index to obtain the cloudified environment after the second scaling; The dynamic filtering and threat marking module is used to perform multi-level cleaning on the network high-speed traffic packets and mark the threat levels; The data aggregation and synchronization module is used to efficiently aggregate multi-node data and achieve low-latency and high-throughput standardized data stream output.
[0027] Embodiment 3 A network high-speed traffic collection device based on a cloudified environment, on which a program is stored, and when the program is executed by a processor, it is used to implement the above-mentioned network high-speed traffic collection method based on a cloudified environment.
[0028] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0029] The preferred embodiments of the invention disclosed above are only used to help explain the invention. The preferred embodiments do not exhaust all the details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the invention, so that those skilled in the art can well understand and utilize the invention.
Claims
1. A method for collecting high-speed network traffic based on a cloudified environment, characterized in that It includes the following steps: S1. Build a cloudified environment and set relevant parameters of the cloudified environment; S2. Mirror the real-time network high-speed traffic to the cloudified environment and perform feature extraction to obtain real-time network high-speed traffic features; S3. Build an FCN prediction model, improve the FCN prediction model using historical network high-speed traffic feature data and an optimization algorithm to obtain an optimized FCN network high-speed traffic prediction model, and combine it with real-time network high-speed traffic features to obtain predicted network high-speed traffic; Decide whether to perform the first cloudified environment scaling based on the predicted network high-speed traffic to obtain the cloudified environment after the first scaling; S4. Collect the load index data of the cloudified environment after the first scaling to obtain real-time load indexes and determine whether to perform the second cloudified environment scaling to obtain the cloudified environment after the second scaling; S5. Collect the network high-speed traffic packets of the cloudified environment after the second scaling and perform dynamic filtering processing to obtain a preprocessed network high-speed traffic data set; S6. Aggregate and synchronize the preprocessed network high-speed traffic data set to obtain a standard network high-speed traffic data stream.
2. The network high-speed traffic collection method based on a cloudified environment according to claim 1, characterized in that, The S1 includes the following steps: S11. Create a virtualized resource pool in the cloud platform, build a collection node cluster based on Kubernetes; Configure the elastic scaling policy of the resource pool; The elastic scaling policy of the resource pool changes according to the size of the predicted network high-speed traffic and the change of the real-time load rate; Allocate a dedicated virtual network interface to each collection node of the resource pool and enable the SR-IOV technology to improve the network card performance; S12. Deploy an SDN controller to the cloud gateway, configure the network high-speed traffic mirroring rule, mirror the target network high-speed traffic to the collection system, and set the mirror network high-speed traffic bandwidth limit; S13. Configure an initial collection policy in the control center, where the initial collection policy includes the network high-speed traffic distribution weight and resource threshold of each node. The network high-speed traffic distribution weight dynamically adjusts the weight according to the node latency, and the resource threshold limits the maximum and minimum numbers of the node cluster; S14. Obtain the cloudified environment through S11, S12, and S13.
3. A network high-speed traffic collection method based on a cloudified environment according to claim 1, characterized in that The S2 includes the following steps: S21. Mirror the real-time network high-speed traffic to the cloudified environment to obtain real-time mirrored network high-speed traffic; S22. Extract the five-tuple feature data set and application layer content features of the real-time mirrored network high-speed traffic through lightweight DPI to obtain real-time network high-speed traffic features; The five-tuple feature data set includes the source IP address, destination IP address, protocol type, source port, and destination port. The application layer content features reflect the priority level of the network high-speed traffic data; Use the consistent hashing algorithm to map the real-time network high-speed traffic to the collection node cluster according to the hash value of the five-tuple feature data set.
4. A network high-speed traffic collection method based on a cloudified environment according to claim 1, characterized in that, The S3 includes the following steps: S31. Construct an FCN prediction model and set the initial learning rate of the FCN prediction model q ; S32. Collect the feature data of the historical network high-speed traffic of the cloudified environment to obtain historical network high-speed traffic feature data; S33. Set the maximum number of training iterations and the training accuracy threshold; input the historical network high-speed traffic feature data into the FCN prediction model for training to obtain the training accuracy; adjust the learning rate of the FCN prediction model according to the training accuracy; when the training accuracy ≥ the training accuracy threshold or the maximum number of training iterations is reached, obtain the FCN network high-speed traffic prediction model; S34. Use an optimization algorithm to find the learning rate of the FCN network high-speed traffic prediction model to obtain the optimal solution; use the optimal solution as the learning rate of the FCN network high-speed traffic prediction model to obtain the optimized FCN network high-speed traffic prediction model; S35. Input the real-time network high-speed traffic feature data into the optimized FCN network high-speed traffic prediction model to obtain the predicted network high-speed traffic; S36. Set the prediction network high-speed traffic scaling rule. According to the prediction network high-speed traffic scaling rule, the real-time network high-speed traffic, and the predicted network high-speed traffic, perform the first cloud environment scaling or keep it unchanged to obtain the cloud environment after the first scaling; the prediction network high-speed traffic scaling rule is to set the predicted network high-speed traffic expansion threshold d , the predicted network high-speed traffic reduction threshold f ; when the ratio of the predicted network high-speed traffic to the current network high-speed traffic e ≥ the predicted network high-speed traffic expansion threshold d , expand the collection node group in the cloud environment to times the current collection node group; when the ratio of the predicted network high-speed traffic to the current network high-speed traffic g ≤ the predicted network high-speed traffic reduction threshold f , reduce the collection node group in the cloud environment to times the current collection node group.
5. A method for collecting high-speed network traffic based on a cloudified environment according to claim 4, characterized in that The step of using an optimization algorithm to find the learning rate of the FCN network high-speed traffic prediction model to obtain the optimal solution in S34 includes the following steps: S341. Construct a sunflower population; S342. According to the learning rate of the FCN network high-speed traffic prediction model, randomly set the initial height of the sunflower population to obtain the initial height set of the sunflower population; S343. There are resource limitations in the height growth of sunflowers. Excessive height leads to excessive energy consumption or unstable structure. Set the height gain coefficient and height penalty coefficient, and define the fitness function; S344. Perform iterative operations on the initial height set of the sunflower population. In each round of iteration, update the height of each sunflower in the initial height set of the sunflower population according to the fitness function, and obtain the best sunflower individual and the global best sunflower in the sunflower population in each round of iteration; S345. Repeat S344. When the maximum number of optimization iterations is reached, stop the iteration and use the global best sunflower as the optimal solution.
6. The network high-speed traffic collection method based on a cloudified environment according to claim 1, wherein The S4 includes the following steps: S41. Collect the real-time CPU utilization rate and real-time memory occupancy rate of the collection node group in the cloudified environment after the first scaling, to obtain the real-time load indicators; S42. Set the rules for scaling the cloud environment based on load metrics. According to the rules for scaling the cloud environment based on load metrics and the real-time load metrics, perform the second scaling of the cloud environment (either expand or keep it unchanged) to obtain the cloud environment after the second scaling. The rules for scaling the cloud environment based on load metrics are to set the maximum CPU utilization threshold h 1. The maximum memory occupancy threshold h 2. The minimum CPU utilization threshold h 3 and the minimum memory occupancy threshold h 4; When the real-time CPU utilization k 1 ≥ the maximum CPU utilization threshold h 1 or the real-time memory occupancy k 2 ≥ the maximum load memory occupancy threshold h 2, expand the collection node group in the cloud environment to times the current collection node group; When the real-time CPU utilization k 1 ≤ the minimum CPU utilization threshold h 3 or the real-time load rate k 2 ≤ the maximum load memory occupancy threshold h 4, reduce the collection node group in the cloud environment to times the current collection node group.
7. A network high-speed traffic collection method based on a cloudified environment according to claim 1, characterized in that, The S5 includes the following steps: S51. In the node cluster of the cloudified environment after the second scaling, batch obtain network high-speed traffic data through the DPDK kernel bypass technology to obtain network high-speed traffic packets; S52. Set the dynamic filtering rule engine to discard the ICMP protocol, block the blacklist IP, and match the attack feature library; perform three-level filtering on the network high-speed traffic packets in combination with the dynamic filtering rule engine to obtain the filtered network high-speed traffic packets; S53. Extract the metadata fields of the filtered network high-speed traffic packets, and add threat level labels to the filtered network high-speed traffic packets according to the metadata fields to generate a preprocessed network high-speed traffic data set.
8. A method for collecting high-speed network traffic based on a cloudified environment according to claim 1, characterized in that, The S6 includes the following steps: S61. Set the circular buffer threshold; when the circular buffer in the cloudified environment accumulates to the circular buffer threshold, discard the data with low priority in the preprocessed network high-speed traffic data set, and compress and encrypt it into the Avro format to obtain the processed network high-speed traffic data set; S62. Write the processed network high-speed traffic dataset into the Kafka message queue. The data in the processed network high-speed traffic dataset is partitioned by the source IP hash value to obtain a standard network high-speed traffic data stream.
9. A network high-speed traffic acquisition system based on a cloudified environment, which is used to implement a network high-speed traffic acquisition method based on a cloudified environment as described in any one of claims 1-8. The system includes a cloudified environment construction module, a traffic mirroring and feature processing module, an optimized FCN network high-speed traffic prediction model module, a first cloudified environment scaling decision module, a second cloudified environment scaling decision module, a dynamic filtering and threat marking module, and a data aggregation and synchronization module.
10. A network high-speed traffic collection device based on a cloudified environment, characterized in that, A program is stored thereon, and when the program is executed by a processor, it implements a network high-speed traffic acquisition method based on a cloudified environment as described in any one of claims 1-8.
Citation Information
Patent Citations
A network traffic collection method and system based on cloud environment
CN117294533B
Network traffic acquisition method and system based on clouded environment
CN117294533A
Hybrid elastic scaling method based on multi-dimensional resource prediction
CN117827434A
Network resource scheduling optimization system of cloud environment
CN117931424A
Cloud collaboration container elastic expansion and contraction method and device and electronic equipment
CN118041787A
Cited By
Cloud desktop resource pool dynamic prediction method, system, device, medium and product
CN120929275A