Traffic monitoring method, system, device, computer device, readable storage medium and program product
By acquiring, aggregating, and statistically analyzing traffic characteristic data of cloud-native WAF nodes, a traffic monitoring view is constructed, solving the problem of traffic monitoring in cloud-native WAF systems and enabling real-time multi-dimensional monitoring and efficient operation. This approach is suitable for cloud-native WAF systems with distributed architectures.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA TELECOM CLOUD TECH CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-19
AI Technical Summary
How to monitor traffic in a cloud-native WAF system, detect anomalies in a timely manner, and better operate the cloud-native WAF system and serve customers in the face of the challenges of massive network traffic.
For each target cloud-native network application firewall (WAF) node, traffic characteristic data is acquired, aggregated and stored, statistically analyzed according to a preset time dimension, and a traffic monitoring view is constructed. The distributed data acquisition module, data aggregation center and distributed data processing module are used to process and store the data, and the traffic monitoring view is constructed to realize the monitoring of the cloud-native WAF system.
It enables real-time, multi-dimensional traffic monitoring of cloud-native WAF systems, allowing for targeted monitoring of traffic across dimensions such as tenants, domains, and source IPs. This improves operational efficiency, provides better real-time monitoring and solutions for complex attacks, and has low resource consumption without affecting the core forwarding and detection logic of the WAF.
Smart Images

Figure CN119561761B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing technology, and in particular to a traffic monitoring method, system, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Technology
[0002] Currently, with the rapid development of cloud computing technology, more and more enterprises are migrating their businesses to cloud platforms. Consequently, cybersecurity threats are also increasing, especially application-level attacks such as Structured Query Language (SQL) injection and cross-site scripting (XSS), which seriously threaten enterprise data security. A Web Application Firewall (WAF) is a specific type of application firewall used to filter, monitor, and block Hypertext Transfer Protocol (HTTP) traffic over web services. By monitoring HTTP traffic, it can prevent attacks that exploit known vulnerabilities in web applications, such as SQL injection and XSS. Therefore, cloud-native WAF systems are widely used.
[0003] Faced with massive amounts of network traffic, how to monitor and analyze the traffic, promptly detect anomalies, and better operate cloud-native WAF systems and serve customers has become a pressing issue for current cloud-native WAF systems. Therefore, a traffic monitoring method capable of monitoring the traffic of cloud-native WAF systems is urgently needed. Summary of the Invention
[0004] Therefore, it is necessary to provide a traffic monitoring method, system, device, computer equipment, computer-readable storage medium, and computer program product that can monitor traffic in cloud-native WAF systems in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a traffic monitoring method, including:
[0006] For each target cloud-native network application firewall (WAF) node, obtain the traffic characteristic data of the target cloud-native WAF node;
[0007] Aggregate and store the traffic characteristic data of each of the target cloud-native WAF nodes;
[0008] According to the preset time dimension, the traffic characteristic data of each target cloud-native WAF node under each traffic characteristic data type dimension is statistically analyzed, and the statistical results are stored in the target database;
[0009] Based on the data stored in the target database, the pre-determined traffic monitoring dimensions, and the preset traffic monitoring view construction rules, a traffic monitoring view is constructed; the traffic monitoring view is used to monitor the cloud-native WAF system.
[0010] In one embodiment, the step of statistically analyzing the traffic characteristic data of each target cloud-native WAF node under each traffic characteristic data type dimension according to a preset time dimension, and storing the statistical results in the target database, includes:
[0011] Based on the current business volume of each target cloud-native WAF node, enable multiple target consumers;
[0012] For each target consumer, the traffic feature data received by the target consumer is written into the channel of the target consumer corresponding to each traffic feature data type dimension.
[0013] For each traffic feature data type dimension, the traffic feature data in the channel of the target consumer corresponding to the traffic feature data type dimension is statistically analyzed according to a preset time dimension through the coroutine of the target consumer.
[0014] The statistically analyzed traffic feature data is written into the write channel of the target consumer corresponding to the traffic feature data type dimension;
[0015] If the preset database write conditions are met, the statistically analyzed traffic characteristic data in the write channel will be written into the target database.
[0016] In one embodiment, the method further includes:
[0017] When a preset first time node is reached, the target historical data of the cloud-native WAF system is statistically analyzed and recorded to a temporary file; the target historical data includes the historical request volume, historical attack volume, and the historical number of each status code; a first time interval is between two adjacent first time nodes;
[0018] When the preset second time node is reached, the target historical data between the previous second time node and the current second time node stored in the temporary file is written into the target database; a second time period is spaced between two adjacent second time nodes; the duration of the second time period is an integer multiple of the duration of the first time period.
[0019] In one embodiment, writing the target historical data from the previous second time node to the current second time node stored in the temporary file into the target database includes:
[0020] If a new status code is added when writing to the target database, the database table field corresponding to the status code is modified according to the range of the status code and the new status code.
[0021] Based on the modified database table fields, the historical number of each status code stored in the temporary file between the previous second time node and the current second time node is split into a preset number of status code data tables.
[0022] Write the status code data table and the total request data table into the target database.
[0023] In one embodiment, the method further includes:
[0024] If the current time and the last record time of the temporary file are in different first time periods but in the same second time period, then the time period between the current time and the last record time is taken as the first missing time period.
[0025] The target historical data of the cloud-native WAF system within the first missing time period is statistically analyzed, and the target historical data within the first missing time period is recorded to the temporary file.
[0026] In one embodiment, the method further includes:
[0027] If the current time and the last record time of the temporary file are in different second time periods, then obtain the earliest retention time of the target historical data retained in the target database;
[0028] The latest of the last recorded time and the earliest retained time is used as the start time for statistics;
[0029] The time period between the start time of statistics and the current time is taken as the second missing time period. Starting from the start time of statistics, the target historical data of the cloud-native WAF system within the second missing time period is statistically processed according to the first time node and the second time node.
[0030] Secondly, this application also provides a traffic monitoring system, including:
[0031] At least one distributed data acquisition module is used to collect traffic characteristic data of each target cloud-native network application firewall (WAF) node; each target cloud-native WAF node includes one of the distributed data acquisition modules.
[0032] The data aggregation center is used to aggregate and store the traffic characteristic data of each of the target cloud-native WAF nodes;
[0033] The distributed data processing module is used to statistically analyze the traffic characteristic data of each target cloud-native WAF node under each traffic characteristic data type dimension according to a preset time dimension, and store the statistical results in the target database.
[0034] The storage service module is used to construct a traffic monitoring view based on the data stored in the target database, the pre-determined traffic monitoring dimensions, and the preset traffic monitoring view construction rules; the traffic monitoring view is used to monitor the cloud-native WAF system.
[0035] Thirdly, this application also provides a traffic monitoring device, comprising:
[0036] The first acquisition module is used to acquire traffic characteristic data of each target cloud-native network application firewall (WAF) node.
[0037] The aggregation module is used to aggregate and store the traffic characteristic data of each of the target cloud-native WAF nodes;
[0038] The first statistics module is used to collect traffic characteristic data of each target cloud-native WAF node under each traffic characteristic data type dimension according to a preset time dimension, and store the statistical results in the target database.
[0039] The construction module is used to construct a traffic monitoring view based on the data stored in the target database, the pre-determined traffic monitoring dimensions, and the preset traffic monitoring view construction rules; the traffic monitoring view is used to monitor the cloud-native WAF system.
[0040] Fourthly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps described in the first aspect above.
[0041] Fifthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps described in the first aspect above.
[0042] Sixthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps described in the first aspect above.
[0043] The aforementioned traffic monitoring method, system, device, computer equipment, computer-readable storage medium, and computer program product acquire traffic characteristic data of each target cloud-native network application firewall (WAF) node; aggregate and store the traffic characteristic data of each target cloud-native WAF node; statistically analyze the traffic characteristic data of each target cloud-native WAF node under each traffic characteristic data type dimension according to a preset time dimension, and store the statistical results in a target database; construct a traffic monitoring view based on the data stored in the target database, a predetermined traffic monitoring dimension, and preset traffic monitoring view construction rules; the traffic monitoring view is used to monitor the cloud-native WAF system. In this way, by acquiring the traffic characteristic data of each target cloud-native WAF node, aggregating and storing it, and statistically analyzing it according to a preset time dimension and each traffic characteristic data type dimension, a traffic monitoring view for monitoring the cloud-native WAF system is constructed, enabling traffic monitoring of the cloud-native WAF system. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 This is a schematic diagram of a traffic monitoring system in one embodiment;
[0046] Figure 2 This is a flowchart illustrating a traffic monitoring method in one embodiment;
[0047] Figure 3 This is a flowchart illustrating the steps of statistically analyzing the traffic characteristic data of each target cloud-native WAF node under each traffic characteristic data type dimension according to a preset time dimension, and storing the statistical results in the target database in one embodiment.
[0048] Figure 4 A schematic diagram illustrating the use of multiple consumers and multiple coroutines to process traffic characteristic data in the distributed data processing module 130.
[0049] Figure 5 This is a flowchart illustrating the steps that the traffic monitoring method may further include in one embodiment;
[0050] Figure 6 This is a flowchart illustrating the step of writing target historical data from the previous second time node to the current second time node stored in a temporary file into the target database in one embodiment.
[0051] Figure 7 This is a flowchart illustrating the steps that the traffic monitoring method further includes in another embodiment;
[0052] Figure 8 This is a flowchart illustrating the steps that the traffic monitoring method further includes in another embodiment;
[0053] Figure 9 This is a structural block diagram of a flow monitoring device in one embodiment;
[0054] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0056] The traffic monitoring method provided in this application embodiment can be applied to, for example, Figure 1 The traffic monitoring system 100 shown includes: at least one distributed data acquisition module 110, a data aggregation center 120, a distributed data processing module 130, and a storage service module 140. Wherein:
[0057] At least one distributed data acquisition module 110 is used to collect traffic characteristic data of each target cloud-native WAF node; each target cloud-native WAF node includes a distributed data acquisition module.
[0058] Data aggregation center 120 is used to aggregate and store traffic characteristic data of each target cloud-native WAF node;
[0059] The distributed data processing module 130 is used to statistically analyze the traffic characteristic data of each target cloud-native WAF node under each traffic characteristic data type dimension according to a preset time dimension, and store the statistical results in the target database.
[0060] The storage service module 140 is used to build a traffic monitoring view based on the data stored in the target database, the pre-determined traffic monitoring dimensions, and the preset traffic monitoring view construction rules; the traffic monitoring view is used to monitor the cloud-native WAF system.
[0061] In this embodiment, the traffic monitoring system 100 adopts a distributed architecture of a data processing center and multiple edge acquisition nodes, with the data acquisition process and the WAF service processing process residing on the same WAF node. This allows for smooth expansion of monitoring based on the expansion of WAF product services, enhancing scalability; monitoring begins as soon as the WAF node starts processing services, improving real-time monitoring operational efficiency.
[0062] Both the distributed data acquisition module 110 and the distributed data processing module 130 can be implemented using Go. After development, they are compiled into executable files and do not depend on the runtime environment. In this way, the logic code of the traffic monitoring system is written in Go, which makes it easier to support multiple underlying architectures, such as x86 and ARM architectures. Only the corresponding architecture version needs to be specified during compilation, which meets the requirements of WAF deployment on multiple architectures.
[0063] The distributed data acquisition module 110 uses multiple different coroutines to periodically acquire, process, and report traffic characteristic information such as tenant ID, domain name, source IP, request time, request size, request processing time, and status code from the access logs and attack logs generated by the current target cloud-native WAF node.
[0064] The data aggregation center 120 can provide a remote dictionary server (Redis). Redis is a high-performance data storage service that, leveraging its stream characteristics, can process collected data sequentially while simultaneously supporting multiple consumers processing different data, thus providing strong data processing capabilities. Access data and attack data can be separated by setting different consumer groups and streams, awaiting consumption and processing of the aggregated data by the distributed data processing module 130.
[0065] The distributed data processing module 130 supports multiple consumers processing data simultaneously. This accelerates data processing and prevents data congestion at the data aggregation center. The distributed data processing module 130 employs multiple coroutines, with each consumer using different coroutines for concurrent processing of each monitoring dimension. This improves processing efficiency.
[0066] Storage service module 140 provides a query service interface. This interface is used for data display and analysis. Users can use the query service interface to query traffic monitoring views across various traffic monitoring dimensions.
[0067] In one example, the resource usage of the running distributed data acquisition module 110 is limited by systemd (System daemon) to control the number of Central Processing Units (CPU), memory, and file handles used. This way, by limiting the CPU, memory, and file handle usage of the distributed data acquisition module process through systemd, even in extreme cases, the monitoring service will not affect the WAF product service.
[0068] Understandably, this method can also be applied to terminals or servers, and further to systems that include both terminals and servers, implemented through interaction between them. Terminals can be, but are not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. Servers can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing cloud computing services.
[0069] In one exemplary embodiment, such as Figure 2 As shown, a traffic monitoring method is provided, which can be applied to... Figure 1 Taking the traffic monitoring system 100 in the example, this method includes the following steps:
[0070] Step 201: For each target cloud-native network application firewall (WAF) node, obtain the traffic characteristic data of that target cloud-native WAF node.
[0071] In this embodiment, the cloud-native Web Application Firewall (WAF) system includes at least one target cloud-native WAF node. The cloud-native WAF system may include multiple target cloud-native WAF nodes. The cloud-native WAF system is the system to be monitored for traffic. The target cloud-native WAF nodes are the cloud-native WAF nodes to be monitored for traffic. The cloud-native WAF nodes perform WAF service processing. Traffic characteristic data is used to define the traffic characteristics of the cloud-native WAF system, and may include at least one of the following: tenant identity document (ID), domain name, source Internet Protocol address (IP), request time, request size, request processing time, and status code; and may also include at least one of the following: request volume, attack volume, request length, and response length.
[0072] For each target cloud-native web application firewall (WAF) node, the distributed data acquisition module 110 within that target WAF node collects its traffic logs. Then, the distributed data acquisition module 110 parses the traffic logs to obtain the traffic characteristic data of the target WAF node. Finally, the distributed data acquisition module 110 sends the traffic characteristic data to the data aggregation center 120. The traffic logs are those generated by the WAF engine, including access logs and attack logs.
[0073] In one example, the distributed data acquisition module 110 reads and processes data line by line, using different regular expressions to extract traffic characteristic information from access logs and attack logs. If the log format meets preset format conditions, the distributed data acquisition module 110 sends the traffic characteristic information from the logs to the data aggregation center 120.
[0074] Step 202: Aggregate and store the traffic characteristic data of each target cloud-native WAF node.
[0075] In this embodiment, the data aggregation center 120 aggregates traffic characteristic data from each target cloud-native WAF node. Then, the data aggregation center 120 uses a remote dictionary server (Redis) to store the traffic characteristic data of each target cloud-native WAF node, and uses different streams to distinguish the traffic characteristic data in the access logs from the traffic characteristic data in the attack logs.
[0076] In one example, the data aggregation center 120 can also mount traffic characteristic data to the host machine. This helps prevent data loss.
[0077] Step 203: According to the preset time dimension, statistically analyze the traffic characteristic data of each target cloud-native WAF node under each traffic characteristic data type dimension, and store the statistical results in the target database.
[0078] In this embodiment, the distributed data processing module 130 receives traffic feature data from access logs and attack logs from different streams. Then, the distributed data processing module 130 statistically analyzes the traffic feature data of each target cloud-native WAF node under each traffic feature data type dimension according to a preset time dimension, and stores the statistical results in the target database. The time dimension can be one or more, including at least one of second-level, minute-level, and ten-minute-level. The traffic feature data type refers to the type of traffic feature data; the aforementioned tenant ID, domain name, source IP, request time, request size, request processing time, status code, request volume, attack volume, request length, and response length are all traffic feature data types. The traffic feature data type dimension refers to the dimension of the traffic feature data type. For example, the traffic feature data type dimension can include tenant ID dimension, domain name dimension, source IP dimension, request time dimension, request size dimension, request processing time dimension, and status code dimension. The target database is the database that stores the statistically analyzed traffic feature information. For example, the target database is a PostgreSQL database.
[0079] Step 204: Construct a traffic monitoring view based on the data stored in the target database, the pre-determined traffic monitoring dimensions, and the preset traffic monitoring view construction rules.
[0080] The traffic monitoring view is used to monitor cloud-native WAF systems.
[0081] In this embodiment, the storage service module 140 constructs a traffic monitoring view based on the data stored in the target database, pre-determined traffic monitoring dimensions, and preset traffic monitoring view construction rules. The data stored in the target database includes statistical results and may also include other data. Traffic monitoring dimensions may include at least one of a time dimension and various traffic characteristic data type dimensions, and may also include other dimensions. Other dimensions may be historical data dimensions. The traffic monitoring view may include real-time sub-traffic monitoring views for dimensions such as tenant, domain name, source IP, and status code, within a preset query period, including request volume, attack volume, request length, response length, request latency, and attack latency. It may also include historical sub-traffic monitoring views for dimensions such as historical request volume, historical attack volume, and historical data for various status codes. The traffic monitoring view may include at least one of the following: sub-traffic monitoring views for the top 10 tenants, domain names, and source IPs by access volume; sub-traffic monitoring views for the top 10 dimensions of attack volume, request size, request time, and status code; and sub-traffic monitoring views for the historical processing request volume, attack volume, and status codes of the cloud-native WAF system.
[0082] In one embodiment, request latency and attack latency are averages. For example, request latency, which is the average of request processing times, can be expressed as:
[0083]
[0084] Where Req_avg_delay is the request delay. This is the current request processing time. The number of requests.
[0085] The aforementioned traffic monitoring method acquires traffic characteristic data from each target cloud-native WAF node, aggregates and stores it, and performs statistical analysis according to preset time and data type dimensions of traffic characteristics. This constructs a traffic monitoring view for monitoring the cloud-native WAF system, enabling traffic monitoring of the cloud-native WAF system. Furthermore, this method can perform real-time multi-dimensional traffic monitoring for cloud-native WAF products. It provides a more targeted real-time multi-dimensional traffic monitoring system for dimensions that cloud-native WAF products need to focus on, such as tenants, domains, source IPs, and response status codes. This helps the WAF better monitor the traffic status of each tenant, domain, and source IP in real time, facilitating better operation and providing users with better and more specific WAF usage suggestions. It also offers a method for dealing with complex attacks and real-time monitoring of network traffic issues. Moreover, this method adopts a distributed data collection node architecture, resulting in low resource consumption; it only sends a small amount of data specific to cloud-native WAF product monitoring and reporting, such as tenant ID, domain name, and source IP; the central processing and analysis are isolated from the WAF nodes, so as not to affect the core forwarding and detection logic of WAF; by collecting and analyzing WAF logs, it enables operations and R&D personnel to monitor the current status of each user, domain name, IP, and status code of WAF, thereby helping WAF products to operate better and discover and handle problems.
[0086] In one exemplary embodiment, such as Figure 3 As shown, the specific process of statistically analyzing the traffic characteristic data of each target cloud-native WAF node under each traffic characteristic data type dimension according to a preset time dimension, and storing the statistical results in the target database includes the following steps:
[0087] Step 301: Enable multiple target consumers based on the current business volume of each target cloud-native WAF node.
[0088] In this embodiment, the distributed data processing module 130 determines the number of target consumers based on the current business volume of each target cloud-native WAF node and activates that number of target consumers. The target consumers are those activated to process data in parallel.
[0089] Step 302: For each target consumer, the traffic feature data received by the target consumer is written into the channel of the target consumer corresponding to each traffic feature data type dimension.
[0090] In this embodiment, for each target consumer, the distributed data processing module 130 writes the traffic feature data received by the target consumer into the corresponding channel of that target consumer for each traffic feature data type dimension. Specifically, for each target consumer, the target consumer writes the traffic feature data received by the target consumer into the corresponding channel of that target consumer for each traffic feature data type dimension. Each target consumer corresponds to a set of channels, and each traffic feature data type dimension corresponds to one channel in that set of channels. The number of channels corresponding to each target consumer can be the same as the number of traffic feature data type dimensions.
[0091] Step 303: For each traffic feature data type dimension, through the coroutine of the target consumer corresponding to the traffic feature data type dimension, according to the preset time dimension, the traffic feature data in the channel of the target consumer corresponding to the traffic feature data type dimension is statistically analyzed.
[0092] In this embodiment, each traffic feature data type dimension corresponds to a coroutine for a target consumer. The number of coroutines corresponding to each target consumer can be the same as the number of traffic feature data type dimensions. Channels and coroutines for the same target consumer can correspond one-to-one.
[0093] Step 304: Write the statistically analyzed traffic feature data into the write channel of the target consumer corresponding to the data type dimension of the traffic feature.
[0094] In this embodiment, each traffic feature data type dimension corresponds to a writer channel for a target consumer. The number of writer channels for each target consumer can be the same as the number of traffic feature data type dimensions. Channels and writer channels for the same target consumer can correspond one-to-one.
[0095] Step 305: If the preset database writing conditions are met, the statistical traffic characteristic data in the writing channel is written to the target database.
[0096] In this embodiment, the database write condition can be either that the amount of data to be written in the write channel exceeds a preset threshold for the amount of data to be written, or that a preset write interval has been reached. The threshold for the amount of data to be written can be 100 records. The write interval can be 1 minute.
[0097] In one example, if the amount of data to be written into the database in the writing channel is greater than the preset threshold for the amount of data to be written into the database, or reaches the preset entry interval time, the distributed data processing module 130 will write the statistical traffic characteristic data in the writing channel into the target database.
[0098] In one embodiment, a schematic diagram of the distributed data processing module 130 using multiple consumers and multiple coroutines to process traffic characteristic data is shown below. Figure 4 As shown. The target consumers include Consumer 1 and Consumer 2. Consumer 1 and Consumer 2 receive traffic characteristic data (access logs and attack logs) from different streams. Consumer 1 writes traffic characteristic data of interest, such as tenant, domain name, and source IP, into different channels (chan). Different goroutines process the data in their respective channels, and after processing and statistics, write the data into different writer channels (writers). When the number of data to be written into a writer channel exceeds a set threshold (default 100 records), or reaches a default time (default one minute), the data in the writer channel is written into the target database (DB). The specific process of Consumer 2 processing traffic characteristic data is the same as that of Consumer 1.
[0099] In the aforementioned traffic monitoring method, multiple target consumers are activated based on the current business volume of each target cloud-native WAF node. For each target consumer, the traffic feature data received by that consumer is written into the channel corresponding to each traffic feature data type dimension. For each traffic feature data type dimension, the traffic feature data in the channel corresponding to that target consumer is statistically analyzed according to a preset time dimension using the coroutine of that target consumer. The statistically analyzed traffic feature data is then written into the write channel of that target consumer. If preset database write conditions are met, the statistically analyzed traffic feature data in the write channel is written into the target database. This method employs a multi-consumer, multi-coroutine architecture to process data from the data aggregation center in parallel. Data in a stream can be processed by multiple different consumers, and different independent coroutines are used for each monitoring focus dimension. This approach is more suitable for the distributed architecture of cloud-native WAF systems and can improve data processing efficiency. Furthermore, the distributed data processing module 130, acting as a data processing center, supports multiple consumers to process collected data simultaneously. As the volume of WAF business increases or decreases, the number of consumers can be gradually increased or decreased, smoothly expanding or shrinking to support WAF business monitoring and supporting relatively convenient business expansion. The multi-consumer, multi-coroutine architecture adopted in this method can also easily realize the expansion of new monitoring dimensions, further enhancing scalability.
[0100] In one embodiment, the method further includes the following steps: checking the number of unconsumed data in the stream at preset check time intervals; if the number of unconsumed data exceeds a preset first unconsumed data threshold, determining that the stream has a message backlog and printing an error log; if the number of unconsumed data exceeds a preset second unconsumed data threshold, outputting an alarm message containing the number of unconsumed data; and determining a consumer increase indication message based on the alarm message. The check time interval can be 5 seconds. The consumer increase indication message indicates whether to increase the number of consumers. For example, a traffic monitoring system linked to an alarm system can output an alarm message containing the number of unconsumed data by sending a message via WeChat.
[0101] In one exemplary embodiment, such as Figure 5 As shown, the method also includes the following steps:
[0102] Step 501: When the preset first time node is reached, collect the target historical data of the cloud-native WAF system and record the target historical data to a temporary file.
[0103] The target historical data includes historical request volume, historical attack volume, and the historical number of each status code. A first time interval separates two adjacent first time nodes.
[0104] In this embodiment, the first time node is a preset set of multiple time points. For example, the first time node can be a minute-by-minute time point, and correspondingly, the first time period is 1 minute. Temporary files are used for caching and temporarily recording data.
[0105] Step 502: When the preset second time node is reached, the target historical data between the previous second time node and the current second time node stored in the temporary file is written into the target database.
[0106] There is a second time interval between two adjacent second time points. The duration of the second time interval is an integer multiple of the duration of the first time interval.
[0107] In this embodiment, the second time node is a preset set of multiple time points. In one example, the first time node is a minute-by-minute time point, and the first time period is 1 minute; the second time node is a ten-minute time point, and the second time period is 10 minutes. For example, when the second time node is a ten-minute time point, the second time node can be 0:00, 0:10, 0:20, ..., 23:50, or it can be 0:05, 0:15, 0:25, ..., 23:55.
[0108] In the aforementioned traffic monitoring method, when a preset first time node is reached, the target historical data of the cloud-native WAF system is statistically analyzed and recorded to a temporary file. When a preset second time node is reached, the target historical data stored in the temporary file between the previous second time node and the current second time node is written to the target database. This way, by recording the number of accesses, attacks, and response status codes within the current second time period at each second time interval, the growth of the cloud-native WAF service and the current request volume can be displayed while also considering database performance. Furthermore, data statistics are executed once at each first time interval, and the statistical results are recorded in a temporary file after execution. After multiple statistical analyses, the statistical data is stored in the database, which avoids excessive database resource consumption and balances database processing capacity. Moreover, when summarizing and recording historical data, multiple statistical analyses and multiple table records are used to record historical access data, ensuring the accurate recording of historical data. In addition, historical request data can reflect changes in the WAF product's business volume, requests, attacks, and status codes over a long period, improving the comprehensiveness of traffic monitoring.
[0109] In one exemplary embodiment, such as Figure 6As shown, the specific process of writing the target historical data from the previous second time node to the current second time node stored in the temporary file into the target database includes the following steps:
[0110] Step 601: If a new status code exists when writing to the target database, modify the database table field corresponding to the status code according to the range of the status code and the new status code.
[0111] In this embodiment, the newly added status code is a status code that has never appeared in the target database before. The status code ranges from 100 to 599.
[0112] Step 602: Based on the modified database table fields, split the historical number of status codes stored in the temporary file between the previous second time node and the current second time node into a preset number of status code data tables.
[0113] In this embodiment, the preset quantity can be 5. The status code data table, also known as a dictionary, can be divided into five tables: 100-199, 200-299, 300-399, 400-499, and 500-599.
[0114] Step 603: Write the status code data table and the total request data table into the target database.
[0115] In this embodiment, the distributed data processing module 130 submits all status code data tables and the total request data table as a single transaction to the target database. Specifically, at the start of the transaction, the distributed data processing module 130 dynamically constructs data insertion commands based on column names and statistical values, and executes the write code until all status code data tables have been processed. Then, the distributed data processing module 130 commits the transaction and determines whether there are any errors in the transaction execution. If errors exist, the distributed data processing module 130 performs a rollback and issues an alarm.
[0116] In the aforementioned traffic monitoring method, if a new status code is added when writing to the target database, the database table fields corresponding to the status code are modified according to the range of the status code and the newly added status code. Based on the modified database table fields, the historical number of each status code between the previous second time node and the current second time node, stored in the temporary file, is split into a preset number of status code data tables. Each status code data table and the total request data table are then written to the target database. This approach, considering the real-time addition of status codes and dynamically modifying the table structure, and considering the large volume of status code data, splitting the status code data tables further ensures the accurate recording of historical data.
[0117] In one exemplary embodiment, such as Figure 7As shown, the method also includes the following steps:
[0118] Step 701: If the current time and the last record time of the temporary file are in different first time periods but in the same second time period, then the time period between the current time and the last record time is taken as the first missing time period.
[0119] In this embodiment, when the first time node is a minute interval and the second time node is 0:00, 0:10, 0:20, ..., 23:50, if the current time is 0:01 and the last record time of the temporary file is 0:09, then the current time and the last record time of the temporary file are in different first time periods but in the same second time period. If the current time is 0:09 and the last record time of the temporary file is 0:11, then the current time and the last record time of the temporary file are in different first time periods and in different second time periods. If the current time is 0:01:12 and the last record time of the temporary file is 0:01:34, then the current time and the last record time of the temporary file are in the same first time period and in the same second time period.
[0120] Step 702: Calculate the target historical data of the cloud-native WAF system within the first missing time period, and record the target historical data within the first missing time period to a temporary file.
[0121] In this embodiment, the distributed data processing module 130 statistically analyzes the target historical data of the cloud-native WAF system within a first missing time period. Then, the distributed data processing module 130 records the target historical data within the first missing time period to a temporary file.
[0122] In the above traffic monitoring method, in the scenario where the distributed data processing module 130 is shut down and then restarted, if the current time and the last record time of the temporary file are in different first time periods but in the same second time period, it is only necessary to restore the cached data from the temporary file. This reduces the recovery process while allowing as much target historical data as possible to be statistically recorded, thereby improving the efficiency and comprehensiveness of traffic monitoring.
[0123] In one exemplary embodiment, such as Figure 8 As shown, the method also includes the following steps:
[0124] Step 801: If the current time and the last record time of the temporary file are in different second time periods, then obtain the earliest retention time of the target historical data stored in the target database.
[0125] In this embodiment, the target database can clean up stored data, prioritizing the cleanup of data stored earlier or for longer periods. The earliest retention time is the earliest time corresponding to the target historical data retained in the target database. For example, if the target database retains target historical data from 8:00 to 9:00, the earliest retention time is 8:00, and target historical data before 8:00 can be the data to be cleaned up.
[0126] Step 802: Use the latest time between the last recorded time and the earliest retained time as the start time for statistics.
[0127] Step 803: The time period between the start time of statistics and the current time is taken as the second missing time period. Starting from the start time of statistics, the target historical data of the cloud-native WAF system within the second missing time period is statistically processed according to the first time node and the second time node.
[0128] In the embodiments of this application, it can be understood that the specific process of the distributed data processing module 130 performing statistical processing on the target historical data of the cloud-native WAF system within the second missing time period, starting from the start of the statistical process and according to the first time node and the second time node, is similar to steps 501-502.
[0129] In one embodiment, if the current time and the last record time of the temporary file are within the same first time period, no further operation is required. If the distributed data processing module 130 is running for the first time, steps 501-502 are executed.
[0130] In the above traffic monitoring method, in the scenario where the distributed data processing module 130 is shut down and then restarted, if the current time and the last record time of the temporary file are in different second time periods, the record time in the temporary file is compared with the earliest retention time of the target historical data in the target database. The latest time is used to count the records, and different data recovery logic is executed for the first run, restart within the first time period, restart within the second time period, and long-term restart. While reducing the recovery process, as much access attack historical data as possible is counted and recorded, which can improve the efficiency and comprehensiveness of traffic monitoring.
[0131] In one embodiment, the traffic monitoring system 100 scans the amount of data stored in the target database every preset scanning time interval. If the amount of stored data exceeds a preset configured number of rows, the traffic monitoring system 100 exports the oldest retained data exceeding the preset configured number of rows and saves it in an export file, and deletes the exported data from the target database. If the retention time of data exceeds a preset retention time threshold, the traffic monitoring system 100 exports the data and saves it in an export file, and deletes the exported data from the target database. When a preset integration time is reached, the traffic monitoring system 100 integrates and packages the exported files and places them in a unified storage location. The scanning time interval can be 1 minute, meaning the above scanning operation can be performed once per minute. The configured number of rows can be in the millions, for example, one million. The retention time threshold can be in the hours, for example, 2 hours. The exported file can be a CSV file. There can be a 1-hour interval between two adjacent integration times, for example, the integration time can be every hour on the hour. The unified storage location can be a disk. The integration and packaging process can be aggregation and compression. The traffic monitoring system 100 may also include a cleanup module, and the specific process in this embodiment can also be executed by this cleanup module. In this way, the number of records in the database table is limited to the millions, and in extreme cases, only the specific processing records within one minute are retained. For data exceeding the configured number, a default of one million records will be cleaned up periodically, ensuring system response speed, balancing performance and availability, and providing more real-time cloud-native WAF traffic monitoring.
[0132] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0133] Based on the same inventive concept, this application also provides a traffic monitoring device for implementing the traffic monitoring method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more traffic monitoring device embodiments provided below can be found in the limitations of the traffic monitoring method described above, and will not be repeated here.
[0134] In one exemplary embodiment, such as Figure 9 As shown, a traffic monitoring device 900 is provided, including: a first acquisition module 910, an aggregation module 920, a first statistics module 930, and a construction module 940, wherein:
[0135] The first acquisition module 910 is used to acquire traffic characteristic data of each target cloud-native network application firewall (WAF) node.
[0136] The aggregation module 920 is used to aggregate and store the traffic characteristic data of each of the target cloud-native WAF nodes;
[0137] The first statistics module 930 is used to collect traffic characteristic data of each target cloud-native WAF node under each traffic characteristic data type dimension according to a preset time dimension, and store the statistical results in the target database.
[0138] The construction module 940 is used to construct a traffic monitoring view based on the data stored in the target database, the predetermined traffic monitoring dimensions, and the preset traffic monitoring view construction rules; the traffic monitoring view is used to monitor the cloud-native WAF system.
[0139] Optionally, the first statistics module 930 is specifically used for:
[0140] Based on the current business volume of each target cloud-native WAF node, enable multiple target consumers;
[0141] For each target consumer, the traffic feature data received by the target consumer is written into the channel of the target consumer corresponding to each traffic feature data type dimension.
[0142] For each traffic feature data type dimension, the traffic feature data in the channel of the target consumer corresponding to the traffic feature data type dimension is statistically analyzed according to a preset time dimension through the coroutine of the target consumer.
[0143] The statistically analyzed traffic feature data is written into the write channel of the target consumer corresponding to the traffic feature data type dimension;
[0144] If the preset database write conditions are met, the statistically analyzed traffic characteristic data in the write channel will be written into the target database.
[0145] Optionally, the device 900 further includes:
[0146] The second statistics module is used to collect target historical data of the cloud-native WAF system when a preset first time node is reached, and record the target historical data to a temporary file; the target historical data includes historical request volume, historical attack volume, and historical number of each status code; the interval between two adjacent first time nodes is a first time period.
[0147] The writing module is used to write the target historical data between the previous second time node and the current second time node stored in the temporary file into the target database when a preset second time node is reached; the interval between two adjacent second time nodes is a second time period; the duration of the second time period is an integer multiple of the duration of the first time period.
[0148] Optionally, the writing module is specifically used for:
[0149] If a new status code is added when writing to the target database, the database table field corresponding to the status code is modified according to the range of the status code and the new status code.
[0150] Based on the modified database table fields, the historical number of each status code stored in the temporary file between the previous second time node and the current second time node is split into a preset number of status code data tables.
[0151] Write the status code data table and the total request data table into the target database.
[0152] Optionally, the device 900 further includes:
[0153] The first determining module is used to determine the time period between the current time and the last recording time of the temporary file as the first missing time period if the current time and the last recording time of the temporary file are in different first time periods but in the same second time period.
[0154] The third statistics module is used to collect the target historical data of the cloud-native WAF system during the first missing time period and record the target historical data during the first missing time period to the temporary file.
[0155] Optionally, the device 900 further includes:
[0156] The second acquisition module is used to acquire the earliest retention time of the target historical data stored in the target database if the current time and the last record time of the temporary file are in different second time periods.
[0157] The second determining module is used to take the latest time among the last recorded time and the earliest retention time as the start time for statistics;
[0158] The fourth statistics module is used to take the time period between the start statistics time and the current time as the second missing time period, and perform statistical processing on the target historical data of the cloud-native WAF system within the second missing time period according to the first time node and the second time node, starting from the start statistics time.
[0159] Each module in the aforementioned traffic monitoring device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0160] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 10 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements a traffic monitoring method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0161] Those skilled in the art will understand that Figure 10The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0162] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0163] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0164] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0165] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0166] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0167] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0168] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A traffic monitoring method, characterized in that, The method includes: For each target cloud-native network application firewall (WAF) node, obtain the traffic characteristic data of the target cloud-native WAF node; Aggregate and store the traffic characteristic data of each of the target cloud-native WAF nodes; According to the preset time dimension, the traffic characteristic data of each target cloud-native WAF node under each traffic characteristic data type dimension is statistically analyzed, and the statistical results are stored in the target database; When the preset first time node is reached, the target historical data of the cloud-native WAF system is counted and recorded to a temporary file; the target historical data includes the historical request volume, historical attack volume, and the historical number of each status code; When the preset second time node is reached and a new status code is added while writing to the target database, the database table field corresponding to the status code is modified according to the range of the status code and the new status code. Based on the modified database table fields, the historical number of each status code stored in the temporary file between the previous second time node and the current second time node is split into a preset number of status code data tables. The system dynamically constructs data insertion commands based on column names and statistical values, executes the write code, and continues until all status code data tables are processed, and then commits the transaction; the transaction is constructed from each of the status code data tables and the total request data table. Based on the data stored in the target database, the pre-determined traffic monitoring dimensions, and the preset traffic monitoring view construction rules, a traffic monitoring view is constructed; the traffic monitoring view is used to monitor the cloud-native WAF system.
2. The method according to claim 1, characterized in that, The step of statistically analyzing the traffic characteristic data of each target cloud-native WAF node under each traffic characteristic data type dimension according to a preset time dimension, and storing the statistical results in the target database, includes: Based on the current business volume of each target cloud-native WAF node, enable multiple target consumers; For each target consumer, the traffic feature data received by the target consumer is written into the channel of the target consumer corresponding to each traffic feature data type dimension. For each traffic feature data type dimension, the traffic feature data in the channel of the target consumer corresponding to the traffic feature data type dimension is statistically analyzed according to a preset time dimension through the coroutine of the target consumer. The statistically analyzed traffic feature data is written into the write channel of the target consumer corresponding to the traffic feature data type dimension; If the preset database write conditions are met, the statistically analyzed traffic characteristic data in the write channel will be written into the target database.
3. The method according to claim 1, characterized in that, A first time interval is between two adjacent first time nodes; a second time interval is between two adjacent second time nodes; the duration of the second time interval is an integer multiple of the duration of the first time interval.
4. The method according to claim 1, characterized in that, After the transaction is committed, the method further includes: Determine if any errors occurred during the execution of the transaction; If an error occurs during the execution of the transaction, a rollback will be performed and an alert will be issued.
5. The method according to claim 3, characterized in that, The method further includes: If the current time and the last record time of the temporary file are in different first time periods but in the same second time period, then the time period between the current time and the last record time is taken as the first missing time period. The target historical data of the cloud-native WAF system within the first missing time period is statistically analyzed, and the target historical data within the first missing time period is recorded to the temporary file.
6. The method according to claim 3, characterized in that, The method further includes: If the current time and the last record time of the temporary file are in different second time periods, then obtain the earliest retention time of the target historical data retained in the target database; The latest of the last recorded time and the earliest retained time is used as the start time for statistics; The time period between the start time of statistics and the current time is taken as the second missing time period. Starting from the start time of statistics, the target historical data of the cloud-native WAF system within the second missing time period is statistically processed according to the first time node and the second time node.
7. A traffic monitoring system, characterized in that, The system includes: At least one distributed data acquisition module is used to collect traffic characteristic data of each target cloud-native network application firewall (WAF) node; each target cloud-native WAF node includes one of the distributed data acquisition modules. The data aggregation center is used to aggregate and store the traffic characteristic data of each of the target cloud-native WAF nodes; A distributed data processing module is used to statistically analyze the traffic characteristic data of each target cloud-native WAF node under each traffic characteristic data type dimension according to a preset time dimension, and store the statistical results in the target database; when a preset first time node is reached, the module analyzes the target historical data of the cloud-native WAF system and records the target historical data to a temporary file; the target historical data includes historical request volume, historical attack volume, and the historical number of each status code; when a preset second time node is reached and a new status code is added while writing to the target database, the module modifies the database table fields corresponding to the status code according to the range of the status code and the new status code; based on the modified database table fields, the module splits the historical number of each status code between the previous second time node and the current second time node stored in the temporary file into a preset number of status code data tables; the module dynamically constructs data insertion commands according to column names and statistical values, executes the write code, until all status code data tables are processed, and commits the transaction; the transaction is constructed from each status code data table and the total request data table; The storage service module is used to construct a traffic monitoring view based on the data stored in the target database, the pre-determined traffic monitoring dimensions, and the preset traffic monitoring view construction rules; the traffic monitoring view is used to monitor the cloud-native WAF system.
8. A flow monitoring device, characterized in that, The device includes: The first acquisition module is used to acquire traffic characteristic data of each target cloud-native network application firewall (WAF) node. The aggregation module is used to aggregate and store the traffic characteristic data of each of the target cloud-native WAF nodes; The first statistics module is used to collect traffic characteristic data of each target cloud-native WAF node under each traffic characteristic data type dimension according to a preset time dimension, and store the statistical results in the target database. The second statistics module is used to collect target historical data of the cloud-native WAF system when a preset first time node is reached, and record the target historical data to a temporary file; the target historical data includes historical request volume, historical attack volume, and historical number of each status code. The writing module is specifically used to modify the database table fields corresponding to the status codes when a new status code is added during the writing to the target database at a preset second time node; based on the range of the status codes and the new status codes, modify the database table fields corresponding to the status codes; based on the modified database table fields, modify the historical number of each status code between the previous second time node and the current second time node stored in the temporary file into a preset number of status code data tables; dynamically construct data insertion commands based on column names and statistical values, and execute the writing code until all status code data tables are processed, and then commit the transaction; the transaction is constructed from each status code data table and the total request data table; The construction module is used to construct a traffic monitoring view based on the data stored in the target database, the pre-determined traffic monitoring dimensions, and the preset traffic monitoring view construction rules; the traffic monitoring view is used to monitor the cloud-native WAF system.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.