Traffic data processing method and device, computer equipment and readable storage medium

By gradually adding noise and using a diffusion model to identify flow characteristics, high-quality diffusion flow samples are generated, solving the problem of low diffusion sample quality in existing technologies and improving the threat detection capability of neural network models.

CN120873540APending Publication Date: 2025-10-31PENG CHENG LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510827640.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies generate low-quality diffusion traffic samples that cannot effectively mimic the dynamic changes and complex patterns in actual attacks, resulting in poor performance of neural network models when facing new or disguised attacks.

Method used

Intermediate diffusion samples are generated by gradually adding noise, and the temporal and nonlinear relationships of traffic characteristics are identified using a diffusion model. Gradual noise reduction is then performed, and finally, category alignment verification and data feature distribution verification are conducted to select high-quality diffusion traffic samples.

Benefits of technology

The generated diffusion traffic samples can more accurately simulate the dynamic changes and complex patterns of actual attacks, improving the realism and diversity of the samples and enhancing the threat detection capabilities of neural network models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873540A_ABST
    Figure CN120873540A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a traffic data processing method and device, computer equipment and a readable storage medium. The method comprises the steps of obtaining an original traffic sample; gradually adding noise to the original flow sample according to a plurality of preset time steps to obtain an intermediate diffusion sample; for each time step, identifying a time sequence dependency relationship and a nonlinear relationship of the flow characteristics of the intermediate diffusion sample through a diffusion model to obtain an identification result, and outputting a predicted noise characteristic corresponding to each time step according to the identification result; based on the predicted noise feature corresponding to each time step, performing step-by-step noise reduction processing on the intermediate diffusion sample to obtain a diffusion flow sample; performing category alignment verification and data feature distribution verification on the diffusion flow sample to obtain a verification result; and screening the plurality of diffusion flow samples according to the plurality of verification results to obtain a target diffusion flow sample corresponding to the sample diffusion category. Therefore, the quality of the generated diffusion flow sample can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, computer equipment, and readable storage medium for processing traffic data. Background Technology

[0002] Currently, with increasingly frequent and complex cyberattacks, critical infrastructure faces growing threats, impacting the normal operation of network services. To address this, neural network models are commonly used for traffic detection to counter escalating attack methods and highly covert attack traffic, ensuring accurate network traffic detection and security protection. However, due to the scarcity of anomalous traffic samples in real-world environments, neural network models often perform poorly due to insufficient training data, affecting the accuracy of threat detection. Therefore, it is necessary to effectively diffuse traffic samples to improve the model's ability to identify complex attack patterns.

[0003] In related technologies, rule-based data augmentation methods are typically used to modify or combine original samples to generate new samples. For example, diffusion traffic samples can be created by simply changing fields such as Internet Protocol addresses, port numbers, and timestamps in the original traffic; or by performing splicing and truncation operations on the packet payload to simulate different types of attack behaviors, diffusion traffic samples can be obtained. However, this generation method is based solely on predefined rules and does not deeply simulate the dynamic changes and complex patterns that may occur in actual attacks, which may lead to poor model performance when facing new or disguised attacks. In other words, the quality of the augmented traffic samples obtained when generating diffusion traffic samples using related technologies is relatively low. Summary of the Invention

[0004] This application proposes a flow data processing method, apparatus, computer equipment, and readable storage medium that can improve the quality of generated diffusion flow samples.

[0005] To achieve the above objectives, a first aspect of this application proposes a traffic data processing method, the method comprising:

[0006] Obtain the original traffic samples corresponding to the sample diffusion category;

[0007] Noise is gradually added to the original flow sample according to multiple preset time steps to obtain intermediate diffusion samples;

[0008] For each time step, the intermediate diffusion sample is input into the diffusion model. The diffusion model identifies the temporal dependency and nonlinear relationship of the flow characteristics of the intermediate diffusion sample to obtain the identification result. Based on the identification result, the predicted noise feature corresponding to each time step is output.

[0009] Based on the predicted noise features corresponding to each time step, the intermediate diffusion samples are subjected to stepwise noise reduction processing to obtain diffusion flow samples;

[0010] The diffusion flow samples were subjected to category alignment verification and data feature distribution verification to obtain the verification results;

[0011] Based on multiple verification results corresponding to multiple diffusion flow samples, the multiple diffusion flow samples are filtered to obtain the target diffusion flow sample corresponding to the diffusion category of the sample.

[0012] Accordingly, a second aspect of the embodiments of this application provides a traffic data processing apparatus, the apparatus comprising:

[0013] The acquisition module is used to acquire the original traffic samples corresponding to the sample diffusion category;

[0014] An addition module is used to gradually add noise to the original flow sample according to multiple preset time steps to obtain intermediate diffusion samples;

[0015] The identification module is used to input the intermediate diffusion sample into the diffusion model for each time step, identify the temporal dependency and nonlinear relationship of the flow characteristics of the intermediate diffusion sample through the diffusion model, obtain the identification result, and output the predicted noise feature corresponding to each time step based on the identification result.

[0016] The processing module is used to perform stepwise noise reduction processing on the intermediate diffusion sample based on the predicted noise features corresponding to each time step to obtain the diffusion flow sample;

[0017] The verification module is used to perform category alignment verification and data feature distribution verification on the diffusion flow samples to obtain verification results;

[0018] The filtering module is used to filter the multiple diffusion flow samples based on the multiple verification results corresponding to the multiple diffusion flow samples, so as to obtain the target diffusion flow sample corresponding to the diffusion category of the sample.

[0019] In some implementations, the verification result includes a first verification result and a second verification result, and the verification module is further configured to:

[0020] Multiple feature centers corresponding to multiple sample diffusion categories are obtained, wherein each feature center is obtained by averaging the features of multiple original traffic samples included in each sample diffusion category;

[0021] The similarity is calculated between each diffusion flow sample included in each diffusion category and each feature center to obtain the target similarity.

[0022] Based on the magnitude relationship between multiple target similarities, the first predicted diffusion category corresponding to each diffusion flow sample is determined;

[0023] The sample diffusion category of each diffusion flow sample is aligned with the first predicted diffusion category to obtain a first verification result;

[0024] Based on the first verification result, the plurality of diffusion flow samples are filtered to obtain the first diffusion flow sample;

[0025] By using a pre-trained filtering network, the data feature distribution of multiple first diffusion flow samples is verified to obtain a second verification result.

[0026] In some implementations, the verification module is further configured to:

[0027] By using a pre-trained filtering network, a second predicted diffusion category for each first diffusion flow sample is predicted based on the data feature distribution of each first diffusion flow sample. The filtering network is obtained by learning the data feature distribution of multiple original flow samples corresponding to each sample diffusion category through an initial filtering network.

[0028] By using the sample diffusion category corresponding to each first diffusion flow sample, the data feature distribution of the second predicted diffusion category is verified to obtain the second verification result.

[0029] In some embodiments, the traffic data processing apparatus further includes a determining module for:

[0030] Obtain the sample category ratio among multiple pre-set sample diffusion categories, and the number of sample categories for multiple target diffusion traffic samples corresponding to each sample diffusion category;

[0031] Based on the sample category ratio and the number of sample categories, determine the target diffusion category from the plurality of sample diffusion categories that needs to increase the number of sample categories, and the number to be increased;

[0032] Based on the multiple original traffic samples corresponding to the target diffusion category, generate the same number of target diffusion traffic samples as the increase quantity.

[0033] In some implementations, the adding module is further configured to:

[0034] According to a preset number of time steps, the noise scheduling coefficient corresponding to each time step is determined, wherein the noise scheduling coefficient of each time step increases sequentially according to the order of each time step;

[0035] For the first time step, noise is added to each original traffic sample based on the corresponding noise scheduling coefficient to obtain the initial diffusion sample corresponding to the first time step;

[0036] For the next time step after the first time step, noise is added to the initial diffusion sample based on the corresponding noise scheduling coefficient to obtain the updated initial diffusion sample for the next time step;

[0037] Repeat the update for the next time step, add noise to the updated initial diffusion sample based on the corresponding noise scheduling coefficient, and obtain the updated initial diffusion sample corresponding to the updated next time step, until the updated next time step corresponding to the updated initial diffusion sample is the last time step of the plurality of time steps, and obtain the intermediate diffusion sample based on the updated initial diffusion sample.

[0038] In some embodiments, the processing module is further configured to:

[0039] Based on the predicted noise features corresponding to each time step, the intermediate diffusion sample is denoised to obtain the updated intermediate diffusion sample corresponding to the previous intermediate time step of each time step.

[0040] For the previous intermediate time step, the updated intermediate diffusion sample is input into the diffusion model to obtain the predicted noise features corresponding to the previous intermediate time step, and based on the predicted noise features, the updated intermediate diffusion sample corresponding to the previous intermediate time step is obtained.

[0041] The process of repeatedly inputting the updated intermediate diffusion sample into the diffusion model for the previous intermediate time step to obtain the predicted noise features corresponding to the previous intermediate time step, and obtaining the updated intermediate diffusion sample corresponding to the updated previous intermediate time step based on the predicted noise features, continues until the updated previous time step corresponding to the updated intermediate diffusion sample is the first time step of the plurality of time steps, and the updated intermediate diffusion sample is used as the diffusion flow sample corresponding to the sample diffusion category.

[0042] In some embodiments, the traffic data processing apparatus further includes a removal module for:

[0043] Multiple preset traffic samples are obtained, and the number of unique values ​​in each preset traffic sample and the frequency of each unique value are determined. Discrete traffic samples are then selected from the multiple preset traffic samples.

[0044] From the plurality of preset traffic samples, the discrete traffic samples are removed to obtain a plurality of original traffic samples.

[0045] Accordingly, a third aspect of the present application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the traffic data processing method of any one of the embodiments of the first aspect of the present application.

[0046] Accordingly, a fourth aspect of the embodiments of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the traffic data processing method of any one of the embodiments of the first aspect of this application.

[0047] This application embodiment obtains the original traffic samples corresponding to the sample diffusion category; adds noise to the original traffic samples step by step according to multiple preset time steps to obtain intermediate diffusion samples; for each time step, the intermediate diffusion samples are input into a diffusion model, and the diffusion model identifies the temporal dependency and nonlinear relationship of the traffic features of the intermediate diffusion samples to obtain identification results, and outputs the predicted noise features corresponding to each time step based on the identification results; performs stepwise noise reduction processing on the intermediate diffusion samples based on the predicted noise features corresponding to each time step to obtain diffusion traffic samples; performs category alignment verification and data feature distribution verification on the diffusion traffic samples to obtain verification results; and filters the multiple diffusion traffic samples according to the multiple verification results corresponding to the multiple diffusion traffic samples to obtain the target diffusion traffic samples corresponding to the sample diffusion category. This approach allows for a more accurate simulation of dynamic changes and complex patterns in actual attacks by gradually adding noise. By utilizing the feature extraction components embedded in the diffusion model, it dynamically captures complex temporal patterns (such as the temporal evolution of attack behavior) and nonlinear relationships (such as the coordinated changes in payload features and network layer parameters) in traffic data. This overcomes the limitation of rule-based methods in simulating the dynamic changes of real attacks, enabling the generated diffusion samples to simulate the dynamic behavioral characteristics of real attack traffic. Furthermore, noise reduction is applied to make the generated diffusion traffic samples more realistic and diverse. On the other hand, the generated diffusion traffic samples undergo category alignment verification (ensuring semantic consistency of the attack) and data feature distribution verification (ensuring matching with the statistical distribution of real abnormal traffic), effectively filtering out low-quality or deviation-from-the-real-distribution generated samples, fundamentally improving the proportion and reliability of effective samples. In summary, this application can improve the quality of generated diffusion traffic samples. Attached Figure Description

[0048] Figure 1 This is a schematic diagram of the architecture of the traffic data processing system provided in the embodiments of this application;

[0049] Figure 2 This is a flowchart of the traffic data processing method provided in the embodiments of this application;

[0050] Figure 3 This is a general flowchart of the traffic data processing method provided in the embodiments of this application;

[0051] Figure 4 This is a schematic diagram of the functional modules of the traffic data processing device provided in the embodiments of this application;

[0052] Figure 5 This is a schematic diagram of the hardware structure of the computer device provided in the embodiments of this application. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0054] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0056] Currently, with increasingly frequent and complex cyberattacks, critical infrastructure faces growing threats, impacting the normal operation of network services. To address this, neural network models are commonly used for traffic detection to counter escalating attack methods and highly covert attack traffic, ensuring accurate network traffic detection and security protection. However, due to the scarcity of anomalous traffic samples in real-world environments, neural network models often perform poorly due to insufficient training data, affecting the accuracy of threat detection. Therefore, it is necessary to effectively diffuse traffic samples to improve the model's ability to identify complex attack patterns.

[0057] In related technologies, rule-based data augmentation methods are typically used to modify or combine original samples to generate new samples. For example, diffusion traffic samples can be created by simply changing fields such as Internet Protocol addresses, port numbers, and timestamps in the original traffic; or by performing splicing and truncation operations on the packet payload to simulate different types of attack behaviors, diffusion traffic samples can be obtained. However, this generation method is based solely on predefined rules and does not deeply simulate the dynamic changes and complex patterns that may occur in actual attacks, which may lead to poor model performance when facing new or disguised attacks. In other words, the quality of the augmented traffic samples obtained when generating diffusion traffic samples using related technologies is relatively low.

[0058] Based on this, embodiments of this application provide a traffic data processing method, apparatus, computer equipment, and readable storage medium, which can improve the quality of the generated diffusion traffic samples.

[0059] The traffic data processing method, apparatus, computer equipment, and readable storage medium provided in the embodiments of this application are specifically described through the following embodiments. First, the traffic data processing system in the embodiments of this application is described.

[0060] Please refer to Figure 1 In some embodiments, this application provides a traffic data processing system, including a terminal 11 and a server 12.

[0061] In some implementations, terminal 11 can be a network edge device with basic computing power, such as an industrial firewall, IoT gateway, enterprise-grade router, or local server. Terminal 11 can capture raw network traffic data in the network environment in real time and perform preliminary feature filtering on the raw network traffic data, such as distinguishing between continuous and discrete features and filtering out discrete features. Furthermore, terminal 11 can transmit the preliminarily processed data to the server through a secure channel for further in-depth analysis and processing.

[0062] Furthermore, the server-side component 12 can be a high-performance server cluster, a cloud server, a data center, etc. The server-side component 12 can deploy the diffusion model and perform highly complex calculations, such as generating diffusion traffic samples, performing class alignment verification and data feature distribution verification on the diffusion traffic samples, and dynamically adjusting the number of generated target diffusion traffic samples until class balance is achieved. Furthermore, the server-side component 12 can also train machine models such as the diffusion model to generate high-quality diffusion traffic samples.

[0063] The traffic data processing method in this application can be illustrated through the following embodiments.

[0064] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user will be obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent will the necessary user-related data for the normal operation of the embodiments of this application be obtained.

[0065] In this embodiment, the description will focus on the perspective of a traffic data processing device, which can specifically be integrated into a computer device. See [link to relevant documentation]. Figure 2 , Figure 2 This is a flowchart illustrating the steps of a traffic data processing method provided in this application embodiment. Taking the traffic data processing device specifically integrated into a terminal or server as an example, the specific process when the processor on the terminal or server executes the program instructions corresponding to the traffic data processing method is as follows:

[0066] Step 101: Obtain the original traffic sample corresponding to the sample diffusion category.

[0067] In some implementations, to clearly define the target of diffusion, the original traffic sample corresponding to at least one sample diffusion category to be diffused can be obtained, so that the diffusion model can perform targeted diffusion based on the original traffic sample of the sample diffusion category, thereby improving the efficiency and accuracy of diffusion.

[0068] Among them, the sample diffusion category can be a specific network traffic category that is significantly insufficient in the dataset and needs to be enhanced by the diffusion model. For example, there are relatively rare attack types with few original traffic samples available. In order to improve the generalization performance of the model during training, sample diffusion can be performed on them.

[0069] The original traffic samples can be raw data collected from real network environments that belong to the sample diffusion category. They can serve as the learning basis for the diffusion model to generate diffusion traffic samples that conform to the real distribution.

[0070] For example, sample propagation categories can include distributed denial-of-service attacks, port scanning, and Trojan communication. When obtaining the original traffic samples corresponding to each sample propagation category, they can be obtained using network intrusion detection tools such as Wireshark and Snort, or after simulating an attack environment, or through public datasets, internal enterprise logs, etc.

[0071] For example, the raw traffic sample may include protocol type and version, source and destination Internet Protocol addresses, source and destination port numbers, timestamps, payload content, etc. The specific sample diffusion category and the corresponding content of the raw traffic sample may vary depending on the scenario.

[0072] By obtaining the original traffic samples corresponding to the sample diffusion categories, it is possible to identify the traffic categories that need to be enhanced, thus avoiding the waste of resources caused by ineffective diffusion and providing accurate basic data for subsequent traffic diffusion.

[0073] In some implementations, to avoid invalid generation, the data distribution characteristics of traffic features can be analyzed to identify discrete features (such as protocol type) that cannot be generated through noise perturbation and to filter out these discrete features. This ensures that subsequent diffusion models only perform data augmentation on optimizable continuous features (such as packet size and transmission interval). For example, before step 101, that is, before "obtaining the original traffic samples corresponding to the sample diffusion categories", the following may also be included:

[0074] (A.1) Obtain multiple preset traffic samples, determine the number of unique values ​​in each preset traffic sample, and the frequency of each unique value, and filter out discrete traffic samples from multiple preset traffic samples;

[0075] (A.2) Remove discrete flow samples from multiple preset flow samples to obtain multiple original flow samples.

[0076] The preset traffic sample can be the original network traffic dataset, which contains all traffic samples of the sample diffusion category (such as attack traffic and normal traffic), and its multidimensional features include continuous (numerical variable) and discrete (finite enumeration value) attributes.

[0077] Among them, the unique value can be a value that is not repeated in the feature column of the current preset traffic sample, such as the protocol type feature.

[0078] Among them, discrete flow samples can be preset flow samples that are classified as discrete features.

[0079] For example, the following operations can be performed on each preset traffic sample: calculate the number of unique values, that is, count the number of different values ​​in the preset traffic sample; analyze the frequency distribution, that is, calculate the frequency of each value, and extract the cumulative frequency of the top N high-frequency values ​​(N can be configured to 5 or other values). Then, if any of the following conditions are met, it is marked as a discrete feature:

[0080] The number of unique values ​​is less than or equal to a unique value threshold, for example, less than or equal to 10. The unique value threshold can be determined according to the actual situation. Alternatively, the cumulative frequency of the top N high-frequency values ​​is greater than a preset frequency threshold, for example, the preset frequency threshold can be 80%, which can also be determined according to the actual situation. Or, the frequency of any value is greater than a preset occurrence frequency, for example, the preset occurrence frequency can be 80%, which can be determined according to the actual situation.

[0081] By using the above filtering method, all preset traffic samples marked as discrete features can be removed, and the remaining preset traffic samples can be retained as original traffic samples to be used as input for the diffusion model in the future.

[0082] By eliminating discrete traffic samples, the diffusion model can avoid adding noise to features that cannot be disturbed, thus fundamentally ensuring the logical rationality of the generated data (such as not generating illegal protocol types). The diffusion model only needs to process the high-dimensional space of continuous features, reducing unnecessary computation and accelerating the data generation process.

[0083] Step 102: Add noise to the original flow sample step by step according to multiple preset time steps to obtain intermediate diffusion samples.

[0084] In some implementations, in order to gradually degrade the data distribution, a progressive noise injection strategy can be used to gradually transform the original flow samples into intermediate states that conform to a standard Gaussian distribution, thereby generating intermediate data containing different noise intensities, which means that the generated intermediate diffusion samples have diversity.

[0085] The time step can be a preset discretization iteration stage number in the diffusion process, used to control the gradual degree of noise addition. Each time step corresponds to a specific noise intensity parameter, and the noise variance is preset through a scheduling strategy (such as linear or cosine) to ensure that the data smoothly transitions from the original distribution to the standard Gaussian distribution.

[0086] The intermediate diffusion sample can be the transitional state data generated after multiple time steps of noise addition, obtained from the original flow sample after t time steps of noise addition.

[0087] In some implementations, a random noise matrix conforming to a standard Gaussian distribution can be pre-generated, with its ε~N(0,I) dimensions being the same as the feature dimensions of the original traffic sample X0. Simultaneously, noise scheduling coefficients for T time steps are preset. (Satisfies 0<β1<β2<...<β T <1), to control the noise intensity at each step.

[0088] Furthermore, for each time step t∈{1,2,...,T}, noise can be added iteratively according to the following formula:

[0089]

[0090] Among them, X t-1 This is the output of the previous time step (initial X0 is the original flow sample); ε t This is a noise matrix generated independently for the current time step; A decay factor to preserve the original characteristics of the original flow samples. This is the noise injection factor.

[0091] After T iterations, the original traffic sample X0 gradually degenerates into:

[0092]

[0093] in,

[0094] Final output X T The intermediate diffused sample follows a standard Gaussian distribution, which is completely deviated from the original data distribution but still contains latent feature structures.

[0095] By adding progressive noise over multiple time steps, the original flow samples can be smoothly transitioned to a learnable Gaussian distribution, eliminating the sensitive dependence on initial parameters, avoiding abrupt distortion problems, ensuring the diversity and authenticity of the generated data, and laying the foundation for high-quality data generation in the subsequent reverse diffusion process.

[0096] In some implementations, to improve the authenticity and diversity of the data, the noise scheduling coefficient can be increased sequentially according to the order of each time step, with the noise intensity increasing as the time step increases, to gradually approximate more complex flow patterns and ultimately generate high-quality intermediate diffusion samples. For example, step 102 may include:

[0097] (102.1) According to a preset number of time steps, the noise scheduling coefficient corresponding to each time step is determined, wherein the noise scheduling coefficient of each time step increases sequentially according to the order of each time step.

[0098] (102.2) For the first time step, noise is added to each original flow sample based on the corresponding noise scheduling coefficient to obtain the initial diffusion sample corresponding to the first time step;

[0099] (102.3) For the next time step after the first time step, noise is added to the initial diffusion sample based on the corresponding noise scheduling coefficient to obtain the updated initial diffusion sample for the next time step.

[0100] (102.4) Repeat the next time step update for the next time step, add noise to the updated initial diffusion sample based on the corresponding noise scheduling coefficient, and obtain the updated initial diffusion sample corresponding to the updated next time step, until the updated next time step corresponding to the updated initial diffusion sample is the last time step of multiple time steps, and obtain the intermediate diffusion sample based on the updated initial diffusion sample.

[0101] The noise scheduling coefficient, also known as the noise variance coefficient, can be a preset noise intensity parameter for each time step, used to control the magnitude of the noise variance added in the current step. The noise scheduling coefficient value can strictly increase sequentially according to the time step.

[0102] The first time step can be the time step (t=1) corresponding to the initial operation phase of the diffusion link, with its noise scheduling coefficient being the minimum value, corresponding to the slightest noise addition intensity, in order to preserve the main characteristics of the original traffic sample.

[0103] The initial diffusion sample can be the first intermediate state data generated after adding noise to the original flow sample at the first time step (t=1), denoted as X1.

[0104] The next time step can be a subsequent stage of the current operation time step (e.g., the next time step after t=1 is t=2), and its noise scheduling coefficient is strictly greater than that of the previous time step.

[0105] Updating the initial diffusion sample can be a sample variable that has been reassigned in the current time step operation, such as an intermediate diffusion sample (e.g., X) generated in the previous time step. t-1 The input for the current step is used as the noise input and then updated to X. t .

[0106] The next time step to be updated can be a new subsequent stage after the current operation is completed. For example, after t=2 is completed, the next time step to be updated is t=3), and the iteration continues until the final time step T.

[0107] Understandably, the noise injected at each time step is an independent and identically distributed Gaussian random matrix, with dimensions consistent with the feature dimensions of the original flow sample (e.g., if the original flow sample is a 100-dimensional feature vector, then the injected noise is also 100-dimensional). The increasing noise scheduling coefficient (noise) ensures that the noise intensity gradually increases with each time step, causing the original flow features to gradually degenerate from a clear state to a pure noise state. However, the final intermediate diffused sample must contain the features of the flow sample.

[0108] For example, if a total number of time steps T is preset (e.g., T = 1000), a noise variance coefficient sequence is generated. Satisfying β1<β2<...<β T That is, βt The noise scheduling coefficient increases linearly with each time step. For the first time step, noise can be added to the original traffic sample X0 to obtain X1:

[0109]

[0110] ε1~N(0,I);

[0111] Output X1 as the initial diffusion sample for the first time step. As a noise disturbance term, injecting random disturbances can simulate fluctuations in the actual network environment (such as legitimate traffic jitter and attack payload variation).

[0112] Furthermore, for the next time step after the first time step, new noise can be added to X1 to obtain X2:

[0113]

[0114] ε2~N(0,I);

[0115] The output X2 is the updated initial diffusion sample for the next time step. This process is repeated until t = T, at which point X is output. T As an intermediate diffusion sample.

[0116] By using the above methods, the generated intermediate diffusion samples can carry temporal characteristics of different noise levels, and more naturally approximate the complex distribution of real attack traffic, thereby enhancing the diversity and authenticity of the intermediate diffusion samples and facilitating the subsequent extraction of high-quality diffusion samples.

[0117] Step 103: For each time step, input the intermediate diffusion sample into the diffusion model, identify the temporal dependency and nonlinear relationship of the flow characteristics of the intermediate diffusion sample through the diffusion model, obtain the identification result, and output the predicted noise feature corresponding to each time step based on the identification result.

[0118] In some implementations, in order to accurately capture the complex dynamic changes in network traffic data, the intermediate diffusion samples can be analyzed by diffusion models to effectively uncover the inherent patterns and potential correlations of the original traffic samples at different points in time, thereby accurately predicting the noise residuals and providing a core basis for inverse diffusion denoising.

[0119] The diffusion model can be a generative model based on the improved RNN-KAN framework of this application, which replaces the traditional linear neural network with the RNN-KAN framework. Here, the diffusion model is a pre-trained model.

[0120] Temporal dependencies can be dynamic correlation characteristics across timestamps in intermediate diffusion samples, such as the trend of continuous data packet size changes and transmission delay fluctuations. Temporal dependencies can be long-term dependency patterns captured by recurrent neural networks (RNNs) through the hidden state cyclical propagation mechanism, such as the pulse timing pattern of DDoS attacks.

[0121] Nonlinear relationships can be complex statistical associations between intermediate diffused samples that cannot be described by linear transformations, such as the interaction effect between protocol type and load size. These can be deep feature combinations analyzed by learnable activation networks, such as the Kolmogorov-Arnold Network (KAN), which achieves high-dimensional nonlinear mapping through learnable spline functions.

[0122] The identification results can be a structured representation of the dynamic patterns of traffic features extracted by analyzing intermediate diffusion samples through the RNN-KAN framework of the diffusion model.

[0123] Among them, the predicted noise feature can be the noise residual estimate output by the diffusion model based on the identification results, which represents the predicted distribution of Gaussian noise contained in the current intermediate diffusion sample.

[0124] For example, the intermediate diffused sample X can be... t Input into the diffusion model, X t The dimension is N×D (N is the number of samples, D is the number of continuous features), and then, X is captured by the RNN in the RNN-KAN framework of the diffusion model. t The time-series evolution pattern of traffic characteristics is analyzed, and the processing logic is as follows:

[0125] h t =RNN(X) t ,h t-1 );

[0126] Among them, h t h is the current hidden state vector. t-1 This is the hidden state from the previous step (initial h0 = 0).

[0127] It is understandable that RNNs can retain historical states h through circular connections. t-1 RNNs learn long-term dependencies in traffic characteristics. For example, RNNs can capture the temporal continuity of the Transmission Control Protocol (TCP) connection state transition sequence (SYN→ACK→FIN), and they can also capture the periodic burst patterns of DDoS attack pulses, thereby accurately identifying traffic characteristics.

[0128] Furthermore, after the RNN outputs the identified traffic features, these features can be input into the KAN to resolve the complex nonlinear mappings between the features. The specific processing logic is as follows:

[0129]

[0130] in, This indicates a feature concatenation operation, which involves combining intermediate diffused samples X. t With the output h of the RNN t This allows for the modeling of implicit coupling relationships, or nonlinear relationships, between features. Examples include the logarithmic relationship between payload length and response time, or the nonlinear interaction between the byte entropy value of encrypted traffic and the protocol type.

[0131] In some implementations, temporal and nonlinear recognition results can be fused to output predicted noise features. The recognition result is the hidden state h of the RNN. t With the output φ of KAN t The joint feature vector [h] obtained by concatenation t ;φ t ], as a joint representation of temporal and nonlinearity.

[0132] Furthermore, the predicted noise characteristics can be calculated in the following way:

[0133] ε θ (X t ,t)=LinearLayer([h t ;φ t ]);

[0134] Where, ε θ For the predicted noise features at the current time step, LinearLayer is a fully connected output layer. In this way, it is possible to simultaneously model the temporal continuity of the traffic (handled by RNN) and nonlinear complexity (handled by KAN), ensuring that the physical constraints of the actual traffic flow are met for the noise prediction load. This allows for the accurate reconstruction of the pulse, periodic, and other behavioral characteristics of the attack traffic, ensuring that the traffic data generated by the reverse diffusion retains the true distribution characteristics and improving the quality of the generated diffusion traffic samples.

[0135] By using the above methods, the predicted noise features at each time step can be accurately extracted, which facilitates subsequent stepwise noise reduction of intermediate diffusion samples to obtain diffusion flow samples.

[0136] Step 104: Based on the predicted noise features corresponding to each time step, perform stepwise noise reduction on the intermediate diffusion samples to obtain diffusion flow samples.

[0137] In some implementations, in order to restore noisy intermediate diffused samples to high-quality synthetic traffic data, noise can be removed step by step through reverse iteration. This allows the diffused traffic samples to gradually approach the real traffic distribution while retaining the key characteristics and complex patterns of the original traffic, thereby ensuring a high degree of consistency between the generated diffused traffic samples and the actual network environment, and enhancing the diversity and reliability of the data.

[0138] The diffusion flow sample can be synthetic flow data generated through a complete reverse diffusion process, obtained by denoising intermediate diffusion samples through a time step.

[0139] In some implementations, inverse diffusion can be performed starting from an intermediate diffused sample (i.e., the diffused sample obtained in the last time step after adding noise at multiple time steps) based on the predicted noise characteristics. For example, for each time step t = T, T-1, ..., 1, the following denoising process can be performed iteratively:

[0140]

[0141] in, The cumulative noise attenuation coefficient is z ~ N(0, I), which is the random disturbance term (enabled when T > 1), and ε θ (X t ,t) represents the predicted noise feature at the corresponding time step, σ t Used to control the intensity of random disturbances.

[0142] Furthermore, the calculation stops when t=1, and X0 is output as a sample of diffusion flow.

[0143] By iteratively removing noise at multiple time steps, the generated traffic samples can gradually approximate the distribution characteristics of real network traffic while preserving the original attack patterns. Furthermore, the introduction of a cumulative noise attenuation coefficient and a random perturbation term further enhances the diversity and controllability of the generated samples, avoiding the problems of homogenization or distortion, and facilitating the generation of diverse, high-quality diffusion samples.

[0144] In some implementations, to ensure that the generated diffusion flow samples not only retain the key characteristics and complex patterns of the original flow but also possess higher realism and diversity, the predicted noise features obtained from the diffusion model can be progressively removed in reverse time step order to gradually restore the high-noise intermediate samples into high-fidelity synthetic flow data. For example, step 104 may include:

[0145] (104.1) Based on the predicted noise features corresponding to each time step, the intermediate diffusion samples are denoised to obtain the updated intermediate diffusion samples corresponding to the previous intermediate time step of each time step.

[0146] (104.2) For the previous intermediate time step, the updated intermediate diffusion sample is input into the diffusion model to obtain the predicted noise features corresponding to the previous intermediate time step, and based on the predicted noise features, the updated intermediate diffusion sample corresponding to the previous intermediate time step is obtained.

[0147] (104.3) Repeat the steps of inputting the updated intermediate diffusion sample into the diffusion model for the previous intermediate time step, obtaining the predicted noise features corresponding to the previous intermediate time step, and obtaining the updated intermediate diffusion sample corresponding to the previous intermediate time step based on the predicted noise features, until the updated previous time step corresponding to the updated intermediate diffusion sample is the first time step of multiple time steps, and use the updated intermediate diffusion sample as the diffusion flow sample corresponding to the sample diffusion category.

[0148] The previous intermediate time step can be the stage preceding the current time step of the reverse diffusion operation. For example, if the sample at time t is being processed, then its previous intermediate time step is t-1 (t decreases from T to 1).

[0149] The updated intermediate diffusion sample can be the transitional state data (denoted as X) after noise reduction processing in the current step during the reverse diffusion process. t-1 ), from input sample X t It is obtained through noise prediction updates.

[0150] In some implementations, inverse diffusion can be performed starting from an intermediate diffused sample (i.e., the diffused sample obtained at the last time step after adding noise at multiple time steps) based on the predicted noise features. For example, for the intermediate diffused sample at time step t and the corresponding predicted noise features ε θ (X t The noise reduction of the intermediate diffused sample X at time step t-1 can be achieved using the following formula. t-1 :

[0151]

[0152] Where, α t This represents the noise attenuation coefficient at the current time step. The cumulative noise attenuation coefficient is z ~ N(0, I), which is the random disturbance term (enabled when T > 1), and ε θ (X t ,t) represents the predicted noise feature at the corresponding time step, σ t Used to control the intensity of random disturbances.

[0153] Furthermore, in the next iteration, X can be... t-1 In the input diffusion model, the predicted noise feature ε at step t-1 is obtained. θ (Xt-1 The algorithm iteratively denoises the data based on the newly generated predicted noise features (t-1) to obtain the updated intermediate diffuse sample X at the previous intermediate time step (t-1) and the updated intermediate diffuse sample X at the previous time step (t-2). t-2 :

[0154]

[0155] Then, repeat the above process at each time step, decreasing the time step from t=T to t=1. This iterative process will not be described in detail here, but can be referred to the above. When the iteration reaches time step 1, that is, the first time step, the loop terminates and the output X0 is the diffusion flow sample corresponding to the sample diffusion category.

[0156] By employing the above methods, the temporal continuity of the denoising process can be enforced through RNNs in each iteration (e.g., DDoS pulse periods, APT attack phase transitions), avoiding temporal breakage issues. Simultaneously, KANs can effectively resolve implicit coupling relationships between features (e.g., nonlinear mapping between payload length and protocol type), generating attack samples that conform to actual statistical distributions. In summary, by using a diffusion model to predict noise features of the current intermediate diffusion sample at each time step and progressively backtracking to remove noise, the synthesized data more closely resembles the evolution of real network traffic in terms of temporal logic and statistical distribution. The final output diffusion traffic samples possess both diversity and maintain a high degree of consistency with the original traffic.

[0157] Step 105: Perform category alignment verification and data feature distribution verification on the diffusion flow samples to obtain the verification results.

[0158] In some implementations, to ensure the authenticity and usability of enhanced traffic samples, the spread traffic samples can be double-verified to prevent semantic drift of the spread samples and constrain the generated samples to conform to the statistical characteristics of real attacks, thereby intelligently filtering out low-quality synthetic data and eliminating samples that deviate from reality.

[0159] Among them, category alignment verification can be used to verify the category semantic consistency of the diffused traffic samples. It is used to filter out cases where category semantic drift is introduced during the noise reduction process. For example, a Structured Query Language (SQL) injection attack sample is mislabeled as a Cross-Site Scripting (XSS) attack. This situation can be filtered out by category alignment verification.

[0160] In this context, data feature distribution verification can be achieved by using a pre-trained filtering network, such as a Siamese network, to verify the authenticity of the statistical distribution of the spread traffic samples, thereby eliminating statistical distribution biases in the generated samples. For example, the temporal fluctuations of Advanced Persistent Threat (APT) attack traffic do not conform to the true Poisson distribution. In this case, the filtering network can make predictions based on the sample distribution. If the predicted category of the spread traffic sample is inconsistent with the true category, then the spread traffic sample can be filtered out. This is data feature distribution verification.

[0161] The verification result can be the verification result for each diffusion flow sample. When a diffusion flow sample passes the category alignment verification and data feature distribution verification respectively, it can be used as the target diffusion flow sample; otherwise, it will be screened out.

[0162] In some implementations, after generating diffusion traffic samples, to ensure their quality and effectiveness, a cosine similarity-based intelligent filtering mechanism can be introduced to perform class alignment verification on the diffusion traffic samples, thereby eliminating low-quality samples that are inconsistent with the original data distribution. First, the features of the original training data and the generated data can be standardized to eliminate the influence of data scaling. Then, the feature center, i.e., the center vector, of the diffusion category of each sample is calculated based on the multiple original traffic samples contained in that category, and then normalized to construct a class benchmark in a high-dimensional feature space. For example, the feature center can be obtained by averaging the features of the multiple original traffic samples contained in the diffusion category of each sample.

[0163] Furthermore, for each diffused traffic sample, its target similarity with each feature center can be calculated, and the feature center with the highest similarity to the diffused traffic sample can be used as the first predicted diffused category for that diffused traffic sample. In this way, the sample diffused category of the original traffic sample based on that diffused traffic sample can be compared with the first predicted diffused category to determine whether the first predicted diffused category is consistent with the original sample diffused category label. If they are consistent, it indicates that the diffused traffic sample has passed the category alignment verification; otherwise, it is discarded. For example, the target similarity between each diffused traffic sample and each feature center can be calculated using methods such as cosine similarity, Euclidean distance, and Manhattan distance.

[0164] Furthermore, to further ensure the authenticity and validity of the synthesized data, a data filtering mechanism based on a Siamese network can be constructed to verify the data feature distribution. First, a Siamese network can be trained to accurately predict the class of a sample based on its distribution. Then, the trained Siamese network can be used to predict the second predicted diffusion class of the first diffusion flow sample that has passed the class alignment verification. The second predicted diffusion class is then verified against the base diffusion class of the first diffusion flow sample. If they match, the first diffusion flow sample is retained; otherwise, it is discarded.

[0165] It is understandable that the basic sample diffusion category is the category of the original flow sample before the diffused flow sample is diffused. For example, if the original flow sample a under sample diffusion category A is diffused to obtain the diffused flow sample a1, then the basic sample diffusion category of the diffused flow sample a1 is sample diffusion category A.

[0166] In some implementations, network attack behaviors exhibit time-series dependencies (e.g., the penetration → reconnaissance → attack phase of an APT attack). Therefore, the temporal pattern trajectory of the spread samples (e.g., packet length changes, protocol switching sequences) can be extracted. This trajectory can then be compared with the original attack category template using a Dynamic Time Warping (DTW) algorithm to achieve dynamic trajectory matching verification of the spread traffic samples and obtain the verification results.

[0167] In some implementations, causal discovery algorithms (such as the PC algorithm) can be used to extract causal emergence features (such as "DNS response time → load entropy value") from the raw traffic to verify whether the generated samples satisfy the same causal structure. If they satisfy the same causal structure, the verification result is considered successful. Specifically, firstly, a standard behavioral trajectory template of the original APT attack, Tori = [SYN scan, HTTP long connection, encrypted payload transmission], is constructed, defining the three-stage logical sequence of the attack. After generating the diffusion traffic sample, its behavioral trajectory, Tdiff = [SYN scan, DNS tunnel, encrypted payload transmission], is extracted. The difference distance between the two trajectories, DTW(Tori, Tdiff) = 1.2, is calculated using the Dynamic Time Warping (DTW) algorithm and compared with a preset threshold τ = 1.5 (which can be adjusted according to the actual situation). For example, since the actual distance 1.2 is less than the threshold 1.5, the diffusion traffic sample is determined to have passed verification. In this way, it is possible to capture the dynamic logical essence of attack behavior by flexibly aligning the timing stages (such as accepting DNS tunnels instead of HTTP long connections as covert channels), where HTTP (Hypertext Transfer Protocol) is the Hypertext Transfer Protocol, thus solving the problem of rule-based methods failing in sequence mutation scenarios.

[0168] In some implementations, causal verification can be performed on the diffused traffic samples to obtain verification results. Specifically, a core causal chain can be designed based on the attack categories of different types of diffused traffic samples. For example, a core causal chain can be extracted from the original APT attack traffic. Specifically, after generating diffused traffic samples, the retention of the core causal chain can be verified by probability calculation: the ratio of the conditional probability of payload length on duration (P(length|time)) to the overall probability of payload length (P(length)) in the diffused traffic sample is calculated to obtain the causal strength index (CI = 2.8). If this value is greater than a preset threshold (τcause = 2.0), it is determined that the sample truly retains the causal logic of the APT attack, and the verification passes. In this way, causal strength quantification can ensure that the generated diffused traffic samples not only match numerical features but also deeply restore the logical chain of attack behavior (such as the inevitable impact of duration on the encrypted payload), completely avoiding the situation where the generated samples are "similar in data but distorted in logic".

[0169] In some implementations, the diffusion flow sample can be verified by combining one or more of the above verification methods to obtain the verification results. The specific method can be selected according to the actual situation.

[0170] By constructing a dual intelligent screening system through the above methods, low-quality samples with inconsistent categories and serious distribution deviations can be eliminated, ensuring the rationality of enhanced data in the feature space, strengthening its matching degree with real attack patterns, and constructing a more representative, diverse, credible, and higher-quality dataset.

[0171] In some implementations, to improve the quality of the generated diffusion flow samples, a dual intelligent verification mechanism can be applied to each diffusion flow sample to accurately separate high-value samples from the diffusion generation data, ensuring the authenticity, logical consistency, and distribution matching of the enhanced diffusion flow samples. For example, step 105 may include:

[0172] (105.1) Obtain multiple feature centers corresponding to multiple sample diffusion categories, wherein each feature center is obtained by averaging the features of multiple original traffic samples included in each sample diffusion category;

[0173] (105.2) Calculate the similarity between each diffusion flow sample included in each diffusion category and each feature center to obtain the target similarity;

[0174] (105.3) Based on the magnitude relationship between multiple target similarities, determine the first predicted diffusion category corresponding to each diffusion flow sample;

[0175] (105.4) Perform class alignment verification between the sample diffusion class of each diffusion flow sample and the first predicted diffusion class to obtain the first verification result;

[0176] (105.5) Based on the first verification result, multiple diffusion flow samples are screened to obtain the first diffusion flow sample;

[0177] (105.6) The data feature distribution of multiple first diffusion flow samples is verified by using a pre-trained screening network to obtain the second verification result.

[0178] The feature center can be the reference vector of each sample diffusion category in the feature space, which can be obtained by taking the arithmetic mean of the features of all the original flow samples of that category and normalizing it.

[0179] The target similarity can be the cosine similarity value between the diffusion flow sample and the feature center, or a similarity value calculated by other similarity calculation methods. It can be used to compare the diffusion flow sample with the sample diffusion category.

[0180] The first predicted diffusion category can be determined by comparing the target similarity between the diffusion flow sample and all feature centers, and taking the category corresponding to the feature center with the highest similarity to the diffusion flow sample as the predicted category of the diffusion flow sample.

[0181] The first verification result can be the conclusion obtained from the category alignment verification, or it can be whether the diffusion flow sample passes or fails the verification.

[0182] The first diffusion flow sample can be a subset of the diffusion flow samples that have passed the first verification (class alignment verification) and used as input for subsequent data feature distribution verification.

[0183] The selection network can be a pre-trained Siamese neural network, or a contrastive learning network, a prototype network, etc.

[0184] The second verification result can be a quantitative conclusion of the data feature distribution verification, which can be whether the first diffusion flow sample passes or fails verification.

[0185] For example, the original training dataset S corresponding to each sample category k can be diffused. k The feature center C of each sample diffusion category k is calculated by analyzing all the original flow samples included. k The specific calculation formula is as follows:

[0186]

[0187] Furthermore, by normalizing the diffusion category for each sample, we can obtain the normalized feature centers:

[0188]

[0189] Furthermore, for each diffusion flux sample X, its feature center C relative to each sample diffusion category can be calculated using the following formula. k Target similarity (i.e., cosine similarity):

[0190]

[0191] Furthermore, from the multiple target similarities between each diffusion traffic sample X and multiple feature centers, the target feature center with the highest target similarity can be selected, and the diffusion category corresponding to this target feature center can be used as the first predicted diffusion category of the diffusion traffic sample X. For example, if the target similarity between diffusion traffic sample X1 and feature center a1 is 0.11, the target similarity with feature center a2 is 0.82, and the target similarity with feature center a3 is 0.28, then the highest target similarity is 0.82. In this case, the diffusion category corresponding to feature center a2 (e.g., DDoS attack traffic) can be used as the first predicted diffusion category of diffusion traffic sample X1.

[0192] For example, if the diffused traffic sample X1 is obtained by diffused traffic from the original traffic sample in the sample diffusion category (e.g., port scan traffic), then the true category corresponding to the diffused traffic sample X1 should be port scan traffic. Therefore, the sample diffusion category corresponding to the diffused traffic sample X1 can be compared and verified with the first predicted diffusion category. If the first verification result shows that the two are consistent, it indicates that the pattern of the diffused traffic sample is consistent with the expectation, and the sample is of high quality and reliable. It can be used as the first diffused traffic sample for data feature distribution verification. If the first verification result shows that the two are inconsistent, it indicates that too much noise or other factors were introduced during the generation process, causing the sample to deviate from the original category features, and it should be discarded.

[0193] Furthermore, a specific implementation of verifying the data feature distribution of multiple first diffusion flow samples through a pre-trained screening network (e.g., a Siamese network) to obtain a second verification result will be described below. Please refer to the following text for details.

[0194] The above methods can effectively identify and eliminate low-quality samples with semantic deviations or abnormal patterns, which not only improves the authenticity and consistency of the diffusion traffic samples, but also enhances the overall quality and credibility of the diffusion traffic samples, and has significant practical application value.

[0195] In some implementations, to effectively identify and eliminate low-quality samples that deviate from the true feature distribution, a screening network can be used to perform a second verification on the first diffusion flow samples to further improve the overall quality and representativeness of the generated dataset. For example, (105.6) may include:

[0196] (105.6.1) Through a pre-trained filtering network, the second predicted diffusion category of each first diffusion flow sample is predicted based on the data feature distribution of each first diffusion flow sample. The filtering network learns the data feature distribution of multiple original flow samples corresponding to each sample diffusion category through an initial filtering network.

[0197] (105.6.2) By using the corresponding sample diffusion category of each first diffusion flow sample, the data feature distribution of the second predicted diffusion category is verified to obtain the second verification result.

[0198] The second predicted diffusion category can be the feature prediction category output by the pre-trained filtering network (Siamese) for the first diffusion flow sample.

[0199] The initial screening network can be an untrained Siamese network initial structure consisting of two identical sub-networks. By learning the feature distribution of the input original traffic sample pairs (positive examples are samples of the same class, and negative examples are samples of different classes), the network parameters can be optimized to distinguish the deep feature patterns of different categories, thus obtaining a trained screening network.

[0200] In this context, data feature distribution learning can be the process by which the initial screening network extracts class-discriminative features from the original traffic samples during the training phase. By minimizing the distance between samples of the same class and maximizing the distance between samples of different classes (such as triplet loss), the screening network can learn the key distribution characteristics of the original traffic samples for each sample diffusion class.

[0201] In some implementations, the initial screening network can learn the intra-class distribution compactness and inter-class separability of the original traffic samples to construct the discrimination boundary of the feature space, so that the distance between features of the same type of samples is less than a preset distance threshold, and the features of dissimilar samples are kept as far apart as possible, thereby training the screening network.

[0202] Furthermore, each first diffusion traffic sample can be input into a filtering network. Since the filtering network has pre-learned the data feature distribution of each sample's diffusion category, it can predict the second predicted diffusion category of each first diffusion traffic sample based on the data feature distribution of the first diffusion traffic sample. Then, the second predicted diffusion category is verified based on the actual sample diffusion category corresponding to the first diffusion traffic sample. If the second predicted diffusion category matches the sample diffusion category, it indicates that its inherent data feature distribution highly conforms to the typical pattern of its category. This first diffusion traffic sample successfully preserves the key features and complex patterns of the original traffic sample, is of high quality and reliable, and is suitable for subsequent model training or testing; it can be used as a target diffusion traffic sample. Otherwise, it is considered that too much noise, pattern distortion, or other factors were introduced during the generation process, causing it to deviate from the expected category feature distribution, and the first diffusion traffic sample is discarded.

[0203] For example, if the first diffused traffic sample X1 is obtained by diffused from the original traffic sample in the sample diffusion category (e.g., port scan traffic), then the true sample diffusion category corresponding to the first diffused traffic sample X1 should be port scan traffic. Therefore, the sample diffusion category corresponding to the first diffused traffic sample X1 can be compared and verified with the second predicted diffusion category. If the second verification result shows that the two are consistent, it is retained as the target diffused traffic sample; if the second verification result shows that the two are inconsistent, it is discarded.

[0204] By introducing a pre-trained filtering network to perform secondary verification of the data feature distribution of the first diffusion traffic sample, low-quality samples that have passed the category alignment verification but still deviate in the deeper feature distribution can be effectively identified. This ensures that the final target diffusion traffic samples are not only consistent in semantic labels, but also highly close to the real traffic in terms of data distribution characteristics, further improving the quality control accuracy and reliability of the generated samples.

[0205] Step 106: Based on the multiple verification results corresponding to multiple diffusion flow samples, filter the multiple diffusion flow samples to obtain the target diffusion flow sample corresponding to the sample diffusion category.

[0206] In some implementations, in order to ensure that the final retained target diffusion traffic samples are highly consistent and of high quality in terms of category labels and data feature distribution, the diffusion traffic samples can be screened based on the verification results corresponding to each diffusion traffic sample, thereby obtaining high-quality target traffic samples, which effectively solves the problem of insufficient minority class samples in the original traffic samples.

[0207] Among them, the target diffusion flow sample can be high-value synthetic flow data that has been screened through dual verification, and must pass both category alignment verification and data feature distribution verification.

[0208] For example, if the diffused traffic sample X1 is obtained by diffused traffic from the original traffic sample in the sample diffusion category (e.g., port scan traffic), then the true category corresponding to the diffused traffic sample X1 should be port scan traffic. Therefore, the sample diffusion category corresponding to the diffused traffic sample X1 can be compared and verified with the first predicted diffusion category. If the first verification result shows that the two are consistent, it indicates that the pattern of the diffused traffic sample is consistent with the expectation, and the sample is of high quality and reliable. It can be used as the first diffused traffic sample for data feature distribution verification. If the first verification result shows that the two are inconsistent, it indicates that too much noise or other factors were introduced during the generation process, causing the sample to deviate from the original category features, and it should be discarded.

[0209] For example, if the first diffused traffic sample X1 is obtained by diffused from the original traffic sample in the sample diffusion category (e.g., port scan traffic), then the true sample diffusion category corresponding to the first diffused traffic sample X1 should be port scan traffic. Therefore, the sample diffusion category corresponding to the first diffused traffic sample X1 can be compared and verified with the second predicted diffusion category. If the second verification result shows that the two are consistent, it is retained as the target diffused traffic sample; if the second verification result shows that the two are inconsistent, it is discarded.

[0210] In some implementations, a diffusion flow sample can be used as a target diffusion flow sample after both the first and second verification results are passed.

[0211] This application embodiment obtains the original traffic samples corresponding to the sample diffusion category; adds noise to the original traffic samples step by step according to multiple preset time steps to obtain intermediate diffusion samples; for each time step, the intermediate diffusion samples are input into a diffusion model, and the diffusion model identifies the temporal dependency and nonlinear relationship of the traffic features of the intermediate diffusion samples to obtain identification results, and outputs the predicted noise features corresponding to each time step based on the identification results; performs stepwise noise reduction processing on the intermediate diffusion samples based on the predicted noise features corresponding to each time step to obtain diffusion traffic samples; performs category alignment verification and data feature distribution verification on the diffusion traffic samples to obtain verification results; and filters the multiple diffusion traffic samples according to the multiple verification results corresponding to the multiple diffusion traffic samples to obtain the target diffusion traffic samples corresponding to the sample diffusion category. This approach allows for a more accurate simulation of dynamic changes and complex patterns in actual attacks by gradually adding noise. By utilizing the feature extraction components embedded in the diffusion model, it dynamically captures complex temporal patterns (such as the temporal evolution of attack behavior) and nonlinear relationships (such as the coordinated changes in payload features and network layer parameters) in traffic data. This overcomes the limitation of rule-based methods in simulating the dynamic changes of real attacks, enabling the generated diffusion samples to simulate the dynamic behavioral characteristics of real attack traffic. Furthermore, noise reduction is applied to make the generated diffusion traffic samples more realistic and diverse. On the other hand, the generated diffusion traffic samples undergo category alignment verification (ensuring semantic consistency of the attack) and data feature distribution verification (ensuring matching with the statistical distribution of real abnormal traffic), effectively filtering out low-quality or deviation-from-the-real-distribution generated samples, fundamentally improving the proportion and reliability of effective samples. In summary, this application can improve the quality of generated diffusion traffic samples.

[0212] In some implementations, to specifically address the residual imbalance problem of minority class samples after screening, the current sample quantity can be compared with the target ratio requirement to identify the category and quantity difference that needs to be enhanced, and the gap category can be further diffused to obtain more target diffusion flow samples, thereby constructing a high-quality training dataset that strictly conforms to the preset ratio. For example, after step 106, that is, after "screening multiple diffusion flow samples based on multiple validation results corresponding to multiple diffusion flow samples to obtain target diffusion flow samples corresponding to sample diffusion categories," the following may also be included:

[0213] (B.1) Obtain the sample category ratio among multiple pre-set sample diffusion categories, and the number of sample categories of multiple target diffusion traffic samples corresponding to each sample diffusion category;

[0214] (B.2) Based on the proportion and number of sample categories, determine the target diffusion category from multiple diffusion categories that needs to be increased, and the number to be increased;

[0215] (B.3) Generate the same number of target diffusion flow samples as the increase number, based on the multiple original flow samples corresponding to the target diffusion category.

[0216] The sample category ratio can be a pre-defined target percentage of each traffic category in the training set (e.g., normal traffic: scanning attack: DDoS attack = 70%: 15%: 15%), used to measure data balance.

[0217] The number of sample categories can be the actual number of currently generated target diffusion traffic samples in the corresponding sample diffusion category (e.g., the current number of samples for the scanning attack category is 850).

[0218] Among them, the target diffusion category can be a minority of categories whose current sample count is lower than the required proportion (e.g., the number of targets for scanning attacks should be 1500, but the current sample count is 800, so it is the target diffusion category). Data augmentation should be performed to generate more target diffusion traffic samples.

[0219] The increase in quantity can be the number of target diffusion traffic samples that need further diffusion for the target diffusion category.

[0220] Specifically, the sample category ratio can be set according to the needs of specific application scenarios. For example, in the field of network security, the sample category ratio can be determined based on historical data or expert experience. Simultaneously, the system also needs to count the actual number of sample categories for each sample diffusion category, corresponding to multiple target diffusion traffic samples. For instance, based on business needs or historical data analysis results, the ratio of each sample diffusion category can be set: 40% for DDoS attacks, 30% for port scanning, and 30% for SQL injection attacks.

[0221] Furthermore, the number of diffusion categories for each sample in the existing dataset can be counted to obtain the current number of sample categories. For example, the existing dataset may contain 200 DDoS attack samples, 150 port scan samples, and 100 SQL injection attack samples.

[0222] Furthermore, the current number of sample categories can be compared with the preset ratios. For example, if the total sample size is 450, then there should be 180 DDoS attacks (40%), 135 port scans (30%), and 135 SQL injection attacks (30%). If there are currently 200 DDoS attack samples, this already meets or exceeds the preset ratio; while there are only 100 SQL injection attack samples, significantly lower than the expected 135. Therefore, 35 more SQL injection attack samples are needed to reach the preset ratio.

[0223] At this point, based on multiple original traffic samples (or target diffusion traffic samples) corresponding to the target diffusion category, more target diffusion traffic samples can be generated in a loop, with the same number of samples added, until the target diffusion traffic samples contained in each diffusion category have reached the preset proportion of sample categories, at which point generation can stop. The specific method for obtaining target diffusion traffic samples has been explained above and will not be repeated here.

[0224] In some implementations, the number of samples in each sample diffusion category can be preset. When the number of sample categories containing multiple target diffusion traffic samples in a stored sample diffusion category is less than the number of samples, the number of target diffusion traffic samples that need to be added to that sample diffusion category is determined based on the difference between the number of samples and the number of sample categories. For example, if 500 target diffusion traffic samples need to be generated for sample diffusion category A, but only 200 target diffusion traffic samples have been generated so far, then the increase quantity can be determined to be 500 - 200 = 300, and the increase quantity is 300.

[0225] By using the above methods, the current sample quantity and preset ratio of each category can be clearly identified, thereby addressing the sample imbalance problem in a targeted manner. This not only effectively alleviates the model training bias caused by insufficient minority class samples, but also improves the overall diversity and representativeness of the dataset, providing higher quality and more balanced training data for security applications such as intrusion detection, and has significant practical application value.

[0226] Please see Figure 3 , Figure 3 As a general embodiment of the traffic data processing method, the following is combined with Figure 3 This section provides an overview of traffic data processing methods.

[0227] First, by deeply analyzing and preprocessing the raw network traffic data, key features are extracted and continuous, effective data is selected, laying the foundation for generating high-quality samples. Next, the RNN-KAN framework, with its powerful time-series modeling capabilities, is inserted into the diffusion model to enhance its ability to learn and generate complex traffic patterns. Then, the aforementioned fusion model is used to generate new diffusion traffic samples, enriching the dataset's diversity. To ensure the accuracy and consistency of the generated samples, the cosine similarity between the generated samples and the real samples is calculated to verify the consistency of their category labels. Then, the Siamese network is further used to deeply validate the data feature distribution of the generated samples, eliminating low-quality samples that deviate from reality in the feature space. Finally, through multiple iterative optimizations, the number of samples in each category is dynamically adjusted until a preset balanced ratio is achieved, thus constructing a high-quality, highly representative balanced dataset.

[0228] Please see Figure 4 This application also provides a traffic data processing apparatus that can implement the above-described traffic data processing method. The traffic data processing apparatus includes:

[0229] The acquisition module 41 is used to acquire the original traffic samples corresponding to the sample diffusion category;

[0230] Add module 42 to gradually add noise to the original flow sample according to multiple preset time steps to obtain intermediate diffusion samples;

[0231] The identification module 43 is used to input intermediate diffusion samples into the diffusion model for each time step, identify the temporal dependence and nonlinear relationship of the flow characteristics of the intermediate diffusion samples through the diffusion model, obtain the identification result, and output the predicted noise features corresponding to each time step based on the identification result.

[0232] Processing module 44 is used to perform stepwise noise reduction on intermediate diffusion samples based on the predicted noise features corresponding to each time step to obtain diffusion flow samples;

[0233] Validation module 45 is used to perform category alignment validation and data feature distribution validation on the diffusion flow samples to obtain validation results;

[0234] The filtering module 46 is used to filter multiple diffusion flow samples based on multiple verification results corresponding to multiple diffusion flow samples, and obtain the target diffusion flow sample corresponding to the sample diffusion category.

[0235] The specific implementation of this traffic data processing device is basically the same as the specific embodiment of the traffic data processing method described above, and will not be repeated here. Subject to meeting the requirements of the embodiments of this application, the traffic data processing device may also be equipped with other functional modules to implement the traffic data processing method in the above embodiments.

[0236] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described traffic data processing method. This computer device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0237] Please see Figure 5 , Figure 5 The hardware structure of a computer device according to another embodiment is illustrated. The computer device includes:

[0238] The processor 51 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0239] The memory 52 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 52 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 52 and called by the processor 51 to execute the traffic data processing method of the embodiments of this application.

[0240] Input / output interface 53 is used to implement information input and output;

[0241] The communication interface 54 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0242] Bus 55 transmits information between various components of the device (e.g., processor 51, memory 52, input / output interface 53, and communication interface 54);

[0243] The processor 51, memory 52, input / output interface 53, and communication interface 54 are connected to each other within the device via bus 55.

[0244] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described traffic data processing method.

[0245] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0246] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0247] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0248] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0249] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0250] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0251] It should be understood that in this application, "at least one" and "several" refer to one or more, and "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0252] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0253] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0254] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0255] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0256] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for processing traffic data, characterized in that, The method includes: Obtain the original traffic samples corresponding to the sample diffusion category; Noise is gradually added to the original flow sample according to multiple preset time steps to obtain intermediate diffusion samples; For each time step, the intermediate diffusion sample is input into the diffusion model. The diffusion model identifies the temporal dependency and nonlinear relationship of the flow characteristics of the intermediate diffusion sample to obtain the identification result. Based on the identification result, the predicted noise feature corresponding to each time step is output. Based on the predicted noise features corresponding to each time step, the intermediate diffusion samples are subjected to stepwise noise reduction processing to obtain diffusion flow samples; The diffusion flow samples were subjected to category alignment verification and data feature distribution verification to obtain the verification results; Based on multiple verification results corresponding to multiple diffusion flow samples, the multiple diffusion flow samples are filtered to obtain the target diffusion flow sample corresponding to the diffusion category of the sample.

2. The traffic data processing method according to claim 1, characterized in that, The verification results include a first verification result and a second verification result. The verification results obtained by performing category alignment verification and data feature distribution verification on the diffusion flow samples include: Multiple feature centers corresponding to multiple sample diffusion categories are obtained, wherein each feature center is obtained by averaging the features of multiple original traffic samples included in each sample diffusion category; The similarity is calculated between each diffusion flow sample included in each diffusion category and each feature center to obtain the target similarity. Based on the magnitude relationship between multiple target similarities, the first predicted diffusion category corresponding to each diffusion flow sample is determined; The sample diffusion category of each diffusion flow sample is aligned with the first predicted diffusion category to obtain a first verification result; Based on the first verification result, the plurality of diffusion flow samples are filtered to obtain the first diffusion flow sample; By using a pre-trained filtering network, the data feature distribution of multiple first diffusion flow samples is verified to obtain a second verification result.

3. The traffic data processing method according to claim 2, characterized in that, The second verification result is obtained by verifying the data feature distribution of multiple first diffusion flow samples through a pre-trained screening network, including: By using a pre-trained filtering network, a second predicted diffusion category for each first diffusion flow sample is predicted based on the data feature distribution of each first diffusion flow sample. The filtering network is obtained by learning the data feature distribution of multiple original flow samples corresponding to each sample diffusion category through an initial filtering network. By using the sample diffusion category corresponding to each first diffusion flow sample, the data feature distribution of the second predicted diffusion category is verified to obtain the second verification result.

4. The traffic data processing method according to claim 1, characterized in that, After filtering the multiple diffusion flow samples based on multiple verification results corresponding to the multiple diffusion flow samples to obtain the target diffusion flow sample corresponding to the sample diffusion category, the method further includes: Obtain the sample category ratio among multiple pre-set sample diffusion categories, and the number of sample categories for multiple target diffusion traffic samples corresponding to each sample diffusion category; Based on the sample category ratio and the number of sample categories, determine the target diffusion category from the plurality of sample diffusion categories that needs to increase the number of sample categories, and the number to be increased; Based on the multiple original traffic samples corresponding to the target diffusion category, generate the same number of target diffusion traffic samples as the increase quantity.

5. The traffic data processing method according to claim 1, characterized in that, The step of gradually adding noise to the original flow sample according to multiple preset time steps to obtain intermediate diffusion samples includes: According to a preset number of time steps, the noise scheduling coefficient corresponding to each time step is determined, wherein the noise scheduling coefficient of each time step increases sequentially according to the order of each time step; For the first time step, noise is added to each original traffic sample based on the corresponding noise scheduling coefficient to obtain the initial diffusion sample corresponding to the first time step; For the next time step after the first time step, noise is added to the initial diffusion sample based on the corresponding noise scheduling coefficient to obtain the updated initial diffusion sample for the next time step; Repeat the update for the next time step, add noise to the updated initial diffusion sample based on the corresponding noise scheduling coefficient, and obtain the updated initial diffusion sample corresponding to the updated next time step, until the updated next time step corresponding to the updated initial diffusion sample is the last time step of the plurality of time steps, and obtain the intermediate diffusion sample based on the updated initial diffusion sample.

6. The traffic data processing method according to claim 1, characterized in that, The intermediate diffusion samples are progressively denoised based on the predicted noise features corresponding to each time step to obtain diffusion flux samples, including: Based on the predicted noise features corresponding to each time step, the intermediate diffusion sample is denoised to obtain the updated intermediate diffusion sample corresponding to the previous intermediate time step of each time step. For the previous intermediate time step, the updated intermediate diffusion sample is input into the diffusion model to obtain the predicted noise features corresponding to the previous intermediate time step, and based on the predicted noise features, the updated intermediate diffusion sample corresponding to the previous intermediate time step is obtained. The process of repeatedly inputting the updated intermediate diffusion sample into the diffusion model for the previous intermediate time step to obtain the predicted noise features corresponding to the previous intermediate time step, and obtaining the updated intermediate diffusion sample corresponding to the updated previous intermediate time step based on the predicted noise features, continues until the updated previous time step corresponding to the updated intermediate diffusion sample is the first time step of the plurality of time steps, and the updated intermediate diffusion sample is used as the diffusion flow sample corresponding to the sample diffusion category.

7. The traffic data processing method according to claim 1, characterized in that, Before obtaining the original traffic sample corresponding to the sample diffusion category, the process also includes: Multiple preset traffic samples are obtained, and the number of unique values ​​in each preset traffic sample and the frequency of each unique value are determined. Discrete traffic samples are then selected from the multiple preset traffic samples. From the plurality of preset traffic samples, the discrete traffic samples are removed to obtain a plurality of original traffic samples.

8. A flow data processing device, characterized in that, The device includes: The acquisition module is used to acquire the original traffic samples corresponding to the sample diffusion category; An addition module is used to gradually add noise to the original flow sample according to multiple preset time steps to obtain intermediate diffusion samples; The identification module is used to input the intermediate diffusion sample into the diffusion model for each time step, identify the temporal dependency and nonlinear relationship of the flow characteristics of the intermediate diffusion sample through the diffusion model, obtain the identification result, and output the predicted noise feature corresponding to each time step based on the identification result. The processing module is used to perform stepwise noise reduction processing on the intermediate diffusion sample based on the predicted noise features corresponding to each time step to obtain the diffusion flow sample; The verification module is used to perform category alignment verification and data feature distribution verification on the diffusion flow samples to obtain verification results; The filtering module is used to filter the multiple diffusion flow samples based on the multiple verification results corresponding to the multiple diffusion flow samples, so as to obtain the target diffusion flow sample corresponding to the diffusion category of the sample.

9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the traffic data processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the traffic data processing method according to any one of claims 1 to 7.