A network traffic classification method and system based on spatiotemporal bias polar coordinate coherent field and cross-channel timing deduction

By using a spatiotemporally biased polar coordinate coherent field and cross-channel time series extrapolation method, static and dynamic features are decoupled, solving the distortion problem of encrypted traffic classification models in existing technologies under extremely small sample and dynamic network environments, and achieving high-precision and robust traffic classification.

CN122496430APending Publication Date: 2026-07-31NANJING UNIV OF INFORMATION SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV OF INFORMATION SCI & TECH
Filing Date
2026-05-28
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing deep learning traffic classification models struggle to decouple static spatial attributes from dynamic temporal distortions when dealing with encrypted traffic. Furthermore, feature extraction is prone to distortion in extremely small sample sizes and dynamic network environments, failing to effectively follow temporal causality and resulting in insufficient classification accuracy and anti-interference capabilities.

Method used

The method employs a spatiotemporal bias polar coordinate coherent field and cross-channel time series extrapolation, generating a static capacity polarization field and a dynamic time-delay bias field through polar coordinate mapping. Combined with self-supervised learning and orthogonal prototype projection, it forces the model to follow the unidirectional time evolution law, strips away static and dynamic features, and uses minimal samples for model fine-tuning.

Benefits of technology

It significantly improves the accuracy and anti-interference ability of encrypted traffic classification in scenarios with extremely small sample sizes. The model maintains high accuracy in extreme jitter environments and can converge quickly with only a small amount of labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122496430A_ABST
    Figure CN122496430A_ABST
Patent Text Reader

Abstract

This invention provides a network traffic classification method and system based on spatiotemporal bias polar coordinate coherent fields and cross-channel time-series extrapolation. The method first extracts the packet size and arrival time interval of the original data packets to construct a basic feature sequence. Then, a temporal decay kernel and phase shift are introduced to map the sequence into a static capacity polarization field and a dynamic time-delay bias field, which are orthogonally fused to form a multi-channel tensor. In a label-free environment, an adaptive mask is generated using local spatial differential gradients. Combined with a cross-attention model embedding a temporal penalty matrix and a local gradient divergence penalty loss, burst traffic patterns are extrapolated under strict unidirectional time constraints. Finally, high-precision classification is achieved through fine-tuning with a normalized coherent projection module and a very small amount of labeled data. This method reconstructs the topological mapping and strict temporal constraints of traffic features, effectively corrects the defects of conventional models that violate communication logic, remains stable under extreme network jitter, and is suitable for scenarios with extremely small sample sizes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of cyberspace security and network traffic analysis technology, specifically to a network traffic classification method and system based on spatiotemporal bias polar coordinate coherent field and cross-channel time series extrapolation. Background Technology

[0002] With the widespread adoption of encrypted communication technologies on the Internet, the extensive use of HTTPS, VPNs, and cloud services has led to encrypted data dominating network traffic. Traditional traffic classification methods based on port numbers and deep packet inspection (DPI) exhibit significant limitations when facing encrypted payloads and dynamic network environments. Therefore, deep learning, especially image-mapping-based traffic classification methods, has gradually become the mainstream technology for encrypted traffic classification because it can transform abstract one-dimensional traffic features into visualized two-dimensional images and leverage the powerful spatial feature extraction capabilities of neural networks to uncover complex patterns. To cope with complex network congestion variations and consider the underlying evolutionary characteristics of network protocols, how to effectively decouple traffic data spatiotemporally and extract features with strong anti-interference capabilities has become a key issue in this field.

[0003] Existing technologies have attempted to classify network traffic using image mapping and generative augmentation. For example, the invention "Tor User Website Access Recognition Method and System Based on Gram Corner Field Transform (CN114710417B)" discloses a method for classifying one-dimensional sequences by converting them into two-dimensional feature matrices using Gram Corner Field Transform (GASF). The invention "A Data Protection Method for Federated Learning Intrusion Detection Based on GAN (CN113468521B)" discloses a method for using Generative Adversarial Networks (GANs) to simulate and generate anomalous data to augment the training dataset. The invention "Network Traffic Generation Data Augmentation Method Based on Diffusion Model (CN118282948A)" discloses converting raw traffic into image form and inputting it into a diffusion model to synthesize a classification model with effectively augmented data. While these methods enrich the data format and improve the visual dimension of features to some extent, they do not deeply consider the following two issues:

[0004] (1) Existing flow-image mapping mechanisms generally suffer from severe symmetric redundancy and fail to decouple static spatial attributes from dynamic temporal distortions. Existing technologies (such as CN114710417B mentioned above) typically use standard transformation methods to directly map a single one-dimensional statistical sequence into a two-dimensional image. This direct numerical mapping often generates a highly symmetric feature matrix, completely ignoring the crucial unidirectional and irreversible (i.e., asymmetric) temporal flow during network data packet transmission. At the same time, this simple mapping fails to effectively separate the features representing the static load scale from the time delay features representing dynamic congestion in the network flow, resulting in a high degree of coupling in the feature space. When facing dynamic environments such as network jitter, the model is easily disturbed by surface-fluctuating time delay noise and struggles to capture the underlying stable capacity topology.

[0005] (2) In the process of learning and pre-training complex temporal features, existing methods are difficult to follow strict temporal causality and are prone to overfitting to sudden congestion features. Mainstream deep learning models (including the above-mentioned methods that rely on GAN or diffusion models, as well as conventional self-supervised attention mechanisms) often allow bidirectional flow of global information when extracting and generating features. This means that the model will violate the underlying temporal logic of the network protocol and use the "future" congestion state to infer the "historical" packet features (i.e., temporal information leakage). Especially when dealing with sudden distortions such as high latency and packet loss (high gradient change regions), the model is often dominated by these invalid congestion burst noises and fails to extract the true consistency of local packet jitter. This not only destroys the true physical topology of the data, but also severely limits the model's ability to orthogonally decouple and finely identify different protocol features in the absence of massive artificial labels (extremely small samples).

[0006] In summary, there is an urgent need for an encrypted traffic classification scheme that can explicitly break time symmetry, enforce unidirectional time evolution constraints, and achieve complete separation of dynamic and static features without destroying the real underlying protocol topology. This would overcome the symmetric redundancy limitations and distortion defects caused by the high coupling of spatiotemporal features in traditional two-dimensional image mapping. Summary of the Invention

[0007] Objective: This invention addresses the shortcomings of existing deep learning traffic classification models, which neglect underlying physical laws, easily exploit future states to violate temporal causality, and suffer from feature extraction distortion in extremely small sample sizes and dynamic network congestion environments. It provides a network traffic classification method and system based on spatiotemporal biased polar coordinate coherent fields and cross-channel temporal inference. This solves the problems of existing technologies over-reliance on surface pixel correlations and difficulty in decoupling static and dynamic features. Furthermore, it forces the model to learn the microscopic consistency and unidirectional temporal evolution law of network packet jitter topology under extreme conditions lacking large-scale manual labels, significantly improving the model's classification accuracy and anti-interference generalization ability in extremely small sample scenarios.

[0008] The method includes the following steps:

[0009] Step 1: Preprocess and extract low-level features from the collected raw network traffic data: directly extract the packet size and arrival time interval of network data packets, and construct normalized packet size sequence and normalized time interval sequence that retain the original low-level features after sequence alignment and maximum-minimum normalization.

[0010] Step 2, Spatiotemporal Feature Mapping and Orthogonal Fusion of Multi-channel Tensors: Based on polar coordinate mapping, a temporal decay kernel and a phase shift mechanism are introduced to convert the normalized packet size sequence into a first-channel static capacity polarization field matrix representing the static scale attribute of the network flow, and the normalized time interval sequence into a second-channel dynamic time delay bias field matrix representing the direction of dynamic congestion evolution; the first and second channels are fused according to spatial coordinate alignment to generate a multi-channel image tensor with asymmetric temporal features;

[0011] Step 3, self-supervised pre-training based on original gradient-driven masking and cross-channel temporal constraints: construct a self-supervised learning task using unlabeled multi-channel image tensors; extract the local spatial difference gradient of the first channel matrix to adaptively mask the high gradient burst feature region of the second channel; use a cross-attention model with embedded temporal penalty matrix to perform unidirectional time series cross-channel extrapolation based solely on the static first channel features, and combine local gradient divergence penalty to calculate composite coherent error and update network parameters;

[0012] Step 4, Perform orthogonal prototype projection and adaptive boundary fine-tuning under minimal sample size: After pre-training, the predictor is removed and the context encoder is retained. An orthogonal prototype projection module is connected to the end of the context encoder. Using traffic data with real labels, which accounts for no more than 5% of the total data volume of the target scene, the global features are explicitly decomposed into static and dynamic independent orthogonal components. A loss function that fuses the original data boundary features and the orthogonal decoupling regularization term is used for fine-tuning and updating. Finally, the classification prediction results of network traffic are output.

[0013] Step 1 includes:

[0014] Step 1-1: Capture the raw network data stream, group it into independent streams by 5-tuples, sort the packets within each stream in ascending order by timestamp, and extract the packet size of each packet. and the arrival time interval of adjacent data packets ;

[0015] Steps 1-2 involve truncating or zero-padding the extracted packet size sequence and arrival time interval sequence to a uniform sequence length of N, directly retaining the original physical variables at the lowest level, and constructing the packet size sequence. With time interval sequence ;in, This represents the normalized bag size value at the Nth position in the normalized bag size sequence. This represents the normalized arrival time interval value at the Nth position in the normalized time interval sequence;

[0016] Steps 1-3 employ max-min normalization to map and transform each original value in the packet size sequence S and time interval sequence T according to the following formula:

[0017] ,

[0018] in, This represents the i-th original data value in the sequence; This represents the maximum value of all original data in the sequence; This represents the minimum value of all the original data in the sequence; This represents the i-th normalized value obtained after calculation, whose value range is strictly within the interval [−1, 1]. Through mapping transformation, the values ​​of the packet size sequence S and the time interval sequence T are mapped to [the specified values]. The interval is used to obtain the normalized bag size sequence. With normalized time interval sequence .

[0019] Step 2 includes:

[0020] Step 2-1, Construct the static capacity polarization field of the first channel: Obtain the normalized packet size sequence of length N output from the preprocessing stage. , convert the sequence Element-wise mapping to polar coordinates; for sequences Each normalized packet size value in Calculate and extract using the inverse cosine function Corresponding spatial angle information Subsequently, all data packet pairs in the sequence are traversed, the sum of the angles corresponding to any two time steps i and j is calculated, and combined with the temporal decay kernel based on index distance, a kernel of size i is constructed. The two-dimensional feature matrix; the polarization field pixel value in the i-th row and j-th column of the generated feature matrix. The calculation formula is:

[0021] ,

[0022] in, This is a preset smoothing control coefficient used to adjust the numerical decay rate in the off-diagonal region of the two-dimensional matrix. This represents the size of the j-th normalized bag. Corresponding spatial angle information;

[0023] by As the fundamental polar coordinate eigenvalues, and The multiplier, used as the distance mapping weight, is based on the relative span of the data packet in the original sequence. Calculate and generate the complete static capacity polarization field matrix C pixel by pixel;

[0024] Step 2-2, Construct the second channel dynamic time-delay bias field: Synchronously acquire a normalized time interval sequence of length N. , convert the sequence Map to polar coordinates and extract the angle components corresponding to each time interval. ,in For sequence The i-th normalized time interval; when calculating the pixel association value of the two-dimensional matrix, an directional offset linearly related to the sequence position span is introduced, resulting in a size of... The pixel value in the i-th row and j-th column of the feature matrix The calculation formula is:

[0025] ,

[0026] in, The bias gain constant is set to control the magnitude of the phase shift; Represents the j-th normalized time interval The corresponding angular components;

[0027] During calculation, additional components will be added. The injected cosine function serves as a directional time-delay phase bias. Because the bias term is affected by the sign of the difference between row number i and column number j, the pixel value generated when the row and column indices are interchanged... Thus, the asymmetric dynamic time-delay bias field matrix D is generated point by point;

[0028] Steps 2-3, orthogonal fusion of coherent fields: The static capacity polarization field matrix C, after time decay mapping, is set as the first channel view of the network flow feature, and the dynamic time-delay bias field matrix D, after phase bias mapping, is set as the second channel view of the network flow feature; the first and second channels are strictly aligned according to row and column pixel coordinates in spatial scale, and then superimposed and merged along the channel dimension to finally generate a spatial shape as follows. The multi-channel image tensor is used as the target tensor and input into the downstream model for subsequent processing.

[0029] Step 3 includes the following steps:

[0030] Step 3-1, Cross-channel decoupling and information flow isolation: The first channel static capacity polarization field matrix C, representing the static scale of the network flow, is input separately to the context encoder; the second channel dynamic time delay bias field matrix D, representing dynamic congestion and time evolution, is input to the target encoder containing the exponential moving average (EMA) mechanism.

[0031] Step 3-2, Adaptive temporal masking strategy based on intrinsic gradient: Extract the local spatial difference gradient of the static capacity polarization field matrix C at the input of the context encoder. Because regions with drastic gradient changes correspond to extreme abrupt changes in network packet size and potential precursors to congestion, The absolute value of the probability density function generates an adaptive mask matrix, which cuts off the feature blocks of the dynamic time delay bias field matrix D corresponding to the high gradient region. This strategy makes the model focus on the burst state region of network traffic and infer the dynamic time evolution characteristics that are prone to distortion across modes based solely on the packet size topology during the steady period.

[0032] Step 3-3: Input the unmasked, complete dynamic time-delay bias field matrix D separately into the target encoder to extract the true high-dimensional topological feature vector as the standard answer. .

[0033] Step 3 also includes:

[0034] Steps 3-4: Temporal Cross-Attention Inference Based on Temporal Penalty: The predictor receives incomplete features from the context encoder and performs cross-channel inference using a Transformer architecture that includes a multi-head cross-attention mechanism. To comply with the irreversible temporal order rule of network data packets and prevent the model from using future states to infer historical congestion, a temporal penalty matrix is ​​embedded in the cross-attention formula. The specific formula is as follows:

[0035] ,

[0036] in, To infer the output feature matrix for cross-attention, softmax is the activation function, K' is the transpose of the key feature matrix K, and Q, K, and V represent the query feature matrix, key feature matrix, and value feature matrix, respectively. The query feature matrix Q is derived from the sequence of incomplete features restored by masking, while the key feature matrix K and value feature matrix V are derived from the surrounding context features. The scaling dimension of the feature vector; Temporal penalty matrix The element in the i-th row and j-th column of the attention graph, when the j-th time step corresponding to the element is greater than the i-th time step. ,otherwise ; The extremely large positive penalty scalar restricts the neuron weights strictly to a unidirectional time axis, ultimately resulting in the output predicting time delay features. ;

[0037] Steps 3-5, calculation of spatiotemporal topological composite coherence loss: a local gradient divergence penalty term is introduced to reconstruct the error evaluation mechanism in order to calculate the prediction time delay characteristics. The true features extracted by the target encoder The topology loss error L between them; this composite formula, while aligning the global high-dimensional vector direction, also enforces microscopic consistency constraints on the local packet jitter topology, and the specific formula is as follows:

[0038] ,

[0039] Where B is the batch size for training; the first part of the formula is the global normalized coherence error, and the second part is the local spatial difference gradient divergence penalty term. Let be the prediction time-delay feature vector of the k-th sample; Let be the target true feature vector of the k-th sample; The predicted local spatial difference gradient for the k-th sample; The true local spatial difference gradient of the k-th sample; This represents the sum of the squared 2 norms of the differences between the predicted local spatial difference gradients and the true local spatial difference gradients for all samples in a batch. The weighting coefficients for the topology penalty term;

[0040] Steps 3-6, Asymmetric update of model parameters: Based on the backpropagation algorithm, the network parameters of the context encoder and predictor are updated in real time using the topological loss error L;

[0041] Steps 3-7 involve performing a gradient stopping operation on the target encoder to disable backpropagation updates; then, using the exponential moving average (EMA) algorithm, the parameters of the target encoder are slowly and smoothly updated using the parameters of the context encoder, as shown in the formula:

[0042] ,

[0043] Among them, symbols The parameter assignment update symbol indicates that the calculation result on the right is assigned and updated to the variable on the left. For the target encoder parameters, Here are the parameters for the context encoder, and λ is the smoothing coefficient.

[0044] Step 4 specifically includes:

[0045] Step 4-1, Extract global topological features: Input the complete target traffic multi-channel image tensor with real labels into the trained context encoder to extract high-dimensional global fusion topological features. ;

[0046] Step 4-2, Construct the orthogonal prototype projection classification head: Keep the main parameters of the context encoder frozen, and connect the orthogonal prototype projector to the tail of the context encoder; the orthogonal prototype projector contains a first orthogonal transformation matrix. With the second orthogonal transformation matrix The first orthogonal transformation matrix The transpose of the second orthogonal transformation matrix The product is a zero matrix; using the first orthogonal transformation matrix Global fusion features Projecting onto the static capacity eigenspace yields the static capacity eigencomponents. Simultaneously, using the second orthogonal transformation matrix Global fusion features Projecting onto the dynamic time-delay distortion subspace yields the dynamic time-delay distortion components. This makes the static capacity eigencomponents With dynamic time delay distortion components Mathematically, they are mutually orthogonal and independent; subsequently, the static capacity eigencomponents are calculated separately in the two-manifold subspace. The Mahalanobis distance C1 from the prototype center of each protocol category, and the dynamic time-delay distortion component. The Mahalanobis distance C2 between each protocol category prototype center is calculated, and C1 and C2 are adaptively fused using dynamic uncertainty weights to output the joint probability distribution of each protocol type.

[0047] Step 4-3, Fine-tuning of coherent large-edge decoupling loss: Fine-tuning the model using real label data that accounts for no more than 5% of the total data volume in the target scene, and using large-edge decoupling loss as the objective function. .

[0048] In step 4-3, the objective function The calculation formula is:

[0049] ,

[0050] in, To fine-tune the objective function; To fine-tune the total number of samples; Let i be the true label of the i-th sample; The feature components of the i-th sample and the true protocol category The angle between the centers of the prototypes; is the angle between the feature component and the center of the c-th non-real protocol category prototype; m is the physical coherence edge penalty term, used to constrain and obtain a more compact intra-class distribution during fine-tuning; This is the scaling factor; For orthogonal decoupling regularization terms, where Let be the transpose of the static capacity eigencomponents. This is a dynamic time-delay distortion component; The weighting coefficient is used to force the static and dynamic feature components to maintain absolute mathematical orthogonality during the fine-tuning stage. By calculating this loss function and backpropagating to update the parameters of the projection classification module, the downstream network traffic classification prediction results are output.

[0051] The present invention also provides a network traffic classification system based on spatiotemporal bias polar coordinate coherent field and cross-channel time series extrapolation for implementing the method, comprising:

[0052] The spatiotemporal topology mapping and multi-channel tensor orthogonal fusion module is used to introduce a temporal decay kernel and phase shift mechanism on the basis of polar coordinate mapping. It converts the normalized sequence into the first channel static capacity polarization field matrix representing the static scale attribute of the network flow and the second channel dynamic time delay bias field matrix representing the direction of dynamic congestion evolution. It then generates a multi-channel image tensor with asymmetric temporal characteristics by aligning it with spatial coordinates.

[0053] A self-supervised pre-training module with physical gradient-driven masking and cross-channel temporal constraints is used to extract the local spatial difference gradient of the first channel matrix to adaptively mask the burst feature region of the second channel; and a cross-attention model with embedded temporal penalty matrix is ​​used to unidirectionally deduce the missing dynamic temporal evolution features, and combined with local gradient divergence penalty to calculate composite coherent error and update network parameters.

[0054] The minimal sample orthogonal prototype projection and adaptive boundary fine-tuning module is used to strip the predictor and retain the context encoder after pre-training. An orthogonal prototype projector is connected at the end. It uses a very small amount of data with real labels to decompose the global features into static and dynamic independent orthogonal components. It is then fine-tuned using a loss function that fuses physical coherent edges and orthogonal decoupling regularization terms, and finally outputs the classification prediction results of network traffic.

[0055] The present invention also provides a traffic classification device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method described.

[0056] The present invention also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method described.

[0057] Compared with the prior art, the present invention has the following beneficial effects:

[0058] 1. This invention addresses the difficulty of existing methods in characterizing asymmetric temporal evolution by innovatively proposing a spatiotemporally biased polar coordinate coherent field. By introducing a temporal decay kernel and a phase shift mechanism into the polar coordinate mapping, a one-dimensional sequence is transformed into a static capacity polarization field and an asymmetric dynamic time-delay bias field. This feature space not only preserves the topological scale of the packet size but also perfectly quantifies the temporal irreversibility and evolution direction of network flows, completely solving the defect of traditional two-dimensional image transformation that destroys the underlying physical laws, thus enabling the model to acquire extremely strong resistance to network jitter.

[0059] 2. This invention addresses the problems of overfitting to invalid noise and difficulty in isolating temporal causality in network feature extraction by designing a gradient-driven masking and orthogonal prototype fine-tuning architecture. During the pre-training phase, static intrinsic gradients are used to adaptively mask dynamic burst regions, and a strict temporal penalty matrix is ​​embedded to force the model to follow a unidirectional time flow for prediction, avoiding the logical collapse of "using the future to predict the past." During the fine-tuning phase, orthogonal prototype decomposition completely decouples static attributes from dynamic distortion. This method enables the model to converge quickly with less than 5% of the data, significantly improving the accuracy and robustness of encrypted traffic classification in small-sample scenarios across different environments. Attached Figure Description

[0060] Figure 1 This is a flowchart of the steps of the method described in this invention.

[0061] Figure 2 This is a schematic diagram illustrating the principle of spatiotemporal feature mapping and orthogonal fusion of multi-channel tensors.

[0062] Figure 3This is a diagram of a self-supervised pre-trained network architecture based on gradient masking and temporal constraints.

[0063] Figure 4 This is a flowchart for fine-tuning and real-time classification for extremely small samples.

[0064] Figure 5 A visualization of the coherent field and mask features in spatiotemporally biased polar coordinates.

[0065] Figure 6 The figure shows the results of cross-scenario technology effect verification and comparison experiments. Detailed Implementation

[0066] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.

[0067] This embodiment provides a network traffic classification method based on spatiotemporal bias polar coordinate coherent field and cross-channel time series extrapolation. In this embodiment, the network traffic classification method can be deployed on a server with parallel computing capabilities (such as a hardware environment equipped with a high-performance GPU), and the underlying software environment can be implemented based on Python and mainstream deep learning frameworks.

[0068] like Figure 1 As shown, the overall process flow of the method described in this invention mainly includes the following steps:

[0069] Step 1: Preprocess the collected raw network traffic data and extract its underlying physical features:

[0070] Step 1-1: Deploy probes at network boundary nodes to capture raw network data streams (e.g., pcap files). Group the data into independent flows based on the network 5-tuple (source IP, destination IP, source port, destination port, transport layer protocol). Sort the packets within each flow in ascending order of timestamps and directly extract the raw characteristics of each packet: packet size S (preferably in bytes) and the arrival time interval T between adjacent packets (preferably in milliseconds).

[0071] Steps 1-2 involve standardizing the extracted packet size sequence and arrival time interval sequence. The standardization sequence length is set to N (preferably N=128 in this embodiment). Zero-padding is used when the sequence length is less than N, and truncation is performed when it exceeds N. This constructs a packet size sequence with a uniform length. With time interval sequence .

[0072] Steps 1-3 employ max-min normalization to strictly map the values ​​of the packet size sequence S and the time interval sequence T to the closed interval [-1, 1], thus obtaining the normalized packet size sequence. With normalized time interval sequence The specific mapping formula is as follows:

[0073] ,

[0074] Step 2: Orthogonal fusion of spatiotemporal feature mapping and multi-channel tensor. For example... Figure 2 As shown, the specific process of orthogonal fusion of the spatiotemporal feature mapping and multi-channel tensor is as follows:

[0075] Step 2-1, Construct the first channel (static capacity polarization field): Obtain the normalized packet size sequence of length N output from the preprocessing stage. , convert the sequence Element-wise mapping to polar coordinates; for sequences Each normalized packet size value in Calculate and extract using the inverse cosine function Corresponding spatial angle information Subsequently, all data packet pairs in the sequence are traversed, the sum of the angles corresponding to any two time steps i and j is calculated, and combined with the temporal decay kernel based on index distance, a kernel of size i is constructed. The two-dimensional feature matrix; the polarization field pixel value in the i-th row and j-th column of the generated feature matrix. The calculation formula is:

[0076] ,

[0077] in, The preset smoothing control coefficient (preferably set to 0.1 in this embodiment) is used to adjust the numerical decay rate of the off-diagonal region in the two-dimensional matrix; This represents the size of the j-th normalized bag. Corresponding spatial angle information; with As the fundamental polar coordinate eigenvalues, and The multiplier, used as the distance mapping weight, is based on the relative span of the data packet in the original sequence. The complete static capacity polarization field matrix C is calculated and generated pixel by pixel.

[0078] Step 2-2, Construct the second channel (dynamic time-delay bias field): Synchronously acquire a normalized time interval sequence of length N. Then map to the polar coordinate system and extract the angle components. Introducing a directional offset linearly related to the sequence position span, generating a size of... The feature matrix, the pixel value in the i-th row and j-th column. The calculation formula is:

[0079] ,

[0080] in, This is the bias gain constant (preferably set to 0.05 in this embodiment). Additional component. This allows when row and column indexes are swapped This allows for the generation of an asymmetric dynamic time-delay bias field matrix D point by point.

[0081] Steps 2-3, orthogonal fusion of coherent fields: Set matrix C as the first channel view and matrix D as the second channel view. Maintain strict alignment of the two channels according to row and column pixel coordinates in spatial scale, and superimpose and merge them along the channel dimensions to finally generate a spatial shape as follows. Multichannel image tensors.

[0082] Step 3: Self-supervised pre-training based on the original gradient-driven masking and cross-channel temporal constraints. For example... Figure 3 As shown, the self-supervised pre-trained network architecture based on gradient masking and temporal constraints, and its internal data processing flow specifically include:

[0083] Step 3-1, Cross-channel decoupling and information flow isolation: Isolation is performed at the input of the two-branch network model (such as the two-branch VisionTransformer). The first channel matrix C, representing static scale, is input separately to the context encoder; the second channel matrix D, representing dynamic congestion, is input to the target encoder.

[0084] Step 3-2, Adaptive temporal masking strategy based on intrinsic gradient: Extract the local spatial difference gradient of the static matrix C. The system is based on The absolute value of the probability density function generates an adaptive mask matrix, precisely cutting off the feature blocks of the dynamic matrix D corresponding to high-gradient regions. This strategy allows the model to focus on regions of sudden states, where dynamic features easily become distorted when extrapolated across modalities from static topology during stable periods. The unmasked complete matrix D is then input separately into the target encoder to extract the true high-dimensional topological feature vector. .

[0085] Step 3-3, Cross-Attention Inference Based on Temporal Penalty: The predictor receives incomplete features and performs cross-channel inference using the Transformer architecture. To comply with the rule of temporal irreversibility, a temporal penalty matrix is ​​embedded in the cross-attention function. The specific formula is as follows:

[0086] ,

[0087] in, The cross-attention derivation outputs a feature matrix, with softmax as the activation function. K' is the transpose of the key feature matrix K. The query feature matrix Q is derived from the sequence of incomplete features restored through masking. The key feature matrix K and the value feature matrix V are derived from the surrounding context features. The scaling dimension of the feature vector; Temporal penalty matrix The element in the i-th row and j-th column of the attention graph, when the j-th time step corresponding to the element is greater than the i-th time step. ,otherwise ; The extremely large positive penalty scalar restricts the neuron weights strictly to a unidirectional time axis, ultimately resulting in the output predicting time delay features. .

[0088] Steps 3-4, Spatiotemporal topology composite coherence loss calculation: A local gradient divergence penalty term is introduced to enforce microscopic consistency of the packet jitter topology. The formula for its composite error function is:

[0089] ,

[0090] Where B is the batch size for training; the first part of the formula is the global normalized coherence error, and the second part is the local spatial difference gradient divergence penalty term. Let be the prediction time-delay feature vector of the k-th sample; Let be the target true feature vector of the k-th sample; The predicted local spatial difference gradient for the k-th sample; The true local spatial difference gradient of the k-th sample; This represents the sum of the squared 2 norms of the differences between the predicted local spatial difference gradients and the true local spatial difference gradients for all samples in a batch. The weighting coefficient for the topology penalty term (preferably 0.5).

[0091] Steps 3-6, Asymmetric update of model parameters: Based on the backpropagation algorithm, the network parameters of the context encoder and predictor are updated in real time using the topological loss error L.

[0092] Steps 3-7 involve performing a gradient stopping operation on the target encoder to disable backpropagation updates; then, using the exponential moving average (EMA) algorithm, the parameters of the target encoder are slowly and smoothly updated using the parameters of the context encoder, as shown in the formula:

[0093] ,

[0094] Among them, symbols The parameter assignment update symbol indicates that the calculation result on the right is assigned and updated to the variable on the left. For the target encoder parameters, Here are the parameters for the context encoder, and λ is the smoothing coefficient (preferably starting at 0.996 and gradually increasing to 1 during training).

[0095] Step 4: Orthogonal AC projection and adaptive boundary fine-tuning under minimal sample conditions. For example... Figure 4 As shown, the specific process of the very small sample fine-tuning and real-time classification is as follows:

[0096] Step 4-1, Extract global topological features: Input the complete target traffic multi-channel image tensor with real labels into the trained context encoder to extract high-dimensional global fusion topological features. ;

[0097] Step 4-2, Construct the orthogonal prototype projection classification head: Keep the main parameters of the context encoder frozen, and connect the orthogonal prototype projector to the tail of the context encoder; the orthogonal prototype projector contains a first orthogonal transformation matrix. With the second orthogonal transformation matrix The first orthogonal transformation matrix The transpose of the second orthogonal transformation matrix The product is a zero matrix; using the first orthogonal transformation matrix Global fusion features Projecting onto the static capacity eigenspace yields the static capacity eigencomponents. Simultaneously, using the second orthogonal transformation matrix The global fusion features Projecting onto the dynamic time-delay distortion subspace yields the dynamic time-delay distortion components. This makes the static capacity intrinsic components With the dynamic time-delay distortion component Mathematically, they are mutually orthogonal and independent; subsequently, the static capacity eigencomponents are calculated separately in the two-manifold subspace. The Mahalanobis distance to the prototype center of each protocol category, and the dynamic time-delay distortion component. The Mahalanobis distance between the two sets of distances and the prototype centers of each protocol type is calculated, and the two sets of distances are adaptively fused using dynamic uncertainty weights to output the joint probability distribution of each protocol type.

[0098] Step 4-3, Fine-tuning of coherent large-edge decoupling loss: Fine-tuning the model using real label data that accounts for no more than 5% of the total data volume in the target scene, and using large-edge decoupling loss as the objective function. .

[0099] In step 4-3, the objective function The calculation formula is:

[0100] ,

[0101] in, To fine-tune the objective function; To fine-tune the total number of samples; Let i be the true label of the i-th sample; The feature components of the i-th sample and the true protocol category The angle between the centers of the prototypes; is the angle between the feature component and the center of the c-th non-real protocol category prototype; m is the physical coherence edge penalty term, used to constrain and obtain a more compact intra-class distribution during fine-tuning; This is the scaling factor; For orthogonal decoupling regularization terms, where Let be the transpose of the static capacity eigencomponents. This is a dynamic time-delay distortion component; The weighting coefficient is used to force the static and dynamic feature components to maintain absolute mathematical orthogonality during the fine-tuning stage. By calculating this loss function and backpropagating to update the parameters of the projection classification module, the downstream network traffic classification prediction results are output.

[0102] Accordingly, this embodiment also provides a traffic classification device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the method.

[0103] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method.

[0104] A specific experimental implementation and technical effect verification:

[0105] To further demonstrate the feasibility of this invention and its significant breakthroughs compared to existing technologies, a complete and specific experimental implementation is provided below, using a real encrypted network traffic dataset (taking the ISCX VPN dataset for pre-training and the ISCX Tor dataset for fine-tuning and testing as an example). The overall experimental deployment strictly follows the methodological architecture proposed in this invention: first, a self-supervised learning task is constructed using large-scale unlabeled multi-channel image tensors; then, adaptive boundary fine-tuning and testing are performed using a very small proportion of real labeled data from the target scene.

[0106] In the specific parameter setting and feature mapping stage, this embodiment uniformly truncates the network flow sequence length N=128. For the static capacity polarization field of the first channel, the smoothing control coefficient is set to 0.1; for the dynamic time-delay bias field of the second channel, the bias gain constant is set to 0.05. Visualization of intermediate experimental results (e.g.) Figure 5 As shown, Figure 5 The horizontal and vertical axes of each subgraph represent the sequence position indexes of data packets extracted from the network flow (the value range corresponds to the sequence length), clearly demonstrating that this invention successfully achieves orthogonal fusion and complete decoupling of spatiotemporal features: the static capacity polarization field accurately preserves the symmetry property of the network packet capacity scale; while the dynamic time-delay bias field breaks the time symmetry through directional time-delay phase bias, generating asymmetric dynamic evolution features. In the cross-channel extrapolation stage, the model accurately extracts the local spatial difference gradient of the static field and uses it to generate an adaptive mask matrix, accurately cutting off the corresponding high-gradient burst feature regions. At the same time, a stop-gradient operation is performed on the target encoder to prohibit backpropagation, and the exponential moving average (EMA) algorithm is used to slowly and smoothly update the parameters, ensuring the absolute stability of the composite coherent error calculation and the asymmetric update of model parameters.

[0107] Detailed comparative experimental data (such as) Figure 6 (As shown) further verifies that the present invention effectively overcomes the inherent defects pointed out in the background art. First, in the comparison of feature mapping mechanisms, the spatiotemporal bias polar coordinate coherent field of the present invention achieves a classification accuracy of up to 97.5% on the target test set. In contrast, existing technologies (such as CN114710417B) directly use the standard Gram angle field (GASF) for simple mapping, which has a serious symmetric redundancy and its accuracy is only 68.5%. Second, in the anti-interference capability test for dealing with complex network congestion changes, when the test set is injected with extreme latency and packet loss jitter of up to 60%, existing data augmentation methods that rely on generative adversarial networks (GAN) or diffusion models (such as CN113468521B) are prone to overfitting to sudden congestion features, causing the accuracy to plummet from 88% to below 48%; while the present invention, with its cross-attention model embedding a temporal penalty matrix and strict unidirectional time constraints, still maintains a high accuracy of over 92% under the extreme jitter environment of 60%.

[0108] Furthermore, in scenarios with extremely small sample sizes lacking massive amounts of manually labeled data, the orthogonal AC prototype projection classification module of this invention exhibits extremely strong feature decoupling capabilities. In target classification scenarios, when there are only 10 real label data per class, conventional deep learning models almost fail (accuracy is only about 30%); however, this system, by explicitly decomposing global features into static capacity intrinsic components and dynamic time-delay distortion components, can quickly converge to a usable accuracy of 72% with only 10 samples, and approach 98% with 80 samples, completely solving the problem of feature extraction distortion under small sample sizes. Finally, ablation experiments confirm that: if the "adaptive temporal masking strategy based on intrinsic gradient" is removed, the accuracy drops to 88.2%; if the "temporal penalty matrix" that forces unidirectional inference is removed, the accuracy drops sharply to 82.1%; and if the "orthogonal AC prototype projector" is not used, the accuracy falls to 85.4%. The above real-world implementation data fully demonstrate that this invention possesses extremely stable industrial-grade feasibility and significant technical advantages.

[0109] This invention provides a network traffic classification method and system based on spatiotemporal bias polar coordinate coherent field and cross-channel time series extrapolation. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment of the invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.

Claims

1. A network traffic classification method based on spatio-temporal bias polar coordinate coherent field and cross-channel temporal deduction, characterized in that, Includes the following steps: Step 1: Preprocess and extract low-level features from the collected raw network traffic data: directly extract the packet size and arrival time interval of network data packets, and construct normalized packet size sequence and normalized time interval sequence that retain the original low-level features after sequence alignment and maximum-minimum normalization. Step 2, Spatiotemporal feature mapping and multi-channel tensor orthogonal fusion: Based on polar coordinate mapping, a temporal decay kernel and phase shift mechanism are introduced to convert the normalized packet size sequence into a first-channel static capacity polarization field matrix representing the static scale attribute of the network flow, and the normalized time interval sequence into a second-channel dynamic time delay bias field matrix representing the direction of dynamic congestion evolution. The first and second channels are aligned and fused according to spatial coordinates to generate a multi-channel image tensor with asymmetric temporal characteristics. Step 3, self-supervised pre-training based on original gradient-driven masking and cross-channel temporal constraints: construct a self-supervised learning task using unlabeled multi-channel image tensors; extract the local spatial difference gradient of the first channel matrix to adaptively mask the high gradient burst feature region of the second channel; use a cross-attention model with embedded temporal penalty matrix to perform unidirectional time series cross-channel extrapolation based solely on the static first channel features, and combine local gradient divergence penalty to calculate composite coherent error and update network parameters; Step 4, Perform orthogonal prototype projection and adaptive boundary fine-tuning under minimal sample size: After pre-training, the predictor is stripped and the context encoder is retained. An orthogonal prototype projection module is connected to the end of the context encoder. Using traffic data with real labels, the global features are explicitly decomposed into static and dynamic independent orthogonal components. A loss function that fuses the original data boundary features and the orthogonal decoupling regularization term is used for fine-tuning and updating. Finally, the classification prediction results of network traffic are output.

2. The method of claim 1, wherein, Step 1 includes: Step 1-1, capture original network data stream, group by five-tuple as independent stream, arrange data packet in each stream in ascending order of timestamp, extract packet size of each data packet and arrival time interval of adjacent data packets ; Steps 1-2 involve truncating or zero-padding the extracted packet size sequence and arrival time interval sequence to a uniform sequence length of N, directly retaining the original physical variables at the lowest level, and constructing the packet size sequence. With time interval sequence ;in, This represents the normalized bag size value at the Nth position in the normalized bag size sequence. This represents the normalized arrival time interval value at the Nth position in the normalized time interval sequence; Steps 1-3 employ max-min normalization to map and transform each original value in the packet size sequence S and time interval sequence T according to the following formula: , in, This represents the i-th original data value in the sequence; This represents the maximum value of all original data in the sequence; This represents the minimum value of all the original data in the sequence; This represents the i-th normalized value obtained after calculation. Through mapping transformation, the values ​​of the packet size sequence S and the time interval sequence T are mapped to... The interval is used to obtain the normalized bag size sequence. With normalized time interval sequence .

3. The method according to claim 2, characterized in that, Step 2 includes: Step 2-1, Construct the static capacity polarization field of the first channel: Obtain the normalized packet size sequence of length N output from the preprocessing stage. , convert the sequence Element-wise mapping to polar coordinates; for sequences Each normalized packet size value in Calculate and extract using the inverse cosine function Corresponding spatial angle information Subsequently, all data packet pairs in the sequence are traversed, the sum of the angles corresponding to any two time steps i and j is calculated, and combined with the temporal decay kernel based on index distance, a kernel of size i is constructed. The two-dimensional feature matrix; the polarization field pixel value in the i-th row and j-th column of the generated feature matrix. The calculation formula is: , in, This is the preset smoothing control coefficient; This represents the size of the j-th normalized bag. Corresponding spatial angle information; by As the fundamental polar coordinate eigenvalues, and The multiplier, used as the distance mapping weight, is based on the relative span of the data packet in the original sequence. Calculate and generate the complete static capacity polarization field matrix C pixel by pixel; Step 2-2, Construct the second channel dynamic time-delay bias field: Synchronously acquire a normalized time interval sequence of length N. , convert the sequence Map to polar coordinates and extract the angle components corresponding to each time interval. ,in For sequence The i-th normalized time interval; when calculating the pixel association value of the two-dimensional matrix, an directional offset linearly related to the sequence position span is introduced, resulting in a size of... The pixel value in the i-th row and j-th column of the feature matrix The calculation formula is: , in, The set bias gain constant; Represents the j-th normalized time interval The corresponding angular components; During calculation, additional components will be added. The cosine function is injected internally as a directional time-delay phase bias, so that the generated pixel value is affected when the row and column indices are interchanged. Thus, the asymmetric dynamic time-delay bias field matrix D is generated point by point; Steps 2-3, orthogonal fusion of coherent fields: The static capacity polarization field matrix C, after time decay mapping, is set as the first channel view of the network flow feature, and the dynamic time-delay bias field matrix D, after phase bias mapping, is set as the second channel view of the network flow feature; the first and second channels are strictly aligned according to row and column pixel coordinates in spatial scale, and then superimposed and merged along the channel dimension to finally generate a spatial shape as follows. The multi-channel image tensor is used as the target tensor and input into the downstream model for subsequent processing.

4. The method according to claim 3, characterized in that, Step 3 includes the following steps: Step 3-1, Cross-channel decoupling and information flow isolation: The first channel static capacity polarization field matrix C, which represents the static scale essence of the network flow, is input separately to the context encoder; The second-channel dynamic time-delay bias field matrix D, representing dynamic congestion and time evolution, is input into the target encoder, which includes an exponential moving average (EMA) mechanism. Step 3-2, Adaptive temporal masking strategy based on intrinsic gradient: Extract the local spatial difference gradient of the static capacity polarization field matrix C at the input of the context encoder. ;by The absolute value of the probability density function generates an adaptive mask matrix, which cuts off the feature blocks of the dynamic time-delay bias field matrix D corresponding to the high gradient region. Step 3-3: Input the unmasked, complete dynamic time-delay bias field matrix D separately into the target encoder to extract the true high-dimensional topological feature vector as the standard answer. .

5. The method according to claim 4, characterized in that, Step 3 also includes: Steps 3-4: Temporal cross-attention deduction based on temporal penalty: The predictor receives the incomplete features output by the context encoder and performs cross-channel deduction using a Transformer architecture that includes a multi-head cross-attention mechanism; a temporal penalty matrix is ​​embedded in the cross-attention formula. The specific formula is as follows: , in, The cross-attention derivation output feature matrix is ​​used, where softmax is the activation function, K' is the transpose of the key feature matrix K, and Q, K, and V represent the query feature matrix, key feature matrix, and value feature matrix, respectively. The scaling dimension of the feature vector; Temporal penalty matrix The element in the i-th row and j-th column of the attention graph, when the j-th time step corresponding to the element is greater than the i-th time step. ,otherwise ; The penalty scalar is a positive number, and the final output is the predicted time delay feature. ; Steps 3-5, calculation of spatiotemporal topological composite coherence loss: a local gradient divergence penalty term is introduced to reconstruct the error evaluation mechanism in order to calculate the prediction time delay characteristics. The true features extracted by the target encoder The topological loss error L between them; the specific formula is: , Where B is the batch size for training; Let be the prediction time-delay feature vector of the k-th sample; Let be the target true feature vector of the k-th sample; The predicted local spatial difference gradient for the k-th sample; The true local spatial difference gradient of the k-th sample; This represents the sum of the squared 2 norms of the differences between the predicted local spatial difference gradients and the true local spatial difference gradients for all samples in a batch. The weighting coefficients for the topology penalty term; Steps 3-6, Asymmetric update of model parameters: Based on the backpropagation algorithm, the network parameters of the context encoder and predictor are updated in real time using the topological loss error L; Steps 3-7 involve performing a gradient stopping operation on the target encoder to disable backpropagation updates; then, using the exponential moving average (EMA) algorithm, the parameters of the target encoder are slowly and smoothly updated using the parameters of the context encoder, as shown in the formula: , Among them, symbols The parameter assignment update symbol indicates that the calculation result on the right is assigned and updated to the variable on the left. For the target encoder parameters, Here are the parameters for the context encoder, and λ is the smoothing coefficient.

6. The method according to claim 5, characterized in that, Step 4 specifically includes: Step 4-1, Extract global topological features: Input the complete target traffic multi-channel image tensor with real labels into the trained context encoder to extract high-dimensional global fusion topological features. ; Step 4-2, Construct the orthogonal prototype projection classification head: Keep the main parameters of the context encoder frozen, and connect the orthogonal prototype projector to the tail of the context encoder; the orthogonal prototype projector contains a first orthogonal transformation matrix. With the second orthogonal transformation matrix The first orthogonal transformation matrix The transpose of the second orthogonal transformation matrix The product is a zero matrix; using the first orthogonal transformation matrix Global fusion features Projecting onto the static capacity eigenspace yields the static capacity eigencomponents. Simultaneously, using the second orthogonal transformation matrix Global fusion features Projecting onto the dynamic time-delay distortion subspace yields the dynamic time-delay distortion components. This makes the static capacity eigencomponents With dynamic time delay distortion components Mathematically, they are mutually orthogonal and independent; subsequently, the static capacity eigencomponents are calculated separately in the two-manifold subspace. The Mahalanobis distance C1 from the prototype center of each protocol category, and the dynamic time-delay distortion component. The Mahalanobis distance C2 between each protocol category prototype center is calculated, and C1 and C2 are adaptively fused using dynamic uncertainty weights to output the joint probability distribution of each protocol type. Step 4-3, Fine-tuning of coherent large-edge decoupling loss: Fine-tuning the model using real-label data from the target scene, employing the large-edge decoupling loss as the objective function. .

7. The method according to claim 6, characterized in that, In step 4-3, the objective function The calculation formula is: , in, To fine-tune the objective function; To fine-tune the total number of samples; Let i be the true label of the i-th sample; The feature components of the i-th sample and the true protocol category The angle between the centers of the prototypes; is the angle between the feature component and the center of the c-th non-real protocol category prototype; m is the physical coherence edge penalty term, used to constrain and obtain a more compact intra-class distribution during fine-tuning; This is the scaling factor; For orthogonal decoupling regularization terms, where Let be the transpose matrix of the static capacity eigencomponents. This is a dynamic time-delay distortion component; These are the weighting coefficients.

8. A network traffic classification system based on spatiotemporal bias polar coordinate coherence field and cross-channel time series extrapolation for implementing the method as described in any one of claims 1 to 7, characterized in that, include: The spatiotemporal topology mapping and multi-channel tensor orthogonal fusion module is used to introduce a temporal decay kernel and phase shift mechanism on the basis of polar coordinate mapping. It converts the normalized sequence into the first channel static capacity polarization field matrix representing the static scale attribute of the network flow and the second channel dynamic time delay bias field matrix representing the direction of dynamic congestion evolution. It then generates a multi-channel image tensor with asymmetric temporal characteristics by aligning it with spatial coordinates. A self-supervised pre-training module with physical gradient-driven masking and cross-channel temporal constraints is used to extract the local spatial difference gradient of the first channel matrix to adaptively mask the burst feature region of the second channel; and a cross-attention model with embedded temporal penalty matrix is ​​used to unidirectionally deduce the missing dynamic temporal evolution features, and combined with local gradient divergence penalty to calculate composite coherent error and update network parameters. The minimal sample orthogonal prototype projection and adaptive boundary fine-tuning module is used to strip the predictor and retain the context encoder after pre-training. An orthogonal prototype projector is connected at the end. The global features are decomposed into static and dynamic independent orthogonal components using data with real labels. The loss function that fuses physical coherence edges and orthogonal decoupling regularization terms is used for fine-tuning, and finally the classification prediction results of network traffic are output.

9. A flow classification device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 7.