Heterogeneous flow analysis method, device and system driven by intelligent large model
By collecting the side channel features of encrypted traffic and performing dynamic position coding and multimodal features fusion, the attention mechanism of the Transformer model is used to solve the problem of low accuracy of encrypted traffic detection in traditional methods, and more efficient attack detection is achieved.
Patent Information
- Application Number
- CN202510855235.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-25
AI Technical Summary
The existing heterogeneous traffic analysis methods are difficult to directly analyze encrypted traffic, and cannot effectively extract and fuse the heterogeneous channel characteristics of the traffic for attack detection, resulting in a decrease in detection accuracy.
The side channel features of encrypted traffic are collected, and attack detection is detected through dynamic position coding and multimodal features fusion.
It improves the accuracy and robustness of encrypted traffic detection, enhances the detection sensitivity of hidden threats in encrypted traffic, and solves the problems of lack of features and semantics in traditional methods.
Smart Images

Figure CN120358103A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of heterogeneous traffic analysis, and particularly to an intelligent large model-driven heterogeneous traffic analysis method, device, and system. Background Art
[0002] In the context of digital transformation, network traffic exhibits significant heterogeneous characteristics, that is, traffic data covers multiple protocols, diverse data forms, and multiple sources, with large structural semantic differences and complex dynamic associations. Traditional analysis relies on manually designed features, making it difficult to capture the deep semantic associations hidden in traffic data and having poor adaptability to new protocols or attack patterns. Large models represented by Transformer have powerful context semantic modeling capabilities and can automatically learn long-range dependency relationships of cross-protocol and cross-modal features through self-attention mechanisms without presetting fixed feature templates.
[0003] Existing heterogeneous traffic analysis methods often require decrypting traffic for analysis. Large models are difficult to directly parse encrypted payloads and can only indirectly infer based on traffic meta-features, unable to effectively extract and fuse heterogeneous side-channel features of traffic for attack detection. Moreover, relying only on single-side channel features results in semantic loss and a decrease in traffic detection accuracy. Summary of the Invention
[0004] To solve the above technical problems, the purpose of this application is to provide an intelligent large model-driven heterogeneous traffic analysis method, device, and system, and the specific technical solutions adopted are as follows: In the first aspect, an embodiment of this application provides an intelligent large model-driven heterogeneous traffic analysis method, which includes the following steps: Collect side-channel features of encrypted traffic, where the side-channel features include timing features and structural features; Based on the length, traffic direction of each data packet in a single session, and the arrival timestamp relative to the session start time, correct the static position encoding method of the Transformer model, perform dynamic position encoding on the side-channel features, and obtain the first feature vector of each data packet; Construct a second feature vector based on the network layer IP and transport layer port number in the five-tuple; construct a third feature vector through digital certificates and encryption suites at the application layer; construct a fourth feature vector based on user IDs and access paths in the business system logs of the computer network; Assign learnable bias vectors to each eigenvector for dynamically adjusting the association weights of different modality features; form the bias vectors of all eigenvectors into a modality bias matrix, which serves as a learnable parameter during the training process of the Transformer model, and correct the expression of the attention mechanism in the Transformer model; utilize the corrected attention mechanism to obtain the fused eigenvector of all eigenvectors; Use a lightweight classifier to perform attack detection on the input traffic of the fused eigenvector.
[0005] In one embodiment, the side-channel features include: The session duration of a single session, the timestamp interval of each data packet, the length of each data packet, and the traffic direction identifier of each data packet.
[0006] In one embodiment, the method for modifying the static position encoding of the Transformer model includes: The position encoding formula for the i-th data packet is , where is the arrival timestamp of the i-th data packet relative to the session start time, is the basic time series encoding of the i-th data packet, is the length of the i-th data packet, is a weight coefficient with a preset value between 0 and 1, is the sign encoding of the traffic direction of the i-th data packet; Among them, ; in the formula, is the session start time of the session corresponding to the i-th data packet, is the average value of the timestamp intervals of all data packets in the session corresponding to the i-th data packet, j is the hidden layer dimension index of the Transformer model, is the hidden layer dimension of the Transformer model, C is a preset value, represents the frequency scaling factor, sin() is the trigonometric sine function, and cos() is the trigonometric cosine function.
[0007] In one embodiment, the determination of the second eigenvector includes: Generate a hash value with a fixed length for the IP address of the network layer using a hash algorithm, perform one-hot encoding on the transport layer port number to generate a sparse vector, and splice the hash value and the sparse vector and linearly map them to the hidden layer dimension of the Transformer model to obtain the second eigenvector.
[0008] In one embodiment, the determination of the third eigenvector includes: By using the digital certificate at the application layer and the hashing algorithm, the certificate hash is obtained. The certificate hash and the encryption suite encoding at the application layer are concatenated and linearly mapped to the hidden layer dimension of the Transformer model to obtain the third feature vector.
[0009] In one embodiment, the determination of the fourth feature vector includes: The user ID is used to generate a hash value of a fixed length by the hashing algorithm. The access path is segmented and then numerically sequenced through word embedding. The hash value generated by the user ID and the numerical sequence are linearly mapped to the hidden layer dimension of the Transformer model to obtain the fourth feature vector.
[0010] In one embodiment, the expression of the attention mechanism in the modified Transformer model includes: ; where Q represents the query vector, specifically the first feature vector, K represents the key vector, specifically any one of the second feature vector, the third feature vector, and the fourth feature vector, is the transpose of the key vector, V is the value vector, represents the scaling factor, is the dimension of the key vector, is the modal bias matrix, represents the activation function.
[0011] In one embodiment, the use of a lightweight classifier to perform attack detection on the input traffic of the fusion feature vector includes: The MobileNetV3 network architecture is adopted, with the fusion feature vector as the input, and the output is the attack probability value, representing the confidence that the input traffic is an attack.
[0012] In a second aspect, an embodiment of the present application also provides a heterogeneous traffic analysis device driven by an intelligent large model. The heterogeneous traffic analysis device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the heterogeneous traffic analysis method described in any one of the above is implemented.
[0013] In a third aspect, an embodiment of the present application also provides an intelligent large model-driven heterogeneous traffic analysis system, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0014] The present application has at least the following beneficial effects: This application collects the side-channel features of encrypted traffic, where the side-channel features include timing features and structural features; based on the length, traffic direction, and arrival timestamp relative to the session start time of each data packet in a single session, it corrects the static position encoding method of the Transformer model, performs dynamic position encoding on the side-channel features, and obtains the first feature vector of each data packet; enhances the model's dynamic perception ability of the timing behavior of encrypted traffic. The fixed sine position encoding of traditional Transformer is difficult to adapt to the non-uniform timing characteristics of network traffic. This application introduces traffic direction weights and relative timestamp offsets, enabling the model to accurately capture the spatio-temporal correlation patterns of data packets within a session, avoiding the over-reliance of traditional methods on the continuity assumption of network traffic, and solving the problem of lack of features caused by the invisibility of encrypted traffic payloads. Through dynamic encoding, the Transformer model can more accurately capture the stable patterns of normal traffic and the mutation characteristics of abnormal traffic, improving the accuracy and robustness of subsequent traffic attack recognition; constructs a second feature vector based on the network layer IP and transport layer port number in the five-tuple; constructs a third feature vector through digital certificates and encryption suites in the application layer; constructs a fourth feature vector based on user IDs and access paths in the business system logs of computer networks; through hierarchical fusion of multi-modal features, it enhances the depth of feature representation; assigns learnable bias vectors to each feature vector to form a modal bias matrix, which serves as a learnable parameter during the training process of the Transformer model, and corrects the expression of the attention mechanism in the Transformer model; breaks through the rigid bottleneck of weight allocation in traditional multi-modal fusion, realizes the deep association between timing-sensitive side-channel features and context semantic features, solves the cross-modal semantic confusion problem in traditional fusion methods, provides high-value inputs containing multi-dimensional association information for subsequent detection models, and improves the decision-making robustness of the model in mixed attack scenarios; uses the corrected attention mechanism to obtain the fused feature vector of all feature vectors; the fused feature vector integrates the timing dependence of side-channel features, traffic structure information, and the protocol semantics and business context of multi-modal features, and is used for subsequent traffic attack detection and classification, enhancing the detection sensitivity of the model to hidden threats in encrypted traffic and improving the accuracy of traffic detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0016] Figure 1The flowchart of steps of a heterogeneous traffic analysis method driven by an intelligent large model provided by an embodiment of the present application; Figure 2 It is a flowchart for determining the fusion feature vector. Detailed implementation manners
[0017] In order to further elaborate on the technical means and effects adopted by the present application to achieve the predetermined invention purpose, the following combines the accompanying drawings and preferred embodiments to detail the specific implementation manners, structures, features and their effects of a heterogeneous traffic analysis method, device and system driven by an intelligent large model proposed according to the present application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.
[0019] The following specifically describes the specific solutions of a heterogeneous traffic analysis method, device and system provided by the present application with reference to the accompanying drawings.
[0020] Please refer to Figure 1 , which shows the flowchart of steps of a heterogeneous traffic analysis method driven by an intelligent large model provided by an embodiment of the present application. The method includes the following steps: S1, collecting side-channel features of encrypted traffic, where the side-channel features include timing features and structural features.
[0021] In this embodiment, the side-channel features of encrypted traffic are collected through eBPF kernel probes, where the side-channel features include timing features and structural features. The eBPF kernel probe is a kernel-level dynamic monitoring mechanism based on the extended Berkeley packet filter (eBPF, Extended Berkeley Packet Filter) technology. By mounting lightweight programs at the network protocol stack, system call interface or performance event points of the operating system kernel, kernel-state data is captured in real time, and the data is efficiently transferred to the user state through the zero-copy technology.
[0022] The core feature of the eBPF kernel probe is that it does not require modifying the kernel source code or loading kernel modules, can be dynamically injected and securely isolated, supports the collection of timing data with nanosecond-level precision and the extraction of traffic features independent of the protocol. It is especially suitable for the side-channel analysis scenario of encrypted traffic. By capturing metadata such as the five-tuple, timestamp, and packet length of packets, it realizes fine-grained monitoring of network behavior with low performance loss, providing underlying data support for subsequent feature modeling and security detection.
[0023] Specifically, the eBPF kernel probe obtains the structured side-channel feature data of a single session based on all the original data packets of the single session and the kernel-mode skb buffer pointer. Among them, the timing feature is to form a time interval sequence and the session duration from the time stamp intervals of each data packet; the structural feature is to form a packet length sequence from the lengths of each data packet and the traffic direction identifier. The data format of the side-channel feature data is a binary structure.
[0024] S2. Based on the length, traffic direction of each data packet in a single session, and the arrival time stamp relative to the start time of the session, correct the static position encoding method of the Transformer model, perform dynamic position encoding on the side-channel features, and obtain the first feature vector of each data packet.
[0025] The static position encoding of the traditional Transformer model only depends on fixed position indexes and cannot capture the dynamically changing time intervals in encrypted traffic, such as the differences between burst traffic and periodic communications, and packet length features, such as large packets may carry malicious payloads, and traffic directions, such as the characteristics of client-initiated attacks, resulting in insufficient semantic representation capabilities of the Transformer model for timing patterns and traffic structures. The parsing of encrypted traffic depends on side-channel features. Therefore, this embodiment constructs a dynamic encoding method that can fuse time, structure, and direction information to enhance the sensitivity of the Transformer model to abnormal patterns in encrypted traffic.
[0026] This application improves the position encoding formula of the traditional Transformer model and introduces a side-channel feature weighting term. The specifically improved position encoding formula is: ; where is the position encoding formula of the i-th data packet, denoted as the dynamic position encoding vector, represents the basic timing encoding of the i-th data packet, encodes the time interval sequence using sine and cosine functions, replaces the static position index, and the expression is , where is the start time of the session corresponding to the i-th data packet, is the mean value of the time stamp intervals of all data packets in the session corresponding to the i-th data packet, j is the hidden layer dimension index of the Transformer model, is the hidden layer dimension of the Transformer model, which is 512 in this embodiment. represents the frequency scaling factor, which increases exponentially as the dimension j increases, making the low dimension correspond to low frequency (long interval) and the high dimension correspond to high frequency (short interval), realizing the multi-resolution coding of time intervals. Essentially, it defines the frequency range. C is a preset value, and its value range is from 1000 to 1e6. In this embodiment, the value of C is 10000. sin() is the trigonometric sine function, and cos() is the trigonometric cosine function. is the length of the i-th data packet. is the weight coefficient with a preset value between 0 and 1. is the symbolic coding of the traffic direction of the i-th data packet.
[0027] It should be understood that the absolute time interval of the i-th data packet relative to the session start time directly reflects the temporal distribution of traffic. By eliminating the difference in time units between different sessions, it ensures the scale consistency of coding. The formula of represents the sine / cosine orthogonal basis, that is, the sine and cosine components of the same frequency form a two-dimensional vector. Through the phase difference it provides orthogonal temporal direction information and enhances the spatial distinguishability of position coding. Using the sine / cosine function to perform periodic coding on the time interval of the data packet arrival reflects the time distribution law of traffic. For example, an interval mutation may indicate an attack; mapping the continuous time interval to the periodic space of [-1, 1] enables the Transformer model to capture cyclic patterns; by defining the frequency components of different dimensions j, it allows the Transformer model to learn the dependencies of multiple time scales through cross-dimensional feature interactions, and the additive property of the trigonometric function enables the Transformer model to generalize time intervals that have not been trained, while the absolute position coding of static indexes lacks this ability.
[0028] represents the packet length weighting term for the length of the i-th data packet to perform power scaling and adjust the intensity of the temporal coding. represents the weight coefficient, and its value range is from 0 to 1. In this embodiment takes the value of 0.3, making the temporal change of long packets have a greater impact on the coding because malicious payloads often contain large packets. represents the symbolic coding of the traffic direction, which is used to distinguish the temporal importance of requests and responses because attacks are usually initiated actively by the client.
[0029] Its operation mode enables the encoded vector to contain both temporal dependencies and traffic structures simultaneously, providing richer input features for the Transformer model and solving the problem of lack of features caused by the invisibility of the payload in encrypted traffic. Through dynamic encoding, the Transformer model can more accurately capture the stable patterns of normal traffic and the mutation characteristics of abnormal traffic, improving the accuracy and robustness of subsequent traffic attack recognition.
[0030] Based on this, this embodiment uses 's position encoding method to perform position encoding on the side-channel feature data of a single session, including temporal features and structural features. Obtain the dynamic position encoding vector of each data packet in the session, denoted as the first feature vector, whose dimension is consistent with the hidden layer of the Transformer model. Each dimension is generated by fusing the time interval, packet length, and direction information, serving as the input encoding of the Transformer model for subsequent temporal dependency modeling and feature extraction.
[0031] S3. Construct a second feature vector based on the network layer IP and transport layer port number in the five-tuple; construct a third feature vector through the digital certificate and encryption suite of the application layer; construct a fourth feature vector based on the user ID and access path in the business system log of the computer network.
[0032] Although the position encoding of side-channel features can capture the temporal dependencies and structural features of traffic sequences, such as packet length changes and time interval patterns, it lacks specific business semantics and cross-layer association information, such as the identities of the communication parties, protocol types, and business operation intentions. It is impossible to determine whether the traffic belongs to a malicious encrypted tunnel only through the position encoding of the packet length sequence. Multimodal features such as the IP reputation of the network layer, the TLS certificate fingerprint of the application layer, and the abnormal access frequency in the business log can provide key semantic support for the Transformer model. This embodiment dynamically associates the position encoding with multimodal features through a cross-modal attention mechanism, enabling the Transformer model to not only perceive the temporal anomalies of traffic, such as high-frequency small packet transmission in a short period of time, but also combine semantic features such as suspicious IPs in the network layer and uncertified certificates in the application layer to achieve accurate semantic parsing and threat determination of encrypted traffic, making up for the semantic lack of single side-channel features and improving the traffic detection accuracy and generalization ability in complex scenarios.
[0033] Based on the above analysis, in this embodiment, the IP address of the network layer and the port number of the transport layer are first obtained based on the IP five-tuple. In this embodiment, the IP address is extracted by the packet parsing tool Scapy, and the IP address is hashed by SHA-256 to generate a hash value with a fixed length of 256 bits. Secondly, one-hot encoding is performed on the port number of the transport layer to generate a sparse vector. Each port number is mapped to a binary vector, where only the position corresponding to the port number is 1 and the rest are 0, so as to retain the classification independence of the port number and prevent the Transformer model from mistakenly treating it as a continuous value. Assuming that the maximum value of the port number is N, the dimension of the one-hot vector corresponding to each port number p is N, and only the p-th bit is 1 and the rest are 0. Finally, the hash value and the sparse vector are concatenated and mapped to the dimension to construct a five-tuple feature vector, denoted as the second feature vector.
[0034] Furthermore, a third feature vector is constructed through the digital certificate and encryption suite of the application layer. The application layer features are the certificate hash and the encryption suite. Specifically, the digital certificate is processed in the same way as the IP address to obtain the corresponding hash value as the certificate hash. Secondly, the encryption suite encoding is obtained through the embedding layer mapping. The certificate hash and the encryption suite encoding are concatenated and mapped to the dimension to construct an application layer feature vector, denoted as the third feature vector.
[0035] The user ID and access path are obtained from the business system log file of the computer network. The user ID is processed in the same way as the IP address to obtain the corresponding hash value, denoted as the user ID hash. After the access path is segmented, a numerical sequence is generated through BERT word embedding. Then, the user ID hash and the numerical sequence are concatenated and mapped to the dimension to construct a business log feature vector, denoted as the fourth feature vector.
[0036] S4. Assign learnable bias vectors to each feature vector to form a modal bias matrix, which is used as a learnable parameter in the training process of the Transformer model, and correct the expression of the attention mechanism in the Transformer model; use the corrected attention mechanism to obtain the fused feature vector of all feature vectors.
[0037] Traditional multi-modal fusion methods ignore the semantic gap and distribution differences of different modal data, resulting in modal confusion during cross-modal feature association, that is, it is difficult for the Transformer model to distinguish the feature sources and mistakenly introduce the noise of irrelevant modalities into the decision-making. For example, when directly concatenating the IP address hash and the TLS encryption suite encoding, the Transformer model may wrongly amplify the weights of irrelevant features. Therefore, a mechanism is needed to explicitly encode the modal differences, guide the attention mechanism in the Transformer model to focus on the effective associations within the same modality, and suppress the cross-modal noise interference.
[0038] Based on the above analysis, in this embodiment, by constructing a modal deviation matrix, a learnable deviation vector is assigned to each modality, that is, a learnable deviation vector is assigned to the side channel, quintuple, application layer, and business system log respectively. The modal deviation matrix is composed of the deviation vectors of the side channel, quintuple, application layer, and business system log. The deviation vector is injected into the attention score matrix during the multi-head attention calculation to dynamically adjust the association weights of different modal features. By explicitly modeling the modal identity, the attention preference of the same-modal features is enhanced. For example, when calculating the association between the IP feature in the network layer and the side channel packet length sequence, the modal deviation matrix will add a positive deviation to the key-value pair corresponding to the quintuple modality, making the attention mechanism more inclined to focus on the feature associations within the network layer, and at the same time imposing a negative adjustment on the weakly correlated cross-modal features. This deviation is dynamically optimized through backpropagation to adapt to the modal interaction patterns in different scenarios.
[0039] Specifically, in this embodiment, the dynamic position encoding vectors of all data packets in a single session are first compressed into a 512-dimensional representation vector through average pooling. Subsequently, the representation vector is used as the query vector Q, and the feature vectors of other modalities are linearly transformed through the key matrix and value matrix respectively to obtain the key vector K and value vector V. The linear transformation is a well-known technology, and the key matrix and value matrix are parameters obtained during the training of the Transformer model. The specific steps will not be elaborated here.
[0040] A learnable deviation vector is assigned to each modality, initialized to a random value, which is used to encode the exclusive semantic features of each modality. The dimension of the deviation vector is the dimension of the key vector of a single head in the multi-head attention mechanism, that is, the ratio of the hidden layer dimension of the Transformer model to the number of attention heads. The deviation vector is used as a parameter of the Transformer model and is jointly optimized with parameters such as attention weights and linear layer weights through backpropagation during the training process. The deviation vectors of all modalities are concatenated row by row to form a matrix, that is, the modal deviation matrix. When calculating the attention score, the modal deviation matrix is added to the original attention score to form a score matrix with modal preferences, denoted as the cross-modal attention mechanism.
[0041] The single-head attention expression of the cross-modal attention mechanism is as follows: , where Q represents the query vector, specifically the first feature vector, that is, the vector generated after the side-channel feature undergoes dynamic positional encoding , representing the target information to be attended to, used to calculate the similarity with the key vector K of other modalities to determine the attention weight. K represents the key vector, that is, the feature vector of the five-tuple, application layer, and business system log, calculates the similarity with the query vector, and characterizes the content to be attended to represents the transpose of the key vector, and V represents the value vector, that is, the feature vector from the same modality as the key vector, and performs weighted summation on it according to the attention weight to generate the final output represents the scaling factor, used to scale the dot product similarity to prevent gradient disappearance specifically represents the dimension of the key vector represents the modality bias matrix, which explicitly encodes the semantic differences between different modalities, namely side-channel, five-tuple, application layer, and business system log, and solves the information confusion problem of the traditional attention mechanism in the cross-modal scenario represents the activation function
[0042] After parallel calculation through the multi-head attention mechanism, the outputs of each head are concatenated and residual-connected with the original side-channel representation, that is, the representation vector, and then layer normalization is performed to generate the fused feature vector. This process realizes the deep association between the time-series sensitive side-channel features and the context semantic features, solves the cross-modal semantic confusion problem in the traditional fusion method, and provides a high-value input containing multi-dimensional association information for the subsequent detection model. The fused feature vector integrates the time-series dependence relationship of the side-channel features, the traffic structure information, and the protocol semantics and business context of the multi-modal features, and is used for subsequent traffic attack detection and classification. The flowchart for determining the fused feature vector is as shown in Figure 2 shown
[0043] S5. A lightweight classifier is used to perform attack detection on the input traffic of the fused feature vector
[0044] Further, a lightweight classifier is used to perform attack detection on the fused feature vector. The lightweight classifier adopts the depthwise separable convolution technology, decomposes the standard convolution into depth convolution and point convolution, and greatly reduces the computational amount while maintaining the feature expression ability. Specifically: In this embodiment, based on the MobileNetV3 basic network architecture, first, the fused feature vector output by the cross-modal attention mechanism is adjusted to a tensor format suitable for input, for example, adding a batch dimension to form [1, 512]. After normalization, it is input into the first fully connected layer and mapped to a dimension matching the inverted residual block. Subsequently, feature extraction is performed through multiple inverted residual blocks. Inside each inverted residual block, the number of channels is first expanded through a 1x1 point convolution, then local features are extracted channel by channel through depthwise separable convolution, combined with the Swish activation function to enhance non-linear expression. Finally, a 1x1 point convolution is used to compress the channels and connect to the Squeeze-and-Excitation module, which dynamically adjusts the channel weights according to the global context to focus on key features. After multi-layer feature extraction, the fused feature vector is input into the average pooling layer to reduce the dimension to 128 dimensions, then mapped to the classification dimension through the fully connected layer, and finally, the attack probability value between 0 and 1 is output through the Sigmoid function, representing the confidence that the input traffic is an attack. Among them, the MobileNetV3 network structure is a well-known existing technology, and the specific process will not be elaborated.
[0045] Based on the same inventive concept as the above method, an embodiment of the present application also provides a heterogeneous traffic analysis device driven by an intelligent large model. The heterogeneous traffic analysis device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the heterogeneous traffic analysis method described in any one of the above.
[0046] Based on the same inventive concept as the above method, an embodiment of the present application also provides an intelligent large model-driven heterogeneous traffic analysis system, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above-described intelligent large model-driven heterogeneous traffic analysis methods.
[0047] It should be noted that: the above sequence of embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. In addition, the above specific embodiments of this specification have been described. Also, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0048] Each embodiment in this specification is described in a progressive manner. The same or similar parts between each embodiment can be referred to each other, and the key points of each embodiment are to illustrate the differences from other embodiments.
[0049] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included within the protection scope of the present application.
Claims
1. An intelligent large model-driven heterogeneous traffic analysis method, characterized in that The method includes the following steps: Collect the side-channel features of encrypted traffic, where the side-channel features include timing features and structural features; Based on the length, traffic direction of each data packet in a single session, and the arrival timestamp relative to the session start time, correct the static position encoding method of the Transformer model, perform dynamic position encoding on the side-channel features, and obtain the first feature vector of each data packet; Construct a second feature vector based on the network layer IP and transport layer port number in the five-tuple; construct a third feature vector through the digital certificate and encryption suite of the application layer; construct a fourth feature vector based on the user ID and access path in the business system log of the computer network; Allocate learnable bias vectors to each feature vector for dynamically adjusting the association weights of different modality features; form the bias vectors of all feature vectors into a modality bias matrix, which is used as a learnable parameter during the training process of the Transformer model, and correct the expression of the attention mechanism in the Transformer model; use the corrected attention mechanism to obtain the fused feature vector of all feature vectors; Use a lightweight classifier to perform attack detection on the input traffic of the fused feature vector.
2. The heterogeneous traffic analysis method driven by an intelligent large model according to claim 1, wherein The side-channel features include: The session duration of a single session, the timestamp interval of each data packet, the length of each data packet, and the traffic direction identifier of each data packet.
3. The heterogeneous traffic analysis method driven by an intelligent large model according to claim 1, characterized in that, The correction of the static position encoding method of the Transformer model includes: Encode the position of the i-th data packet as , where is the arrival timestamp of the i-th data packet relative to the session start time, is the basic timing encoding of the i-th data packet, is the length of the i-th data packet, is a preset weight coefficient with a value between 0 and 1, is the symbol encoding of the traffic direction of the i-th data packet; Among them, ; In the formula, is the session start time of the session corresponding to the i-th data packet, is the mean of the timestamp intervals of all data packets in the session corresponding to the i-th data packet, j is the hidden layer dimension index of the Transformer model, is the hidden layer dimension of the Transformer model, C is a preset value, represents the frequency scaling factor, sin() is the trigonometric sine function, and cos() is the trigonometric cosine function.
4. The heterogeneous traffic analysis method driven by an intelligent large model according to claim 1, wherein, The determination of the second feature vector includes: Generate a fixed-length hash value for the IP address in the network layer using the hash algorithm, perform one-hot encoding on the transport layer port number to generate a sparse vector, splice the hash value and the sparse vector, and linearly map them to the hidden layer dimension of the Transformer model to obtain the second feature vector.
5. An intelligent large model-driven heterogeneous traffic analysis method according to claim 1, characterized in that The determination of the third feature vector includes: Through the digital certificate of the application layer, use the hash algorithm to obtain the certificate hash, splice the certificate hash and the encoding of the application layer encryption suite, and linearly map them to the hidden layer dimension of the Transformer model to obtain the third feature vector.
6. The heterogeneous traffic analysis method driven by an intelligent large model according to claim 1, characterized in that, The determination of the fourth feature vector includes: Generate a fixed-length hash value for the user ID using the hash algorithm, generate a numerical sequence through word embedding after segmenting the access path, and linearly map the hash value generated by the user ID and the numerical sequence to the hidden layer dimension of the Transformer model to obtain the fourth feature vector.
7. The heterogeneous traffic analysis method driven by an intelligent large model according to claim 1, characterized in that, The correction of the expression of the attention mechanism in the Transformer model includes: ; where Q represents the query vector, specifically the first eigenvector, K represents the key vector, specifically any one of the second eigenvector, the third eigenvector, and the fourth eigenvector, is the transpose of the key vector, and V is the value vector, represents the scaling factor, is the dimension of the key vector, is the modal deviation matrix, represents the activation function.
8. The heterogeneous traffic analysis method driven by an intelligent large model according to claim 1, wherein, The use of a lightweight classifier to perform attack detection on the fused feature vector includes: Adopt the MobileNetV3 network architecture, use the fused feature vector as the input, and the output is the attack probability value, indicating the confidence that the input traffic is an attack.
9. An intelligent large model-driven heterogeneous traffic analysis device, characterized in that, The heterogeneous traffic analysis device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the heterogeneous traffic analysis method described in any one of claims 1-8 is implemented.
10. An intelligent large model-driven heterogeneous traffic analysis system, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, the steps of the method described in any one of claims 1-8 are implemented.
Citation Information
Patent Citations
Encrypted traffic classification method and device based on multi-modal learning, and storage medium
CN116451138A
Method and device for constructing side channel attack model and side channel attack method and device
CN117081722A
Lightweight Internet of Vehicles intrusion detection method based on improved MobileNetV3 model
CN118101326A
Charging pile peak queuing time prediction method based on big data analysis
CN118966408A
Encrypted traffic classification method based on graph structure and dual-channel sequence feature mixing
CN119357768A
Cited By
Channel reconstruction method and device, computer equipment and storage medium
CN121308884A