Concealed channel communication flow identification method

By constructing a multi-head generator and discriminator through multi-task conditional generative adversarial networks (MT-cGANs), the problems of insufficient diversity of generated samples and generalization ability of discriminators in covert channel communication identification in existing technologies are solved, and efficient identification of complex covert communication traffic is achieved.

CN120856468AActive Publication Date: 2025-10-28GUANGZHOU UNIVERSITY

Patent Information

Application Number
CN202511334078.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-10-28
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

Existing covert channel communication identification methods have deficiencies in generated sample diversity, discriminator generalization ability and training stability, and are unable to effectively identify complex covert communication traffic.

Method used

The multi-task conditional generative adversarial network (MT-cGANs) is used to build a multi-head generator and discriminator, combined with the self-attention mechanism and dynamic weight adjustment, to generate and identify covert communication traffic, thereby improving the model's generation and discrimination capabilities.

Benefits of technology

The model's adaptability and robustness to various covert communication patterns have been improved, enhancing detection efficiency and stability, and enabling effective identification of complex covert communication traffic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856468A_ABST
    Figure CN120856468A_ABST
Patent Text Reader

Abstract

The invention provides a covert channel communication flow identification method. The method comprises the following steps: acquiring a covert communication flow sample set; extracting data features and converting the data features to obtain data feature values, vectorizing the data feature values to obtain feature vector representation, extracting time sequence features of the covert communication flow sample set, identifying interaction mode features of the covert communication flow sample set, and fusing the interaction mode features to generate context feature vectors; splicing the context feature vector and the feature vector representation to obtain an enhanced feature vector; performing label distribution according to the communication behavior type of the covert communication flow to obtain a condition label; constructing a multi-head generator and a multi-head discriminator; training is carried out according to a set loss function to obtain a trained generator and a discriminator, a dynamic weight adjustment mechanism is set, and the discriminator is used for judging whether the flow data to be recognized is the real covert communication flow or not. By applying the method, the diversity, the accuracy and the calculation efficiency of covert communication flow detection are improved, and the method has obvious advantages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, and in particular to a method for identifying covert channel communication traffic. Background Technology

[0002] With the rapid development of network communication and information technology, the Internet has become the core infrastructure for information dissemination and data transmission. However, this has also provided network attackers with avenues for covert communication, particularly covert channel communication that uses legitimate protocols to transmit illegal data. Covert channel communication is a method of using legitimate network protocols (such as DNS and HTTP) to transmit unauthorized data. Attackers disguise legitimate traffic to evade detection by traditional network security monitoring tools, creating significant security vulnerabilities.

[0003] Currently, existing research and technologies have attempted to use Generative Adversarial Networks (GANs) to generate realistic fake covert communication traffic, thereby assisting discriminators in distinguishing between real and spoofed traffic. However, existing implementations have several problems and limitations, specifically in the following aspects: While existing cGANs models generate different types of covert communication traffic by introducing conditional variables, the diversity of their generated samples remains insufficient. Due to the high complexity of network communication traffic content, existing models cannot fully capture the features within it. Consequently, the samples generated by the generators often only conform to covert communication patterns on a superficial level, lacking the fine-grained temporal dependencies found in real traffic. This results in weak generation capabilities, making it impossible to comprehensively simulate real covert communication traffic patterns.

[0004] Most existing discriminators employ simple binary classification networks to determine the authenticity of samples and simultaneously classify covert communication types. However, when dealing with multiple covert communication types, these discriminators fail to adequately distinguish the characteristics of different types of covert traffic, resulting in insufficient generalization ability when facing multi-task discrimination. Classification performance deteriorates significantly, especially when dealing with complex covert communication traffic.

[0005] In existing technologies, the generator and discriminator cannot dynamically adjust their weights according to the complexity of the current task when processing different tasks, resulting in unstable training. This is especially true when multiple covert communication types coexist, where the model converges slowly and is prone to getting stuck in local optima.

[0006] Therefore, it is necessary to provide a covert channel communication traffic identification method that is highly adaptable to various covert communication modes, has high robustness and training stability. Summary of the Invention

[0007] The purpose of this invention is to provide a method for identifying covert channel communication traffic, in order to solve the problem that covert channel communication behavior is difficult to identify effectively under conventional protocols.

[0008] In a first aspect, the covert channel communication traffic identification method provided by the present invention includes: acquiring a covert communication traffic sample set; Data features of network traffic in the covert communication traffic sample set are extracted and standardized to obtain data feature values. The data feature values ​​are then vectorized to obtain feature vector representations. Temporal features of the covert communication traffic sample set are extracted, and interaction mode features of the covert communication traffic sample set are identified. The temporal features and interaction mode features are fused to generate context feature vectors. The context feature vectors are then concatenated with the feature vector representations to obtain enhanced feature vectors. Conditional labels are obtained by assigning labels to the communication traffic in the covert communication traffic sample set based on the communication behavior type of the covert communication traffic. Construct a multi-head generator to generate covert communication traffic pseudo-samples based on the enhanced feature vector, the conditional label, and the randomly generated noise vector; A multi-head discriminator is constructed, which adopts a multi-head discriminator structure and a self-attention mechanism to determine the authenticity of covert communication traffic and identify the type of covert communication behavior based on data characteristics. The multi-head generator and multi-head discriminator are trained according to the set loss function to obtain the trained generator and discriminator. A dynamic weight adjustment mechanism is set to adjust the task weights of the generator and the discriminator. The discriminator is used to determine whether the traffic data to be identified is real covert communication traffic, and the covert communication message type is determined for the traffic data determined to be covert communication traffic.

[0009] The beneficial effects of the covert channel communication traffic identification method provided by this invention are as follows: It introduces a learning strategy of multi-task conditional generative adversarial networks (MT-cGANs), enabling the discriminator to learn simultaneously on classification and true / false detection tasks, significantly improving detection efficiency in multi-task scenarios. By analyzing the temporal dependencies and interaction patterns between network traffic data packets and integrating them into contextual features, the model's ability to capture the temporal and contextual features of covert communication is enhanced. The generator and discriminator are designed with a multi-head structure, improving the model's flexibility and robustness, enabling it to exhibit better performance when facing various covert communication patterns. A dynamic weight adjustment mechanism based on game equilibrium theory dynamically adjusts the weights of the generator and discriminator in multi-task learning, improving the model's training efficiency and stability. By designing a two-stage prediction output strategy, the computational burden of the model is optimized, improving classification efficiency.

[0010] In one possible embodiment, the data feature values ​​include traffic metadata feature values ​​and protocol content feature values; the data feature values ​​are vectorized to obtain feature vector representations, including: performing one-hot encoding on the traffic metadata feature values ​​to obtain corresponding feature vectors, and performing bag-of-words model processing or word embedding processing on the protocol content feature values ​​to obtain corresponding feature vectors.

[0011] In another possible embodiment, the covert communication traffic sample set includes T data packets, where T is a positive integer. The data characteristics of the traffic in the covert communication traffic sample set include the source IP and the destination IP. The temporal features of the covert communication traffic sample set are extracted, and the interaction pattern features of the covert communication traffic sample set are identified. The temporal features and the interaction pattern features are then fused to generate a context feature vector. This includes: calculating the time interval between each data packet and the previous data packet as a temporal feature; identifying the interaction pattern between the source IP and the destination IP, and constructing an interaction matrix as an interaction pattern feature; and fusing the temporal features and the interaction pattern features to generate a context feature vector that satisfies the following formula. ,in, Represents the context feature vector. Representing temporal characteristics, denoted by , where f represents the interaction mode feature and f represents the feature fusion function.

[0012] In other possible embodiments, a multi-head generator is constructed to generate covert communication traffic pseudo-samples based on the enhanced feature vector, the conditional label, and the randomly generated noise vector, including: A multi-head generator consists of multiple generator heads; The generator head receives a random noise vector and a conditional label, and generates a pseudo sample of covert communication traffic based on the random noise vector and the conditional label, combined with the enhanced feature vector. The output of each generator head satisfies ,in, This represents a pseudo-sample of covert communication traffic. Indicates random noise in the application. Indicates a conditional label. Indicates the type of covert communication. Indicates the generator parameters.

[0013] The multi-head generator and multi-head discriminator are trained according to a set loss function, including: the multi-head generator and multi-head discriminator are trained with the goal of minimizing the loss function; the loss function of the multi-head generator satisfies: ,in, This represents the loss function of the multi-head generator. This represents the calculation of the expected value of a random variable. Indicates random noise. Indicates a conditional label. Indicates generator parameters, Indicates the discriminator parameters. This indicates that the discriminator is in response to a given condition label. The following is the result of judging the authenticity of the generated samples; the loss function of the multi-head discriminator satisfies: ,in, , , This represents the loss function of the multi-head discriminator. The discriminant loss represents the loss of the real sample. This represents the discriminant loss of the discriminator on the generated samples. Represents a real sample. Let x represent the probability distribution of the true sample x. This represents the discriminator's judgment of whether an input sample x is true or false under a given condition label y.

[0014] A dynamic weight adjustment mechanism is set up to adjust the task weights of the generator and the discriminator, including: defining the total loss function for dynamic weight adjustment as: ,in, This represents the total loss function for dynamic weight adjustment. This indicates the number of covert communication traffic types. Indicates the generator at the 1st... The weight of each task, Indicates the generator at the 1st... Losses in each task Indicates the discriminator at the 1st The weight of each task, Indicates the discriminator at the 1st Losses on individual tasks; and Dynamic changes, specifically satisfying: , ,in, This represents the rate of change of the generator's loss in the current task. This represents the rate of change of the discriminator's loss in the current task. Indicates the adjustment parameter. Represents the natural constant.

[0015] Secondly, the present invention also provides a covert channel communication traffic identification device, comprising: a sample acquisition unit for acquiring a covert communication traffic sample set; The data processing unit is used to extract data features of network traffic in the covert communication traffic sample set and perform data standardization transformation to obtain data feature values, perform vectorization processing on the data feature values ​​to obtain feature vector representation, extract the temporal features of the covert communication traffic sample set, identify the interaction mode features of the covert communication traffic sample set, fuse the temporal features and the interaction mode features to generate a context feature vector, and concatenate the context feature vector with the feature vector representation to obtain an enhanced feature vector. The tag allocation unit is used to allocate conditional tags to the communication traffic in the covert communication traffic sample set according to the communication behavior type of the covert communication traffic. A generator construction unit is used to construct a multi-head generator that generates covert communication traffic pseudo-samples based on the enhanced feature vector, the conditional label, and the randomly generated noise vector. The discriminator construction unit is used to construct a multi-head discriminator. The multi-head discriminator adopts a multi-head discriminator structure and a self-attention mechanism to determine the authenticity of covert communication traffic and identify the type of covert communication behavior based on data characteristics. The training and optimization unit is used to train the constructed multi-head generator and multi-head discriminator according to the set loss function to obtain the trained generator and discriminator. A dynamic weight adjustment mechanism is set to adjust the task weights of the generator and the discriminator. The discriminator is used to determine whether the traffic data to be identified is real covert communication traffic, and the covert communication message type is determined for the traffic data determined to be covert communication traffic.

[0016] Thirdly, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described method for identifying covert channel communication traffic.

[0017] Fourthly, the present invention also provides an electronic device, comprising: a processor and a memory; the memory being used to store a computer program; the processor being used to execute the computer program stored in the memory, so that the electronic device performs the above-described covert channel communication traffic identification method.

[0018] For the beneficial effects of the second to fourth aspects mentioned above, please refer to the description of the first aspect mentioned above. Attached Figure Description

[0019] Figure 1 A flowchart illustrating a method for identifying covert channel communication traffic according to an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating an implementation example of a covert channel communication traffic identification method provided in this invention. Figure 3 A schematic diagram of a covert channel communication traffic identification device provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of an electronic device structure provided in an embodiment of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention. Unless otherwise defined, the technical or scientific terms used herein should have the ordinary meaning understood by those skilled in the art. The terms "comprising" and similar expressions used herein mean that the element or object preceding the word covers the element or object listed following the word and its equivalents, but do not exclude other elements or objects.

[0021] This embodiment provides a method for identifying covert channel communication traffic. See also... Figure 1 and Figure 2 The method includes: S101: Obtain a sample set of covert communication traffic.

[0022] Network traffic data includes two categories: normal network traffic and covert communication traffic. Normal network traffic includes normal data packets of conventional network protocols (such as DNS, HTTP, etc.); covert communication traffic includes various covert communication tunnel traffic (such as DNS tunnels, HTTP tunnels, etc.), which are similar to normal traffic in terms of surface characteristics.

[0023] In a specific embodiment, the traffic samples in the data covert communication traffic sample set obtained from network traffic are composed of the following two types of data features: traffic metadata: such as source IP, destination IP, packet size, timestamp, etc.; protocol content features: such as DNS query fields, HTTP request headers, and other specific protocol field content.

[0024] S102: Extract data features of network traffic from the covert communication traffic sample set and perform data standardization transformation to obtain data feature values. Perform vectorization processing on the data feature values ​​to obtain feature vector representation. Extract the temporal features of the covert communication traffic sample set. Identify the interaction mode features of the covert communication traffic sample set. Fuse the temporal features and interaction mode features to generate a context feature vector. Concatenate the context feature vector with the feature vector representation to obtain an enhanced feature vector.

[0025] In one possible embodiment, the data feature values ​​include traffic metadata feature values ​​and protocol content feature values; the data feature values ​​are vectorized to obtain feature vector representations, including: performing one-hot encoding on the traffic metadata feature values ​​to obtain the corresponding feature vectors, and performing bag-of-words model processing or word embedding processing on the protocol content feature values ​​to obtain the corresponding feature vectors.

[0026] In one possible embodiment, the covert communication traffic sample set includes T data packets, where T is a positive integer. The data characteristics of the traffic in the covert communication traffic sample set include the source IP and destination IP. Temporal features of the covert communication traffic sample set are extracted, and interaction pattern features of the covert communication traffic sample set are identified. The temporal features and interaction pattern features are fused to generate a context feature vector, including: calculating the time interval between each data packet and the previous data packet as a temporal feature; identifying the interaction pattern between the source IP and the destination IP, constructing an interaction matrix as an interaction pattern feature; and fusing the temporal features and interaction pattern features to generate a context feature vector satisfying the following formula. ,in, Represents the context feature vector. Representing temporal characteristics, denoted by , where f represents the interaction mode feature and f represents the feature fusion function.

[0027] In a specific embodiment, to eliminate the differences between different feature value ranges and prevent gradient vanishing or exploding during model training due to some feature values ​​being too large or too small, it is necessary to standardize the traffic data in the covert communication traffic sample set. Specifically, the standardization operation involves scaling each feature value to a standard normal distribution with a mean of 0 and a variance of 1. The specific operation is as follows: ,in, This represents the raw feature value obtained by extracting data features from network traffic. The mean of the feature. The standard deviation of a feature This represents the standardized feature value.

[0028] The data in the covert communication traffic sample set is transformed into numerical vectors to serve as input for subsequent steps. This includes network traffic metadata vectorization and protocol content feature vectorization. Network traffic metadata vectorization specifically involves one-hot encoding of categorical features such as source IP and destination IP, and converting timestamps into relative time intervals while maintaining continuous numerical values. Protocol content feature vectorization specifically involves processing textual features such as DNS query requests and HTTP request header fields using a bag-of-words model or word embedding, transforming them into fixed-length numerical vectors.

[0029] For example, the sample set of covert communication traffic captured within a preset time period includes T data packets, and the traffic sequence is represented as follows: Each data packet Includes the following features: Network traffic metadata: timestamps Protocol type Source IP Target IP Data packet size ...; Protocol content characteristics: DNS query fields HTTP request header fields ...

[0030] The above features are processed into data vectors to obtain numerical vectors. ,in, Indicates the first Numerical representation of each feature Indicates the number of features.

[0031] Calculate the time interval between each data packet and the previous data packet. This is then incorporated as part of the temporal characteristics. Interaction patterns between source and target IPs are identified, and an interaction matrix is ​​constructed. Record the number and patterns of data packet interactions between different IP pairs. Incorporate time-series characteristics. and interaction mode features Fusion generates contextual feature vectors Specifically, it is expressed as: Where f represents the feature fusion function. The context features are... With original features (referring to original characteristics) Features extracted from the covert communication traffic dataset (input form is a standardized and vectorized numerical vector) are concatenated to form the enhanced feature vector. .

[0032] Context reconstruction analyzes the dependencies between each data packet, extracts and merges features such as time intervals and interaction patterns into the original features, generating a context feature vector. The final enhanced feature vector... The temporal dependencies between data packets were extracted, and the communication traffic was serialized. This enabled the enhanced feature vector to not only contain the detailed features of individual data packets, but also incorporate the temporal dependencies and interaction pattern information in the traffic sequence, thus providing richer and more semantic input data for subsequent generative adversarial network models.

[0033] S103: Based on the communication behavior type of the covert communication traffic, assign labels to the communication traffic in the covert communication traffic sample set to obtain conditional labels.

[0034] In one possible implementation, a unique label is assigned to each type of covert communication behavior. For example, normal traffic is labeled as DNS covert tunnel communication is tagged as HTTP covert tunnel communication is tagged as ,wait.

[0035] S104: Construct a multi-head generator that generates covert communication traffic pseudo-samples based on enhanced feature vectors, conditional labels, and randomly generated noise vectors.

[0036] In one possible embodiment, a multi-head generator is constructed to generate covert communication traffic pseudo-samples based on enhanced feature vectors, conditional labels, and randomly generated noise vectors. The multi-head generator includes multiple generator heads; each generator head receives a random noise vector and a conditional label, and generates covert communication traffic pseudo-samples based on the random noise vector and the conditional label, combined with the enhanced feature vector; the output of each generator head satisfies… ,in, This represents a pseudo-sample of covert communication traffic. Indicates random noise in the application. Indicates a conditional label. Indicates the type of covert communication. Indicates the generator parameters.

[0037] In one specific embodiment, the generator's task is to receive a random noise vector. and conditional tags The algorithm combines enhanced feature vectors to generate fake samples with covert communication characteristics. The generator structure is a multi-head generator, with each generator head corresponding to a covert communication type.

[0038] The input to the multi-head generator includes random noise. and conditional tags ,in ( (This is the dimension of the noise vector). This is a one-hot vector, representing the type of covert communication.

[0039] A multi-head generator consists of multiple generator heads, each of which is dedicated to generating a specific type of covert communication traffic sample. ,in Indicates the type of covert communication. Indicates the generator parameters.

[0040] The output of the multi-head generator is the samples generated under given conditional labels: The generated sample contains features of covert communication and combines them with a feature vector of protocol context features. This generates more realistic fake samples.

[0041] S105: Construct a multi-head discriminator. The multi-head discriminator adopts a multi-head discrimination structure and a self-attention mechanism to determine the authenticity of covert communication traffic and identify the type of covert communication behavior based on data characteristics.

[0042] In one specific embodiment, the discriminator's task is to determine the authenticity of input samples and classify their covert communication types. The discriminator employs a multi-head discrimination structure and a self-attention mechanism to enhance its ability to handle complex multi-task processing.

[0043] The input to the multi-head discriminator includes enhanced feature vector samples after standardization and protocol context feature fusion. and conditional tags , and the pseudo-samples of covert communication traffic generated by the generator.

[0044] The multi-head discriminator's output layer is designed with a dual, multi-task output structure, including a true / false determination node and a covert communication type classification node. The true / false determination node refers to the determination of the authenticity of the output sample. ,in Represents a real sample. This indicates the generated sample. The covert communication type classification node indicates the covert communication type to which the output sample belongs. ,in This indicates the number of covert communication types.

[0045] The total output of the multi-head discriminator is: Through a multi-task output structure, the discriminator can simultaneously handle the tasks of determining authenticity and classifying communication types.

[0046] In one possible embodiment, the covert channel communication traffic identification of the present invention is based on multi-task conditional generative adversarial networks (MT-cGANs), and the model used for covert channel communication traffic identification consists of a generator and a discriminator.

[0047] S106: Train the constructed multi-head generator and multi-head discriminator according to the set loss function to obtain the trained generator and discriminator. Set a dynamic weight adjustment mechanism to adjust the task weights of the generator and the discriminator. Use the discriminator to determine whether the traffic data to be identified is real covert communication traffic, and determine the covert communication message type for the traffic data determined to be covert communication traffic.

[0048] In one possible embodiment, the constructed multi-head generator and multi-head discriminator are trained according to a set loss function, including: the multi-head generator and multi-head discriminator are trained with the goal of minimizing the loss function; the loss function of the multi-head generator satisfies: ,in, This represents the loss function of the multi-head generator. This represents the calculation of the expected value of a random variable. Indicates random noise. Indicates a conditional label. Indicates generator parameters, Indicates the discriminator parameters. This indicates that the discriminator is in response to a given condition label. The following are the results of judging the authenticity of the generated samples; the loss function of the multi-head discriminator satisfies: ,in, , , This represents the loss function of the multi-head discriminator. The discriminant loss represents the loss of the real sample. This represents the discriminant loss of the discriminator on the generated samples. This represents the true sample, that is, data sampled from the real data distribution. Let x represent the probability distribution of the true sample x. This represents the discriminator's judgment of whether an input sample x is true or false under a given condition label y.

[0049] In one possible embodiment, a dynamic weight adjustment mechanism is set to adjust the task weights of the generator and discriminator, including: defining the total loss function for dynamic weight adjustment as: ,in, This represents the total loss function for dynamic weight adjustment. This indicates the number of covert communication traffic types. Indicates the generator at the 1st... The weight of each task, Indicates the generator at the 1st... Losses in each task Indicates the discriminator at the 1st The weight of each task, Indicates the discriminator at the 1st Losses on individual tasks; and Dynamic changes, specifically satisfying: , ,in, This represents the rate of change of the generator's loss in the current task. This represents the rate of change of the discriminator's loss in the current task. Indicates the adjustment parameter. Represents the natural constant.

[0050] For example, the loss function of a multi-head generator The aim is to maximize its ability to deceive the discriminator, so that the discriminator can detect a given condition label. In this case, it is difficult to correctly distinguish between generated samples and real samples. The goal of a multi-head generator is to enable the discriminator to correctly distinguish between generated samples and real samples. The output of the multi-head generator is considered to be as close as possible to the output of the real sample, i.e., it is judged as real. The formula for the loss function of the multi-head generator satisfies: ,in, and These are the parameters for the multi-head generator and the multi-head discriminator, respectively. This indicates that the multi-head discriminator is in response to a given condition label. The following is a judgment of the authenticity of the generated samples.

[0051] The goal of a multi-head discriminator is to distinguish between real and generated samples, and simultaneously classify the covert communication type of the input samples based on the input condition y. The discriminator's loss function consists of two parts: one part is used to determine the authenticity of the input sample, and the other part is used to classify the covert communication type.

[0052] Discriminator loss function It consists of two parts: the discriminant loss for real samples and the discriminant loss for generated samples. The discriminant loss for real samples maximizes the probability that the discriminator correctly classifies the real samples. The discriminator aims to achieve the correct classification probability of the real samples given certain conditions. The following correctly identifies the true sample as genuine. The loss function specifically satisfies the formula: .

[0053] The loss function of the discriminative loss component for generated samples is used to maximize the probability that the discriminator correctly classifies the generated samples. The discriminator aims to identify and classify fake samples generated by the generator as fake.

[0054] The total loss function of the multi-head discriminator is the sum of the two parts mentioned above: The total loss function considers both the authenticity of real samples and the fakeness of generated samples, ensuring that the discriminator can continuously improve its ability to distinguish between real and fake samples during adversarial training.

[0055] Within the framework of multi-task learning, the generator and discriminator each undertake multiple distinct tasks: the generator not only needs to generate realistic covert communication traffic but also ensures the diversity of samples generated under various covert communication patterns; the discriminator, on the other hand, needs to simultaneously perform both real / fake detection and covert communication type classification tasks. To reasonably balance the loss functions of each task during training, this invention introduces a dynamic weight adjustment mechanism, adaptively adjusting task weights based on the complexity and loss state of the current task. The total loss function for dynamic weight adjustment is defined as: .in: This indicates the number of covert communication traffic types. Indicates the generator at the 1st... The weight of each task is used to represent the importance of the generator in that task. Indicates the discriminator at the 1st The weight of each task, and These represent the generator and discriminator at the th... Losses on individual tasks.

[0056] Weight and The generator and discriminator dynamically adjust based on game equilibrium theory, maintaining a dynamic balance across different tasks. The specific adjustment method satisfies: .in, and These represent the rate of change of loss for the generator and discriminator, respectively, in the current task. This indicates the adjustment parameter. This mechanism ensures that different tasks do not overemphasize any one type of task during training, thus guaranteeing the balance of multi-task learning.

[0057] An adversarial learning model for identifying covert communication traffic is optimized using generators and discriminators. Generator The goal is to generate forged samples with covert communication characteristics, making the discriminator... It becomes difficult to distinguish between real and generated samples, while the discriminator continuously learns to improve its ability to differentiate between real and generated samples. Through this adversarial training mechanism, a dynamic game is formed between the generator and the discriminator, thereby continuously optimizing the model. By designing specific loss functions for the generator and discriminator, and introducing a dynamic weight adjustment mechanism and a multi-task loss balancing strategy, the model's performance is further improved.

[0058] After setting up a dynamic weight adjustment mechanism for the generator and discriminator on the training set, the discriminator can then distinguish the traffic data to be identified. To improve classification efficiency and reduce computational burden, the discriminator adopts a two-stage prediction output strategy. This two-stage strategy consists of two phases: Phase 1: The discriminator determines whether all input traffic is genuine or fake, i.e., whether the input sample is real covert communication traffic. At this stage, the discriminator primarily relies on its real / fake determination node. The first stage outputs a binary classification result. The second stage involves further classification of the covert communication type only for samples deemed "suspicious" in the first stage. This stage is based on the classification nodes of the discriminator. This two-stage strategy performs multi-class classification to determine the specific covert communication type of the input traffic samples. This significantly reduces computational load, especially when dealing with large traffic samples, by minimizing the classification of normal traffic and substantially improving classification efficiency.

[0059] The covert channel communication traffic identification method of this invention introduces a learning strategy of Multi-Task Conditional Generative Adversarial Networks (MT-cGANs), enabling the generator to generate fake traffic samples that meet specific conditions under different covert communication modes. The discriminator can simultaneously optimize in both true / false detection and multi-task classification, thereby improving the model's generalization ability and adaptability. Conditional labels are introduced during the generation process, and the commonalities and differences of covert communication traffic are learned through shared features, ensuring the model's stability and robustness when handling multiple covert communication modes. This strategy allows the model to learn the features of multiple covert communication traffic simultaneously and guides the generator to generate specific types of data through conditional labels. Compared with existing technologies, MT-cGANs, through conditional generation and discrimination processes, enables the discriminator to learn simultaneously on classification and true / false detection tasks, significantly improving detection efficiency in multi-task scenarios.

[0060] By analyzing the temporal dependencies and interaction patterns among network traffic packets, a protocol context reconstruction mechanism is proposed. This mechanism integrates this information into contextual features, which are then used in conjunction with the original packet features in generative adversarial networks (GANs). This enhances the model's ability to capture the temporal and contextual features of covert communication, significantly improving its detection capabilities, especially its adaptability to complex covert communication traffic.

[0061] The generator and discriminator employ a multi-head structure. The generator generates independent fake samples for each type of covert communication traffic, while the discriminator provides independent true / false decision nodes and multi-classification nodes for each covert communication type. This structure ensures that the generator and discriminator can be optimized and processed separately for different types of covert communication patterns. By introducing a multi-head structure, this invention significantly improves the model's flexibility and robustness, enabling it to exhibit better performance when facing multiple covert communication patterns.

[0062] Based on game equilibrium theory, a dynamic weight adjustment mechanism dynamically adjusts the weights of the generator and discriminator in multi-task learning. This allows the loss functions for different tasks to adaptively balance according to the current training situation, improving the model's training efficiency and stability. This enables the model to learn and optimize more robustly in various covert communication traffic patterns.

[0063] By designing a two-stage prediction output strategy, the discriminator first determines whether the input traffic is real or fake, and only performs further classification of covert communication types on suspicious samples. This design optimizes the computational burden of the model and improves classification efficiency. Staged processing greatly reduces redundant classification calculations for normal traffic, improving overall detection efficiency.

[0064] In summary, this invention proposes innovative designs in several key technical aspects, which have significant advantages in effectively improving the diversity, accuracy and computational efficiency of covert communication traffic detection, and provide a more robust and practical solution compared to existing technologies.

[0065] See the instruction manual appendix Figure 3 This embodiment also provides a covert channel communication traffic identification device, which is used to implement the above method embodiment. The device includes: The sample acquisition unit 201 is used to acquire a sample set of covert communication traffic.

[0066] The data processing unit 202 is used to extract data features of network traffic in the covert communication traffic sample set and perform data standardization transformation to obtain data feature values, perform vectorization processing on the data feature values ​​to obtain feature vector representation, extract the temporal features of the covert communication traffic sample set, identify the interaction mode features of the covert communication traffic sample set, fuse the temporal features and the interaction mode features to generate a context feature vector, and concatenate the context feature vector with the feature vector representation to obtain an enhanced feature vector.

[0067] The label allocation unit 203 is used to allocate conditional labels to the communication traffic in the covert communication traffic sample set according to the communication behavior type of the covert communication traffic.

[0068] Generator building unit 204 is used to build a multi-head generator that generates covert communication traffic pseudo samples based on enhanced feature vectors, conditional labels and randomly generated noise vectors.

[0069] The discriminator construction unit 205 is used to construct a multi-head discriminator. The multi-head discriminator adopts a multi-head discrimination structure and a self-attention mechanism to determine the authenticity of covert communication traffic and identify the type of covert communication behavior based on data characteristics.

[0070] The training and optimization unit 206 is used to train the constructed multi-head generator and multi-head discriminator according to the set loss function to obtain the trained generator and discriminator. A dynamic weight adjustment mechanism is set to adjust the task weights of the generator and discriminator. The discriminator is used to determine whether the traffic data to be identified is real covert communication traffic, and the covert communication information type is determined for the traffic data determined to be covert communication traffic.

[0071] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0072] In other embodiments of this application, an electronic device is disclosed, such as... Figure 4 As shown, the electronic device 300 may include: one or more processors 301; a memory 302; a display 303; one or more application programs (not shown); and one or more computer programs 304. These devices can be connected via one or more communication buses 305. The one or more computer programs 304 are stored in the memory and configured to be executed by the one or more processors 301. The one or more computer programs 304 include instructions that can be used to perform actions such as... Figure 1 , Figure 3 And the various steps in the corresponding embodiments.

[0073] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0074] In the embodiments of this application, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0075] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.

[0076] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.

Claims

1. A method for identifying covert channel communication traffic, characterized in that, include: Obtain a sample set of covert communication traffic; Data features of network traffic in the covert communication traffic sample set are extracted and standardized to obtain data feature values. The data feature values ​​are then vectorized to obtain feature vector representations. Temporal features of the covert communication traffic sample set are extracted, and interaction mode features of the covert communication traffic sample set are identified. The temporal features and interaction mode features are fused to generate context feature vectors. The context feature vectors are then concatenated with the feature vector representations to obtain enhanced feature vectors. Conditional labels are obtained by assigning labels to the communication traffic in the covert communication traffic sample set based on the communication behavior type of the covert communication traffic. Construct a multi-head generator to generate covert communication traffic pseudo-samples based on the enhanced feature vector, the conditional label, and the randomly generated noise vector; A multi-head discriminator is constructed, which adopts a multi-head discriminator structure and a self-attention mechanism to determine the authenticity of covert communication traffic and identify the type of covert communication behavior based on data characteristics. The multi-head generator and multi-head discriminator are trained according to the set loss function to obtain the trained generator and discriminator. A dynamic weight adjustment mechanism is set to adjust the task weights of the generator and the discriminator. The discriminator is used to determine whether the traffic data to be identified is real covert communication traffic, and the covert communication message type is determined for the traffic data determined to be covert communication traffic.

2. The method according to claim 1, characterized in that, The data feature values ​​include traffic metadata feature values ​​and protocol content feature values; The data feature values ​​are vectorized to obtain feature vector representations, including: The traffic metadata feature values ​​are one-hot encoded to obtain the corresponding feature vectors, and the protocol content feature values ​​are processed by bag-of-words model or word embedding to obtain the corresponding feature vectors.

3. The method according to claim 1, characterized in that, The covert communication traffic sample set includes T data packets, where T is a positive integer. The data characteristics of the traffic in the covert communication traffic sample set are the source IP and destination IP. Extracting the temporal features of the covert communication traffic sample set, identifying the interaction pattern features of the covert communication traffic sample set, and fusing the temporal features and the interaction pattern features to generate a context feature vector, including: Calculate the time interval between each data packet and the previous data packet as a timing feature; Identify the interaction patterns between source IPs and target IPs, and construct an interaction matrix as interaction pattern features; The temporal features and the interaction pattern features are fused to generate a context feature vector that satisfies the following formula. ,in, Represents the context feature vector. Representing temporal characteristics, denoted by , where f represents the interaction mode feature and f represents the feature fusion function.

4. The method according to claim 1, characterized in that, Constructing a multi-head generator to generate covert communication traffic pseudo-samples based on the enhanced feature vector, the conditional label, and a randomly generated noise vector includes: A multi-head generator consists of multiple generator heads; The generator head receives a random noise vector and a conditional label, and generates a pseudo sample of covert communication traffic based on the random noise vector and the conditional label, combined with the enhanced feature vector. The output of each generator head satisfies ,in, This represents a pseudo-sample of covert communication traffic. Indicates random noise in the application. Indicates a conditional label. Indicates the type of covert communication. Indicates the generator parameters.

5. The method according to claim 1, characterized in that, The multi-head generator and multi-head discriminator are trained according to the set loss function, including: The multi-head generator and multi-head discriminator are trained with the goal of minimizing the loss function; The loss function of the multi-head generator satisfies: ,in, This represents the loss function of the multi-head generator. This represents the calculation of the expected value of a random variable. Indicates random noise. Indicates a conditional label. Indicates generator parameters, Indicates the discriminator parameters, This indicates that the discriminator is in response to a given condition label. The following is a judgment of the authenticity of the generated samples; The loss function of the multi-head discriminator satisfies: ,in, , , This represents the loss function of the multi-head discriminator. This represents the discriminant loss for real samples. This represents the discriminant loss of the discriminator on the generated samples. Represents a real sample. Let x represent the probability distribution of the true sample x. This represents the discriminator's judgment of whether an input sample x is true or false under a given condition label y.

6. The method according to claim 1, characterized in that, A dynamic weight adjustment mechanism is set up to adjust the task weights of the generator and the discriminator, including: Define the total loss function for dynamic weight adjustment as follows: ,in, This represents the total loss function for dynamic weight adjustment. This indicates the number of covert communication traffic types. Indicates the generator at the 1st... The weight of each task, Indicates the generator at the 1st... Losses in each task The discriminator is in the first... The weight of each task, The discriminator is in the first... Losses on individual tasks; and Dynamic changes, specifically satisfying: , ,in, This represents the rate of change of the generator's loss in the current task. This represents the rate of change of the discriminator's loss in the current task. Indicates the adjustment parameter. Represents the natural constant.

7. A covert channel communication traffic identification device, characterized in that, The device includes: The sample acquisition unit is used to acquire a sample set of covert communication traffic. The data processing unit is used to extract data features of network traffic in the covert communication traffic sample set and perform data standardization transformation to obtain data feature values, perform vectorization processing on the data feature values ​​to obtain feature vector representation, extract the temporal features of the covert communication traffic sample set, identify the interaction mode features of the covert communication traffic sample set, fuse the temporal features and the interaction mode features to generate a context feature vector, and concatenate the context feature vector with the feature vector representation to obtain an enhanced feature vector. The tag allocation unit is used to allocate conditional tags to the communication traffic in the covert communication traffic sample set according to the communication behavior type of the covert communication traffic. A generator construction unit is used to construct a multi-head generator that generates covert communication traffic pseudo-samples based on the enhanced feature vector, the conditional label, and the randomly generated noise vector. The discriminator construction unit is used to construct a multi-head discriminator. The multi-head discriminator adopts a multi-head discriminator structure and a self-attention mechanism to determine the authenticity of covert communication traffic and identify the type of covert communication behavior based on data characteristics. The training and optimization unit is used to train the constructed multi-head generator and multi-head discriminator according to the set loss function to obtain the trained generator and discriminator. A dynamic weight adjustment mechanism is set to adjust the task weights of the generator and the discriminator. The discriminator is used to determine whether the traffic data to be identified is real covert communication traffic, and the covert communication message type is determined for the traffic data determined to be covert communication traffic.

8. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the covert channel communication traffic identification method according to any one of claims 1 to 6.

9. An electronic device, characterized in that, include: Processor and memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory to cause the electronic device to perform the covert channel communication traffic identification method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-class unbalanced network traffic data enhancement method based on improved generative adversarial network

    CN117892125A

Cited By

  • Behavior mixed interference resistant SSH behavior fine-grained identification method and device

    CN122160195A