Anti-confusion traffic generation method and device, storage medium and program product
Through the adversarial obfuscated traffic generation method based on reinforcement learning, an adversarial interference sequence is generated and network traffic data packets are embedded, which solves the problem of network traffic behavior recognition that the existing technology cannot effectively resist the deep learning model identification, and effectively protects network traffic privacy.
Patent Information
- Application Number
- CN202510195445.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-30
AI Technical Summary
Existing traffic obfuscation technologies cannot effectively resist network traffic behavior recognition based on deep learning models, resulting in an increased risk of network traffic privacy leakage.
Adversarial obfuscation traffic generation method based on reinforcement learning is adopted to intercept the original network traffic packets, generate adversarial interference sequences, and embed them into the load area of the data packets to generate obfuscation network traffic packets to affect the representation ability of the deep learning model.
It significantly changes the statistical and semantic characteristics of network traffic packets, reduces the accuracy of network behavior recognition based on packet representation, and enhances the privacy protection of network traffic.
Smart Images

Figure CN120074901A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular, to an adversarial obfuscated traffic generation method, device, storage medium, and program product based on reinforcement learning. Background Art
[0002] With the booming development of the Internet, its usage rate is getting higher and higher, bringing great convenience to users. However, network traffic data hides rich user behavior information. Once this information is intercepted and utilized by attackers, it will bring unpredictable losses to users. Therefore, the concealment protection of network traffic has become a key requirement.
[0003] First of all, traditional network traffic protection focuses on the security of data payloads in traffic. An important method is to ensure the confidentiality, integrity, etc. of data through encryption, digital signatures, etc. However, the development of decryption and other technologies has brought new challenges to traffic data protection. Moreover, traditional protection methods cannot prevent attackers from analyzing network traffic, that is, attackers can still analyze users' behavior patterns based on encrypted network traffic.
[0004] Secondly, traffic obfuscation technology has solved the above problems to a certain extent. At present, existing traffic obfuscation schemes include randomization, tunneling, and shaping. Randomization randomizes information such as feature fields in the target traffic through various methods; shaping reshapes the target traffic into sample traffic by using various methods such as modifying target traffic data packets and filling virtual data packets, so that it has the characteristics of sample traffic; tunneling transmits the target traffic through a tunnel.
[0005] However, due to the development of artificial intelligence, the above-mentioned traffic obfuscation technologies have been cracked. On the one hand, attackers can use machine learning methods to model network traffic behavior recognition as a classification problem, and realize the recognition of network traffic behavior by modeling the statistical features of traffic (such as traffic packet interval time, traffic packet direction information, traffic packet length distribution). On the other hand, with the development of deep learning models, attackers can use a packet-based representation model to realize traffic behavior recognition, and the recognition accuracy of this scheme is extremely high, reaching more than 90%, which further exacerbates the risk of network traffic privacy leakage.
[0006] When the inventor was researching traffic obfuscation, it was found that existing traffic obfuscation technologies are ineffective against the recognition method of traffic behavior based on traffic packet representation, and existing schemes cannot reduce its recognition accuracy. Summary of the Invention
[0007] In view of the deficiencies of the prior art, the present invention proposes an adversarial obfuscated traffic generation method, device, storage medium, and program product based on reinforcement learning. This method can effectively change the characteristics of the original network traffic data packets, affect the representation ability of the deep learning model for the data packets, and further affect the behavior recognition of the original network traffic data packets, reducing the recognition accuracy.
[0008] To achieve the above object, on the one hand, the present invention provides an adversarial obfuscated traffic generation method based on reinforcement learning, including the following steps:
[0009] Intercept the original network traffic data packets to be protected;
[0010] Generate an adversarial interference sequence based on the reinforcement learning model;
[0011] Embed the adversarial interference sequence into the payload area of the original network traffic data packet to generate an obfuscated network traffic data packet;
[0012] Verify the adversarial effectiveness of the obfuscated network traffic data packet against the network traffic behavior classification model;
[0013] If the verification fails, iteratively adjust the adversarial interference sequence until the preset adversarial condition is met or the maximum filling length is reached.
[0014] In an embodiment of the present invention, the adversarial interference sequence is embedded into the head and / or tail of the payload area.
[0015] In an embodiment of the present invention, generating the adversarial interference sequence based on the reinforcement learning model includes:
[0016] Use the payload area of the original network traffic data packet as the input state of the reinforcement learning model;
[0017] Output an action value through the policy network of the reinforcement learning model, and the action value corresponds to the interference byte;
[0018] Generate the adversarial interference sequence according to the action value and embed the adversarial interference sequence into the payload area.
[0019] In an embodiment of the present invention, the reinforcement learning model adopts the proximal policy optimization algorithm, including a policy network and an evaluation network;
[0020] The policy network is used to generate the action value according to the input state;
[0021] The evaluation network is used to evaluate the long-term reward of the action value and optimize the policy of the policy network.
[0022] In an embodiment of the present invention, the byte sequence of the original network traffic packet is input into the policy network, and the policy network adds a preset tag to the head of the byte sequence when inputting;
[0023] The policy network extracts the representation of the preset tag and generates a high-dimensional vector through multiple linear transformations;
[0024] Perform probability normalization processing on the high-dimensional vector to generate a probability distribution of action values;
[0025] Determine the target action value according to the maximum probability index of the probability distribution, and convert the target action value into interference bytes of a fixed length.
[0026] In an embodiment of the present invention, the verification of the adversarial effectiveness of the obfuscated network traffic packet against the network traffic behavior classification model includes:
[0027] Input the obfuscated network traffic packet into the pre-trained network traffic behavior classification model, and the network traffic behavior classification model outputs the corresponding category information or the probability distribution of the category;
[0028] Construct the reward function of the network traffic behavior classification model from the corresponding category information or the probability distribution of the category output by the network traffic behavior classification model;
[0029] Set the adversarial condition according to the reward function;
[0030] If the adversarial condition is satisfied, it is determined that the adversary is effective and the verification passes;
[0031] If not, it is determined that the adversary is ineffective and the verification fails.
[0032] In an embodiment of the present invention, the reward function of the network traffic behavior classification model is defined as:
[0033] If the category information output by the network traffic behavior classification model is inconsistent with the initial category of the original network traffic packet, the reward value is the first preset value, otherwise it is the second preset value; or,
[0034] Then calculate the divergence distance between the probability distributions of the categories of the obfuscated network traffic packet and the original network traffic packet, and generate a composite reward value in combination with the Euclidean distance between the representation vectors of the two.
[0035] In an embodiment of the present invention, the network traffic behavior classification model uses a black-box network traffic behavior classification model and outputs the corresponding category information; or,
[0036] The network traffic behavior classification model adopts a gray-box network traffic behavior classification model or a white-box network traffic behavior classification model, and outputs the probability distribution of the corresponding category.
[0037] In an embodiment of the present invention, the training of the network traffic behavior classification model includes:
[0038] Collect a network traffic data set, and extract the data packet payload area or continuous data packet sequences as training samples;
[0039] Using the training samples, construct the network traffic behavior classification model based on a deep learning model, and the deep learning model includes a time series feature extraction network or a context encoding architecture.
[0040] On the other hand, the present invention also provides an adversarial obfuscated traffic generation device based on reinforcement learning, including:
[0041] An interception module for intercepting the original network traffic data packets to be protected;
[0042] A sequence generation module for generating an adversarial interference sequence based on a reinforcement learning model;
[0043] A filling module for embedding the adversarial interference sequence into the payload area of the original network traffic data packet to generate an obfuscated network traffic data packet;
[0044] A verification module for verifying the adversarial effectiveness of the obfuscated network traffic data packet against the network traffic behavior classification model; and
[0045] If the verification fails, iteratively adjust the adversarial interference sequence until a preset adversarial condition is met or the maximum filling length is reached.
[0046] On yet another aspect, the present invention provides a computer-readable storage medium storing a computer program, and the computer program is executed by a processor to perform the steps of the above method.
[0047] On yet another aspect, the present invention provides a computer program product including a computer program, and when the computer program is executed by a processor, it implements the steps of the above method.
[0048] From the above solutions, the advantages of the present invention are as follows:
[0049] The method for generating adversarial obfuscated traffic based on reinforcement learning provided by the present invention intercepts the original network traffic packets to be protected; generates an adversarial interference sequence based on a reinforcement learning model; embeds the adversarial interference sequence into the payload area of the original network traffic packets to generate obfuscated network traffic packets; verifies the adversarial effectiveness of the obfuscated network traffic packets against a network traffic behavior classification model; if the verification fails, iteratively adjusts the adversarial interference sequence until a preset adversarial condition is met or the maximum filling length is reached. This method modifies the original network traffic packets based on the pre-fill and post-fill methods, significantly changing the statistical and semantic features of the original packets, and at the same time affecting the model's representation ability for network traffic packets, and can effectively resist attacks that identify network behaviors based on packet representations and network behavior identification attacks based on network statistical features. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 FIG. is a schematic flowchart of a method for generating adversarial obfuscated traffic provided by an embodiment of the present invention;
[0051] Figure 2 FIG. is a schematic diagram of an application scenario of a method for generating adversarial obfuscated traffic based on reinforcement learning;
[0052] Figure 3 FIG. is a flowchart of training a reinforcement learning model;
[0053] Figure 4 FIG. is a schematic diagram of the structure of a policy network;
[0054] Figure 5 FIG. is a schematic diagram of the structure of an adversarial obfuscated traffic generation device provided by an embodiment of the present invention.
[0055] Among them, the reference numerals are:
[0056] 10: Client;
[0057] 20: Server;
[0058] 400: Adversarial obfuscated traffic generation device;
[0059] 410: Interception module;
[0060] 420: Sequence generation module;
[0061] 430: Filling module;
[0062] 440: Verification module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] In order to make the above features and effects of the present invention more clearly and understandably described, specific embodiments are hereinafter given and detailed descriptions are made in conjunction with the accompanying drawings of the specification as follows.
[0064] When conducting research on traffic obfuscation, the inventor found that existing traffic obfuscation technologies are ineffective against the identification methods of traffic behaviors based on traffic packet representations, and existing solutions cannot reduce their identification accuracy. Through in-depth research, the inventor found that the reason for this situation is that among all the traffic bytes contained in the packet, the preceding bytes contain significant features of traffic behaviors. Therefore, modifying the packet based on existing methods cannot achieve the effect of countering, and the model's ability to identify the behaviors of traffic packets is still very strong, and the stealthy transmission of traffic cannot be realized.
[0065] In response to this, to address the deficiencies in the defense performance of the currently proposed traffic obfuscation scheme when identifying network behaviors based on packet representations, this method is based on the reinforcement learning method for network traffic. It generates traffic packet adversarial samples by filling interference byte sequences in the payloads of original network traffic packets (the data payloads of IP packets or TCP packets), confuses the semantic and statistical features of the traffic, and successfully resists attacks on identifying network behaviors based on packet representations, enhancing the security of network traffic stealth.
[0066] Specifically, referring to Figure 1 as shown in Figure 1 Fig. shows a schematic flowchart of an adversarial obfuscated traffic generation method based on reinforcement learning;
[0067] An adversarial obfuscated traffic generation method based on reinforcement learning includes the following steps:
[0068] Step S1, intercept the original network traffic packet to be protected.
[0069] In the actual usage process, typical usage scenarios are as Figure 2 shown. During the communication between two client nodes 10 in the network, they may communicate directly or may need to communicate across multiple nodes through the server 20. It is necessary to install and deploy node traffic interception and forwarding software in the nodes using adversarial obfuscated traffic. The role of the node traffic interception and forwarding software is, on the one hand, to intercept the original network traffic packet. If there is an obfuscated traffic part, the obfuscated byte sequence is removed and then handed over to the upper-layer information analysis and processing application for use. On the other hand, for the intercepted traffic, traffic obfuscation operations are performed, and then the obfuscated traffic is sent to the network.
[0070] When an original network traffic packet needs to be protected, through the traffic data interception scheme, the original network traffic packet to be protected is intercepted. And according to actual requirements, the network traffic filling length is determined to control the network load.
[0071] Step S2, generate an adversarial interference sequence based on the reinforcement learning model.
[0072] How to obtain an effective reinforcement learning model is crucial. The training process of the reinforcement learning model is as Figure 3 shown, which includes three important components: state, agent, and environment. In the reinforcement learning framework, the agent learns by interacting with the environment. At each step, the agent selects an action according to the policy π and the current environmental state s, and then the environment will feedback a new state and a reward Reward to the agent based on this action. The goal of the agent is to learn an optimal policy through this interaction process, that is, a series of actions that maximize the long-term cumulative reward.
[0073] In the reinforcement learning framework of the present invention, the state is the target traffic to be obfuscated, that is, the byte sequence of the target traffic; the agent is analogous to the human brain, the input is the byte sequence of the target traffic, and the output is an action, where the action represents the adversarial interference sequence to be added. Here, we define the action as an integer between 0 and 65535, that is, each time an interference byte of a given length is output and filled into the payload area of the original network traffic packet to obtain a generated obfuscated network traffic packet (state); the environment is the reaction to the action made by the agent. The present invention defines the environment as a network traffic behavior classification model to be countered. The obfuscated network traffic packet is input into the network traffic behavior classification model, and the category information of the obfuscated network traffic packet and the probability distribution of the category are output.
[0074] In one embodiment, the reinforcement learning model adopts the proximal policy optimization algorithm, including a policy network and an evaluation network. The policy network is the policy π of the agent, which is used to generate the action value according to the input state; the evaluation network is used to evaluate the long-term reward of the action value and optimize the policy of the policy network. In one embodiment, the evaluation network is implemented by an LSTM model, the input is network traffic, and the output is a specific value.
[0075] In the process of generating the adversarial interference sequence by using the reinforcement learning model, the payload area of the original network traffic packet is used as the input state of the reinforcement learning model; the action value is output through the policy network of the reinforcement learning model, and the action value corresponds to the interference byte; the adversarial interference sequence is generated according to the action value, and the adversarial interference sequence is embedded into the payload area.
[0076] In one embodiment, the policy network is the policy π of the agent, and the structure of the policy network is as Figure 4As shown, the byte sequence of the original network traffic packet is input into the policy network. When inputting, the policy network adds a preset tag, such as [CLS], to the head of the byte sequence; extracts the representation of the preset tag through the policy network, and generates a high-dimensional vector after multiple linear transformations; performs probability normalization processing on the high-dimensional vector based on the Softmax() function to generate the probability distribution of the action value; determines the target action value according to the maximum probability index of the probability distribution, and converts the target action value into interference bytes with a fixed length (for example, 2-byte length) to generate an adversarial interference sequence to be added to the payload area of the original network traffic packet S, obtaining a new obfuscated network traffic packet S new Then, the obfuscated network traffic packet is input into the environment, and the category information or the category to which the original network traffic packet belongs is obtained through the network traffic behavior classification model. This method can find the optimal byte sequence for filling. The optimal byte sequence can affect the characteristics of the packet with fewer bytes, more significantly, and more effectively reduce the accuracy of identifying network behavior based on the packet representation.
[0077] Step S3: Embed the adversarial interference sequence into the payload area of the original network traffic packet to generate an obfuscated network traffic packet.
[0078] In one embodiment, embedding the adversarial interference sequence into the payload area of the original network traffic packet to generate an obfuscated network traffic packet can effectively change the characteristics of the original network traffic packet, affect the representation ability of the deep learning model for the packet, and further affect the behavior recognition of the original network traffic packet, reducing the recognition accuracy.
[0079] In one embodiment, the adversarial interference sequence is separately embedded into the head of the payload area by pre-padding, separately embedded into the tail of the payload area by post-padding, or simultaneously embedded into the head and tail of the payload area by a combination of pre-padding and post-padding to generate an obfuscated network traffic packet. Mask the characteristics of the original packet and protect the privacy and security of anonymous traffic users, so that the generated obfuscated traffic can not only resist attacks on identifying network behavior based on packet representation, but also resist network behavior recognition attacks based on network statistical characteristics.
[0080] Step S4: Verify the adversarial effectiveness of the obfuscated network traffic packet against the network traffic behavior classification model.
[0081] There are usually two types of network traffic behavior classification models. One is an existing trained model, which can be directly used; the other is a model that needs to be retrained. For this type of network traffic behavior classification model that needs to be pre-trained, a network traffic data set is collected, the payload area of the data packet is extracted, or the traffic data packet is converted into a burst form. The burst refers to a sequence of consecutive data packets in the same direction, and the effective information of the data packet in the burst form is the concatenation of the payloads of each data packet. The payload area of the data packet or the sequence of consecutive data packets is extracted as a training sample, and the network traffic behavior classification model is constructed based on the deep learning model using the training sample.
[0082] Verify the adversarial effectiveness of the obfuscated network traffic data packet against the network traffic behavior classification model. Specifically, input the obfuscated network traffic data packet into the pre-trained network traffic behavior classification model, and the network traffic behavior classification model outputs the corresponding category information or the probability distribution of the category. Construct the reward function of the network traffic behavior classification model according to the category information or the probability distribution of the category output by the network traffic behavior classification model; set the adversarial condition based on the reward function. If the adversarial condition is satisfied, it is determined that the adversarial is effective and the verification passes; if not, it is determined that the adversarial is ineffective and the verification fails.
[0083] In an embodiment, the category information or probability distribution of the original network traffic data packet and the obfuscated network traffic data packet is obtained through the traffic behavior recognition model. The reward function of the network traffic behavior classification model is defined as follows: if the category information output by the network traffic behavior classification model is inconsistent with the initial category of the original network traffic data packet, the reward value is the first preset value (for example, 1), otherwise it is the second preset value (for example, 0). Or, calculate the divergence distance between the probability distributions of the categories of the obfuscated network traffic data packet and the original network traffic data packet, and generate a composite reward value in combination with the Euclidean distance between the feature vectors of the two.
[0084] In one embodiment, when the network traffic behavior classification model adopts a black-box network traffic behavior classification model, only the corresponding category information can be output, and the corresponding reward function is as follows: if the category information output by the network traffic behavior classification model is inconsistent with the initial category of the original network traffic data packet, the reward value is a first preset value (e.g., 1), otherwise it is a second preset value (e.g., 0). When the network traffic behavior classification model adopts a gray-box network traffic behavior classification model or a white-box network traffic behavior classification model, the probability distribution of the corresponding category is output, and then the divergence distance between the probability distribution of the confused network traffic data packet and the probability distribution of the original network traffic data packet is calculated, and a composite reward value is generated in combination with the Euclidean distance between the two representation vectors. In addition, if the network traffic behavior recognition model can only obtain the traffic category result, a white-box model similar to the black-box network traffic behavior recognition model can be constructed by means of model approximation, so as to obtain the traffic category distribution, and then, consistent with the above scheme, the reward value can be calculated based on the KL divergence.
[0085] In this way, an adversarial condition can be constructed using the reward function. The adversarial condition refers to the standard for verifying whether the confused network traffic data packet successfully deceives the classification model after it is generated. For example, the adversarial condition can be set as follows: when the reward value is the first preset value (i.e., the classification results are inconsistent), the verification passes; otherwise, it fails. Or, the adversarial condition is set according to the composite reward value, and it is set that the difference between the probability distribution of the confused network traffic data packet and the probability distribution of the original network traffic data packet needs to meet a preset threshold, and at the same time, the representation vectors are far enough apart in the vector space. When (D 散度 (P 原始 ,P 混淆 )≥θ 1 )∧(D 欧式 (v 原始 ,v 混淆 )≥θ 2 ), the adversarial is effective, where: D 散度 is the distribution difference measure such as KL divergence; D 欧式 is the Euclidean distance; θ 1 , θ 2 are preset thresholds, P 原始 , P 混淆 are probability distributions, v 原始 , v 混淆 are vectors. The composite reward value is: R = α * D 散度 + β * D 欧式 , where α and β are weight coefficients.
[0086] Combining the above steps, (the original network traffic data packet S, action value Action, reward value Reward, confused network traffic data packet S new)The quadruple can be used to train the policy network and the evaluation network through the standard PPO algorithm. When an Internet traffic data packet is input, an interference sequence can be calculated based on the policy network and added to the traffic load area, significantly changing the traffic characteristics and affecting the recognition accuracy of the Internet traffic behavior recognition model.
[0087] In one embodiment, in order to accelerate the training speed of the model, for the black-box and gray-box Internet traffic behavior classification models, it is expected to obtain an Internet traffic behavior classification model similar to the policy model as shown in Figure 4 . Its recognition of Internet traffic is also calculated based on the representation E of [CLS]. [CLS] Specifically, the representation of [CLS] is processed through multiple (for example, 2) linear layers to calculate two vectors E [1] = F(E [CLS] ), E [2] = F(E [1] ). Then, based on the Softmax() function, a probability distribution E [3] = F(E [2] ) with a numerical range between 0 and 1 is calculated. At this time, the definition of Reward is divided into two components. One is the KL divergence F [KL] obtained based on the class distributions of the original Internet traffic data packet and the obfuscated Internet traffic data packet, and the other is the Euclidean distance between the vectors E [1] obtained from the original Internet traffic data packet and the obfuscated Internet traffic data packet. Since the Euclidean distance can ensure that the representation distances of the two traffic flows are as far as possible, and the KL divergence can ensure that the distributions of the two traffic flow representations are as different as possible, that is, belonging to different classes, so as to achieve the adversarial purpose.
[0088] Step S5: If the verification fails, iteratively adjust the adversarial interference sequence until the preset adversarial condition is met or the maximum filling length is reached.
[0089] Input the obfuscated Internet traffic data packet into the Internet traffic behavior classification model to verify whether the adversary is effective. If the verification is effective, end. If it is not successful, further verify whether the maximum filling length is reached. If the maximum length is reached, end. If not, go to step S3 and iteratively adjust the adversarial interference sequence again until the preset adversarial condition is met.
[0090] In summary, the method for generating adversarial obfuscated traffic based on reinforcement learning provided in this embodiment intercepts the original network traffic packets to be protected; generates an adversarial interference sequence based on a reinforcement learning model; embeds the adversarial interference sequence into the payload area of the original network traffic packets to generate obfuscated network traffic packets; verifies the adversarial effectiveness of the obfuscated network traffic packets against a network traffic behavior classification model; if the verification fails, iteratively adjusts the adversarial interference sequence until a preset adversarial condition is met or the maximum filling length is reached. This method modifies the original network traffic packets based on the pre- and post-filling methods, significantly changing the statistical and semantic features of the original network traffic packets, and at the same time also affecting the model's representation ability for the original network traffic packets, resulting in misclassification of traffic behavior recognition attacks; at the same time, the method based on reinforcement learning generates byte sequences that are as short as possible and can significantly affect the target packets, enhancing the interference performance of the filling scheme, and only filling a small number of bytes into the traffic packets to avoid a large amount of bandwidth overhead caused by overly long filled traffic, enhancing the deployability. The method of the present invention can also be combined with existing methods for generating traffic packet adversarial samples such as byte sequence filling and adding virtual packets to obtain more robust data packets, which can effectively resist attacks on network behavior recognition based on packet representation and network behavior recognition attacks based on network statistical features at the same time.
[0091] Referring to Figure 5 , Figure 5 FIG. shows an adversarial obfuscated traffic generation device 400 based on reinforcement learning, which can implement each process realized by the adversarial obfuscated traffic generation method as shown in Figure 1 .
[0092] An adversarial obfuscated traffic generation device 400 based on reinforcement learning includes at least
[0093] An interception module 410 for intercepting the original network traffic packets to be protected.
[0094] A sequence generation module 420 for generating an adversarial interference sequence based on a reinforcement learning model.
[0095] A filling module 430 for embedding the adversarial interference sequence into the payload area of the original network traffic packets to generate obfuscated network traffic packets.
[0096] A verification module 440 for verifying the adversarial effectiveness of the obfuscated network traffic packets against a network traffic behavior classification model; and
[0097] If the verification fails, iteratively adjust the adversarial interference sequence until a preset adversarial condition is met or the maximum filling length is reached.
[0098] In addition, it should be understood that in the device according to the embodiments of the present application, only the above-mentioned division of each functional module is used for illustration. In practical applications, the above-mentioned functions can be allocated to different functional modules according to needs, that is, the device can be divided into functional modules different from the above-mentioned illustrated modules to complete all or part of the functions described above.
[0099] According to an embodiment of the present invention, the present invention also provides a machine-readable medium. The machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. When the computer program is executed by a processor, it implements the steps of the above-mentioned method for generating anti-confusion traffic.
[0100] In addition, the machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the above. More specific examples of the machine-readable storage medium will include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0101] The present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a readable storage medium. When the computer program is executed by a processor, the computer can execute the method for generating anti-confusion traffic provided by the above-mentioned various methods.
[0102] In addition, it should be noted that the method for generating anti-confusion traffic disclosed in the present invention is not limited to being executed in the order of the above-mentioned steps S1-S5. It can be reordered, steps can be added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. The present invention does not limit this here.
[0103] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for generating adversarial obfuscation traffic based on reinforcement learning, characterized in that: The following steps are involved: Intercept the original network traffic data packets to be protected; Generate adversarial perturbation sequences based on reinforcement learning models; Embedding the adversarial interference sequence into the payload area of the original network traffic data packet to generate an obfuscated network traffic data packet; Verify the effectiveness of the obfuscated network traffic data packets against the network traffic behavior classification model; If the verification fails, the adversarial interference sequence is iteratively adjusted until a preset adversarial condition is met or a maximum padding length is reached.
2. The method according to claim 1, characterized in that The countermeasure interference sequence is embedded into the head and / or tail of the load region.
3. The method according to claim 2, characterized in that The generating of the adversarial interference sequence based on the reinforcement learning model comprises: Using the load area of the original network traffic data packet as an input state of the reinforcement learning model; Outputting an action value through a policy network of the reinforcement learning model, wherein the action value corresponds to an interference byte; The adversarial interference sequence is generated according to the action value, and the adversarial interference sequence is embedded into the load area.
4. The method according to claim 3, characterized in that The reinforcement learning model adopts a proximal strategy optimization algorithm, including a strategy network and an evaluation network; The policy network is used to generate the action value according to the input state; The evaluation network is used to evaluate the long-term reward of the action value and optimize the strategy of the policy network.
5. The method according to claim 4, characterized in that Inputting the byte sequence of the original network traffic data packet into the policy network, wherein the policy network adds a preset label to the head of the byte sequence when inputting; Extracting the representation of the preset label through the strategy network, and generating a high-dimensional vector through multiple linear transformations; Performing probability normalization processing on the high-dimensional vector to generate a probability distribution of the action value; A target action value is determined according to a maximum probability index of the probability distribution, and the target action value is converted into interference bytes of a fixed length.
6. The method according to claim 1, characterized in that The verifying the effectiveness of the obfuscated network traffic data packet against the network traffic behavior classification model comprises: Inputting the obfuscated network traffic data packet into the pre-trained network traffic behavior classification model, the network traffic behavior classification model outputting corresponding category information or probability distribution of the category; Constructing a reward function of the network traffic behavior classification model based on the corresponding category information or the probability distribution of the category output by the network traffic behavior classification model; Set the adversarial conditions based on the reward function; If the confrontation conditions are met, the confrontation is determined to be effective and the verification is passed; If not, the confrontation is deemed invalid and the verification fails.
7. The method according to claim 6, characterized in that The reward function of the network traffic behavior classification model is defined as: If the category information output by the network traffic behavior classification model is inconsistent with the initial category of the original network traffic data packet, the reward value is a first preset value, otherwise it is a second preset value; or, The divergence distance between the probability distributions of the categories to which the obfuscated network traffic data packet and the original network traffic data packet belong is calculated, and a composite reward value is generated by combining the Euclidean distance between the representation vectors of the two.
8. The method according to claim 6, characterized in that The network traffic behavior classification model adopts a black box network traffic behavior classification model to output corresponding category information; or The network traffic behavior classification model adopts a gray box network traffic behavior classification model or a white box network traffic behavior classification model to output the probability distribution of the corresponding category.
9. The method according to claim 1, characterized in that: The training of the network traffic behavior classification model includes: Collect network traffic data sets, extract data packet load areas or continuous data packet sequences as training samples; use the training samples to build the network traffic behavior classification model based on a deep learning model.
10. A device for generating anti-obfuscation traffic based on reinforcement learning, characterized in that: Include: An interception module, used to intercept the original network traffic data packets to be protected; A sequence generation module, used to generate adversarial interference sequences based on the reinforcement learning model; A filling module, used for embedding the adversarial interference sequence into the payload area of the original network traffic data packet to generate an obfuscated network traffic data packet; A verification module, used to verify the effectiveness of the obfuscated network traffic data packet against the network traffic behavior classification model; as well as If the verification fails, the adversarial interference sequence is iteratively adjusted until a preset adversarial condition is met or a maximum padding length is reached.
11. A computer-readable storage medium storing a computer program, characterized in that: The computer program is used by a processor to execute the steps of the method according to any one of claims 1 to 9.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 9 are implemented.