Confrontation-resistant traffic generation method based on reinforcement learning

By reinforcing the learning model, obfuscated bytes are inserted into the network traffic to generate traffic data packets that are adversarial to the obfuscation. This solves the problem that the existing technology cannot effectively resist the recognition ability of the recognition model, achieves better traffic concealment transmission effect, and enhances the security of network traffic.

CN120729616APending Publication Date: 2025-09-30INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511093759.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing network traffic obfuscation technology cannot effectively resist traffic behavior identification models based on machine learning and deep learning, resulting in a high risk of traffic privacy leakage. Traditional methods cannot achieve adversarial effects under limited disturbances and cannot achieve hidden transmission of traffic.

Method used

A reinforcement learning model is used to insert obfuscated bytes into the original traffic data packets multiple times until the traffic behavior recognition model cannot accurately classify them. The intelligent agent in the reinforcement learning model selects obfuscation actions, inserts bytes and specifies positions, and generates adversarially obfuscated traffic data packets with the optimization goal of maximizing the cumulative reward of the adversarial generation process.

Benefits of technology

Effectively reduce the accuracy of traffic behavior identification models, enhance the hidden security of network traffic, resist identification attacks based on graph neural networks, and achieve better traffic transmission protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120729616A_ABST
    Figure CN120729616A_ABST
Patent Text Reader

Abstract

The invention provides an anti-confusion traffic generation method based on reinforcement learning, and the method comprises the steps: taking a traffic behavior recognition model as an environment, carrying out the multiple times of anti-confusion learning based on an original traffic data packet in the environment through employing a reinforcement learning model, and obtaining a final confused traffic data packet, the first traffic data packet is an original traffic data packet, and the other traffic data packets are traffic data packets mixed in the previous time; the agent selects the confusion action of this time from a preset confusion action space according to the current state by taking maximization of the accumulated reward as an optimization target, and the confusion action comprises bytes needing to be inserted and a specified position; and inserting bytes needing to be inserted into the specified position of the traffic data packet, generating a traffic data packet after confusion, obtaining a classification result of the traffic data packet after confusion and an award positively correlated with a classification result error through the environment, and if the classification result is different from the category label, executing the award if the classification result is not the same as the category label. And constructing a final traffic data packet based on the traffic data packet generated this time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, specifically, to the field of adversarial generation of neural network recognition models, and more specifically, to a method for generating adversarial obfuscation traffic based on reinforcement learning. Background Art

[0002] With the rapid development of information technology, the internet has become deeply integrated into all areas of society, and its high-frequency application scenarios have significantly improved the efficiency of information exchange. It is worth noting that traffic data generated during network transmission, as a key digital asset, fully records users' online behavior and operational characteristics. When this sensitive data is illegally intercepted and deeply mined, it can lead to multi-dimensional security threats such as user privacy leakage and digital identity fraud. Therefore, effective network traffic anonymization has become a key technical issue in ensuring the security of the digital ecosystem.

[0003] First, traditional network traffic protection focuses on the security of the data payload within traffic. One key approach involves using encryption and digital signatures to ensure data confidentiality and integrity. However, the development of decryption technologies has brought new challenges to traffic data protection. Furthermore, traditional protection methods cannot prevent attackers from analyzing network traffic; attackers can still analyze user behavior patterns based on encrypted network traffic.

[0004] Secondly, traffic obfuscation technology has solved the problems faced by traditional network traffic to a certain extent. At present, existing traffic obfuscation schemes include randomization, tunneling, and shaping. For example, the randomization scheme in reference [1] uses various methods to randomize information such as characteristic fields in the target traffic; the shaping scheme in reference [2] uses various methods such as modification of target traffic data packets and virtual data packet filling to reshape the target traffic into sample traffic, so that it has the characteristics of sample traffic; tunneling means transmitting the target traffic in a tunneling manner. The tunneling mechanism establishes an encrypted or encapsulated transmission channel between the client and the server, and encapsulates the original target traffic in the channel for transmission. This tunneling method makes it impossible for intermediate nodes to identify the specific content or true destination of the original data, thereby achieving the purpose of traffic obfuscation, hiding, protocol disguise, or crossing network restrictions.

[0005] Finally, with the iterative evolution of artificial intelligence technology, existing network traffic obfuscation technology is facing the risk of systemic failure. First, as in Reference [3], it uses a machine learning method to extract traffic statistical features (such as traffic packet interval time, traffic packet direction information, and traffic packet length distribution) to build a classification decision model to achieve the identification of network traffic behavior characteristics; second, as in Reference [4], it uses a deep learning framework to directly encode the features of the original data packet to form a traffic behavior classification model with strong discrimination. Third, based on graph neural networks, attackers convert network traffic data into graph data and then train a graph neural network traffic behavior classification model with few parameters and high accuracy, such as Reference [5] and Reference [6]. Empirical data shows that the recognition accuracy of the latter two types of intelligent algorithms has exceeded the 90% threshold, which further exacerbates the risk of network traffic privacy leakage.

[0006] Therefore, based on the existing traffic obfuscation perturbation method to modify the data packet, the existing traffic behavior recognition model still has a strong ability to identify the behavior of the modified traffic data packet. Under the condition of limited disturbance, it is impossible to achieve the confrontation effect, the hidden transmission of traffic cannot be achieved, and the traffic transmission cannot be well protected.

[0007] It should be noted that this background information is intended solely to introduce relevant information related to the present invention to facilitate understanding of the present invention's technical solution. It does not necessarily constitute prior art. Relevant information submitted and disclosed together with the present invention's solution should not be considered prior art unless there is evidence that the relevant information was disclosed prior to the filing date of the present invention.

[0008] References are as follows:

[0009] [1] Yao Zhongjiang, Ge Jingguo, Zhang Xiaodan, et al. A review of traffic obfuscation technology and corresponding identification and tracking technology [J]. Journal of Software, 2018, 29(10): 3205-3222.

[0010] [2]Meier R, Lenders V, Vanbever L. ditto: WAN Traffic Obfuscation atLine Rate[C] / / NDSS. 2022.

[0011] [3]Rimmer V, Preuveneers D, Juarez M, et al. Automated websitefingerprinting through deep learning[J]. arXiv preprint arXiv:1708.06376,2017.

[0012] [4]Lin

[0013] [5]Zhang, H., Yu, L., Xiao, X., Li, Q., Mercaldo, F., Luo,

[0014] [6]Zhang, H., Xiao, Summary of the Invention

[0015] Therefore, the purpose of the present invention is to overcome the above-mentioned defects of the prior art and provide a method for generating anti-obfuscation traffic based on reinforcement learning.

[0016] The purpose of the present invention is achieved through the following technical solutions:

[0017] According to a first aspect of the present invention, a method for generating adversarial obfuscation traffic based on reinforcement learning is provided, comprising: using a trained traffic behavior recognition model that an original traffic data packet needs to contend with as an environment, and using a reinforcement learning model in the environment to perform one or more adversarial learning based on the original traffic data packet to generate an obfuscated traffic data packet, each time comprising: S1, taking the current traffic data packet as the current state, using the original traffic data packet for the first time, and using the previous adversarial obfuscated traffic data packet for the other times; S2, with maximizing the cumulative reward of the adversarial generation process as the optimization goal, selecting the current obfuscation action from a preset obfuscation action space based on the header data and the payload data in the current state through the intelligent agent of the reinforcement learning model; S3, inserting the bytes to be inserted into the designated position of the current traffic data packet to generate the current adversarial obfuscated traffic data packet; S4, inputting the traffic data packet generated in S3 into the traffic behavior recognition model in the environment to obtain a classification result, and generating a reward positively correlated with the difference between the classification result and the original traffic category; wherein, when the classification result is different from the category label of the original traffic data packet, constructing a final obfuscated traffic data packet based on the current obfuscated traffic data packet.

[0018] In some embodiments of the present invention, in S2, the preset obfuscation action space includes an optional integer range at each position in an optional position range, and the position range includes the position before each byte of multiple bytes of payload data of the traffic data packet; the bytes to be inserted in the obfuscation action are generated based on integers selected from the optional integer range of the preset obfuscation action space.

[0019] In some embodiments of the present invention, in S2, the cumulative reward is the accumulation of the rewards for each time in multiple adversarial generation processes, and the method of selecting this obfuscation action includes: converting a preset number of bytes of header data in the current state into a numerical value within a preset numerical range to obtain a first sub-vector; converting a preset number of bytes of payload data in the current state into a numerical value within a preset numerical range to obtain a second sub-vector; splicing the first sub-vector and the second sub-vector to obtain a spliced ​​vector, and processing the spliced ​​vector using an intelligent agent constructed based on a neural network to obtain the bytes to be inserted and the specified position selected from the preset action space.

[0020] In some embodiments of the present invention, the intelligent agent includes an embedding layer, a feature extraction module, a first output layer and a second output layer, and the method of processing the splicing vector using the intelligent agent constructed based on the neural network includes: using the embedding layer to embed the splicing vector to obtain an embedded vector, and using the feature extraction module to extract a feature vector from the embedded vector; inputting the feature vector into the first output layer for processing to obtain the probability of each integer in the optional integer range being selected; inputting the feature vector into the second output layer for processing to obtain the probability of each position in the optional position range being selected; and determining the bytes to be inserted and the specified position based on the probability of each integer being selected and the probability of each position being selected.

[0021] In some embodiments of the present invention, the model is a traffic behavior recognition model based on a graph neural network; in S4, when the classification result of this time is the same as the category of the original traffic data packet but the total length of the bytes inserted in the traffic data packet after this adversarial obfuscation is equal to the preset byte length, the final obfuscated traffic data packet is constructed based on the traffic data packet after this obfuscation.

[0022] In some embodiments of the present invention, in S4, the method of constructing the final obfuscated traffic data packet includes: filling the total number of bytes inserted cumulatively and the position offset of the specified position when inserting bytes each time relative to the first byte position of the payload data into the position after the last byte in the payload data of this traffic data packet to obtain the final obfuscated traffic data packet.

[0023] According to the second aspect of the present invention, there is provided an anti-obfuscation traffic processing device, which includes: a node traffic interception and forwarding module, which is used to intercept traffic data packets that need to be protected, and send the final obfuscated traffic data packets to the next network node, and directly forward traffic data packets that do not need to be protected to the next network node; a traffic intelligent anti-obfuscation module, which is used to process the traffic data packets that need to be protected using the method described in the first aspect of the present invention to generate the final obfuscated traffic data packets; an information processing and analysis module, which is used to restore the received obfuscated traffic data packets, including removing the inserted bytes.

[0024] According to the third aspect of the present invention, a network communication system is provided, which includes a plurality of network nodes communicating with each other, and each network node is deployed with an anti-obfuscation traffic processing device based on the second aspect of the present invention. In this system, the communication method between any two network nodes includes: for traffic data packets that need to be protected, one network node uses its anti-obfuscation traffic processing device to generate a final obfuscated traffic data packet, and sends the obfuscated traffic data packet to another network node to achieve communication between the two network nodes; and for traffic data packets that do not need to be protected, one network node directly sends the traffic data packet to another network node to achieve communication between the two network nodes.

[0025] According to the fourth aspect of the present invention, an electronic device is provided, comprising: one or more processors; and a memory, wherein the memory is used to store executable instructions; the one or more processors are configured to implement the steps of the method of the first aspect of the present invention by executing the executable instructions.

[0026] Compared with the prior art, the advantages of the present invention are:

[0027] The method of the present invention uses a traffic behavior recognition model as an environment and employs a reinforcement learning model to repeatedly insert obfuscated bytes into the original traffic data packets until the traffic behavior recognition model in the environment is unable to accurately classify the traffic data packets to be protected, achieving a good adversarial effect and thus better protecting traffic transmission. Furthermore, in the reinforcement learning model, the reward is set to be positively correlated with the difference between the classification result and the original traffic category. Furthermore, each time an obfuscated byte is inserted, the cumulative reward of the adversarial generation process is optimized to maximize the cumulative reward. The bytes to be inserted and the designated location are then searched for. This allows for a more significant obfuscation effect on the characteristics of the original data packet with fewer obfuscated bytes, more effectively reducing the accuracy of the traffic behavior recognition model in identifying network behavior based on data packet representation. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The embodiments of the present invention are further described below with reference to the accompanying drawings, in which:

[0029] Figure 1 Schematic diagram of a flow chart of a method for generating anti-obfuscation traffic based on reinforcement learning according to an embodiment of the present invention;

[0030] Figure 2 A schematic diagram of the result of inserting bytes into the payload data of a data packet according to an embodiment of the present invention;

[0031] Figure 3 Schematic diagram of the structure principle of the Actor module of an intelligent agent according to an embodiment of the present invention;

[0032] Figure 4A schematic diagram of an adversarial learning process of a reinforcement learning model according to an embodiment of the present invention;

[0033] Figure 5 Schematic diagram of a complete execution flow of a method for generating anti-obfuscation traffic based on reinforcement learning according to an embodiment of the present invention;

[0034] Figure 6 Schematic diagram of a communication process between two network nodes in a network communication system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0035] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0036] As mentioned in the background technology section, based on the existing traffic obfuscation perturbation method to modify the data packet, the existing traffic behavior recognition model still has a strong ability to identify the behavior of the modified traffic data packet. Under the condition of limited disturbance, it is impossible to achieve the countermeasure effect, cannot realize the hidden transmission of traffic, and still cannot protect the traffic transmission well.

[0037] To address the above issues, a reinforcement learning-based adversarial obfuscation traffic generation method is proposed. This method uses the trained traffic behavior recognition model that the original traffic data packet needs to contend with as the environment. When a network traffic data packet needs to be protected, a reinforcement learning model is used in the environment to insert obfuscated bytes into the original traffic data packet multiple times until the traffic behavior recognition model in the environment is unable to accurately classify the category of the traffic data packet to be protected, thereby achieving a good adversarial effect. In addition, in the reinforcement learning model, the traffic behavior recognition model identifies the traffic data packet after each obfuscation byte is inserted, obtains a classification result, and obtains a reward that is positively correlated with the difference between the classification result and the original traffic category. Moreover, each time an obfuscated byte is inserted, the bytes to be inserted and the specified position are found with the optimization goal of maximizing the cumulative reward of the adversarial generation process. This can achieve a more effective obfuscation effect on the characteristics of the original data packet with a smaller number of obfuscated bytes, more effectively reduce the accuracy of the recognition model in identifying network behavior based on data packet representation, and thus better protect traffic transmission.

[0038] The inventors have found that the existing traffic behavior recognition method based on graph neural network will model traffic byte data into graph data, which improves the generalization ability of the model. However, based on the existing method to modify the data packet, the traffic behavior recognition model based on graph neural network still has a strong ability to recognize the behavior of the traffic data packet, and still cannot achieve a good confrontation effect, and cannot realize the hidden transmission of traffic. Therefore, according to one embodiment of the present invention, the trained traffic behavior recognition model in the environment can adopt a traffic behavior recognition model based on graph neural network, such as the network behavior recognition classifier based on graph neural network in reference [5] or reference [6]. The technical solution of this embodiment can at least achieve the following beneficial technical effects: using the traffic behavior recognition model based on graph neural network with strong recognition ability as the environment, to counter the generated obfuscated traffic data packet, to achieve a good confrontation effect, can better resist the attack based on the model recognition network behavior of the graph neural network, and enhance the network traffic concealment security.

[0039] According to one embodiment of the present invention, the traffic behavior identification model can use an existing trained model. This type of model is a trained traffic behavior identification model and can be used directly; a model that needs to be retrained can also be used. When retraining the model, it is necessary to collect network traffic, extract data packets from the network traffic, and then input the payload data of the obtained data packets into the model training, or convert the traffic data packets into a burst form, and the valid information of the data packets in the burst form is the splicing of the payload data of each data packet and input into the model for training, so as to obtain a trained traffic behavior identification model, which can classify the behavior of the traffic more accurately. Among them, a burst refers to a sequence of continuous data packets in the same transmission direction.

[0040] According to one embodiment of the present invention, see Figure 1 , which is a flow chart of a reinforcement learning-based method for generating adversarial obfuscated traffic. This method includes using a trained traffic behavior recognition model against which the original traffic packets are to be subjected as an environment, and using a reinforcement learning model within this environment to perform one or more adversarial learning exercises based on the original traffic packets to generate obfuscated traffic packets. Each adversarial learning exercise includes steps S1, S2, S3, and S4. To better understand the present invention, each step is described in detail below in conjunction with specific embodiments.

[0041] Step S1: Take the current traffic data packet as the current state, use the original traffic data packet for the first time, and use the traffic data packet after the previous anti-obfuscation for the other times.

[0042] According to an embodiment of the present invention, the original traffic data packet that needs to be protected in the network can be obtained by intercepting the original traffic data packet through a traffic data interception solution.

[0043] Step S2: With the optimization goal of maximizing the cumulative reward of the adversarial generation process, the intelligent agent of the reinforcement learning model selects the current obfuscation action from the preset obfuscation action space based on the header data and payload data in the current state, including the bytes to be inserted and the specified position.

[0044] Step S2 is described in the following three aspects:

[0045] 1) Accumulated rewards

[0046] According to one embodiment of the present invention, the cumulative reward is the sum of the rewards for each of the multiple confrontation generation processes. For example, if the first reward is -1, the second is -1, and the third is 0, the cumulative reward is (-1) + (-1) + 0 = -2.

[0047] 2) Preset confusion action space

[0048] According to one embodiment of the present invention, step S2 includes obtaining a preset obfuscation action space and then selecting the current obfuscation action from the preset obfuscation action space. The process of constructing the preset obfuscation action space is as follows:

[0049] The inventors have found that among all the traffic bytes contained in the data packet, the bytes at the front contain significant characteristics of the traffic behavior, especially the payload data part of the data packet. The payload data is a sequence of effective payload bytes that carry business data. Therefore, according to one embodiment of the present invention, in each adversarial learning, it is possible to specify the insertion of bytes in the payload data part of the traffic data packet. The technical solution of this embodiment can at least achieve the following beneficial technical effects: inserting bytes in the payload data can effectively change the characteristics of the network traffic data packet, affect the recognition model's ability to characterize the data packet, and thereby reduce the recognition model's accuracy in identifying the behavior of the network traffic data packet.

[0050] According to one embodiment of the present invention, see Figure 2 , which is a schematic diagram of the result of inserting bytes into the payload data of the data packet. In the figure, the blue parts are all byte sequences of the payload data of the original traffic data packet, and the other colors are all inserted bytes. There are many ways to insert bytes in the payload data part of the traffic data packet. Among them, if it is specified to insert the byte sequence after the last byte in the byte sequence of the data packet payload data, it is called the post-fill method; if it is specified to insert the byte sequence before the first byte in the byte sequence, it is called the pre-fill method; if it is specified to insert a single byte before any byte in the byte sequence, it is called the insert fill method; the insert fill method and the post-fill method can also be combined to insert bytes, and the pre-fill method and the post-fill method can also be combined to insert bytes.

[0051] According to one embodiment of the present invention, in step S2, a preset obfuscation action space is constructed by inserting and filling. The preset obfuscation action space includes an optional integer range at each position in an optional position range, and the optional position range includes the position before each byte in the multiple bytes of the payload data of the traffic data packet; the bytes to be inserted in the obfuscation action are generated based on integers selected from the optional integer range of the preset obfuscation action space. The technical solution of this embodiment can at least achieve the following beneficial technical effects: by inserting a single byte at any middle position or head position of the payload data, it is possible to achieve a more significant change in the characteristics of the affected data packet with a smaller number of bytes, thereby effectively reducing the recognition accuracy of the traffic behavior recognition model.

[0052] According to one embodiment of the present invention, the optional integer range can be set to [0, 255]. Combined with the insertion and padding method described in the above embodiment, the final optional position range can be set to [0, N]. Each position is also considered to be a position offset relative to the first byte position of the payload data. In [0, N], 0 represents the position before the first byte of the payload data (also indicating a position offset of 0 relative to the first byte position of the payload data), and N represents the position before the N+1th byte of the payload data (also indicating a position offset of N relative to the first byte position of the payload data). If the specified position is N, it indicates a position between bytes N and N+1. If the corresponding position offset is also N, then a byte needs to be inserted between bytes N and N+1. Therefore, N is set to be less than or equal to the byte sequence length of the payload data. Since the traffic behavior recognition model generally selects the first 150 bytes of the payload data as input and extracts representations of the input bytes for recognition, the typical value of N is 150. It should be understood that this is for illustrative purposes only and N can also be set based on the actual input data length selected by the model, such as 100, 120, etc.

[0053] 3) Select the current confusion action from the preset confusion action space

[0054] According to one embodiment of the present invention, the method of selecting the current obfuscation action includes steps a1, a2, and a3:

[0055] Step a1: Convert a preset number of bytes of header data in the current state into numerical values ​​within a preset numerical range to obtain a first sub-vector.

[0056] According to one embodiment of the present invention, the first M bytes of the current traffic packet header data corresponding to the current state are converted into integer values ​​within a preset range of [0, 255] to obtain an M-dimensional first sub-vector. If the current traffic packet header is less than M bytes, the integer 256 is padded, where the preset number M can be set to 50, 60, etc. For example, if M is 50 and the current traffic packet header data is only 40 bytes, the last 10 of the 50 elements in the first sub-vector are padded with 256.

[0057] Step a2: Convert a preset number of bytes of the payload data in the current state into numerical values ​​within a preset numerical range to obtain a second sub-vector.

[0058] According to one embodiment of the present invention, the first N bytes of the payload data of the current traffic data packet corresponding to the current state are converted into integer values ​​within a preset numerical range of [0, 255] to obtain an N-dimensional second sub-vector. N is set to be less than or equal to the byte sequence length of the payload data, and N is generally 150. If the payload data is less than N bytes, it is padded with the integer 256. If N is 150, and the payload data of the current traffic data packet is only 145 bytes, the last five of the 150 elements in the second sub-vector are padded with 256.

[0059] Step a3: Concatenate the first sub-vector and the second sub-vector to obtain a concatenated vector, and process the concatenated vector using an agent built based on a neural network to obtain the bytes to be inserted and the specified positions selected from the preset action space.

[0060] According to one embodiment of the present invention, the dimension of the splicing vector is M+N. If M=50 and N=150, a 200-dimensional splicing vector V is obtained.

[0061] According to one embodiment of the present invention, the reinforcement learning model may adopt an Actor-Critic algorithm (such as the PPO algorithm). The reinforcement learning model includes an intelligent agent constructed by a neural network, which includes an Actor module and a Critic module. The Actor module is constructed based on the transformer encoder network. The input of the Actor module is the concatenated vector of the target traffic data packet, and the output is the obfuscated action. The Critic module is implemented using an LSTM network. The input is the original network traffic data packet, and the output is a numerical value used to evaluate the quality of the obfuscated action output by the Actor module. At the same time, this specific value is used for gradient backpropagation to adjust the parameters of the Actor module with the optimization goal of maximizing the cumulative reward of the adversarial generation process, so as to make the action more obfuscated.

[0062] According to one embodiment of the present invention, see Figure 3, which is a schematic diagram of the structural principle of the agent's Actor module. The Actor module includes an embedding layer, a feature extraction module, a first output layer, and a second output layer. The first and second output layers are both composed of two linear layers (such as fully connected neural network layers), and the parameters of the first and second output layers are not shared. The method of processing the splicing vector using the neural network-based agent includes processing through the Actor module, and the processing method is as follows: steps b1-b4:

[0063] Step b1: Use the embedding layer to embed the concatenated vector to obtain an embedded vector, and use the feature extraction module to extract the feature vector from the embedded vector.

[0064] According to one embodiment of the present invention, the embedding process includes: performing token embedding and position encoding on the concatenated vector (such as vector (45, 23, 128, ...)) through the embedding layer to obtain the token encoding (form 、 ...) and sequence position encoding (e.g., 0, 1, 2, ...). The token encoding and sequence position encoding are concatenated to generate an embedding vector, which is then input into the feature extraction module to generate a feature vector. The feature extraction module can utilize the encoder in the transformer network module. The Actor module also includes a pooling layer. The encoder's output feature vector is mean-pooled through the pooling layer to produce the pooled feature vector F.

[0065] Step b2: Input the feature vector into the first output layer for processing to obtain the probability of each integer being selected in the optional integer range.

[0066] According to one embodiment of the present invention, the pooled feature vector F is input to the first output layer, and the first output layer outputs the representation , then Apply the Softmax activation function to calculate the first probability distribution with a value range between 0 and 1. When the optional integer range is set to [0, 255], the corresponding representation The first probability distribution is 255-dimensional, and includes the probability of each integer in the optional integer range [0, 255] being selected.

[0067] Step b3: Input the feature vector into the second output layer for processing to obtain the probability of each position being selected in the optional position range.

[0068] According to one embodiment of the present invention, the pooled feature vector F is input to the second output layer, and the output layer outputs the representation , then Apply the Softmax activation function to calculate the second probability distribution with a value range between 0 and 1. When the optional position range is set to [0,150], the corresponding representation The second probability distribution has 150 dimensions and includes the probability of each position being selected in the optional position range [0, 150].

[0069] Step b4: Determine the bytes to be inserted and the designated positions based on the probabilities of each integer being selected and the probabilities of each position being selected.

[0070] According to one embodiment of the present invention, the integer B corresponding to the maximum probability value and the specified position index position L corresponding to the maximum probability value are respectively taken from the first probability distribution and the second probability distribution, and the integer B is used as the byte to be inserted. That is, the selected obfuscation action is (B, L). Assuming that the integer range is [0, 255] and the position range is [0, N], each optional obfuscation action in the preset obfuscation action space is a tuple, the first element of the tuple is an integer byte selected from the range [0, 255], and the second element is a position selected from the range [0, N]. In addition, the 255-dimensional representation obtained by the first output layer is After the Softmax function is processed, the 180th element has the highest probability; the 150-dimensional representation obtained by the second output layer After being processed by the Softmax function, the 93rd element has the highest probability, and the final selected confusion action is (180,93).

[0071] Step S3: insert the bytes to be inserted into the specified position of the current traffic data packet to generate the traffic data packet after the current anti-obfuscation.

[0072] According to one embodiment of the present invention, after obtaining the value of the obfuscation action, since the integer B in the obfuscation action is a number between 0 and 255, this decimal number is converted into a 1-byte hexadecimal number (0x00-0xFF), which is consistent with the format of the traffic byte sequence. That is, the integer B is converted into a 1-byte value, and then the 1 byte corresponding to B is inserted into the payload data portion of the current traffic data packet S according to the value of L (i.e., the specified position or the specified position offset), to obtain the traffic data packet S after this anti-obfuscation. new For example, the obfuscation action (180,93) means inserting the byte 0xB4 between the 93rd and 94th elements of the traffic packet payload data.

[0073] Step S4: input the traffic data packet generated by S3 into the traffic behavior recognition model in the environment to obtain a classification result, and generate a reward that is positively correlated with the difference between the classification result and the original traffic category. When the classification result is different from the category label of the original traffic data packet, a final obfuscated traffic data packet is constructed based on the obfuscated traffic data packet.

[0074] According to one embodiment of the present invention, the traffic data packets generated by S3 are input into the environment, and a classification result is obtained using the traffic behavior recognition model in the environment, which is a probability distribution of traffic categories or traffic categories. A reward is then calculated based on the probability distribution of traffic categories or traffic categories. If the black box network traffic behavior recognition model used can only obtain traffic categories, a reward can be obtained by determining whether the category label of the original traffic data packet and the classification result corresponding to the obfuscated traffic data packet are consistent. If they are consistent, the reward is -1; if they are inconsistent, the reward is 0.

[0075] According to one embodiment of the present invention, in step S4, if the classification result is the same as the category of the original traffic data packet, but the total length of the bytes inserted in the traffic data packet after the current adversarial obfuscation is equal to the preset byte length, a final obfuscated traffic data packet is constructed based on the traffic data packet after the current adversarial obfuscation. The technical solution of this embodiment can at least achieve the following beneficial technical effect: based on actual needs, the network traffic padding length is determined to not exceed the preset byte length, thereby controlling network load.

[0076] According to one embodiment of the present invention, in step S4, the method for constructing a final obfuscated traffic data packet based on the traffic data packet after the current anti-obfuscation includes: using a post-fill method to fill the position after the last byte of the payload data of the current traffic data packet with the total number of bytes inserted and the position offset of the designated position of each byte insertion relative to the first byte position of the payload data, thereby obtaining the final obfuscated traffic data packet. For example, if a total of three single-byte insertions are performed, and the position offsets of the designated positions relative to the first byte position of the payload data are 5, 7, and 66, respectively, then the total number of bytes inserted (3) and the position offsets (5, 7, and 66) are converted into a 1-byte hexadecimal number (0x00-0xFF), respectively, to obtain 0x03, 0x05, 0x07, and 0x42, and these four bytes are filled in the position after the last byte of the payload data, thereby obtaining the final obfuscated traffic data packet.

[0077] The technical solution of the above-mentioned post-filling embodiment can at least achieve the following beneficial technical effects: on the one hand, it is convenient for the receiving end to reproduce the original traffic data packet based on the information at the end of the payload data. On the other hand, the post-filling method is used to fill the total number of bytes inserted and the corresponding position offset when each byte is inserted at the position after the last byte of the payload data. Since some identification methods will count the head and / or tail features for identification, the present invention can fill bytes at the head, middle and tail of the network traffic payload data in multiple adversarial generation. The generated obfuscated traffic data packet can not only resist attacks based on data packet characterization to identify network behavior, but also resist attacks based on traffic behavior identification based on network traffic statistical characteristics, further enhancing traffic transmission security.

[0078] According to one embodiment of the present invention, see Figure 4 , which is a schematic diagram of the adversarial learning process of the reinforcement learning model. It includes three important components: state, agent, and environment. In the reinforcement learning model, the agent learns by interacting with the environment. In each adversarial learning process, the agent selects an obfuscation action (this action includes which bytes to insert at a specified position) based on the preset obfuscation action space and the header and payload data in the current state. The obfuscated traffic data packet is obtained based on this action. The environment then feeds the obfuscated traffic data packet back to the agent as a new state based on this action. The environment then uses its traffic behavior recognition model to identify the category and category label of the obfuscated traffic data packet and obtains a reward. The reward is then fed back to the agent to represent the impact of the inserted bytes on the traffic. The agent's goal is to learn an optimal strategy through this interactive process, that is, a series of obfuscation actions that maximize the cumulative reward.

[0079] Combining the above steps, we get the four-tuple: (current state (S), action (action), reward (Reward), obfuscated traffic data packet (S new The Actor and Critic modules can be trained using an Actor-Critic algorithm. When a network traffic packet is input, the Actor module gradually calculates the byte to be inserted and the specified position offset. The byte is then inserted into the payload of the traffic packet, significantly changing the traffic characteristics and improving the recognition accuracy of the traffic behavior recognition model.

[0080] According to one embodiment of the present invention, see Figure 5 , which is a complete execution flow chart of the present invention's method for generating anti-obfuscation traffic based on reinforcement learning. The complete execution process is as follows:

[0081] Step 1: At the beginning, obtain the trained traffic behavior recognition model;

[0082] Step 2. Enter the original traffic data packet to be obfuscated;

[0083] Step 3: limit the padding byte length (i.e., preset byte length);

[0084] Step 4: Based on the agent in the reinforcement learning model, the obfuscation action is selected, including the bytes to be inserted and the specified position;

[0085] Step 5: Insert the bytes to be inserted into the specified position of the traffic data packet to obtain the obfuscated traffic data packet;

[0086] Step 6: Identify the obfuscated traffic data packets using the traffic behavior recognition model in the environment to obtain classification results.

[0087] Step 7: Determine whether the inserted bytes are valid. If the classification result is consistent with the category label of the original traffic data packet, it means that the inserted bytes are invalid, and then execute step 8. Otherwise, it means that the inserted bytes are valid, and then the final obfuscated traffic data packet is constructed.

[0088] Step 8: Determine whether the total length of the bytes inserted in the obfuscated traffic data packet is greater than the preset byte length (i.e., the set maximum byte length). If so, end the byte insertion and construct the final obfuscated traffic data packet. Otherwise, return to step 4 to continue execution.

[0089] According to one embodiment of the present invention, there is provided an anti-obfuscation traffic processing device, which includes: a node traffic interception and forwarding module, which is used to intercept traffic data packets that need to be protected, and send the final obfuscated traffic data packets to the next network node, and directly forward traffic data packets that do not need to be protected to the next network node; a traffic intelligent anti-obfuscation module, which is used to process the traffic data packets that need to be protected using the method described in the above embodiment to generate the final obfuscated traffic data packets; an information processing and analysis module, which is used to restore the received obfuscated traffic data packets, including removing the inserted bytes.

[0090] According to one embodiment of the present invention, a network communication system is provided, which includes a plurality of network nodes communicating with each other, and each network node is deployed with an anti-obfuscation traffic processing device based on the above-mentioned embodiment. In the system, the communication method between any two network nodes includes: for traffic data packets that need to be protected, one network node uses its anti-obfuscation traffic processing device to generate a final obfuscated traffic data packet, and sends the obfuscated traffic data packet to another network node to achieve communication between the two network nodes; and for traffic data packets that do not need to be protected, one network node directly sends the traffic data packet to another network node to achieve communication between the two network nodes.

[0091] Schematically, see Figure 6 , which is a schematic diagram of the communication process between two network nodes in a network communication system. Client 1, Client 2, and the server in the figure are all network nodes. During the communication process between two network nodes, communication may occur directly or across multiple network nodes. An anti-obfuscation traffic processing device can be deployed in each network node, or it can be installed and deployed only in network nodes that require anti-obfuscation traffic. Network nodes deployed with anti-obfuscation traffic processing devices serve as concealed network traffic generation and processing nodes.

[0092] When two network nodes communicate, the nodes generating and processing hidden network traffic will analyze the traffic data packets transmitted. The analysis process includes: for traffic data packets that need to be protected, the node traffic interception and forwarding module in the device intercepts the traffic data packets that need to be protected, and then the traffic intelligent anti-obfuscation module uses the method described in the above embodiment to process the traffic data packets that need to be protected, generate the final obfuscated traffic data packets, and send the final obfuscated traffic data packets to the next network node through the node traffic interception and forwarding module to achieve hidden transmission of network traffic. Traffic data packets that do not need to be protected are directly forwarded to the next network node.

[0093] It should be noted that although the above describes the various steps in a specific order, it does not mean that the steps must be performed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order as long as the required functions can be achieved.

[0094] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0095] A computer-readable storage medium may be a tangible device that holds and stores instructions used by an instruction execution device. Computer-readable storage media may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove having instructions stored thereon, and any suitable combination thereof.

[0096] While various embodiments of the present invention have been described above, the above descriptions are intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A reinforcement learning-based adversarial obfuscation traffic generation method, comprising: The trained traffic behavior recognition model that the original traffic data packet needs to confront is used as an environment, and a reinforcement learning model is used in the environment to perform one or more adversarial learning based on the original traffic data packet to counter the generated obfuscated traffic data packet, each time including: S1: Take the current traffic data packet as the current state, use the original traffic data packet for the first time, and use the traffic data packet after the previous adversarial obfuscation for the rest of the time; S2: With the goal of maximizing the cumulative reward of the adversarial generation process, the agent of the reinforcement learning model selects the current obfuscation action from the preset obfuscation action space based on the header data and payload data in the current state, including the bytes to be inserted and the specified position; S3. Insert the bytes to be inserted into the specified position of the current traffic data packet to generate the traffic data packet after the current anti-obfuscation; S4. Input the traffic data packet generated by S3 into the traffic behavior recognition model in the environment to obtain a classification result, and generate a reward that is positively correlated with the difference between the classification result and the original traffic category; wherein, when the classification result is different from the category label of the original traffic data packet, a final obfuscated traffic data packet is constructed based on the obfuscated traffic data packet.

2. The method according to claim 1, characterized in that In S2, the preset obfuscation action space includes an optional integer range at each position in an optional position range, and the position range includes a position before each byte of multiple bytes of payload data of the traffic data packet; The bytes to be inserted into the obfuscation action are generated based on integers selected from a selectable integer range of a preset obfuscation action space.

3. The method according to claim 2, characterized in that In S2, the cumulative reward is the accumulation of rewards for each of the multiple adversarial generation processes. The method of selecting the current obfuscation action includes: Converting a preset number of bytes of the header data in the current state into values ​​within a preset value range to obtain a first sub-vector; Converting a preset number of bytes of the load data in the current state into a value within a preset value range to obtain a second sub-vector; The first sub-vector and the second sub-vector are concatenated to obtain a concatenated vector, and the concatenated vector is processed by an agent built based on a neural network to obtain bytes to be inserted and designated positions selected from a preset action space.

4. The method according to claim 3, characterized in that The intelligent agent includes an embedding layer, a feature extraction module, a first output layer, and a second output layer. The method of processing the splicing vector using the intelligent agent constructed based on the neural network includes: The concatenated vector is embedded using an embedding layer to obtain an embedded vector, and a feature extraction module is used to extract a feature vector from the embedded vector. The feature vector is input into the first output layer for processing to obtain the probability of each integer being selected in the optional integer range; The feature vector is input into the second output layer for processing to obtain the probability of each position being selected in the optional position range; The bytes to be inserted and the designated positions are determined based on the probabilities of each integer being selected and the probabilities of each position being selected.

5. The method according to claim 1, wherein The model is a traffic behavior recognition model based on graph neural network; In S4, when the classification result is the same as the category of the original traffic data packet but the total length of the bytes inserted in the traffic data packet after this adversarial obfuscation is equal to the preset byte length, the final obfuscated traffic data packet is constructed based on the traffic data packet after this obfuscation.

6. The method according to claim 1 or 5, characterized in that In the above S4, the final obfuscated traffic data packet is constructed in the following manner: The total number of bytes inserted and the position offset of the specified position when each byte is inserted relative to the first byte position of the payload data are filled in the position after the last byte in the payload data of this traffic data packet to obtain the final obfuscated traffic data packet.

7. A device for processing traffic against obfuscation, characterized in that: The device includes: The node traffic interception and forwarding module is used to intercept traffic data packets that need to be protected, and send the final obfuscated traffic data packets to the next network node, and forward traffic data packets that do not need to be protected directly to the next network node; A traffic intelligent anti-obfuscation module, configured to process the traffic data packet to be protected using the method described in any one of claims 1 to 6 to generate a final obfuscated traffic data packet; The information processing and analysis module is used to restore the received obfuscated traffic data packets, including removing the inserted bytes.

8. A network communication system, comprising a plurality of network nodes communicating with each other, each of which is equipped with the anti-obfuscation traffic processing device according to claim 7, wherein the communication mode between any two network nodes includes: For a traffic data packet that needs to be protected, a network node uses its anti-obfuscation traffic processing device to generate a final obfuscated traffic data packet, and sends the obfuscated traffic data packet to another network node to achieve communication between the two network nodes; as well as For traffic data packets that do not require protection, one network node directly sends the traffic data packets to another network node to achieve communication between the two network nodes.

9. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the method according to any one of claims 1 to 6.

10. An electronic device, characterized in that: include: one or more processors; as well as a memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method of any one of claims 1 to 6 by executing the executable instructions.