A clustering method for unknown protocol text based on reinforcement learning
By applying reinforcement learning algorithms in the field of network security to process network traffic texts of unknown protocols, the problem that the existing technology is difficult to distinguish the differences in text structures of unknown protocols is solved, and more efficient clustering effect and accuracy are achieved.
Patent Information
- Application Number
- CN202210848560.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-07-19
AI Technical Summary
Existing clustering algorithms are difficult to effectively deal with network traffic texts of unknown protocols, especially when facing a wide variety of texts with different structures, it is difficult to distinguish structural differences, and there are problems of time loss and clustering performance impacts.
Using reinforcement learning-based method, traffic text is processed through CNN, combined with the full connection layer to output Q values, decision-making is made using ε-greedy strategy, and training is carried out through dual Q networks and Bellman equations to realize clustering of traffic texts of unknown protocols.
This method can continuously use previous experience to provide reference for later stages in the process of self-learning, effectively distinguish text field structure, improve clustering effect and accuracy, and is more suitable for processing unknown protocol text than the algorithm based on European distance.
Smart Images

Figure CN115309896B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a clustering method for network traffic text using an unknown protocol, in particular to an unknown protocol text clustering method based on reinforcement learning. Background Art
[0002] Network protocols define the rules for communication between hosts on the Internet. As an indispensable means of information transmission in the network, the structure and definition of the protocol are directly related to communication security. A large number of network security and privacy security issues are caused by the widespread use of protocols. However, with the development of communication technology, most communications currently use private unknown protocols. There are no public protocol specifications, and most network security strategies and methods are helpless in the face of unknown protocols. At the same time, the results of parsing unknown protocols are widely used in a wide range of applications, including malicious code analysis, intrusion detection, network security policy formulation, software security analysis, application session replay, etc. Therefore, parsing network traffic using unknown protocols is of great significance to the research and development of the network security field.
[0003] Unknown protocol reverse parsing is the process of inferring the protocol structure in the absence of formal specifications. The current mainstream methods are divided into static analysis based on network tracing and dynamic analysis based on execution tracing. Execution tracing is a piece of code or tool executed during a single run of an application between two or more communicating hosts. By capturing the internal process information of a binary executable file, the structure of the protocol field is identified by using the call status of various memory data reflected by the instruction execution trajectory during the protocol processing. A notable feature of this method is that platform-specific instructions make it difficult to transplant the tracing code or tool, which has a strong platform dependency, and the method requires access to the executable file of the application, which requires a high level of professional knowledge in protocol reverse engineering, and has high complexity and difficulty in implementation. In contrast, the static analysis method based on network tracing is simple to operate, has strong timeliness, has no dependency on applications and platform tools, and is suitable for situations where there are many types of network traffic samples and protocols used. It is very suitable for processing traffic information intercepted during the communication between two or more hosts. The present invention uses a static analysis method based on network tracing.
[0004] At present, the commonly used means in existing research is to use clustering algorithms such as KMeans, Agnes, and neural networks to classify network traffic texts containing known public protocols such as TCP, IP, and ICMP, and good results have been achieved. However, in the face of a wide variety of unknown protocol texts with different structures, no further solutions have been given. In addition, algorithms such as KMeans and Agnes, which rely on the Euclidean distance of the sample space for clustering, are difficult to distinguish structural differences when the unknown protocol network traffic performs similarly in each digit. Even in the later stage of the operation, it is almost impossible to use the previous calculation results, which results in time loss. Problems such as improper selection of the initial clustering center point and improper selection of the division method will have a significant impact on the clustering performance.
[0005] Based on the above problems, the present invention proposes an unknown protocol text clustering algorithm based on reinforcement learning. Reinforcement learning has the ability of autonomous learning and learns knowledge by interacting with the environment. It is mainly used for strategic problems and shines in the fields of games and unmanned driving. Combining reinforcement learning with text representation and related text structure feature extraction to improve the effect of text clustering algorithm is a problem with explorability and room for improvement. Reinforcement learning can continuously use previous experience to provide reference for the later stage during the self-learning process. At the same time, its method of classification learning according to the given target clustering results is more suitable for the logic of clustering of various unknown protocol texts than the algorithm based on Euclidean distance division, and can better achieve the expected effect of distinguishing from the text field structure. Summary of the invention
[0006] The technical problem to be solved by the present invention is to provide a reinforcement learning-based unknown protocol traffic text clustering method (for the convenience of description, the present invention method is referred to as RLPA) to classify the source or type of network traffic using unknown protocols and achieve basic classification. The technical solution is as follows:
[0007] A method for clustering unknown protocol texts based on reinforcement learning, characterized in that it comprises the following steps:
[0008] Step 1: Start the application that needs to capture traffic on the host, start the packet capture tool or code to capture traffic, and pre-process the obtained traffic text;
[0009] Step 2: Set the action code to the number 0-4 corresponding to the category, regard each traffic text as an environmental state, process the input traffic text into a feature vector through CNN, and then input it into the fully connected layer, and output it as 5 Q values;
[0010] For each action A adopted in each environment state s, the position corresponding to the number of action A is set to 1, and the other 4 positions are set to 0 to form a one-dimensional action vector. The one-dimensional action vector is multiplied by the 5 Q values and the sum is calculated to obtain Q(s, A) of action A adopted in each environment state s. Q(s, A) is the cumulative reward of taking action A in a certain state s following a certain strategy. Two identical models are established with traffic text and one-dimensional action vector as input and all Q(s, A) as output, namely model 1 for decision-making and model 2 for training.
[0011] Step 3: The decision-making process uses the ε-greedy strategy, with an ε probability of selecting an unknown action and a 1-ε probability of selecting the action with the highest reward in experience, ensuring that each state-action pair has a probability of being visited;
[0012] Step 4: Set a function as the system environment. In each cycle, randomly select a text from the input sample to be classified and throw it to the model1 responsible for making decisions. The text is the current state. Combined with the ε-greedy strategy described in step 3, the probability of ε randomly selects an action and returns a Q value of 0. The probability of 1-ε makes model1 predict the text category and calculate the Q value. The predicted result is compared with the actual result and reward and punishment feedback is given. If it is correct, it will increase points, and if it is wrong, it will decrease points. The system environment randomly throws the next text as the next state. The returned Q value and the corresponding state s, action A, the obtained score, and the next state are recorded in the experience memory.
[0013] Step 5: Randomly sample a batch of data from the experience memory, input it into the Bellman equation to calculate the new Q value, and update the new Q value to model2; at regular intervals, copy the parameters of model1 to model2;
[0014] Step 6: Repeat the above training steps until the specified number of training times is completed, or the total score reaches the specified threshold, then the training is completed; the trained model must be able to accurately distinguish the source or type of the newly input traffic text.
[0015] Furthermore, the preprocessing of the traffic text in step 1 is specifically as follows:
[0016] Step 1.1: Decode the captured source data into a hexadecimal string, then convert the hexadecimal characters into corresponding decimal numbers, fill in the null places, and each hexadecimal character is a processing unit;
[0017] Step 1.2: Unify the length of data with different lengths obtained from different data sources
[0018] By using fit_transform to fit part of the data, find out the mean and variance indicators; transform the data, including standardization, normalization, and PCA dimensionality reduction to retain valid information, obtain the average value x of the data lengths from different data sources after dimensionality reduction, and re-perform PCA dimensionality reduction with parameter x on all data so that all data lengths are uniformly reduced to x, and then all data lengths are uniformly cut.
[0019] Furthermore, in step 2, CNN uses two convolutional layers and two pooling layers, the step sizes are all set to 1, the activation function uses the relu function, and the loss function uses the cross entropy loss function.
[0020] Furthermore, the calculation formula of ε in step 3 is as follows:
[0021]
[0022] Among them, εmin is the minimum value of ε set by yourself, εmax is the maximum value of ε set by yourself, step is the current number of training times, and total is a constant set by yourself.
[0023] Furthermore, in step 5, the parameters of model1 are copied to model2, the action corresponding to the maximum Q value is first found in model1, and then the action is used to calculate the target Q value in the network of model2. The calculation formula is as follows:
[0024] y j =R j +γQ`(Φ(S` j ),arg maxQ(Φ(S` j ),a,ω),ω`) (2)
[0025] Among them, y j is the current target Q value; Q represents the current Q network of model1, Q` represents the target Q network of model2; R j is the reward corresponding to the current state-action; γ is the attenuation degree, which indicates the degree of dependence on the future, and its value range is between 0 and 1; S` j is the current state S j The new state after executing action A, Φ(S` j ) is state S` j The characteristic vector of ; a is the learning rate; ω is the network parameter of the current Q network; ω` is the network parameter of the target Q network.
[0026] Furthermore, the Bellman equation in step 5 is as follows:
[0027] Q(S0,A i)←(1-ɑ)Q(S0,A i )+ɑR(S0,A i ) (3)
[0028] Among them, Q(S0,A i ) means to select action A in state S0 i The updated value of the cumulative return, Q(S0,A i ) is the old value of the cumulative return; R(S0,A i ) means to select action A in state S0 i The reward is ; ɑ is the learning rate.
[0029] Compared with the existing research results, the characteristics of the present invention are:
[0030] 1) The present invention applies reinforcement learning to the field of unknown protocol parsing in network security, a field in which there are currently not many solutions and related research based on reinforcement learning.
[0031] 2) The self-learning characteristics of reinforcement learning enable it to maintain a relatively stable performance when facing unknown protocol traffic texts of different structures and a wide variety. At the same time, compared with the partition-based clustering algorithm, it can more effectively learn the field structure rules and characteristics of the text, and the classification effect is more accurate than the partition-based method. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 This is the core flow chart of the model of RLPA method of the present invention.
[0033] Figure 2 is a preprocessing step for raw hexadecimal network traffic text.
[0034] Figure 3 This is the basic construction process of the DQN model.
[0035] Figure 4 This is the training flow chart of the DQN model using a dual Q network.
[0036] Figure 5 Figure 2 shows the effect of source classification using 5 video playback application traffic datasets.
[0037] Figure 6 This is a diagram showing the effect of using more than 1,500 Android app traffic data sets to classify normal / advertising / malicious traffic types. DETAILED DESCRIPTION
[0038] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. Five application source categories are also taken as an example for the sake of description.
[0039] The overall process of the method of the present invention is as follows Figure 1 As shown in the figure, the method is built based on the reinforcement learning DQN algorithm, and the network traffic text is preprocessed to convert it into one-dimensional text data that can be used by the model. The text data is input into the built algorithm framework, and the environment continuously throws text data to the model for classification and gives reward feedback. The model continuously interacts with the environment and learns from experience until the accumulated rewards reach the set threshold.
[0040] Step 1: Start the application that needs to capture traffic on the host, start the packet capture tool or code to capture the traffic, and obtain the original data stream after transcoding. The text length and structure of the traffic of different applications and the traffic of different functions of the same application vary greatly, and need to be processed uniformly. First, decode the source data into a hexadecimal string. Since the input of model training only supports numbers, convert the hexadecimal characters into corresponding decimal numbers, fill in the null places, and each hexadecimal character is a processing unit. The data lengths obtained from different data sources are different, and the lengths need to be unified. Use fit_transform to fit part of the data, find out the mean, variance and other indicators, and then transform the data, including standardization, normalization, etc., and perform PCA dimensionality reduction to retain most of the valid information. Get the average value x of the data lengths of different data sources after dimensionality reduction, and re-perform PCA dimensionality reduction with parameter x on all data, so that all data lengths are uniformly reduced to x, and then all data lengths are uniformly cut. The whole process is like Figure 2 Shown
[0041] Step 2: The action is encoded as a number 0-4 corresponding to the category. Each traffic text is regarded as an environmental state. The neural network is used to encode the features first and then connect to the fully connected layer.
[0042] The input text is processed into a feature vector through a built CNN. The specific structure of CNN is 2 layers of convolutional layers and 2 layers of pooling layers alternately. The classic framework Lenet-5 convolutional neural network model of image classification algorithm is referenced. Four common activation functions, tanh, sigmoid, relu, and elu, are considered in the selection of activation functions. Sigmoid is similar to tanh and has the problem of soft saturation leading to gradient disappearance; relu is similar to elu and does not have the limitations of tanh and sigmoid. The performance is similar, but the computational complexity of relu is lower, so the relu function is finally used. The loss function considers the mean square error loss MSE and the cross entropy loss function. MSE is sensitive to outliers and will reduce the overall model performance. It is not suitable for unknown protocol text datasets with more special samples; the cross entropy loss function is often used for classification problems. It has a high degree of adaptability to the softmax function used in the fully connected layer, and the gradient is larger and the optimization is faster, so the cross entropy loss function is used. CNN outputs a one-dimensional feature vector to the fully connected layer, and the output is 5 Q values. For each action A adopted in each state s, the one-dimensional action vector whose 4 bits are 0 except the position corresponding to the action A number is 1 is multiplied with the 5 Q values and the sum is calculated to obtain Q(s,A) of the action A adopted in each state s. The whole process is as follows: Figure 3 As shown. With traffic text and one-dimensional action vector as input and all Q(s,A) as output, two completely identical models are established, model1 for decision-making and model2 for training. If there is only one Q network, the update and calculation of the target Q value and the Q network parameters need to use each other. The mutual dependence of the two blocks the convergence of the algorithm. Therefore, the improved DQN by predecessors uses two networks with the same structure but different parameters, one for selecting actions and the other for calculating the target Q value. Modal1 used to select actions uses the latest parameters, and modal2 used to estimate the target Q value uses the previous parameters. After a certain number of steps, the latest modal1 parameters are copied to mod2 to reduce the correlation between the target Q value calculation and the Q network parameter update. The overall process of the dual-network DQN is as follows Figure 4 shown.
[0043] In the original Q-learning, the agent's optimal strategy always chooses the best behavior in a given state. The background assumption of this action is that the best action has the largest estimated Q value. However, the actual situation is that the agent knows nothing about the environment at first and needs to estimate and update the Q value. In fact, it is impossible to determine whether the action with the largest estimated Q value is the best action. In fact, in most cases, the Q value of the best behavior is not the maximum Q value. Although the optimal strategy method adopted by the agent can make the Q value quickly approach the target, it will lead to overestimation of the Q value, resulting in suboptimal strategy problems. Therefore, the dual-network DQN algorithm improved by predecessors in reinforcement learning is used to establish two networks with the same structure, one for decision-making actions, namely model1, and the other for calculating the target Q value, namely model2. The parameters of model1 are copied to model2 every certain rounds to reduce the correlation between Q value calculation and network parameter update. Instead of directly finding the maximum Q value of each action in the model2 network, the action corresponding to the maximum Q value is first found in model1, and then the action is used to calculate the target Q value in the model2 network. The calculation formula is as follows:
[0044] y j =R j +γQ`(Φ(S` j ),arg maxQ(Φ(S` j ),a,ω),ω`) (1)
[0045] Among them, y j is the current target Q value; Q represents the current Q network of model1, Q` represents the target Q network of model2; R j is the reward corresponding to the current state-action; γ is the attenuation degree, which indicates the degree of dependence on the future, and its value range is between 0 and 1; S` j is the current state S j The new state after executing action A, Φ(S` j ) is state S` j The characteristic vector of ; a is the learning rate; ω is the network parameter of the current Q network; ω` is the network parameter of the target Q network.
[0046] Step 3: The decision-making process uses the ε-greedy strategy, with a probability of ε to select unknown actions and a probability of 1-ε to select the action with the highest reward in experience, ensuring that each state-action pair has a certain probability of being visited, avoiding being limited to the known maximum value action and no longer making new attempts; which may lead to the problem of not being able to achieve the global optimum. At the beginning of the learning process, a larger ε is required to be able to try as many unselected actions as possible. As the agent accumulates learning experience and becomes more and more certain about the Q value, ε should decrease. Therefore, a separate ε calculation function is set to calculate the ε value for each round of learning. Passing the number of learning rounds step as a parameter into the calculation formula can ensure its decrease. The minimum value of ε is set to 0.01 and the maximum value is set to 1.
[0047] The calculation of ε in the ε-greedy strategy described above requires a larger ε at the beginning of the learning process so that it can try as many actions as possible that have not been selected. As the agent accumulates learning experience and becomes more and more certain about the Q value, ε should decrease, making it more likely that the agent will use experience. Therefore, the calculation formula for ε is as follows:
[0048]
[0049] Among them, εmin is the minimum value of ε set by yourself, εmax is the maximum value of ε set by yourself, step is the current number of training times, and total is a constant set by yourself.
[0050] Step 4: Set an env function as the "system environment". In each cycle, randomly select one from the input samples to be classified and throw it to the model1 responsible for making decisions. The text is the current state s. Combined with the ε-greedy strategy described in step 3, the probability of ε randomly selects an action and returns a Q value of 0; the probability of 1-ε makes model1 predict the text category and calculate the Q value, compare its prediction result with the actual result and give reward and punishment feedback, adding points if correct and deducting points if wrong; the system environment randomly throws the next text as the next state; the returned Q value and the corresponding state s, action A, the obtained score, and the next state are recorded in the experience memory
[0051] Let modal1 classify the text and compare it with the actual result and give reward feedback, because the reward signal can only be used to convey the goal, not how to achieve the goal, and because text clustering is a task with a simple purpose and a small state-action space, the reward mechanism is very simple. If it is the initial state, the reward is 0, 2 points are added for correctness, and 1 point is subtracted for errors. The plus point is slightly larger than the minus point, because the use of penalties may cause the agent to avoid penalties and remain motionless, causing the training process to have insufficient search problems. Since the clustering task is simple, the agent can clarify the goal through the ε-greedy strategy described in step 3, that is, to correctly complete the text classification, ensuring that the effective transfer toward the goal can occupy a certain proportion, so that the algorithm can eventually converge.
[0052] Step 5: Randomly sample a batch of data from the experience memory, input it into the Bellman equation to calculate the new Q value, and update the new Q value to model2; every certain round, copy the parameters of model1 to model2.
[0053] A batch of data is randomly sampled from the experience memory and input into the Bellman equation to calculate the new Q value. Experience replay is a memory replay technology in reinforcement learning. It stores the (St, At, Rt, St+1) experienced by the agent at each step into an array called memory, and then randomly samples a batch of experience for learning. The key reason for using experience replay is that in general reinforcement learning problems, the state St and the next state St+1 are considered to be strongly correlated, and the update of Q(S0, A0) in the Bellman equation depends on Q(Si, Ai). Experience replay can break the correlation between continuous samples and avoid the increase in the variance of parameter updates caused by continuous samples. However, since the classification task of the invention is discrete; each flow data is independent of each other, and there is no certain correlation between the front and the back, and there is no need for experience replay to solve the problem of strong correlation of data. Therefore, the part of the Bellman equation that uses St+1 to calculate the Q value of St is deleted to simplify the overall structure and reduce the amount of calculation.
[0054] The Bellman equation is as follows:
[0055] Q(s,a)←Q(s,a)+ɑ[R(S0,A i )+γmax{Q(S i ,A0),Q(S i ,A1)...}-Q(s,a)] (3)
[0056] Among them, ɑ is the learning rate, and γ is the decay degree, which indicates the degree of dependence on the future. The larger the value, the higher the attention to the future. Otherwise, only the current state is considered. Since the state St and the next state St+1 are considered to be strongly correlated in general reinforcement learning problems, it is reflected in the fact that the update of Q(S0, A0) in the Bellman equation depends on Q(Si, Ai). The random sampling from the experience memory mentioned above is an "experience replay" strategy, which is used to break the correlation between continuous samples and avoid the increase in the variance of parameter updates caused by continuous samples. The neural network requires the correlation of training data to be as low as possible. Therefore, experience replay is used in the reinforcement learning framework based on the continuous state background to maintain the independence of data.
[0057] However, since the present invention is a classification problem, which is a discrete problem rather than a continuous problem, each flow text is independent of each other, and there is no certain correlation between the previous and the next. Our agency does not need to solve the problem of strong data correlation. Therefore, the algorithm framework of the present invention removes the Q value update calculation of integrating the next state St+1 into the current state St, simplifies the overall structure, reduces the amount of calculation, and thus simplifies the Bellman equation as follows:
[0058] Q(S0,A0)←(1-ɑ)Q(S0,A0)+ɑR(S0,A i ) (4)
[0059] Among them, Q(S0,A i ) means to select action A in state S0 i The updated value of the cumulative return, Q(S0,A i ) is the old value of the cumulative return; R(S0,A i ) means to select action A in state S0 i The reward is ; ɑ is the learning rate.
[0060] Step 6: Repeat the above training steps. Generally, the training epoch is set to 1500-2000. The algorithm of the present invention has basically converged when the number of training times reaches 1500, and the cumulative reward exceeds 600 points, so it can be set to 1500.
[0061] Reinforcement learning is to provide the agent with a clear target classification number for learning. The model can eventually accurately output the number of the category to which each text data belongs. It can directly obtain the four data of the number of positive and negative samples that are correctly or incorrectly predicted, and thus calculate the following evaluation indicators:
[0062] 1) Precision: The proportion of all positive samples that are correctly predicted
[0063] 2) Accuracy: The proportion of all samples that are correctly predicted
[0064] 3) Recall: The proportion of samples predicted to be positive that are actually correct
[0065] Figure 5 and Figure 6 The effects of the algorithm of the present invention on source classification of 5 video playback application traffic data sets and classification of normal / advertising / malicious traffic types of more than 1,500 Android app traffic data sets are respectively demonstrated. The algorithm of the present invention is carried out around 5 video playback application traffic data sets in the process of building and training the model, and in the experimental verification and comparison stage, a ready-made Android app traffic data set that has been pre-processed by a third party is used for testing. Therefore, all evaluation indicators have declined, indicating that the video traffic data set processing and model structure of this project may have a certain overfitting phenomenon. The model structure and the parameters used in the construction process are selected around the characteristics of a data set, which results in the model not being able to play an absolutely stable parsing effect on unfamiliar data sets with different structures, but it still maintains a high level of more than 85%, indicating that the algorithm of the present invention performs relatively well in the unknown protocol text classification task with unfamiliar data set input and completely different classification methods, and can play an effective classification role.
Claims
1. A method for clustering unknown protocol texts based on reinforcement learning, characterized in that: The following steps are involved: Step 1: Start the application that needs to capture traffic on the host, start the packet capture tool or code to capture traffic, and pre-process the obtained traffic text; Step 2: Set the action code to the number 0-4 corresponding to the category, regard each traffic text as an environmental state, first process the input traffic text into a feature vector through CNN, then input it into the fully connected layer, and output it as 5 Q values; for the action A adopted in each environmental state s, make the position corresponding to the number of action A 1, and the other 4 positions 0 to form a one-dimensional action vector, dot multiplication of the one-dimensional action vector and the 5 Q values and calculate the sum, and obtain Q(s,A) of action A adopted in each environmental state s. Q(s,A) is the cumulative reward of taking action A in a certain state s according to a certain strategy; use traffic text and one-dimensional action vector as input, and all Q(s,A) as output to establish two completely identical models, model1 for decision-making and model2 for training; Step 3: The decision-making process uses the ε-greedy strategy, with an ε probability of selecting an unknown action and a 1-ε probability of selecting the action with the highest reward in experience, ensuring that each state-action pair has a probability of being visited; Step 4: Set a function as the system environment. In each cycle, randomly select a text from the input sample to be classified and throw it to the model1 responsible for making decisions. The text is the current state. Combined with the ε-greedy strategy described in step 3, the probability of ε randomly selects an action and returns a Q value of 0. The probability of 1-ε makes model1 predict the text category and calculate the Q value. The predicted result is compared with the actual result and reward and punishment feedback is given. If it is correct, it will increase points, and if it is wrong, it will decrease points. The system environment randomly throws the next text as the next state. The returned Q value and the corresponding state s, action A, the obtained score, and the next state are recorded in the experience memory. Step 5: Randomly sample a batch of data from the experience memory, input it into the Bellman equation to calculate the new Q value, and update the new Q value to model2; at regular intervals, copy the parameters of model1 to model2; Step 6: Repeat the training steps until the specified number of training times is completed, or the total score reaches the specified threshold, then the training is completed; the trained model must be able to accurately distinguish the source or type of the newly input traffic text; In step 5, the parameters of model1 are copied to model2. The action corresponding to the maximum Q value is first found in model1, and then the action is used to calculate the target Q value in the network of model2. The calculation formula is as follows: y j =R j +γQ`(Φ(S` j ),arg maxQ(Φ(S` j ),a,ω),ω`) (2) Among them, y j is the current target Q value; Q represents the current Q network of model1, Q` represents the target Q network of model2; R j is the reward corresponding to the current state-action; γ is the decay degree, which indicates the degree of dependence on the future, and its value range is between 0 and 1; S` j is the current state S j The new state after executing action A, Φ(S` j ) is state S` j The characteristic vector of ; a is the learning rate; ω is the network parameter of the current Q network; ω` is the network parameter of the target Q network.
2. The unknown protocol text clustering method based on reinforcement learning according to claim 1 is characterized in that: The specific steps of preprocessing the traffic text in step 1 are as follows: Step 1.1: Decode the captured source data into a hexadecimal string, then convert the hexadecimal characters into corresponding decimal numbers, fill in the null places, and each hexadecimal character is a processing unit; Step 1.2: Unify the length of data with different lengths obtained from different data sources. By using fit_transform to fit part of the data, find out the mean and variance indicators; transform the data, including standardization, normalization, and PCA dimensionality reduction to retain valid information, obtain the data length of different data sources after dimensionality reduction and take the average value x, re-perform PCA dimensionality reduction with parameter x on all data, so that all data lengths are uniformly reduced to x, and then all data lengths are uniformly cut.
3. The unknown protocol text clustering method based on reinforcement learning according to claim 1 is characterized in that: In step 2, CNN uses two convolutional layers and two pooling layers, the step sizes are all set to 1, the activation function uses the relu function, and the loss function uses the cross entropy loss function.
4. The unknown protocol text clustering method based on reinforcement learning according to claim 1 is characterized in that: The calculation formula of ε in step 3 is as follows: Among them, εmin is the minimum value of ε set by yourself, εmax is the maximum value of ε set by yourself, step is the current number of training times, and total is a constant set by yourself.
5. The unknown protocol text clustering method based on reinforcement learning according to claim 1 is characterized in that: The Bellman equation in step 5 is as follows: Q(S0,A i )←(1-ɑ)Q(S0,A i )+ɑR(S0,A i )(3) Among them, Q(S0,A i ) means to select action A in state S0 i The updated value of the cumulative return, Q(S0,A i ) is the old value of the cumulative return; R(S0,A i ) means to select action A in state S0 i The reward is ; ɑ is the learning rate.
Citation Information
Patent Citations
Polarized SAR image classification method based on reinforcement learning
CN113627480A
Traffic classification method for unknown network protocol of application layer
CN114666273A