GB35114 traffic classification and grading method and device based on reinforcement learning, medium and product

By classifying and grading the GB35114 traffic data based on reinforcement learning, the problems of low accuracy and efficiency in the existing technology are solved, and high accuracy and high efficiency traffic grading are achieved, and the information security protection capability of the video surveillance system is improved.

CN120180236AActive Publication Date: 2025-06-20BEIJING TIANFANG SECURITY TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510637452.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-20
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately classify and classify traffic under the GB35114 standard, and face the impact of subjectivity of rule matching, the accuracy of cluster analysis, and the quality of feature sets of machine learning methods.

Method used

Using reinforcement learning-based method, by obtaining the original video traffic data, extracting GB35114 traffic data, performing feature extraction and preprocessing, and using the reinforcement learning model of the hierarchical structure Critic network to output traffic grading results.

Benefits of technology

It realizes effective classification and grading of GB35114 traffic data, improves the accuracy and efficiency of traffic grading, reduces the workload of security personnel, and improves the information security protection capabilities of the public safety video surveillance system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180236A_ABST
    Figure CN120180236A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of traffic grading, in particular to a GB35114 traffic classification and grading method and device based on reinforcement learning, a medium and a product. The method comprises the following steps: acquiring original video traffic data, and classifying GB35114 traffic data from the original video traffic data; performing feature extraction and feature preprocessing on the GB35114 traffic data to obtain preprocessed feature data; inputting the feature data into a pre-trained reinforcement learning model to obtain a traffic grading result for the GB35114 traffic data output by the reinforcement learning model; wherein the Critic network in the reinforcement learning model is of a layered structure, and each layer corresponds to one traffic level. According to the method and the device, accurate classification and grading of the GB35114 traffic can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of traffic classification, and particularly to a GB35114 traffic classification and grading method, device, medium and product based on reinforcement learning. Background Art

[0002] With the rapid development of the information society, video surveillance systems are increasingly widely used in fields such as urban public security, traffic management, and financial security. The data processed and transmitted by these systems often contains a large amount of sensitive information. How to ensure that this information is not illegally obtained, tampered with, or misused has become an urgent problem to be solved. For this reason, the "GB35114-2017 Technical Requirements for Information Security of Public Security Video Surveillance Networking" has emerged. This standard is led by the Ministry of Public Security and jointly formulated by multiple scientific research institutions, aiming to provide comprehensive technical protection for information security in the field of public security video surveillance. Through the GB35114 standard, the information security protection ability of the security video surveillance system can be strengthened, the confidentiality, integrity, and availability of video data can be ensured, and the safety of the country and the people can be guaranteed.

[0003] GB35114 divides into three levels according to different encryption levels, namely Level A, Level B, and Level C. Level A conducts two-way device authentication based on digital certificates to confirm the legitimacy of the identity. That is, a new GB35114 certificate verification logic is added to the traditional GB / T28181 identity registration and authentication process. Based on the certificate and key, the legitimacy of the user and the access device is authenticated to ensure that only legitimate devices can access the management platform. Level B conducts video source authenticity verification based on the video data signature ability of digital certificates to prevent video data from being tampered with during transmission. Level C is the highest level. On the basis of Level B, the video is encrypted and transmitted to effectively prevent video data leakage.

[0004] With the implementation of the GB35114 standard, more and more users will deploy terminal devices that meet the GB35114 standard at important points according to national policies and their own security requirements, which can effectively prevent security problems such as illegal access, video tampering, and video leakage. The asset ratio of terminal devices that meet the GB35114 standard in the user network is usually very low. How to quickly identify the traffic that meets the GB35114 standard from video traffic and automatically classify its security level is convenient for evaluating the number of devices that meet the GB35114 standard in the user network and for realizing the security ability certification and rating of the video surveillance system.

[0005] In the related art, the methods for classifying and grading according to the GB35114 standard mainly include techniques such as rule matching, clustering analysis, and machine learning. However, these methods face many challenges in practical applications. First, it relies on expert experience and knowledge to formulate rules, which not only has a certain degree of subjectivity, but also different rules may lead to inconsistent classification results. Second, in GB35114, its cascading, interconnection, etc. still adopt the GB / T28181 standard, but in the process of network formation, according to the device access requirements, the security algorithm mechanism of GB35114 is added. This will lead to difficulty in distinguishing between GB / T28181 and GB35114 during clustering analysis, thereby affecting the accuracy of the classification results. In addition, although machine learning methods can automatically learn features, researchers still need to manually construct a feature set based on expert experience, and the quality of the feature set directly affects the classification effect. At the same time, the marking work of network traffic is time-consuming and laborious, further increasing the complexity of the classification task.

[0006] Therefore, how to effectively and accurately classify and grade the traffic under the GB35114 standard has become an urgent technical problem to be solved. Summary of the Invention

[0007] In order to solve the problem that the prior art cannot achieve accurate classification and grading of GB35114 traffic, the present application provides a GB35114 traffic classification and grading method, device, medium, and product based on reinforcement learning.

[0008] In the first aspect, the present application provides a GB35114 traffic classification and grading method based on reinforcement learning, adopting the following technical solutions: A GB35114 traffic classification and grading method based on reinforcement learning includes: Obtain the original video traffic data, and classify the GB35114 traffic data from the original video traffic data; Extract and preprocess the features of the GB35114 traffic data to obtain preprocessed feature data; Input the feature data into a pre-trained reinforcement learning model to obtain the traffic grading result of the GB35114 traffic data output by the reinforcement learning model; wherein, the Critic network in the reinforcement learning model is a hierarchical structure, and each layer corresponds to a traffic level.

[0009] By adopting the above technical solution, the original video traffic data is obtained and the GB35114 traffic data is classified. Then, feature extraction and preprocessing are performed on it. Finally, the reinforcement learning model of the hierarchical structure Critic network outputs the traffic classification result, realizing the effective classification and grading of the GB35114 traffic data. Through the Critic network of the hierarchical structure, different traffic levels can be more accurately corresponded, improving the accuracy and efficiency of traffic classification and grading.

[0010] In a preferred example of the present application, it can be further configured that the construction process of the reinforcement learning model includes: Initializing the Actor network, Critic network, target network, and experience replay pool; Defining the reward function and the target function; Obtaining training sample data, using the Actor network and the reward function to generate quintuple data corresponding to the training sample data, and storing the quintuple data in the experience replay pool; Repeatedly executing the model training step until the model training stop condition is met.

[0011] By adopting the above technical solution, initializing each network and the experience replay pool to build a basic framework, defining the reward and target functions to provide a direction and optimization goal for model training, generating quintuple data through the Actor network combined with the reward function and storing it in the experience replay pool to accumulate training materials, and repeatedly training until the stop condition is met. This process enables the reinforcement learning model to effectively learn data features, optimize network parameters, and finally achieve the ability to accurately classify and grade the GB35114 traffic data.

[0012] In a preferred example of the present application, it can be further configured that the model training step includes: Randomly sampling experience data from the experience replay pool, calculating the target Q value based on the target network and the experience data, and updating the predicted Q value of the Critic network; Calculating the loss function value of the Critic network based on the predicted Q value and the target Q value, and updating the parameters of the Critic network based on the loss function value; Calculating the function value of the target function, and updating the parameters of the Actor network; Updating the parameters of the target network.

[0013] By adopting the above technical solution, randomly sampling experience data from the experience replay pool can effectively utilize past experience and reduce data correlation. The target Q value is calculated based on the target network and experience data to update the predicted Q value of the Critic network, making the prediction more accurate; by calculating the loss function value to update the parameters of the Critic network, the network's estimation of value can be optimized; calculating the objective function value to update the parameters of the Actor network enables the Actor network to learn a better strategy; updating the parameters of the target network helps to stabilize the training process and improve the convergence speed and stability of the model.

[0014] In a preferred example of the present application, it can be further configured that: updating the predicted Q value of the Critic network includes: For each layer of the Critic network, obtain the Q value of the layer, and determine the matching degree between the experience data and the preset layer features of the layer. Take the sum of the Q value of the layer and the matching degree as the sub-predicted Q value of the layer; Calculate the predicted Q value of the Critic network based on the sub-predicted Q values of each layer in the Critic network.

[0015] By adopting the above technical solution, when updating the predicted Q value of the Critic network, first calculate the sum of its Q value and the matching degree between the experience data and the preset layer features for each layer as the sub-predicted Q value. This can more carefully evaluate the value by combining the characteristics of each layer and its fit with the experience data. Then, calculate the overall predicted Q value based on the sub-predicted Q values of each layer, which can be dynamically adjusted according to the security level of the current state, avoiding the fuzzy processing of complex security constraints by a single Q value.

[0016] In a preferred example of the present application, it can be further configured that: calculating the function value of the objective function includes: Determine the security index and communication delay index based on the experience data; Calculate the function value of the objective function based on the predicted Q value, the security index, and the communication delay index.

[0017] By adopting the above technical solution, an innovative multi-objective optimization strategy design is carried out in the optimization strategy. In addition to the value estimation of traffic data, additional objective functions (security and communication delay) are introduced. Through the design of an optimal objective function with weights for fusion calculation, multi-objective optimization is realized to balance performance and security.

[0018] In a preferred example of the present application, it can be further configured that: classifying the GB35114 traffic data from the original video traffic data includes: Perform a feature classification step on each piece of traffic data in the original video traffic data, and classify the GB35114 traffic data from the original video traffic data based on the classification results of each piece of traffic data; Among them, the feature classification step includes: matching the current traffic data with the relevant features of GB / T28181 traffic to obtain a comprehensive feature matching degree, and judging whether the current traffic data is the GB35114 traffic data based on the comprehensive feature matching degree.

[0019] By adopting the above technical solution, perform the feature classification step on each piece of the original video traffic data one by one. By matching the current traffic data with the relevant features of GB / T28181 traffic to obtain the comprehensive feature matching degree, and thus judge whether it is the GB35114 traffic data. This method can carefully and accurately screen out the GB35114 traffic data from a large amount of original video traffic data, improving the accuracy and efficiency of data classification.

[0020] In a preferred example of the present application, it can be further configured as: performing feature extraction and feature preprocessing on the GB35114 traffic data to obtain preprocessed feature data, including: Performing multi-dimensional feature extraction on the GB35114 traffic data to obtain initial feature data; Successively performing outlier removal and standardization processing on the initial feature data to obtain intermediate feature data; Extracting the time series features of the intermediate feature data; Performing format normalization processing on the intermediate feature data and the time series features to obtain preprocessed feature data.

[0021] By adopting the above technical solution, first perform multi-dimensional feature extraction on the GB35114 traffic data to comprehensively obtain data features. Then, obtain intermediate feature data through outlier removal and standardization processing, which can improve data quality and model training efficiency. Then, extract the time series features and perform format normalization processing together with the intermediate feature data to make the data format unified and regular. The finally obtained preprocessed feature data can better be applied to subsequent model training and analysis, improving the accuracy and stability of the model.

[0022] In a second aspect, the present application provides an electronic device, adopting the following technical solution: One or more processors; A memory; At least one application program, where at least one application program is stored in the memory and is configured to be executed by at least one processor. The at least one application program is configured to: execute the GB35114 traffic classification and grading method based on reinforcement learning as described in any item of the first aspect.

[0023] In a third aspect, the present application provides a computer-readable storage medium, adopting the following technical solution: A computer-readable storage medium stores a computer program, which, when executed on a computer, causes the computer to execute the GB35114 traffic classification and grading method based on reinforcement learning according to any one of the first aspect.

[0024] In a fourth aspect, the present application provides a computer program product, adopting the following technical solution: A computer program product includes a computer program, which, when executed by a processor, implements the GB35114 traffic classification and grading method based on reinforcement learning according to any one of the first aspect.

[0025] In summary, the present application includes the following beneficial technical effects: By acquiring the original video traffic data and classifying the GB35114 traffic data, then extracting features and preprocessing it, and finally using the reinforcement learning model of the hierarchical structure Critic network to output the traffic grading result, the present application realizes the effective classification and grading of the GB35114 traffic data. Through the Critic network with a hierarchical structure, it can more accurately correspond to different traffic levels, improving the accuracy and efficiency of traffic grading. Description of the Drawings

[0026] Figure 1 is a schematic structural diagram of a GB35114 traffic classification and grading system based on reinforcement learning provided by an embodiment of the present application; Figure 2 is a schematic flowchart of a GB35114 traffic classification and grading method based on reinforcement learning provided by an embodiment of the present application; Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0027] The following will Figure 1 - be Figure 3 further described in detail with reference to the accompanying

[0028] This specific embodiment is only an interpretation of the present application and does not limit the present application. Those skilled in the art can make modifications without creative contributions to this embodiment according to their needs after reading this specification, but as long as they are within the scope of the claims of the present application, they are protected by the patent law.

[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0030] In addition, the term "and / or" in this document is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, unless otherwise specified.

[0031] It should be noted that in the optional embodiments of this application, for relevant data such as object information, when the embodiments in this application are applied to specific products or technologies, object permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. That is to say, if the embodiments of this application involve data related to objects, it needs to be obtained under the authorization and consent of the objects, the authorization and consent of relevant departments, and compliance with the relevant laws, regulations, and standards of the country and region. If personal information is involved in the embodiments, the acquisition of all personal information needs to obtain the consent of the individual. If sensitive information is involved, the separate consent of the information subject needs to be obtained, and the embodiments also need to be implemented under the authorization and consent of the objects.

[0032] The objective of this application is to automatically extract traffic data that conforms to the GB35144 standard from video traffic data, effectively and accurately automatically classify the A, B, and C levels of traffic in GB35144, and facilitate the implementation of the security capability assessment and rating of video surveillance systems. It can not only improve the accuracy and efficiency of traffic classification and grading under the GB35114 standard, but also greatly reduce the time and effort consumed by security personnel in classifying and grading traffic, thus significantly enhancing the information security protection capability of public security video surveillance systems.

[0033] Based on this, the embodiments of this application provide a GB35114 traffic classification and grading system based on reinforcement learning. This system can be loaded on an electronic device, such as Figure 1 As shown, this system includes: a data collection and coarse-grained classification module 101, a traffic data preprocessing module 102, a reinforcement learning model construction module 103, a model training and optimization module 104, and a real-time fine-grained grading module 105.

[0034] The data acquisition and coarse-grained classification module 101 is used to classify the recognizable features extracted from the original video traffic data in a coarse-grained manner by constructing a string matching mechanism, so as to classify the GB35114 traffic data from the original video traffic data.

[0035] The traffic data preprocessing module 102 is used to extract traffic behavior features from the GB35114 traffic data, construct a multi-dimensional feature vector, clean and preprocess the extracted features to remove outliers and noise interference, and ensure the high quality of the data.

[0036] The reinforcement learning model construction module 103 is used to construct a deep neural network model, and let the model automatically learn the complex mapping relationship between the traffic behavior features and the A, B, and C level standards of GB35144.

[0037] The model training and optimization module 104 is used to train the model with a large number of labeled traffic sample data, continuously optimize the model parameters, and improve the generalization ability and classification accuracy of the model.

[0038] The real-time fine-grained classification module 105 is used to input the collected real-time video traffic data into the trained reinforcement learning model in actual applications, and receive the classification result of the GB35114 traffic output by the model, which can significantly improve the accuracy and efficiency of traffic classification and grading under the GB35114 standard.

[0039] The embodiment of the present application provides a GB35114 traffic classification and grading method based on reinforcement learning, as Figure 1 shown. In the method provided in the embodiment of the present application, it is applied to a GB35114 traffic classification and grading system based on reinforcement learning, which is executed by an electronic device. The electronic device can be a server or a terminal device. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods. The embodiment of the present application does not make any restrictions here. The method includes steps S201 - step S203, where: S201. Obtain the original video traffic data, and classify the GB35114 traffic data from the original video traffic data.

[0040] Specifically, traffic data in the video surveillance network can be collected through a probe or a network interface to obtain the original video traffic data of the video surveillance network. Since GB / T28181 and GB35114 are highly similar and it is easy to make misjudgments, this application proposes a method of coarse-grained clustering to achieve rapid classification of GB / T28181 and GB35114 traffic. The original video traffic data may contain multiple traffic data, and each traffic data can be subjected to coarse-grained clustering analysis to determine whether the traffic data is GB / T28181 or GB35114 traffic.

[0041] Specifically, relevant features of GB / T28181 traffic are defined in advance based on the characteristics of GB / T28181. There can be multiple relevant features, and a matching degree threshold is defined. Any traffic data in the original video traffic data is used as the current traffic data, and the current traffic data is matched with the relevant features of GB / T28181 traffic to obtain a comprehensive feature matching degree. The larger the comprehensive feature matching degree, the higher the similarity between the current traffic data and the relevant features of GB / T28181 traffic. Furthermore, the comprehensive feature matching degree is compared with the matching degree threshold. If the comprehensive feature matching degree does not exceed the matching degree threshold, it is determined that the current traffic data is GB35114 traffic. If the comprehensive feature matching degree exceeds the matching degree threshold, it is determined that the current traffic data is GB / T28181 traffic.

[0042] S202. Extract features and perform feature preprocessing on the GB35114 traffic data to obtain preprocessed feature data.

[0043] Specifically, for each traffic data determined to be GB35114 traffic data, multi-dimensional features are extracted as the initial feature data. Among them, the feature dimensions include: transmission behavior features, connection behavior features, data packet features, and encryption features, etc. Specific features include: the average data packet rate and the standard deviation of rate fluctuation in the transmission behavior feature dimension, the connection duration in the connection behavior feature dimension, the data packet size distribution entropy value in the data packet feature dimension, the encryption algorithm identifier, key negotiation frequency, signaling security, identity authentication, signature control, encryption control, data encryption, etc. in the encryption feature dimension.

[0044] Furthermore, outlier removal and standardization processing are sequentially performed on the initial feature data to obtain intermediate feature data. The time series features of the intermediate feature data are extracted, and the intermediate feature data and the time series features are subjected to format normalization processing to obtain preprocessed feature data. The preprocessed feature data is a 16-dimensional feature vector.

[0045] S203. Input the feature data into a pre-trained reinforcement learning model to obtain the traffic classification result for GB35114 traffic data output by the reinforcement learning model. Among them, the Critic network in the reinforcement learning model is a hierarchical structure, and each layer corresponds to a traffic level.

[0046] This embodiment realizes the classification of two protocols with high similarity, namely GB / T28181 and GB35114, accurately extracts the GB35114 traffic data, and the reinforcement learning model realizes the classification of the real-time traffic into three levels, namely A, B, and C in GB35114. The system monitors the accuracy of traffic classification and the security status of the system in real time. If a classification error or an abnormal situation occurs in the system, the relevant information is fed back to the reinforcement learning model, and the model performs online learning and adjustment according to the new feedback information and real-time traffic data, continuously optimizing the classification strategy to adapt to the dynamic changes of the network environment and ensuring the effective protection of the information security of the video surveillance system at all times.

[0047] A possible implementation manner of the embodiment of the present application for classifying GB35114 traffic data from the original video traffic data includes: Performing a feature classification step on each traffic data in the original video traffic data, and classifying the GB35114 traffic data from the original video traffic data based on the classification result of each traffic data; Among them, the feature classification step includes: matching the current traffic data with the relevant features of GB / T28181 traffic to obtain a comprehensive feature matching degree, and judging whether the current traffic data is GB35114 traffic data based on the comprehensive feature matching degree.

[0048] In this embodiment, the relevant features of GB / T28181 traffic are specific SIP methods defined to implement the functions of the video surveillance system, including: the INVITE method is used to establish a video call session, the ACK method is used to confirm the session establishment, the BYE method is used to end the session, the OPTIONS method is used to query device capabilities, REGISTER is used for device registration, a specific 401 Unauthorized response, a 20-digit digital coding rule for the device id, and the content of the application / sdp type is included in the message body, etc. If these specific SIP methods are found in the traffic and the message content is related to the video surveillance service, then it can be determined as GB / T28181 traffic.

[0049] Construct a fusion coarse-grained classification algorithm for GB / T28181 and GB35114: Among them, represents the comprehensive matching degree of the relevant features of the current traffic data and the GB / T28181 traffic; A is whether the INVITE method is matched; B is whether the ACK method and the BYE method are matched; C is whether the OPTIONS method and the REGISTER method are matched; D is whether a specific 401 Unauthorized response is matched; E is whether a 20-digit digital coding rule conforming to the device id is matched; F is whether the message body contains content of the application / sdp type. For A, B, C, D, E, and F, if the matching result is a match, the corresponding value is set to 1, and if not, the corresponding value is set to 0. are the respective parameter weights and can be set according to actual experience.

[0050] Furthermore, the obtained comprehensive feature matching degree is compared with the preset comprehensive feature matching degree threshold. When the comprehensive feature matching degree exceeds the preset comprehensive feature matching degree threshold, it indicates that the matching degree of the relevant features of the current traffic data and the GB / T28181 traffic is too high, and the current traffic data is determined to be GB / T28181 traffic; when the matching degree of the comprehensive features does not exceed the preset comprehensive feature matching degree threshold, the current traffic data is determined to be GB35114 traffic data.

[0051] This embodiment can classify the original video traffic by integrating the coarse-grained matching algorithm, accurately extract the GB35114 traffic data therein, so as to facilitate the next analysis and processing of the GB35114 traffic data for classification of Class A, Class B, and Class C.

[0052] A possible implementation manner of the embodiment of the present application is to perform feature extraction and feature preprocessing on the GB35114 traffic data to obtain preprocessed feature data, including: Performing multi-dimensional feature extraction on the GB35114 traffic data to obtain initial feature data; Successively performing outlier removal and normalization processing on the initial feature data to obtain intermediate feature data; Extracting the time series features of the intermediate feature data; Performing format normalization processing on the intermediate feature data and the time series features to obtain preprocessed feature data.

[0053] In this embodiment, outlier removal is performed on the extracted initial feature data, including: adopting the IsolationForest algorithm to identify and remove abnormal traffic, , where E(h(x)) is the average value of the path lengths of the data point x in all trees, n is the sample size of the training data, and c(n) is the average path length of the trees, which is used for normalization. Then, the data is normalized, and Z-score normalization is performed on each dimension. The normalization formula: , where Z is the standardized data point, X is the original data point before standardization, μ is the mean of the original data, and σ is the standard deviation of the original data.

[0054] Extract the time series features of the intermediate feature data, including: obtaining the time series features by processing the intermediate feature data through a sliding window.

[0055] Furthermore, the format planning and processing includes dimensionality reduction and format integration. Perform dimensionality reduction operations on the intermediate feature data and the time series features. Apply PCA to compress the original 32-dimensional features to 16 dimensions, and then use the destination IP address as the key (key) to integrate the dimensionality-reduced features into the JSON format to obtain the preprocessed feature data, which is convenient for subsequent data processing and analysis.

[0056] In this embodiment, multi-dimensional feature extraction is first performed on the GB35114 traffic data to comprehensively obtain data features. Then, intermediate feature data is obtained through outlier removal and standardization processing, which can improve data quality and model training efficiency. Next, time series features are extracted and format normalization processing is performed together with the intermediate feature data to make the data format unified and regular. Finally, the preprocessed feature data obtained can better be applied to subsequent model training and analysis, improving the accuracy and stability of the model.

[0057] The Deep Deterministic Policy Gradient (DDPG) algorithm is an algorithm based on deep reinforcement learning, aiming to optimize a deterministic policy through gradient descent rather than a random policy. In the field involved in this application, network traffic can be divided into different levels according to the GB35114 standard. There are many features and high dimensions of traffic. In most cases, a single traffic feature cannot play a decisive role in the classification result, but multiple features need to be fused and calculated, which will make the process of reinforcement learning very unstable and the convergence speed very slow. After multiple experiments, this embodiment selects the deep deterministic policy gradient algorithm, which has advantages in dealing with problems of high-dimensional states, and has good stability and convergence, and can effectively learn complex policies.

[0058] A possible implementation manner of the embodiment of this application, the construction process of the reinforcement learning model includes: Initialize the Actor network, Critic network, target network, and experience replay pool; Define the reward function and the target function; Obtain the training sample data, use the Actor network and the reward function to generate the five-tuple data corresponding to the training sample data, and store the five-tuple data in the experience replay pool; Repeat the model training steps until the model training stop condition is met.

[0059] In this embodiment, the Deep Deterministic Policy Gradient (DDPG) algorithm adopts a dual-network architecture. After dimensionality reduction of traffic features, a 16-dimensional core feature vector is extracted, covering key security metrics such as traffic payload, connection timing, protocol features, signaling security, identity authentication, signature control, encryption control, and data encryption. The resulting state space S is a 16-dimensional feature vector. The Actor network is responsible for generating actions, and the Critic network is responsible for evaluating the quality of actions. The Critic network of traditional DDPG outputs a single Q value and cannot perform differential evaluation for different security levels. In the field of GB35114 traffic classification, Level A includes certificate verification logic, Level B adds signature verification logic on the basis of Level A, and Level C encrypts data on the basis of Level B.

[0060] In this embodiment, the action space A of the Actor network is innovatively discretized into three security level decisions {0, 1, 2}, corresponding to Level A, Level B, and Level C of GB35114 traffic respectively. The Critic network is designed as a hierarchical structure with three layers, corresponding to Level A, Level B, and Level C of GB35114 traffic respectively. This can not only evaluate the action quality more precisely, optimize the action selection strategy for different security levels, but also improve the adaptability of the model to complex environments. The target network includes the target Actor network corresponding to the Actor network and the target Critic network corresponding to the Critic network.

[0061] Based on the characteristics of GB35114 traffic classification, a reward function R is defined to better match the requirements of GB35114 traffic classification, including: the reward score for correct classification is +100, the reward score for incorrect classification is -50, the reward score for detecting an adversarial network is -20 (an isolation forest or autoencoder can be used to detect the adversarial network), and the reward score for action delay exceeding the threshold is -10.

[0062] Initialize the relevant parameters of the reinforcement learning model, including the weights and biases of the neural network. The training sample data has the same data format as the feature data after preprocessing the GB35114 traffic data in this application. The training sample data includes training traffic data and corresponding classification labels. For each piece of training sample data, it is input into the reinforcement learning model. The constructed model selects an action according to the current policy and classifies the traffic. Calculate the reward value according to the selected action and the actual labeled GB35114 classification result. Store the current state, action, reward value, and the next state in the experience replay pool. Specifically, at state take a deterministic behavior strategy μ, generate an action through the Actor network θ, execute the action to obtain a new state and a reward , and whether it is in a termination state , store the five-tuple data in the experience replay pool.

[0063] The conventional optimization strategy is to use the predicted Q-value as the objective function and maximize the predicted Q-value of the Critic network for optimization: , is the performance metric of the Actor network, representing the average predicted Q-value evaluated by the Critic network under the parameters . The goal is to maximize this performance metric. Then, use the gradient backpropagation of the neural network to update all the parameters of the Actor network.

[0064] This embodiment innovatively adopts a multi-objective optimization strategy. The objective function not only includes maximizing the Q-value but also introduces additional security metrics and communication delay metrics, and balances performance and security through multi-objective optimization. The objective function is: .

[0065] Among them, is the performance metric of the Actor network, and the goal is to maximize this performance metric; are the parameters of the Actor network; represents the th sample; represents the th sample state; represents the output action of the Actor network in the state; represents the predicted Q-value of the Critic network; represents the parameters of the Critic network; N is the batch size, that is, the number of samples of the empirical data sampled from the experience replay pool; represents the security metric; represents the communication delay metric; is the weight parameter used to balance the importance of each objective.

[0066] After the number of samples in the experience replay pool reaches a certain threshold, a small batch of experience data is randomly sampled from the experience replay pool, and after each sampling, a round of model training is performed using the sampled experience data. The neural network parameters of the model are updated by minimizing the loss function (such as the mean squared error loss function), and at the same time, hyperparameters of the neural network such as the learning rate, the number of neurons in the hidden layer, and the batch size are adjusted. The optimal combination of hyperparameters is selected through methods such as cross-validation, so that the Q-value function is closer to the true state-action value function. The above process is continuously repeated. As the training progresses, the policy is gradually optimized, so that the model can more often select the optimal action based on the Q-value, improving the accuracy of classification. Until the model training stop condition is met, such as reaching the maximum number of training rounds.

[0067] This embodiment introduces the reinforcement learning technology into the security classification scenario of GB35114 video surveillance traffic, breaking through the traditional passive classification mode based on rule matching or static feature analysis (such as the prior art relying on manually preset thresholds or fixed risk indicators), and realizing the automated and adaptive determination of the risk level through the reward and punishment mechanism of the reinforcement learning algorithm model and the dynamic interaction with the traffic environment, providing a new technical path for high-dynamic and multi-dimensional security assessment.

[0068] Aiming at the problem that the single layer of the value evaluation network in traditional reinforcement learning cannot finely match complex security levels, the present invention innovatively designs the Critic network of the deep deterministic policy gradient algorithm in layers, with each layer corresponding to Class A / Class B / Class C of GB35114, and dynamically adjusts according to the security level of the current state. It can not only more finely evaluate the action quality, optimize the action selection strategy for different security levels, but also improve the adaptability of the model to complex environments.

[0069] A possible implementation manner of the embodiment of the present application, the model training steps include: Randomly sample experience data from the experience replay pool, calculate the target Q-value based on the target network and the experience data, and update the predicted Q-value of the Critic network; Calculate the loss function value of the Critic network based on the predicted Q-value and the target Q-value, and update the parameters of the Critic network based on the loss function value; Calculate the function value of the target function, and update the parameters of the Actor network; Update the parameters of the target network.

[0070] In this embodiment, for each training sample in the empirical data, the training sample is input into the Actor network to generate an action, and the next state is input into the target Actor network. The target Actor network receives the next state in the training sample and outputs a target action. The target Critic network evaluates the value of the current state-action pair of the next state and the target action. The target Q-value is the true value estimate of the current state-action pair. The target Q-value of the Critic network is updated using the Bellman equation: ; where y is the target Q-value, and R is the function value of the reward function; is the target Actor network, are the parameters of the target Actor network; is the target Critic network, are the parameters of the target Critic network; is the discount factor, indicating the degree of attenuation of future rewards, with a value range of (0, 1). The closer it is to 1, the more importance is attached to long-term benefits; is the next state; is the target action.

[0071] For each layer of the Critic network, obtain the Q-value of the layer and determine the matching degree between the training sample and the preset layer features of the layer. Take the sum of the Q-value of the layer and the matching degree as the sub-predicted Q-value of the layer; Calculate the predicted Q-value of the Critic network based on the sub-predicted Q-values of each layer in the Critic network.

[0072] Traverse each training sample in the empirical data to obtain the target Q-value and the predicted Q-value corresponding to each training sample. Calculate the loss function value using the mean squared error based on the target Q-value and the predicted Q-value corresponding to each training sample in the empirical data, , is the loss function value of the Critic network; are the parameters of the Critic network; is the target Q-value; N is the batch size, that is, the number of training samples of the empirical data sampled from the experience replay pool; represents the th state of the training sample; represents the th action of the training sample. Based on the calculated loss function value, use the gradient backpropagation of the neural network to update all the parameters of the Critic network.

[0073] Based on the defined objective function and the predicted Q-value, safety index, and communication delay index of the current round of training, calculate the function value of the objective function. Use the soft update mechanism to update the parameters of the target network, ; is the soft update coefficient.

[0074] In this embodiment, by randomly sampling experience data from the experience replay pool, past experience can be effectively utilized and data correlation can be reduced. The target Q value is calculated based on the target network and the experience data to update the predicted Q value of the Critic network, making the prediction more accurate. By calculating the loss function value to update the Critic network parameters, the network's estimation of the value can be optimized. By calculating the objective function value to update the Actor network parameters, the Actor network can learn a better strategy. Updating the target network parameters helps to stabilize the training process and improve the convergence speed and stability of the model.

[0075] A possible implementation of the embodiment of the present application for updating the predicted Q value of the Critic network includes: For each layer of the Critic network, obtain the Q value of the layer, determine the matching degree between the experience data and the preset layer features of the layer, and use the sum of the Q value of the layer and the matching degree as the sub-predicted Q value of the layer; Calculate the predicted Q value of the Critic network based on the sub-predicted Q values of each layer in the Critic network.

[0076] The original Critic network outputs a single Q value: , The modified hierarchical Critic network in this embodiment outputs multiple Q values, corresponding to different security levels respectively. Each layer of the network can learn the decision preferences under specific security levels. Level A matches the Authorization feature of the GB35114 traffic and calculates the similarity , represents the multi-dimensional feature vector represented by the training sample, represents the Authorization feature vector. Level B matches the VideoSignatureControl feature of the GB35114 traffic and calculates the similarity , represents the multi-dimensional feature vector represented by the training sample, represents the VideoSignatureControl feature. Level C matches the VideoEncryptionControl of the GB35114 traffic, the VideoPlayEncryption of the GB35114 traffic, and the encryption confusion degree, and calculates the similarity , represents the multi-dimensional feature vector represented by the training sample, represents the VideoEncryptionControl feature, represents the VideoPlayEncryption feature, The correlation coefficient representing the degree of encryption and obfuscation, is the preset weight.

[0077] Calculate the decision preference degree for each layer to obtain the sub-prediction Q value corresponding to each layer: , where is the sub-prediction Q value of the k-th layer, represents the Q value function of the k-th layer, 3 is the total number of security levels, is the similarity of the k-th layer.

[0078] The predicted Q value of the Critic network is obtained by weighted summation of the sub-prediction Q values of each layer, , is the weight of the k-th layer.

[0079] In this embodiment, when updating the predicted Q value of the Critic network, first calculate the sum of its Q value and the matching degree of the empirical data and the preset layer features for each layer as the sub-prediction Q value, so as to more carefully evaluate the value by combining the characteristics of each layer itself and its fit with the empirical data. Then, calculate the overall predicted Q value based on the sub-prediction Q values of each layer, which can be dynamically adjusted according to the security level of the current state, avoiding the fuzzy processing of complex security constraints by a single Q value.

[0080] A possible implementation manner of the embodiment of the present application, calculating the function value of the objective function, includes: Determine the security index and communication delay index based on the empirical data; Calculate the function value of the objective function based on the predicted Q value, security index, and communication delay index.

[0081] In this embodiment, the encryption feature and transmission behavior feature can be extracted from the empirical data first. The security index is determined based on the encryption feature, and the communication delay index is determined based on the transmission behavior feature.

[0082] For the security index, set the scoring rule. The encryption feature may include: encryption control field, encryption algorithm identifier, key length, certificate validity, etc. Set the corresponding encryption strength scores for unencrypted, encrypted and using a certain encryption algorithm respectively, set the validity score for the certificate validity, and perform weighted summation on the encryption strength and validity scores to obtain the security index.

[0083] For the communication delay index, set the scoring rule. The transmission behavior feature may include: packet arrival time interval, network congestion mark, protocol type, etc. Set the delay baseline based on the protocol type, and adjust the delay baseline based on network congestion. For example, when network congestion occurs, the delay baseline increases, indicating that the maximum allowed delay becomes larger. Calculate the difference between the packet arrival time interval (current delay) and the delay baseline, and use the ratio of the difference to the delay baseline as the security index.

[0084] The objective function is as follows: Based on the predicted Q value, safety index, communication delay index, and their respective weights, calculate the function value of the objective function.

[0085] In this embodiment, a multi-objective optimization strategy design is innovatively carried out in the optimization strategy. By introducing additional objective functions (safety and communication delay), and through the design of an optimal objective function with weights for fusion calculation, multi-objective optimization is achieved to balance performance and safety.

[0086] An electronic device is provided in an embodiment of the present application, as Figure 3 shown. Figure 3 The electronic device 300 shown in the figure includes: a processor 301 and a memory 303. Among them, the processor 301 and the memory 303 are connected, such as through a bus 302. Optionally, the electronic device 300 may further include a transceiver 304. It should be noted that in practical applications, the transceiver 304 is not limited to one, and the structure of the electronic device 300 does not constitute a limitation to the embodiments of the present application.

[0087] The processor 301 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the disclosure of the present application. The processor 301 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0088] The bus 302 may include a path for transmitting information between the above components. The bus 302 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 302 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 3 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0089] The memory 303 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0090] The memory 303 is used to store the application program code for executing the solution of this application, and is controlled by the processor 301 for execution. The processor 301 is used to execute the application program code stored in the memory 303 to implement the content shown in the foregoing embodiments of the GB35114 traffic classification and grading method based on reinforcement learning.

[0091] Figure 3 The illustrated electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of this application.

[0092] The embodiments of this application provide a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, it enables the computer to execute the content shown in the foregoing embodiments of the GB35114 traffic classification and grading method based on reinforcement learning.

[0093] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially according to the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order restriction, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0094] An embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the content shown in the foregoing embodiment of the GB35114 traffic classification and grading method based on reinforcement learning.

[0095] The above are only partial embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A GB35114 traffic classification and grading method based on reinforcement learning, characterized in that: include: Acquire original video traffic data, and classify GB35114 traffic data from the original video traffic data; Performing feature extraction and feature preprocessing on the GB35114 flow data to obtain preprocessed feature data; The feature data is input into a pre-trained reinforcement learning model to obtain a traffic classification result for the GB35114 traffic data output by the reinforcement learning model; wherein the Critic network in the reinforcement learning model is a hierarchical structure, and each layer corresponds to a traffic level.

2. According to the GB35114 traffic classification and grading method based on reinforcement learning according to claim 1, it is characterized in that: The construction process of the reinforcement learning model includes: Initialize the Actor network, Critic network, target network and experience replay pool; Define reward function and objective function; Acquire training sample data, generate quintuple data corresponding to the training sample data using the Actor network and the reward function, and store the quintuple data in the experience replay pool; Repeat the model training steps until the model training stop condition is met.

3. The GB35114 traffic classification and grading method based on reinforcement learning according to claim 2 is characterized in that: The model training step comprises: Randomly sampling experience data from the experience replay pool, calculating a target Q value based on the target network and the experience data, and updating the predicted Q value of the Critic network; Calculating a loss function value of the Critic network based on the predicted Q value and the target Q value, and updating parameters of the Critic network based on the loss function value; Calculate the function value of the objective function and update the parameters of the Actor network; The parameters of the target network are updated.

4. The GB35114 traffic classification and grading method based on reinforcement learning according to claim 3 is characterized in that: The updating of the predicted Q value of the Critic network includes: For each layer of the Critic network, obtain the Q value of the layer, determine the matching degree between the empirical data and the preset layer features of the layer, and use the sum of the Q value of the layer and the matching degree as the sub-prediction Q value of the layer; The predicted Q value of the Critic network is calculated based on the sub-predicted Q values ​​of each layer in the Critic network.

5. The GB35114 traffic classification and grading method based on reinforcement learning according to claim 3 is characterized in that: The function value of the calculated objective function includes: Determining a security index and a communication delay index based on the empirical data; Based on the predicted Q value, the safety index and the communication delay index, a function value of the objective function is calculated.

6. The GB35114 traffic classification and grading method based on reinforcement learning according to claim 1 is characterized in that: The classifying GB35114 traffic data from the original video traffic data includes: Performing a feature classification step on each piece of traffic data in the original video traffic data, and classifying the GB35114 traffic data from the original video traffic data based on the classification result of each piece of traffic data; Among them, the feature classification step includes: matching the current flow data with relevant features of GB / T28181 flow to obtain a comprehensive feature matching degree, and judging whether the current flow data is the GB35114 flow data based on the comprehensive feature matching degree.

7. The GB35114 traffic classification and grading method based on reinforcement learning according to claim 1 is characterized in that: The feature extraction and feature preprocessing of the GB35114 flow data to obtain the preprocessed feature data includes: Perform multi-dimensional feature extraction on the GB35114 traffic data to obtain initial feature data; The initial feature data is sequentially subjected to outlier elimination and standardization processing to obtain intermediate feature data; Extracting time series features of the intermediate feature data; The intermediate feature data and the time series feature are subjected to format normalization processing to obtain preprocessed feature data.

8. An electronic device, characterized in that: include: at least one processor; Memory; At least one application, wherein at least one application is stored in a memory and configured to be executed by at least one processor, and the at least one application is configured to: execute the GB35114 traffic classification and grading method based on reinforcement learning as described in any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed in a computer, the computer is caused to execute the GB35114 traffic classification and grading method based on reinforcement learning as described in any one of claims 1 to 7.

10. A computer program product, characterized in that It includes a computer program, which, when executed by a processor, implements the steps of the GB35114 traffic classification and grading method based on reinforcement learning described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for realizing monitoring video security encryption function, processor and computer readable storage medium thereof

    CN113965381A

  • Audio and video data processing method and system based on GB35114

    CN114554286A

  • Welding abnormity real-time diagnosis method based on Actor-Critic reinforcement learning model

    CN115673596A

  • Data encryption method and device, equipment and storage medium

    CN119272294A

  • Method and device to provide a security level for communication

    WO2022171657A1