Network session behavior analysis method, firewall policy optimization method and system

Optimizing firewall policies through network session behavior analysis and Q-learning algorithms has solved the problem that existing firewall technologies cannot adjust their policies in real time when facing complex attacks, and achieved dynamic adaptability and high-precision detection.

CN120200828APending Publication Date: 2025-06-24BEIJING CATHAY INTERNET INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510506965.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The lack of dynamic learning and optimization capabilities of existing firewall technologies leads to high false positive rates and underreporting problems, especially in the face of complex attacks such as advanced persistent threats (APTs), and the inability to adjust strategies in real time.

Method used

The network session behavior analysis method is adopted to identify abnormal sessions through session monitoring and behavior analysis, and use machine learning models such as random forests or convolutional neural networks to optimize firewall strategies in combination with Q learning algorithms.

Benefits of technology

It realizes dynamic adaptability, automatically adapts to new network attacks, reduces false alarm rates and missed alarm rates, and improves the intelligence level and response speed of the firewall.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200828A_ABST
    Figure CN120200828A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network security, and discloses a network session behavior analysis method and a firewall policy optimization method and system, and the behavior analysis method comprises the following steps: S1, session monitoring: capturing and analyzing network session behavior data, and extracting session features; s2, behavior analysis: classifying session features, and analyzing whether network session behaviors are abnormal or not. According to the invention, the problems of false alarm rate or missing alarm and the like in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular, to a network session behavior analysis method, a firewall policy optimization method, and a system. Background Art

[0002] Most of the existing firewall technologies rely on statically configured rule sets. This traditional method is difficult to cope with the dynamically changing network environment and emerging attack patterns. Although modern firewalls can use deep packet inspection (DPI) and traffic analysis to identify known threats, these technologies lack the ability of dynamic learning and optimization, resulting in high false positive rates and missed alarm problems. Especially when facing complex attacks such as advanced persistent threats (APT), the existing firewalls lack the adaptive ability and cannot adjust the strategy in real time.

[0003] Therefore, how to achieve dynamic firewall policy generation and optimization, automatically learn and cope with new attack behaviors, has become a major challenge in the current technology. Summary of the Invention

[0004] To overcome the deficiencies of the prior art, the present invention provides a network session behavior analysis method, a firewall policy optimization method, and a system, which solve the problems such as false positive rate or missed alarm existing in the prior art.

[0005] The technical solution adopted by the present invention to solve the above problems is as follows:

[0006] A network session behavior analysis method includes the following steps:

[0007] S1, Session monitoring: Capture and parse network session behavior data, and extract session features;

[0008] S2, Behavior analysis: Classify the session features and analyze whether the network session behavior is abnormal.

[0009] As a preferred technical solution, in step S2, the random forest algorithm is used to classify the session features, including the following steps:

[0010] SA21, Data preprocessing: Preprocess the session features;

[0011] SA22, Labeling data: Label the session features as normal sessions or abnormal sessions;

[0012] SA23, Training the random forest model: Use the labeled data as the training set of the random forest algorithm for training, and classify the session features by constructing multiple decision trees; among them, each tree independently judges whether a session is abnormal, and finally determines the classification result by voting;

[0013] SA24, Analyze network traffic using the random forest model.

[0014] As a preferred technical solution, in step S2, a convolutional neural network is used to classify session features, including the following steps:

[0015] SB21, data input: Convert the session features into an input format acceptable to the convolutional neural network and input them into the convolutional neural network;

[0016] SB22, training: Use a public dataset containing normal sessions and labeled attack samples as the training set of the convolutional neural network for training, and then convert the output of the convolutional neural network into a probability distribution;

[0017] Among them, the expression of the classification result is:

[0018]

[0019] In the formula: i, j represent the numbers of two categories, n represents the number of categories, i = 1, 2,..., n, j = 1, 2,..., n, σ(z) i represents the probability of the i-th category, represents the natural constant, z i represents the score of the i-th category, z j represents the score of the j-th category.

[0020] As a preferred technical solution, in step S1, the network session behavior data is sharded and stored in the cache, and the cache adopts a circular queue structure; parallel computing is used when extracting session features.

[0021] A method for optimizing network session firewall policies includes the above-mentioned method for analyzing network session behaviors, and further includes the following steps:

[0022] S3, policy generation: Generate firewall policies according to the classification results obtained in step S2;

[0023] S4, policy optimization: Use the Q-learning algorithm to implement firewall policy optimization.

[0024] As a preferred technical solution, in step S3, the types of firewall policies include one or more of allowing, blocking, and traffic limiting; the firewall policies are sorted by priority based on the threat level, and the features affecting the sorting include one or more of attack severity, historical occurrence frequency, and affected network area.

[0025] As a preferred technical solution, in step S4, the Q-value update formula of the Q-learning algorithm is:

[0026] Among them, s represents the current state, a represents the action taken in the current state s, Q(s, a) represents the expected cumulative reward that can be obtained after taking action a from state s, α represents the learning rate, R represents the immediate reward obtained after taking action a in the current state s, γ represents the discount factor, s' represents the next state, a′ represents the possible action that may be taken in the next state s', and Q(s′, a′) represents the Q value of the possible action a' in the next state s'. Represents the maximum Q value among all possible actions a' in the next state s'.

[0027] As a preferred technical solution, in step S4, in the Q-learning algorithm, R = W1 × (1 - False alarm rate) + W2 × Attack interception rate - W3 × Delay penalty;

[0028] Among them, W1 represents the weight parameter for balancing and adjusting the priority of false alarms, W2 represents the weight parameter for adjusting the attack interception ability, and W3 represents the weight parameter for restricting the level of delay.

[0029] As a preferred technical solution, it further includes the following steps:

[0030] S5, Log recording and feedback: Real-time record network session data, and use the network session data in the log to execute step S2.

[0031] A network session firewall policy optimization system for implementing the described network session firewall policy optimization method, including the following modules connected in sequence:

[0032] S1, Session monitoring module: Used to execute step S1;

[0033] S2, Behavior analysis module: Used to execute step S2;

[0034] S3, Policy generation module: Used to execute step S3;

[0035] S4, Policy optimization module: Used to execute step S4.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] (1) Dynamic adaptability: Through reinforcement learning and real-time feedback mechanism, it can automatically adapt to new network attacks and dynamically adjust the firewall policy;

[0038] (2) High-precision detection: The behavior analysis module combining multiple models can efficiently identify normal and abnormal sessions, significantly reducing the false alarm rate and the missed alarm rate;

[0039] (3) Self-optimization ability: It can automatically optimize and deploy strategies, reduce manual intervention, and improve the intelligence level of the firewall;

[0040] (4) Data security: It adopts encrypted storage and strict access control mechanisms to ensure the privacy and security of the system and data. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is the overall flowchart of a network session firewall policy optimization method according to the present invention;

[0042] Figure 2 is the flowchart of session detection;

[0043] Figure 3 is the flowchart of the behavior analysis module;

[0044] Figure 4 is the flowchart of policy generation;

[0045] Figure 5 is the flowchart of policy optimization;

[0046] Figure 6 is the flowchart of log recording and behavior feedback;

[0047] Figure 7 is an example diagram of the convolutional neural network structure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0049] Embodiment 1

[0050] As Figures 1 to 7 shown, the present invention discloses a system and method for network session behavior analysis and automated firewall policy optimization by combining machine learning and reinforcement learning technologies. Specifically, the present invention automatically generates and optimizes firewall policies by real-time monitoring and analyzing network traffic, combining multiple machine learning models and reinforcement learning algorithms, thereby improving the intelligence level of the firewall and optimizing the deployment of security policies. The present invention has the capabilities of adaptive learning and dynamic optimization, can respond to new network attacks in real time, and significantly improves the response speed and protection accuracy of the firewall.

[0051] The system structure of the present invention includes the following modules:

[0052] 1. Session monitoring module;

[0053] 2. Behavior analysis module;

[0054] 3. Policy generation module;

[0055] 4. Policy Optimization Module;

[0056] 5. Logging and Feedback Module.

[0057] More specifically:

[0058] 1. Session Monitoring Module

[0059] The session monitoring module is responsible for capturing and parsing network packets in real time. It achieves efficient packet capture through DPDK technology and supports real-time monitoring of network traffic from gigabit to ten-gigabit levels. To ensure the efficient collection and processing of network traffic, the following technologies are used for optimization:

[0060] Traffic Sharding and Caching Mechanism: To avoid packet loss in high-traffic scenarios, packets are sharded and stored in the cache. The cache adopts a circular queue structure to ensure data real-time.

[0061] Feature Extraction: The module extracts key features of each session, including source / destination IP, port, protocol type, session duration, data traffic, etc., as the input for subsequent analysis. For real-time requirements, feature extraction uses a parallel computing framework to ensure real-time processing of high-concurrency traffic.

[0062] 2. Behavior Analysis Module

[0063] The behavior analysis module performs network traffic analysis through a multi-model combination strategy, combining Random Forest (RF) and Convolutional Neural Network (CNN). The main processes include:

[0064] Initial Screening: Use the random forest model to perform preliminary classification on the captured session features (features include: traffic size, transmission rate, request frequency, IP address, port number, protocol type, session duration, etc.).

[0065] The specific steps of the random forest include:

[0066] 1. Data Preprocessing: including collecting features, data cleaning, and feature selection;

[0067] 2. Labeling Data: Labeled as normal session or abnormal session;

[0068] 3. Training the Random Forest Model: Model selection, using the training set (labeled data) to train the model. During the training process, the random forest model classifies session features by constructing multiple decision trees. Each tree independently judges whether a session is abnormal, and finally determines the classification result through a voting method;

[0069] 4. Using the random forest model to analyze network traffic.

[0070] In-depth analysis: For the recognition of complex behaviors, a convolutional neural network (CNN) is used for in-depth analysis. The advantage of CNN in processing traffic pattern recognition lies in its ability to automatically learn complex network attack patterns, which is particularly suitable for dealing with attacks with long-duration and complex behaviors (such as DDoS, scanning, etc.).

[0071] The specific steps of CNN include:

[0072] Data input: Convert the network traffic characteristics into an input format acceptable to the CNN model. For example, convert the traffic time series into a matrix representation and input it into the CNN model.

[0073] Model structure: The CNN network includes two convolutional layers, two max-pooling layers, and one fully connected layer. Finally, the classification result is output through the softmax output layer.

[0074] Figure 7 One CNN model structure is listed. The implementation of the present invention is not limited to the specific structure listed here. Figure 7 The symbol explanations are as follows:

[0075] Feature extraction layer (C layer):

[0076] Another name for the feature extraction layer: The feature extraction layer is also called the convolutional layer.

[0077] Function of the feature extraction layer: Perform convolution processing on the network traffic data to extract traffic features, such as protocol patterns, traffic behavior patterns, or abnormal communication patterns.

[0078] Input data form of the feature extraction layer: The network traffic data can be represented as a two-dimensional matrix or a time series segment. For example, the rows represent time or the packet order, and the columns represent different features of the traffic, such as byte distribution, protocol fields, traffic direction, etc.

[0079] Neuron connection of the feature extraction layer: The input of each neuron is connected to the local receptive field of the previous layer to extract the local feature pattern.

[0080] Feature mapping layer (S layer):

[0081] Another name for the feature mapping layer: The feature mapping layer is also called the pooling layer, downsampling layer, or computational layer.

[0082] Function of the feature mapping layer: Reduce the dimension of the traffic feature map output by the convolutional layer through pooling operations, compress the data scale, extract more representative key features, and at the same time enhance the translational invariance and robustness of the features.

[0083] Composition of the feature mapping layer: Each feature mapping layer consists of multiple feature maps. Each feature map is a two-dimensional plane, and all neurons on the plane share weights. For example: the statistical distribution of traffic feature patterns.

[0084] Activation function of the feature mapping layer: A non-linear activation function is used to enhance the model's expressive power. The activation function makes the feature map more adaptable to complex traffic pattern analysis tasks.

[0085] Function of the pooling layer:

[0086] Dimensionality reduction: The pooling layer compresses the data size and reduces the number of parameters through max pooling or average pooling, thereby alleviating overfitting.

[0087] Feature enhancement: Improve the model's adaptability to the translation and scale invariance of local patterns.

[0088] Precautions for the pooling layer:

[0089] Pooling size: Select an appropriate pooling window size to ensure dimensionality reduction while retaining key features as much as possible.

[0090] Number of pooling layers: Too many pooling layers may cause the loss of important traffic features. Therefore, it is necessary to balance the relationship between feature retention and dimensionality reduction.

[0091] Relationship between the feature extraction layer and the feature mapping layer:

[0092] Connection method:

[0093] In network traffic analysis, each feature extraction layer is usually followed by a pooling layer to gradually extract key features and reduce dimensionality.

[0094] Advantages of the connection method:

[0095] It can capture the features of local traffic patterns. Tolerate a certain range of traffic data distortion or inconsistency, such as timing jitter or noise interference.

[0096] Common point: Each layer consists of multiple feature planes, and each plane is a feature map of traffic patterns. This structured extraction method can effectively retain key features and eliminate redundant information.

[0097] Operation process of the convolutional neural network in network traffic analysis:

[0098] Convolution process: The input network traffic feature data is convolved with multiple convolutional kernels to extract local feature patterns and generate feature maps.

[0099] Pooling process: Perform pooling operations on the feature maps output by each convolutional layer to extract the core information of the features while reducing dimensionality.

[0100] Loop iteration: Use the output of the previous pooling layer as the input of the next convolutional layer, and continuously perform feature extraction and dimensionality reduction.

[0101] Feature map rasterization: Flatten the feature map of the last layer into a vector representation as the input of the classifier.

[0102] Classification process: Use a fully connected neural network to further analyze the feature vector, complete tasks such as traffic classification and anomaly detection, and obtain the final output.

[0103] Fully connected neural network NN: After convolutional and pooling operations, the network usually flattens the features and further processes them through fully connected layers. Each neuron in the fully connected layer is connected to all neurons in the previous layer, and it is usually used to map the extracted high-level features to the final output space (such as classification labels).

[0104] The Softmax output layer is a common output layer in deep learning, usually used for multi-classification tasks. Its main function is to convert the output of the network (logits) into a probability distribution, where the probability value of each class is between 0 and 1, and the sum of the probability values of all classes is 1.

[0105] Among them, the expression of the Softmax function is:

[0106]

[0107] In the formula: i and j represent the numbers of two classes, n represents the number of classes, i = 1, 2,..., n, j = 1, 2,..., n, and σ(z) i represents the probability of the i-th class, represents the natural constant, and z i represents the score of the i-th class, and z j represents the score of the j-th class, and are to exponentiate the original scores as an intermediate step in Softmax.

[0108] Training method:

[0109] Use a publicly available dataset (such as the CICIDS2017 dataset) containing normal sessions and labeled attack samples. Manually label different attack types (such as DDoS, scanning attacks, etc.). The training set accounts for 70%, and the test set accounts for 30%.

[0110] Model integration: Based on the outputs of multiple models, use the soft voting method to fuse the model results and further improve the classification accuracy. To optimize the effect of multi-model integration, a weighted voting strategy is adopted, and weights are assigned to each model according to its historical performance.

[0111] 3. Policy Generation Module

[0112] The policy generation module automatically generates firewall policies based on the results output by the behavior analysis module. This module realizes efficient and flexible policy generation in the following ways:

[0113] Rule Template Generation: According to the results of behavior analysis, firewall policy templates are automatically generated, and the policy types include allow, block, traffic limiting, etc. Each policy template can be flexibly configured according to characteristics such as source / destination IP, port, protocol, session duration, etc.

[0114] Policy Priority Sorting: This module supports sorting the priorities of firewall policies based on threat levels.

[0115] Feature Definition: The features affecting sorting include attack severity, historical occurrence frequency, affected network areas, etc.

[0116] Algorithm Description: A decision tree is constructed based on information gain, nodes are divided, and priorities are output.

[0117] Example: For an attack frequency > 10 times / hour, the decision tree marks its severity as "high", and the priority is increased to the top 10%.

[0118] API Interface Support: The policy generation module communicates with the firewall device through an open RESTful API interface, supports real-time policy issuance and updates as needed.

[0119] 4. Policy Optimization Module

[0120] The policy optimization module uses the Q-learning reinforcement learning algorithm to realize dynamic optimization of firewall policies, including the following steps:

[0121] Definition of State Space and Action Space:

[0122] State Space: Session characteristics (frequency, protocol, duration, etc.).

[0123] Action Space: Policy adjustment (allow, block, traffic limiting, etc.).

[0124] Reward Mechanism: R = W1×(1 - False Alarm Rate) + W2×Attack Interception Rate - W3×Delay Penalty

[0125] In the above formula:

[0126] R represents the immediate reward, W1 represents the weight parameter used to balance the priority of adjusting false alarms, W2 represents the weight parameter used to adjust the attack interception ability, and W3 represents the weight parameter used to limit the level of delay.

[0127] By adjusting the weights, the system can be flexibly made to pay more attention to reducing false positives (increasing W1) or improving the interception rate (increasing W2), or even taking into account performance (appropriately reducing W3).

[0128] Q-value (reward) update formula based on reinforcement learning:

[0129]

[0130] Parameter settings:

[0131] Learning rate α: Set relatively high initially (e.g., 0.1) and gradually decrease later.

[0132] Discount factor γ: Set the long-term reward to 0.9 - 0.99 and the quick response to 0.5 - 0.8.

[0133] Update process: Dynamically adjust the Q-value after the policy is executed. For example, the corresponding reward increases after the false positive rate is reduced.

[0134] The Q-value update formula is the core formula in reinforcement learning (especially Q-learning), used to iteratively update the values in the Q-table, thereby helping the agent gradually learn the optimal policy. The following is a detailed explanation of the formula and its symbols:

[0135] Meaning of symbols in the Q-value update formula:

[0136] Q(s,a): The expected cumulative Q-value of taking action a in the current state s, representing the expected cumulative reward that can be obtained after taking action a from state s under the current policy.

[0137] α: Learning rate, with a value range of 0 ≤ α ≤ 1.

[0138] Function: Control the balance between the old and new Q-values. The larger the value, the higher the degree of emphasis on new information; the smaller the value, the more dependent on past experience.

[0139] Typical value: Usually set to 0.1 or gradually decay to improve learning stability.

[0140] R: Immediate reward, representing the immediate reward obtained after taking action a in the current state s.

[0141] Source: Defined by the system, usually designed according to the task objective reward mechanism.

[0142] Example: In network firewall optimization, R may be calculated based on the false positive rate, attack interception rate, latency, etc.

[0143] γ: Discount factor, with a value range of 0 ≤ γ ≤ 10.

[0144] Function: Weigh the importance of short-term rewards and long-term benefits.

[0145] When γ = 0: Only consider the current reward (myopic).

[0146] When γ → 1: Focus on long-term rewards.

[0147] Typical value: Generally set to 0.9 - 0.99.

[0148]

[0149] Meaning: The maximum Q value among all possible actions a' in the next state s'.

[0150] Function: Represents the expected return of taking the optimal action in the next state s', and is used to estimate future rewards.

[0151] s and s':

[0152] s: The current state, representing the state of the system before performing an action.

[0153] s': The next state, representing the state the system enters after performing an action.

[0154] a and a′:

[0155] a: The current action, representing the action taken by the agent in the current state s.

[0156] a': The next action, representing the possible action in the next state s'.

[0157] The core idea of the Q-value update formula is to revise the expectation of future rewards based on current experience. By introducing parameters such as the learning rate and discount factor, the formula can respond quickly to changes while ensuring long-term stability.

[0158] 5. Logging and Feedback Mechanism

[0159] The logging and feedback mechanism provides a basis for the system to continuously learn, ensuring the effectiveness and accuracy of firewall policies during long-term operation:

[0160] Real-time logging and feedback: The system real-time logs data such as session behaviors, policy execution effects, false alarm rates, and missed alarm rates, and feeds them back to the behavior analysis module and policy optimization module through log analysis. Through this closed-loop feedback mechanism, the system can continuously improve itself.

[0161] Incremental learning and retraining: The data in the logs will be regularly used for the retraining of the behavior analysis module to enhance the adaptability and accuracy of the model. The retraining process adopts incremental learning methods and only trains on new data to ensure the efficiency of the training process.

[0162] Encryption and access control: All recorded data is stored encrypted, and strict access control mechanisms are set up to ensure data privacy and security.

[0163] The method steps are as follows:

[0164] 1. Session monitoring:

[0165] Real-time capture of network traffic through DPDK technology, extraction of key features, and preprocessing operations such as denoising and normalization.

[0166] 2. Behavior analysis:

[0167] Through a multi-model combination strategy, use random forest to classify the preliminary data, use CNN to deeply analyze complex attack patterns, and integrate the model results through the soft voting method.

[0168] 3. Policy generation:

[0169] Automatically generate firewall policy templates according to the behavior analysis results, and issue policies through the API interface, supporting flexible rule priority configuration.

[0170] 4. Policy optimization:

[0171] Based on the Q-learning reinforcement learning algorithm for policy optimization, use the ε-greedy policy to balance exploration and exploitation, and adaptively adjust the firewall policy through the reward mechanism.

[0172] 5. Log recording and feedback:

[0173] Provide a new dataset for the behavior analysis module for incremental training to ensure that the system continuously adapts to new attack patterns and network environments.

[0174] As described above, the present invention can be preferably implemented.

[0175] All features disclosed in all embodiments in this specification, or all steps in the methods or processes implicitly disclosed, except for mutually exclusive features and / or steps, can be combined and / or extended, replaced in any way.

[0176] The above is only a preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Based on the technical essence of the present invention, any simple modifications, equivalent replacements, and improvements made to the above embodiments within the spirit and principles of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A network conversation behavior analysis method, characterized in that: The following steps are involved: S1, session monitoring: capturing and parsing network session behavior data and extracting session features; S2, Behavior Analysis: Classify session features and analyze whether network session behavior is abnormal.

2. A network conversation behavior analysis method according to claim 1, characterized in that: In step S2, the session features are classified using a random forest algorithm, including the following steps: SA21, data preprocessing: preprocessing session features; SA22, labeled data: labeling session features as normal sessions or abnormal sessions; SA23, training random forest model: Use the labeled data as the training set of the random forest model to classify session features by building multiple decision trees. Each tree independently determines whether a session is abnormal or not, and finally decides the classification result by voting. SA24, Analyzing Network Traffic Using Random Forest Model.

3. A network conversation behavior analysis method according to claim 1, characterized in that: In step S2, a convolutional neural network is used to classify the session features, including the following steps: SB21, data input: convert the session features into an input format acceptable to the convolutional neural network and input them into the convolutional neural network; SB22, training: Use a public dataset containing normal conversations and labeled attack samples as the training set of the convolutional neural network, and then convert the output of the convolutional neural network into a probability distribution; The expression of the classification result is: Where: i and j represent the numbers of two categories, n represents the number of categories, i = 1, 2, ..., n, j = 1, 2, ..., n, σ(z) i represents the probability of the i-th category, e represents a natural constant, z i represents the score of the i-th category, z j represents the score of the j-th category.

4. A network conversation behavior analysis method according to any one of claims 1 to 3, characterized in that: In step S1, the network session behavior data is segmented and stored in a cache, which adopts a circular queue structure; parallel computing is used to extract session features.

5. A network session firewall policy optimization method, characterized in that: A network conversation behavior analysis method according to any one of claims 1 to 4, further comprising the following steps: S3, policy generation: Generate a firewall policy based on the classification results obtained in step S2; S4, policy optimization: Use Q learning algorithm to achieve firewall policy optimization.

6. A network session firewall policy optimization method according to claim 5, characterized in that: In step S3, the firewall policy type includes one or more of release, blocking, and flow control; the firewall policies are prioritized based on the threat level, and the characteristics that affect the ranking include one or more of the severity of the attack, the historical frequency of occurrence, and the affected network area.

7. A network session firewall policy optimization method according to claim 6, characterized in that: In step S4, the Q value update formula of the Q learning algorithm is: Where s represents the current state, a represents the action taken in the current state s, Q(s,a) represents the expected cumulative reward after taking action a from state s, α represents the learning rate, R represents the immediate reward after taking action a in the current state s, γ represents the discount factor, s' represents the next state, a' represents the possible action in the next state s', Q(s',a') represents the Q value of possible action a' in the next state s', It represents the maximum Q value among all possible actions a' in the next state s'.

8. A network session firewall policy optimization method according to claim 7, characterized in that: In step S4, in the Q learning algorithm, R = W1×(1-false alarm rate)+W2×attack interception rate-W3×delay penalty; Among them, W1 represents a weight parameter for balancing and adjusting the priority of false alarms, W2 represents a weight parameter for adjusting the attack interception capability, and W3 represents a weight parameter for limiting the delay.

9. A network session firewall policy optimization method according to any one of claims 5 to 8, characterized in that: The following steps are also included: S5, log recording and feedback: record the network session data in real time, and use the network session data in the log to execute step S2.

10. A network session firewall policy optimization system, characterized in that: A network session firewall policy optimization method for implementing any one of claims 5 to 9, comprising the following modules connected in sequence: S1, session monitoring module: used to execute step S1; S2, behavior analysis module: used to execute step S2; S3, strategy generation module: used to execute step S3; S4, strategy optimization module: used to execute step S4.