Method and system for generating and evaluating security policy of network / security equipment based on large language model
By introducing LLM collaboration mechanism, multi-dimensional policy evaluation and dynamic adaptation mechanism into the network security protection system, the shortcomings of existing systems in policy generation and evaluation are solved, the effectiveness and reliability of equipment security policies are improved, and faster and more accurate network security response is achieved.
Patent Information
- Application Number
- CN202510462450.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing network security protection system based on large language models lacks adaptability to dynamic changes in the network environment when generating defense strategies and responds slowly; the strategy evaluation mechanism mostly relies on a single certain quantity indicator and lacks multi-dimensional comprehensive evaluation; the decision logic of LLM in complex network environments needs to be optimized to improve its accuracy and reliability in practical applications.
By introducing LLM collaboration mechanism, multi-dimensional strategy evaluation mechanism and feedback-based dynamic adaptation mechanism, we will improve the effectiveness and reliability of equipment security policies. Specific methods include: fine-tuning threat analysis LLM and policy generation LLM, outputting threat information quadruples through threat analysis LLM, policy generation LLM generates device security policies, and evaluates policy accuracy, effectiveness and adaptability by evaluating the Agent.
It improves the response speed and adaptability of equipment security strategies, realizes multi-dimensional strategy evaluation, improves the decision-making accuracy and reliability of LLM in complex network environments, and solves the problems of insufficient timeliness of threat handling and high operation and maintenance costs.
Smart Images

Figure CN119996085A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of traffic anomaly detection, and in particular to a method and system for generating and evaluating network / security device security policies based on a large language model. Background Art
[0002] In network threat or daily operation and maintenance scenarios, security operation and maintenance personnel often face the problems of insufficient threat handling timeliness and high operation and maintenance costs, especially for internal networks with high control, strong isolation, and high security levels. In recent years, researchers have begun to explore the application of artificial intelligence technology to network security protection, especially security response and operation and maintenance systems based on machine learning and deep learning. The above-mentioned traditional security protection methods still have limitations in dealing with complex attack patterns and changing attack behaviors, especially in policy generation and dynamic adaptation. Moreover, security policy generation mostly relies on manual work, which usually has the problems of insufficient handling timeliness and high operation and maintenance costs. The emergence of large language models (LLMs) provides new ideas for solving these problems. Its powerful semantic understanding and generation capabilities enable it to extract key information from complex intranet attack logs, including node information and attack behavior information, and generate targeted internal / dedicated device defense strategies.
[0003] Despite this, the existing LLM-based network security protection system still has some problems that need to be solved. First, LLM often lacks adaptability to dynamic changes in the network environment when generating defense strategies, resulting in a slow response when facing new attacks. Second, the existing strategy evaluation mechanism mostly relies on a single quantitative indicator and lacks a multi-dimensional comprehensive evaluation of the strategy execution effect. In addition, the decision-making logic of LLM in complex network environments still needs to be further optimized to improve its accuracy and reliability in practical applications. Summary of the invention
[0004] In view of this, the present application provides a network / security device security policy generation and evaluation method and system based on a large language model. Through the LLM collaboration mechanism, multi-dimensional policy evaluation mechanism and feedback-based dynamic adaptive mechanism, it aims to improve the effectiveness and reliability of device security policies, promote the intelligent level of network security protection, and solve the problems of insufficient timeliness of threat disposal and high operation and maintenance costs.
[0005] The present application discloses a network / security device security policy generation and evaluation method based on a large language model, which includes: Step 1: Use the threat analysis dataset and strategy generation dataset to fine-tune the base model to obtain the threat analysis LLM and strategy generation LLM; Step 2: Threat analysis LLM performs threat analysis on the abnormal logs output by the log collection module and outputs a threat information quadruple; Step 3: The policy generation LLM generates the device security policy by calling the knowledge base Agent based on the threat information quadruple output by the threat assessment LLM; Step 4: The policy generation LLM calls the evaluation agent to evaluate the device security policy from three aspects: policy correctness evaluation, policy effectiveness evaluation, and policy adaptability evaluation.
[0006] Furthermore, the step 1 comprises: According to the application service scenarios of the target internal / private network, multiple attack paths under various application scenarios are designed to launch attacks on the target server or terminal. The relevant log data is collected through the abnormal log collection module, and different types of log data from different devices are sorted out. Finally, a threat analysis data set is constructed. The fields of the threat analysis data set are: {time, log, attacker IP, attacker port, victim IP, victim port, attack stage, threat level}; Design threat / abnormal scenarios covering the target network according to the target network environment and asset conditions. Obtain log data by launching multiple rounds of attacks. Combined with historical abnormal data, identify the attack stage through threat analysis LLM and locate the problem node. Combined with the relevant information provided by the knowledge base agent, establish defense strategy logic for the attack path and logic, perform annotation, and finally obtain the threat strategy data set. The fields of the threat strategy data set are: {threat event, strategy}; related information includes disposal nodes; A base model is selected, and LoRA+ based on human feedback reinforcement learning is used as the fine-tuning technology. The threat analysis dataset and strategy generation dataset are used as fine-tuning datasets to fine-tune the base model and obtain the threat analysis LLM and strategy generation LLM.
[0007] Furthermore, the abnormal log collection module collects relevant log data and sorts out different types of log data from different devices, including: Download the filebeat installation package, edit the filebeat configuration file, and run the . / filebeat script on the target host where the filebeat configuration file is deployed; the filebeat configuration file includes the type and path of the logs collected locally and the information pushed to Logstash; Download the Logstash installation package, edit the Logstash configuration file, run the . / Logstash script, configure different Logstash server clusters for different types of logs on different devices, and start the Logstash service on the server through the Logstash configuration file; the Logstash configuration file includes the log source port, log content filtering rules, and the configuration of pushing logs to ES-Kibana; Store various log entries from different devices as a search engine; integrate various log data from different devices, count and visualize logs; log in to each host to view log indexes and specific data.
[0008] Furthermore, the step 2 comprises: The threat analysis LLM performs threat analysis on the abnormal logs received from the log collection module and outputs a threat information quadruple, which is: {attack stage, threat source node, victim node, security event}. The threat quadruple data is used as the input of the policy generation LLM, which assists the policy generation LLM in generating the device security policy under the current threat scenario.
[0009] Furthermore, the step 3 comprises: Fix the errors in the process of generating threat information quadruple, encapsulate the relevant information into tools in advance, and use the langchain framework to call the policy generation LLM to call the tool, call the knowledge base Agent for knowledge retrieval, and collect the response of each step, and continue to determine whether the tool needs to be used again. Repeat this iterative cycle until the answer is given. The answer includes network topology knowledge, disposal node device information and security event information; the relevant information includes the target network topology information and device information.
[0010] Furthermore, the method for searching for the processing node device information is: if there is a threat source node, then query the nearest disposable node device information that can process the threat source node; if there is no threat source node, then query the nearest disposable node device information that can process the victim node. The policy generates LLM to output security policy binary information based on the threat analysis LLM analysis result. The security policy binary information is: {threat event, security policy}.
[0011] Furthermore, the construction process of the knowledge base includes: Construct a node attribute table and a connection relationship table, import the node attribute table and the connection relationship table into the graph database to form a network topology knowledge graph; the fields of the node attribute table are: {node serial number, host name, node IP address, node port, node operating system, network segment, network area, log type, node grouping, devices that can handle the node, node device type}; the fields of the connection relationship table are: {starting point serial number, end point serial number, connection type}.
[0012] Furthermore, the step 4 comprises: Policy correctness assessment: The device security policy output by the policy generation LLM is retrieved and compared with the reference policy set, and the confidence ranking calculation of the semantic similarity is performed. If the similarity between the output device security policy and any reference policy in the reference policy set exceeds the preset threshold, the device security policy output by the policy generation LLM is considered correct; then the command configuration of the corresponding network / security device is matched, and the policy is issued. If the command configuration of the corresponding network / security device is not matched, the result is fed back to the policy generation LLM; the reference policy set refers to the set of all possible security policies for devices in the target network when the target network faces a threat scenario.
[0013] Furthermore, the step 4 also includes: Policy effectiveness evaluation: The policy effectiveness evaluation is completed by the policy effectiveness module. The traffic collection module pushes the abnormal traffic data it collects to the simulated attack module in the evaluation agent. The simulated attack module replays the attack in the isolated area to simulate the launch of the attack. In the simulated network, the device security policy generated by the policy LLM output is sent to the virtual network / security device terminal for policy execution. The policy effectiveness module monitors the host log to observe the effect of policy implementation. If the attack fails after a period of execution of the device security policy, the output device security policy is judged to be valid, and the time from the device security policy being sent to the defense taking effect is recorded; otherwise, the device security policy is judged to be invalid, and the result is fed back to the threat assessment model.
[0014] Furthermore, the step 4 also includes: Policy adaptability assessment: Policy adaptability assessment is completed by the policy adaptability module in combination with the human-computer interaction mechanism. The device security policy is sent to the virtual network / security device terminal in the simulated network for policy execution. The policy adaptability module monitors the traffic fluctuations of the simulated network and the impact on the important business flows of the device security policy intranet. According to the preset network traffic fluctuation threshold, the policy adaptability results are output and pushed to network security experts for comprehensive analysis. If it fails, the policy adaptability results are fed back to the policy generation LLM.
[0015] The present application also discloses a network / security device security policy generation and evaluation system based on a large language model, which implements the above-mentioned network / security device security policy generation and evaluation method based on a large language model, including a log collection module, a traffic collection module, a threat analysis LLM, a policy generation LLM, a knowledge base Agent and an evaluation Agent; the evaluation Agent includes a policy validity module, a policy correctness module and a policy adaptability module; the log collection module is connected to the policy generation LLM through the threat analysis LLM; the policy generation LLM is connected to the knowledge base Agent and the evaluation Agent respectively; the policy validity module is connected to the threat analysis LLM, the simulated attack module in the policy validity module is connected to the traffic collection module, the policy correctness module and the policy adaptability module are connected to the policy generation LLM respectively, the policy validity module is used for policy validity evaluation; the policy adaptability module is used for policy adaptability evaluation; the policy correctness module is used for policy correctness evaluation.
[0016] Due to the adoption of the above technical solution, the present application has the following advantages: 1. For the target network, the LLM collaboration mechanism is introduced to realize information sharing and strategy optimization among different large models, optimize the training of a pair of large model groups with high-quality strategy generation capabilities, and based on the multi-agent calling idea, improve the decision-making efficiency and reliability of the overall system.
[0017] 2. Based on multi-dimensional and multi-level evaluation indicators, introduce feedback loops to dynamically adjust evaluations and “human-computer interaction” mechanisms to adapt to the ever-changing network environment and improve the effectiveness and reliability of internal / dedicated security policies. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0019] Figure 1 A flowchart of a network / security device security policy generation and evaluation method based on a large language model according to an embodiment of the present application; Figure 2 This is a block diagram of a network / security device security policy generation and evaluation system based on a large language model according to an embodiment of the present application. DETAILED DESCRIPTION
[0020] The present application is further described in conjunction with the accompanying drawings and embodiments, and the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field should fall within the scope of protection of the embodiments of the present application.
[0021] This application mainly aims at the problem of high false negative rate of unknown threats in internal / private networks, and proposes a network / security device security policy generation and evaluation method and system based on a large language model. Specifically, it solves the following problems: (1) How to use the semantic understanding capabilities of LLM to parse anomalies and determine their related attack information, and then generate preliminary defense strategies that can be sent to internal / dedicated devices.
[0022] (2) How to construct a strategy evaluation space and design a multi-dimensional comprehensive evaluation method to comprehensively evaluate the execution effect of the LLM generated strategy.
[0023] (3) How to improve the adaptability of LLM in complex scenarios and the quality and reliability of defense strategy generation.
[0024] See also Figure 1 , a network / security device security policy generation and evaluation method based on a large language model is proposed in an embodiment of the present application, which includes: Step 1: Use the threat analysis dataset and strategy generation dataset to fine-tune the base model to obtain the threat analysis LLM and strategy generation LLM; Abnormal data collection is a key network security task, which aims to identify and prevent potential security threats by recording and analyzing abnormal behaviors in the network. Unlike the Internet environment, the internal / dedicated environment has special functions and deploys specific application services, so the attacks it faces are different from those in the Internet environment. In addition, most internal / dedicated environments have higher security requirements and special functional area divisions, which ensure the reliable operation of the intranet while also making the attacks against the intranet special.
[0025] This application designs the attack scenarios that the target network environment may face based on the specific environment and service applications, sorts out possible attack paths, launches simulated attacks, and collects relevant log data. At the same time, it sorts out different types of log data from different devices, and screens appropriate data features and fields according to the needs of subsequent analysis. In addition, in order to unify the standards for subsequent threat analysis and response disposal, it is planned to divide the data window. This application sets 10 seconds as the data window, and divides and annotates the attack data according to the attack stages of Adversarial Tactics, Techniques, and Common Knowledge (ATT&CK) to finally construct a threat analysis data set, as follows: 1) Attack scenario and path design Specifically, this application designs multiple attack paths in various application scenarios based on the application service scenarios of the target internal / private network, and uses multiple attack tools and techniques to attack the target server or terminal. The data set fields are set as: {time, log, attacker IP, attacker port, victim IP, victim port, attack stage, threat level}. In order to clearly describe the construction of the data set, one of the several attack paths against the intranet File Transfer Protocol (FTP) server is used as an example to explain. The attack scenario is that the attacker infiltrates the intranet through a virtual private network (VPN) and steals or destroys data on the intranet FTP server. The attack path is as follows: (1) The attacker steals the VPN credentials of the remote user and connects to the DMZ area through the VPN server proxy; (2) The attacker launches a scanning attack (passive attack) in the DMZ area and finds that port 21 is open, judging it as an FTP server; (3) Based on the backdoor of vsftpd 2.3.4 version, the attacker launches a "smiley face" attack on the FTP server and finally obtains the root permission of the FTP server; (4) Weak password blasting is used to obtain permissions; (5) Cobalt Strike is used to generate an exe Trojan horse, disguised as a normal file, and put it to the target terminal through the FTP server; (6) Connect to the Trojan horse and rebound the shell to communicate remotely with the remote attacker; (7) Use the vulnerability script in MSF to elevate the target terminal to obtain root permissions; (8) The attacker uses the root permissions of the target terminal to steal private information or damage data / system. According to the above attack path, the corresponding attack stages and log descriptions are divided as follows: (1) Reconnaissance: VPN connection log, from the router's vpnlog, represents the attacker accessing the intranet through the proxy; (2) Reconnaissance: Port scanning log, from the FTP server's ufwl.og, represents the attacker scanning the server's open ports to obtain the application service list; (3) Credential access: FTP blasting log, from the FTP server's vsftpd.log, represents the attacker blasting the server to obtain FTP connection permissions; (4) Execution: FTP file log tampering log, from the FTP server's vsftpd.log, represents the attacker's process of uploading malicious files; (5) Command and control: User-side control log, from the user terminal's syslog, represents the user executing the malicious file; (6) Leakage: User rebound Sell log, from the user terminal's syslog, represents the user executing the malicious file and then getting a virus; (7) Termination: VPN disconnection log, from the router's vpnlog, represents the attacker disconnecting the proxy to end the attack.
[0026] Similarly, based on multiple paths of multiple preset attack methods, the "threat analysis" data set is constructed. Taking the three attack scenarios of FTP attack, MySQL database attack and intranet forum attack as examples, 100 rounds of FTP attack were launched, generating more than 104,200 abnormal logs, 110 rounds of MySQL database attack were launched, generating more than 133,800 abnormal logs, and 700 rounds of intranet forum injection attack were launched, generating more than 191,400 abnormal logs.
[0027] 2) Abnormal log collection Based on the aforementioned attack scenarios and paths, relevant logs need to be collected, and abnormal log collection is completed by the log collection module. After extensive preliminary research and testing, for abnormal log data, this application uses Elasticsearch in the Elastic Stack technology stack to provide persistence services, uses Filebeat-Logstash for log collection and preprocessing, and uses Kibana for visualization. The overall log collection process is as follows: ① Download the filebeat installation package, edit the filebeat configuration file, mainly including the type and path of the local log collection and the information pushed to Logstash, and run the . / filebeat script on the target host where filebeat is deployed; ② Download the Logstash installation package, edit the Logstash configuration file, mainly including the log source port, log content filtering rules, and the configuration of pushing logs to ES-Kibana, run the . / Logstash script, configure different Logstash server clusters for different types of logs, and start the Logstash service on the server through the above configuration; ③Store various log entries as a search engine; ④Integrate various log data, count and visualize logs; ⑤Log in to each host to view the log index and specific data.
[0028] Regarding the construction of the strategy generation dataset: Since the disposal strategy needs to be aligned with the original input log data in time to achieve real-time analysis and disposal, therefore, consistent with the previous threat assessment data set, this application sets a 10-second data window, aggregates the attack phase and problem node output of the previous threat assessment LLM, the network topology information provided by the knowledge base agent, the importance of the equipment, the area where the problem node is located, and the disposal node information associated with the problem node, designs the disposal strategy logic according to the actual online scenario, and manually annotates the data set.
[0029] It is worth noting that this application is mainly aimed at internal / private network environments, so the threat / abnormal scenarios covering the target network can be designed according to the target network environment and asset conditions. Then, by launching multiple rounds of attacks, log data is obtained, combined with historical abnormal data, the attack stage is determined through the previous threat analysis LLM, and the problem node is located. Combined with the disposal nodes and other related information provided by the knowledge base Agent, the defense strategy logic is established for the attack path and logic, and manual labeling is carried out.
[0030] Specifically, based on the attack scenarios and paths when constructing the previous threat assessment dataset, a threat strategy dataset is constructed, and the dataset fields are: {threat event, strategy}. To clearly describe the construction of the dataset, we will take the aforementioned attack on the intranet FTP server as an example to explain the construction logic of the dataset as follows: (1) Threat event: At ROUTER-1, it is determined that there is an attack from 10.1.40.X during the reconnaissance phase. Strategy: Handle node ip: 192.168.1.1, and issue an alarm. (2) Threat event: At SERVER-4, it is determined that there is an attack from 10.1.40.X during the reconnaissance phase. Strategy: Handle node ip: 192.168.1.1, and issue an alarm. (3) Threat event: At SERVER-4, it is determined that there is an attack from 10.1.40.X during the reconnaissance phase and an attack from 10.1.40.X during the credential access phase. Strategy 1: Handle node ip: 192.168.1.1, and issue an alarm. (4) Threat event: At SERVER-4, it is determined that there is an attack from 10.1.40.X during the reconnaissance phase. 10.1.40.X's VPN access credentials are deleted. (5) Threat event: On SERVER-4, it is determined that there is an attack from 10.1.40.X during the execution phase. Policy: Disposal node ip: 192.168.0.1. Create a new security rule on the firewall to prohibit 10.1.40.X from accessing the trust zone and DMZ zone. (6) Threat event: On SERVER-4, it is determined that there is an attack from 10.1.10 / 20.X during the execution phase. Policy: Disposal node ip: 192.168.0.1. Alarm. (7) Threat event: On TERMINAL-1, it is determined that there is an attack from 10.1.10 / 20.X during the command and control phase. Policy: Disposal node ip: 192.168.0.1. Create a new security rule on the firewall to prohibit 10.1.10 / 20.X from accessing the trust zone and DMZ zone. Access to the trust zone and dmz zone; (8) Threat event: On SERVER-4, it is determined that there is an attack at the leakage stage from 10.1.10 / 20.X. Strategy: Disposal node: 192.168.0.1, block port 4444 of 10.1.10 / 20.X.
[0031] Similarly, based on the preset FTP attack, MySQL database attack and intranet forum attack scenarios, the “threat analysis” data set is constructed.
[0032] About the Threat Analysis LLM and Strategy Generation LLM Fine-tuning: Based on the existing research foundation and comprehensive consideration of training and reasoning cost control, this application adopts a technical solution of fine-tuning a smaller-scale large model with a higher-quality dataset, and compares the performance of the large model through task scenario-oriented actual measurements. It uses Llama3-8B as the base model, LoRA+ based on human feedback reinforcement learning (RLHF) as the fine-tuning technology, and threat analysis datasets and strategy generation datasets as fine-tuning datasets to carry out model fine-tuning, and finally obtain threat analysis LLM and strategy generation LLM.
[0033] The embodiment of the present application also includes the collection of abnormal traffic: In this application, it is necessary to collect abnormal traffic and output it to the simulated attack module, which replays the abnormal traffic to achieve simulated reproduction of the attack. The abnormal traffic collection is completed by the traffic collection module. For abnormal traffic data, this application uses Wireshark to collect network traffic data and analyze each data packet in the network. Wireshark is a common network data packet analysis tool that can intercept various network packets online, display detailed information of network packets, and analyze existing message data, such as message data collected by tcpdump / Win Dump, Wireshark, etc. Wireshark provides a variety of filtering rules for message filtering. It has seven major functions: packet capture, protocol analysis, traffic analysis, filtering, packet decoding, marking and annotation, and report generation.
[0034] Step 2: Threat analysis LLM performs threat analysis on the abnormal logs output by the log collection module and outputs a threat information quadruple; In this application, after receiving the input abnormal data, the threat analysis LLM outputs a four-tuple information, which is defined as: {attack stage, threat source node, victim node, security event}, and this four-tuple data is used as the input of the policy generation LLM to assist the policy generation LLM in generating the device security policy under the current threat scenario.
[0035] Step 3: The policy generation LLM generates the device security policy by calling the knowledge base Agent based on the threat information quadruple output by the threat assessment LLM; The policy generation LLM receives the four-tuple information generated by the threat analysis LLM. First, it performs preliminary data processing to repair some errors in the LLM generation process, such as output word errors, network information errors or null values, etc.; then, the target network topology information and device information are pre-encapsulated into tools. The policy generation LLM implements tool calls through the langchain framework, calls the knowledge base Agent for knowledge retrieval, and collects responses at each step, and continues to determine whether the tool needs to be used again. This iterative cycle continues until the answer is finally given. The answer includes network topology knowledge, disposal device node information, and security event information. It is worth noting that the logic of searching for disposable node devices preset in this application is: if there is a threat source node, query the nearest disposable node device information that can handle the threat source node; if there is no threat source node, query the nearest disposable node device information that can handle the victim node. Finally, the policy generation LLM outputs security policy binary information based on the threat analysis LLM analysis results, and the output is defined as: {threat event, security policy}.
[0036] Specifically, the present application constructs a knowledge base agent based on the network topology of the target network. In the system of the present application, the knowledge base agent only involves the network topology, so a graph database is used to establish the knowledge base agent in the present application. The construction steps of the graph database are: (1) constructing a node attribute table; (2) constructing a connection relationship table; (3) importing the node attribute table and the connection relationship table into the neo4j graph database to form a network topology knowledge graph. Among them, the fields of the node attribute table are: {node serial number, host name, node IP address, node port, node operating system, network segment, network area, log type, node grouping, devices that can handle the node, node device type}; the fields of the connection relationship table are: {starting point serial number, end point serial number, connection type}.
[0037] Step 4: The policy generation LLM calls the evaluation agent to evaluate the device security policy from three aspects: policy correctness evaluation, policy effectiveness evaluation, and policy adaptability evaluation.
[0038] Specifically, the evaluation agent in this application consists of three parts: strategy correctness evaluation, strategy effectiveness evaluation, and strategy adaptability evaluation. The evaluation results will serve as the benchmark for the feedback fine-tuning framework LLM. The fine-tuning method adopts the Proximal Policy Optimization (PPO) reinforcement learning algorithm.
[0039] (1) Policy correctness evaluation: The policy correctness evaluation is completed based on a reference policy set. Since it is aimed at internal / private networks, this application presets a reference policy set based on the target network environment and the device. First, the security policy output by the policy generation LLM is retrieved and compared with the reference policy set, and the confidence ranking calculation of the semantic similarity is performed by importing the pre-trained model "cross-encoder / stsb-distilroberta-base" in sentence_transformers. If the output policy has a very high similarity with one of the reference policies, it can be determined that the policy generation of the model is correct; then, it is accurately matched to the specific CLI command configuration of the corresponding network / security device, and the policy automation / manual intervention is performed. If the correctness match fails, the result is fed back to the policy generation LLM. It is worth noting that the reference policy set in this application refers to the set of all possible security policies for the devices in this network when facing threat scenarios for the target network.
[0040] (2) Policy effectiveness evaluation: The policy effectiveness evaluation is completed by the policy effectiveness module. The traffic collection module pushes the abnormal traffic data to the simulated attack module, which replays the attack in the isolated area to simulate the initiation of the attack. In the simulated network, the security policy output by the policy generation LLM is sent to the virtual network / security device terminal for policy execution. The module will continuously monitor the host log to observe the effect of policy implementation. If the attack fails after a period of policy execution, the output network / security device security policy is judged to be "valid" and the time from the policy being issued to the defense taking effect is recorded; otherwise, the security policy is judged to be "invalid" and the result is fed back to the attack analysis model. Specifically, the policy effectiveness module uses the network / security device image and the terminal device image to build a simulated test network, in which one terminal acts as a virtual attacker to receive the abnormal traffic data from the traffic collection module, and the simulated attack module of the policy effectiveness module implements a certain stage of attack behavior. After the policy generation model outputs the correct network / security device security policy, the instruction will be sent to the virtual network / security device terminal for policy execution.
[0041] (3) Policy adaptability evaluation: The policy adaptability evaluation is completed by the policy adaptability module. The output network / security device security policy is sent to the virtual network / security device terminal in the simulated network for policy execution. The policy adaptability module continuously monitors the traffic fluctuations of the simulated network and the impact on important business flows in the policy intranet. According to the preset network traffic fluctuation threshold, the policy adaptability results are output. It is worth noting that, considering that simple threshold settings are not sufficient to fully reflect the actual business situation of the network, this application introduces a human-computer interaction mechanism, sets up a network security expert interface, and pushes the policy adaptability results to network security experts, who can conduct a comprehensive assessment. If they fail, the results will be fed back to the policy generation model.
[0042] In the above embodiments, device security policies refer to a series of rules, configurations, and practices developed by devices such as firewalls, routers, and switches to protect the devices themselves and the information assets they process, in order to ensure the confidentiality, integrity, and availability of the devices and prevent unauthorized access, attacks, and other threats. These policies typically include mechanisms to block terminal devices (such as blacklists and whitelists), strong authentication mechanisms (such as multi-factor authentication), access control lists (ACLs) to limit network traffic, enable logging and monitoring to detect abnormal activities, and implement encryption technology to protect data transmission. In addition, physical security measures are also involved, such as limiting physical access to hardware to fully protect the safe operation of network / security devices.
[0043] See also Figure 2 The embodiment of the present application also provides a network / security device security policy generation and evaluation system based on a large language model, which is used to implement the above-mentioned network / security device security policy generation and evaluation method based on a large language model, which includes a log collection module, a traffic collection module, a threat analysis LLM, a policy generation LLM, a knowledge base Agent and an evaluation Agent; the evaluation Agent includes a policy validity module, a policy correctness module and a policy adaptability module; the log collection module is connected to the policy generation LLM through the threat analysis LLM; the policy generation LLM is connected to the knowledge base Agent and the evaluation Agent respectively; the policy validity module is connected to the threat analysis LLM, the simulated attack module in the policy validity module is connected to the traffic collection module, the policy correctness module and the policy adaptability module are connected to the policy generation LLM respectively, the policy validity module is used for policy validity evaluation; the policy adaptability module is used for policy adaptability evaluation; the policy correctness module is used for policy correctness evaluation.
[0044] Optionally, in the training phase, for the target network, the present application constructs a threat analysis data set and a policy generation data set, and completes the fine-tuning of the base model. In the testing phase, first, the abnormal data enters the threat analysis LLM after preprocessing, and the model outputs four-tuple information related to the attack; then, the policy generation LLM obtains the disposal node device information based on the four-tuple information by calling the knowledge base Agent, and generates the corresponding security response strategy; next, in order to measure the quality of the strategy generated by the LLM, the present application designs an evaluation Agent to perform a quality evaluation on the generated security strategy, and feeds back the evaluation results to the corresponding LLM according to the preset feedback mechanism, optimizes the iterative strategy, and stops iterating to output the strategy until the preset conditions are met. The preset conditions can be set according to user needs, such as when the number of iterations reaches the required number, and when the network fluctuation reaches a certain value.
[0045] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present application rather than to limit it. Although the present application has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present application can still be modified or replaced by equivalents, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present application should be included in the scope of protection of the claims of the present application.
Claims
1. A network / security device security policy generation and evaluation method based on a large language model, characterized in that: include: Step 1: Use the threat analysis dataset and strategy generation dataset to fine-tune the base model to obtain the threat analysis LLM and strategy generation LLM; Step 2: Threat analysis LLM performs threat analysis on the abnormal logs output by the log collection module and outputs a threat information quadruple; Step 3: The policy generation LLM generates the device security policy by calling the knowledge base Agent based on the threat information quadruple output by the threat analysis LLM; Step 4: The policy generation LLM calls the evaluation agent to evaluate the device security policy from three aspects: policy correctness evaluation, policy effectiveness evaluation, and policy adaptability evaluation.
2. The method for generating and evaluating network / security device security policies based on a large language model according to claim 1, characterized in that: The step 1 comprises: According to the application service scenarios of the target internal / private network, multiple attack paths under various application scenarios are designed to launch attacks on the target server or terminal. The relevant log data is collected through the abnormal log collection module, and different types of log data from different devices are sorted out. Finally, a threat analysis data set is constructed. The fields of the threat analysis data set are: {time, log, attacker IP, attacker port, victim IP, victim port, attack stage, threat level}; Design threat / abnormal scenarios covering the target network according to the target network environment and asset conditions. Obtain log data by launching multiple rounds of attacks. Combined with historical abnormal data, identify the attack stage through threat analysis LLM and locate the problem node. Combined with the relevant information provided by the knowledge base agent, establish defense strategy logic for the attack path and logic, perform annotation, and finally obtain the threat strategy data set. The fields of the threat strategy data set are: {threat event, strategy}; related information includes disposal nodes; A base model is selected, and LoRA+ based on human feedback reinforcement learning is used as the fine-tuning technology. The threat analysis dataset and strategy generation dataset are used as the fine-tuning datasets to fine-tune the base model and obtain the threat analysis LLM and strategy generation LLM.
3. The network / security device security policy generation and evaluation method based on a large language model according to claim 2 is characterized in that: The abnormal log collection module collects relevant log data and sorts out different types of log data from different devices, including: Download the filebeat installation package, edit the filebeat configuration file, and run the . / filebeat script on the target host where the filebeat configuration file is deployed; the filebeat configuration file includes the type and path of the logs collected locally and the information pushed to Logstash; Download the Logstash installation package, edit the Logstash configuration file, run the . / Logstash script, configure different Logstash server clusters for different types of logs on different devices, and start the Logstash service on the server through the Logstash configuration file; the Logstash configuration file includes the log source port, log content filtering rules, and the configuration of pushing logs to ES-Kibana; Store various log entries from different devices as a search engine; integrate various log data from different devices, count and visualize logs; log in to each host to view log indexes and specific data.
4. The network / security device security policy generation and evaluation method based on a large language model according to claim 2 is characterized in that: The step 2 comprises: The threat analysis LLM performs threat analysis on the abnormal logs received from the log collection module and outputs a threat information quadruple, which is: {attack stage, threat source node, victim node, security event}. The threat quadruple data is used as the input of the policy generation LLM, which assists the policy generation LLM in generating the device security policy under the current threat scenario.
5. The network / security device security policy generation and evaluation method based on a large language model according to claim 1 is characterized in that: The step 3 comprises: Fix the errors in the process of generating threat information quadruple, encapsulate the relevant information into tools in advance, and use the langchain framework to call the policy generation LLM to call the tool, call the knowledge base Agent for knowledge retrieval, and collect the response of each step, and continue to determine whether the tool needs to be used again. Repeat this iterative cycle until the answer is given. The answer includes network topology knowledge, disposal node device information and security event information; the relevant information includes the target network topology information and device information.
6. The network / security device security policy generation and evaluation method based on a large language model according to claim 5 is characterized in that: The method for searching the processing node device information is: if there is a threat source node, query the nearest disposable node device information that can process the threat source node; if there is no threat source node, query the nearest disposable node device information that can process the victim node. The policy generates LLM to output security policy binary information based on the threat analysis LLM analysis result. The security policy binary information is: {threat event, security policy}.
7. The method for generating and evaluating network / security device security policies based on a large language model according to claim 5, characterized in that: The construction process of the knowledge base includes: Construct a node attribute table and a connection relationship table, import the node attribute table and the connection relationship table into the graph database to form a network topology knowledge graph; the fields of the node attribute table are: {node serial number, host name, node IP address, node port, node operating system, network segment, network area, log type, node grouping, devices that can handle the node, node device type}; the fields of the connection relationship table are: {starting point serial number, end point serial number, connection type}.
8. The network / security device security policy generation and evaluation method based on a large language model according to claim 1 is characterized in that: The step 4 comprises: Policy correctness assessment: The device security policy output by the policy generation LLM is retrieved and compared with the reference policy set, and the confidence ranking calculation of the semantic similarity is performed. If the similarity between the output device security policy and any reference policy in the reference policy set exceeds the preset threshold, the device security policy output by the policy generation LLM is considered correct; then the command configuration of the corresponding network / security device is matched, and the policy is issued. If the command configuration of the corresponding network / security device is not matched, the result is fed back to the policy generation LLM; the reference policy set refers to the set of all possible security policies for devices in the target network when the target network faces a threat scenario.
9. The network / security device security policy generation and evaluation method based on a large language model according to claim 1 is characterized in that: The step 4 also includes: Policy effectiveness evaluation: The policy effectiveness evaluation is completed by the policy effectiveness module. The traffic collection module pushes the abnormal traffic data it collects to the simulated attack module in the evaluation agent. The simulated attack module replays the attack in the isolated area to simulate the launch of the attack. In the simulated network, the device security policy generated by the policy LLM output is sent to the virtual network / security device terminal for policy execution. The policy effectiveness module monitors the host log to observe the effect of policy implementation. If the attack fails after a period of execution of the device security policy, the output device security policy is judged to be valid, and the time from the device security policy being sent to the defense taking effect is recorded; otherwise, the device security policy is judged to be invalid, and the result is fed back to the threat assessment model.
10. The network / security device security policy generation and evaluation method based on a large language model according to claim 1, characterized in that: The step 4 also includes: Policy adaptability assessment: Policy adaptability assessment is completed by the policy adaptability module in combination with the human-computer interaction mechanism. The device security policy is sent to the virtual network / security device terminal in the simulated network for policy execution. The policy adaptability module monitors the traffic fluctuations of the simulated network and the impact on the important business flows of the device security policy intranet. According to the preset network traffic fluctuation threshold, the policy adaptability results are output and pushed to network security experts for comprehensive analysis. If it fails, the policy adaptability results are fed back to the policy generation LLM.
11. A network / security device security policy generation and evaluation system based on a large language model, implementing the network / security device security policy generation and evaluation method based on a large language model according to any one of claims 1 to 10, characterized in that: It includes log collection module, traffic collection module, threat analysis LLM, strategy generation LLM, knowledge base agent and evaluation agent; the evaluation agent includes strategy validity module, strategy correctness module and strategy adaptability module; The log collection module is connected to the policy generation LLM through the threat analysis LLM; the policy generation LLM is connected to the knowledge base agent and the evaluation agent respectively; the policy effectiveness module is connected to the threat analysis LLM, the simulated attack module in the policy effectiveness module is connected to the traffic collection module, and the policy correctness module and the policy adaptability module are connected to the policy generation LLM respectively; the policy effectiveness module is used for policy effectiveness evaluation; the policy adaptability module is used for policy adaptability evaluation; and the policy correctness module is used for policy correctness evaluation.
Citation Information
Patent Citations
LLM-driven industrial network intrusion detection method and response system
CN118381627A
Network security threat perception identification response method based on security knowledge graph
CN119011251A
Network security alarm automatic studying and judging method, device, equipment and medium
CN119402282A
Adaptive network security policy dynamic adjustment method
CN119766555A
Interactive cyber-security user-interface for cybersecurity components that cooperates with a set of llms
US20240414191A1
Cited By
Implementation method of security Agent system oriented to operating system
CN120180452A
Evaluation method and device of Internet of Things interface security policy, equipment and medium
CN120785638A
Method, device and medium for evaluating security policy of internet of things interface
CN120785638B
Flow collection system, threat analysis method and strategy generation method
CN120785652A
Network security decision-making method and device based on large language model
CN120856385A