Method and System for Generating and Evaluating Security Policies of Network / Security Devices Based on Large Language Models

By introducing LLM collaboration mechanism, multi-dimensional policy evaluation and dynamic adaptation mechanism into the network security protection system, the shortcomings of existing systems in strategy generation and evaluation are solved, the effectiveness and reliability of strategies are improved, and the intelligence level of network security protection is promoted.

CN119996085BActive Publication Date: 2025-06-10NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510462450.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-06-10
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing network security protection system based on large language models lacks adaptability to dynamic changes in the network environment when generating defense strategies and responds slowly; the strategy evaluation mechanism mostly relies on a single certain quantity indicator and lacks multi-dimensional comprehensive evaluation; the decision logic of LLM in complex network environments needs to be further optimized to improve its accuracy and reliability in practical applications.

Method used

By introducing LLM collaboration mechanism, multi-dimensional strategy evaluation mechanism and feedback-based dynamic adaptive mechanism, we will improve the effectiveness and reliability of equipment security policies and promote the intelligence level of network security protection. Specifically, it includes: fine-tuning threat analysis LLM and policy generation LLM, outputting threat information quadruples through threat analysis LLM, policy generation LLM generates device security policies, and evaluates policy accuracy, effectiveness and adaptability by evaluating the Agent.

Benefits of technology

It improves the efficiency of equipment security policies generation and evaluation, enhances the adaptability to changes in the network environment, improves the effectiveness and reliability of the strategies, and reduces the problems of insufficient timeliness of threat handling and high operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996085B_ABST
    Figure CN119996085B_ABST
Patent Text Reader

Abstract

The present application discloses a method and system for generating and evaluating security policies of network / security devices based on large language models. The method includes: fine-tuning a base model using a threat judgment data set and a policy generation data set respectively to obtain a threat judgment LLM and a policy generation LLM; the threat judgment LLM performs threat judgment on the abnormal logs output by the log collection module and outputs a threat information quadruple; the policy generation LLM generates a device security policy by calling a knowledge base Agent according to the threat information quadruple output by the threat judgment LLM; the policy generation LLM calls an evaluation Agent to evaluate the device security policy from three aspects: policy correctness evaluation, policy effectiveness evaluation, and policy adaptability evaluation. The present application improves the effectiveness and reliability of device security policies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of traffic anomaly detection, and in particular to a method and system for generating and evaluating network / security device security policies based on a large language model. Background Art

[0002] In network threat or daily operation and maintenance scenarios, security operation and maintenance personnel often face the problems of insufficient threat handling timeliness and high operation and maintenance costs, especially for internal networks with high control, strong isolation, and high security levels. In recent years, researchers have begun to explore the application of artificial intelligence technology to network security protection, especially security response and operation and maintenance systems based on machine learning and deep learning. The above-mentioned traditional security protection methods still have limitations in dealing with complex attack patterns and changing attack behaviors, especially in policy generation and dynamic adaptation. Moreover, security policy generation mostly relies on manual work, which usually has the problems of insufficient handling timeliness and high operation and maintenance costs. The emergence of large language models (LLMs) provides new ideas for solving these problems. Its powerful semantic understanding and generation capabilities enable it to extract key information from complex intranet attack logs, including node information and attack behavior information, and generate targeted internal / dedicated device defense strategies.

[0003] Despite this, the existing LLM-based network security protection system still has some problems that need to be solved. First, LLM often lacks adaptability to dynamic changes in the network environment when generating defense strategies, resulting in a slow response when facing new attacks. Second, the existing strategy evaluation mechanism mostly relies on a single quantitative indicator and lacks a multi-dimensional comprehensive evaluation of the strategy execution effect. In addition, the decision-making logic of LLM in complex network environments still needs to be further optimized to improve its accuracy and reliability in practical applications. Summary of the invention

[0004] In view of this, the present application provides a network / security device security policy generation and evaluation method and system based on a large language model. Through the LLM collaboration mechanism, multi-dimensional policy evaluation mechanism and feedback-based dynamic adaptive mechanism, it aims to improve the effectiveness and reliability of device security policies, promote the intelligent level of network security protection, and solve the problems of insufficient timeliness of threat disposal and high operation and maintenance costs.

[0005] The present application discloses a network / security device security policy generation and evaluation method based on a large language model, which includes:

[0006] Step 1: Use the threat analysis dataset and strategy generation dataset to fine-tune the base model to obtain the threat analysis LLM and strategy generation LLM;

[0007] Step 2: Threat Assessment The LLM for threat assessment conducts threat assessment on the abnormal logs output by the log collection module and outputs a quadruple of threat information;

[0008] Step 3: Policy Generation The LLM for policy generation generates device security policies by calling the knowledge base Agent according to the quadruple of threat information output by the LLM for threat assessment;

[0009] Step 4: Policy Evaluation The LLM for policy generation calls the evaluation Agent to evaluate the device security policies from three aspects: policy correctness evaluation, policy effectiveness evaluation, and policy adaptability evaluation.

[0010] Further, the said Step 1 includes:

[0011] According to the application service scenarios of the target internal / special network, design multiple attack paths under various application scenarios, launch attacks on the target server or terminal, collect relevant log data through the abnormal log collection module, sort out different types of log data of different devices, and finally construct a threat assessment data set. The fields of the threat assessment data set are: {time, log, attacker IP, attacker port, victim IP, victim port, attack stage, threat level};

[0012] Design threat / abnormal scenarios covering the target network according to the target network environment and asset conditions. Through multiple rounds of attacks, obtain log data, combine historical abnormal data, judge the attack stage through the LLM for threat assessment, locate the problem nodes, combine the relevant information provided by the knowledge base Agent, establish a defense policy logic for the attack paths and logics, and conduct annotation. Finally, obtain a threat policy data set. The fields of the threat policy data set are: {threat event, policy}; the relevant information includes the handling nodes;

[0013] Select a base model, use LoRA + reinforcement learning based on human feedback as the fine-tuning technology, and use the threat assessment data set and the policy generation data set as the fine-tuning data sets to carry out fine-tuning of the base model to obtain the LLM for threat assessment and the LLM for policy generation.

[0014] Further, the collection of relevant log data through the abnormal log collection module and the sorting out of different types of log data of different devices include:

[0015] Download the filebeat installation package, edit the filebeat configuration file, and run the. / filebeat script on the target host where the filebeat configuration file is deployed; the filebeat configuration file includes the types and paths of logs collected locally and the information pushed to Logstash;

[0016] Download the Logstash installation package, edit the Logstash configuration file, run the. / Logstash script, configure different Logstash server clusters for different types of logs of different devices, and start the Logstash service on the server through the Logstash configuration file; the Logstash configuration file includes the log source port, the log content filtering rules, and the configuration for pushing the logs to ES-Kibana;

[0017] Store various log entries of different devices as a search engine; integrate various log data of different devices, count and visualize the logs; log in to each host to view the log index and specific data.

[0018] Further, the step 2 includes:

[0019] The threat judgment LLM conducts threat judgment on the abnormal logs output by the log collection module and outputs a threat information quadruple. The threat information quadruple is: {attack stage, threat source node, victim node, security event}. The threat quadruple data is used as the input of the policy generation LLM to assist the policy generation LLM in generating device security policies under the current threat scenario.

[0020] Further, the step 3 includes:

[0021] Repair the errors in the process of generating the threat information quadruple, pre-encapsulate the relevant information into tools, the policy generation LLM realizes tool calls through the langchain framework, calls the knowledge base Agent for knowledge retrieval, and collects the responses of each step, and continues to judge whether it is necessary to use the tools again, and so on in an iterative loop until the final answer is given. The answer includes network topology knowledge, disposal node device information, and security event information; the relevant information includes the target network topology information and device information.

[0022] Further, the method for searching the disposal node device information is: if there is a threat source node, query the information of the nearest disposable node device that can dispose of the threat source node; if there is no threat source node, query the information of the nearest disposable node device that can dispose of the victim node. The policy generation LLM outputs security policy binary tuple information for the judgment result of the threat judgment LLM. The security policy binary tuple is: {threat event, security policy}.

[0023] Further, the construction process of the knowledge base includes:

[0024] Construct a node attribute table and a connection relationship table, and import the node attribute table and the connection relationship table into the graph database to form a network topology knowledge graph; among them, the fields of the node attribute table are: {node serial number, host name, node IP address, node port, node operating system, affiliated network segment, affiliated network area, log type, node group, devices that can handle this node, node device type}; the fields of the connection relationship table are: {starting point serial number, ending point serial number, connection type}.

[0025] Further, step 4 includes:

[0026] Policy correctness evaluation: Retrieve and compare the device security policy output by the policy generation LLM with the reference policy set, and calculate the confidence ranking of semantic similarity. If the similarity between the output device security policy and any reference policy in the reference policy set exceeds the preset threshold, it is considered that the device security policy output by the policy generation LLM is correct; then match the command configuration of the corresponding network / security device and perform the policy distribution operation. If the command configuration of the corresponding network / security device is not matched, the result is fed back to the policy generation LLM; the reference policy set refers to the set of all possible security policies of the devices in the target network when the target network faces threat scenarios.

[0027] Further, step 4 also includes:

[0028] Policy effectiveness evaluation: Policy effectiveness evaluation is completed by the policy effectiveness module. The traffic collection module pushes the abnormal traffic data it collects to the simulation attack module in the evaluation Agent. The simulation attack module replays the attack in the isolation area to simulate the initiation of the attack, and in the simulation network, the device security policy output by the policy generation LLM is distributed to the virtual network / security device terminal for policy execution. The policy effectiveness module monitors the host logs to observe the policy implementation effect; if the attack fails after the device security policy has been executed for a period of time, it is determined that the output device security policy is effective, and the time from when the device security policy is distributed to when the defense takes effect is recorded; otherwise, it is determined that the device security policy is ineffective, and the result is fed back to the threat judgment model.

[0029] Further, step 4 also includes:

[0030] Policy adaptability evaluation: Policy adaptability evaluation is completed by the policy adaptability module in combination with the human-computer interaction mechanism; the device security policy is distributed to the virtual network / security device terminal in the simulation network for policy execution. The policy adaptability module monitors the traffic fluctuation situation of the simulation network and the impact on the important business flows in the intranet of the device security policy. According to the preset network traffic fluctuation threshold, the policy adaptability result is output, and the policy adaptability result is pushed to the network security expert. The network security expert conducts a comprehensive judgment. If it fails, the policy adaptability result is fed back to the policy generation LLM.

[0031] The present application also discloses a network / security device security policy generation and evaluation system based on large language models, which implements the above-mentioned network / security device security policy generation and evaluation method based on large language models, including a log collection module, a traffic collection module, a threat judgment LLM, a policy generation LLM, a knowledge base Agent, and an evaluation Agent; the evaluation Agent includes a policy effectiveness module, a policy correctness module, and a policy adaptability module; the log collection module is connected to the policy generation LLM through the threat judgment LLM; the policy generation LLM is respectively connected to the knowledge base Agent and the evaluation Agent; the policy effectiveness module is connected to the threat judgment LLM, and the simulation attack module in the policy effectiveness module is connected to the traffic collection module, the policy correctness module and the policy adaptability module are respectively connected to the policy generation LLM, and the policy effectiveness module is used for policy effectiveness evaluation; the policy adaptability module is used for policy adaptability evaluation; the policy correctness module is used for policy correctness evaluation.

[0032] Due to the adoption of the above technical solutions, the present application has the following advantages:

[0033] 1. For the target network, introduce the LLM collaboration mechanism to achieve information sharing and policy optimization between different large models, optimize and train a pair of large model groups with excellent policy generation capabilities, and based on the idea of multi-Agent invocation, improve the decision-making efficiency and reliability of the overall system.

[0034] 2. Based on multi-dimensional and multi-level evaluation indicators, introduce a feedback loop to dynamically adjust the evaluation and "human-computer interaction" mechanism to adapt to the changing network environment and improve the effectiveness and reliability of internal / special security policies. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0036] Figure 1 It is a schematic flowchart of a method for generating and evaluating network / security device security policies based on large language models according to an embodiment of the present application;

[0037] Figure 2 It is a block diagram of a network / security device security policy generation and evaluation system based on large language models according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The present application will be further described in conjunction with the accompanying drawings and embodiments. The described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art shall fall within the scope of protection of the embodiments of the present application.

[0039] The present application mainly aims at the problem of high false negative rates of unknown threats in internal / special networks, and proposes a method and system for generating and evaluating network / security device security policies based on large language models. Specifically, the following problems are solved:

[0040] (1) How to utilize the semantic understanding ability of the LLM to analyze anomalies and judge their relevant attack information, and then generate preliminary defense policies that can be issued to internal / special devices.

[0041] (2) How to construct a policy evaluation space and design a comprehensive evaluation method with multiple dimensions to comprehensively evaluate the execution effect of the policies generated by the LLM.

[0042] (3) How to improve the adaptability of the LLM in complex scenarios and the quality and reliability of the generated defense policies.

[0043] See Figure 1 , a method for generating and evaluating network / security device security policies based on large language models proposed by the embodiments of the present application includes:

[0044] Step 1: Fine-tune the base model using the threat judgment data set and the policy generation data set respectively to obtain the threat judgment LLM and the policy generation LLM;

[0045] Anomaly data collection is a key network security task aimed at identifying and preventing potential security threats by recording and analyzing abnormal behaviors in the network. Different from the Internet environment, due to its special functions and deployment of specific application services, the internal / special environment faces different attacks from those in the Internet environment. In addition, most internal / special environments have high security requirements and special functional area divisions, which not only ensure the reliable operation of the internal network but also make the attack behaviors against the internal network have particularities.

[0046] Based on specific environments and service applications, this application designs the attack scenarios that the target network environment may face, sorts out the possible attack paths, launches simulated attacks, and collects relevant log data. At the same time, it sorts out different types of log data from different devices, and filters appropriate data features and fields according to the requirements of subsequent research and judgment. In addition, in order to unify the standards for subsequent threat research and judgment and response handling, it is planned to divide data windows. This application sets 10 seconds as the data window, divides and labels the attack data according to the attack phases of Adversarial Tactics, Techniques, and Common Knowledge (ATT&CK), and finally constructs a threat research and judgment data set, which is specifically as follows:

[0047] 1) Design of attack scenarios and paths

[0048] Specifically, according to the application service scenarios of the target internal / special network in this application, multiple attack paths in various application scenarios are designed, and multiple attack tools and technologies are used to launch attacks on the target server or terminal. The dataset fields are set as: {time, log, attacker IP, attacker port, victim IP, victim port, attack stage, threat level}. To clearly describe the construction of the dataset, here, one of the several attack paths targeting an intranet File Transfer Protocol (FTP) server is taken as an example for elaboration. The attack scenario is that the attacker penetrates into the intranet through a Virtual Private Network (VPN) and realizes data theft or damage to the intranet FTP server. The attack path is as follows: (1) The attacker steals the VPN credentials of a remote user and connects to the DMZ area through the VPN server proxy; (2) The attacker launches a scanning attack (passive attack) in the DMZ area, discovers that port 21 is open, and determines it as an FTP server; (3) A "smiley face" attack is launched against the FTP server due to a backdoor in the vsftpd 2.3.4 version, and finally, the root permission of the FTP server is obtained; (4) Weak password brute force to obtain permission; (5) Use Cobalt Strike to generate an exe trojan, disguise it as a normal file, and upload it to the target terminal through the FTP server using the put command; (6) Connect to the trojan and bounce a shell to communicate remotely with the remote attacker; (7) Use the vulnerability script in MSF to escalate privileges on the target terminal and obtain root permission; (8) The attacker steals privacy information using the root permission of the target terminal or damages data / systems. According to the above attack path, the corresponding attack stages and log descriptions are divided as follows: (1) Reconnaissance: VPN connection log, vpnlog from the router, representing that the attacker accesses the intranet through the proxy; (2) Reconnaissance: Port scanning log, ufwl.og from the FTP server, representing that the attacker scans the open ports of the server to obtain the application service list; (3) Credential Access: FTP brute force log, vsftpd.log from the FTP server, representing that the attacker brute forces the server to obtain FTP connection permission; (4) Execution: FTP file log tampering log, vsftpd.log from the FTP server, representing the process of the attacker uploading malicious files; (5) Command and Control: Client-side control log, syslog from the user terminal, representing that the user executes a malicious file; (6) Exfiltration: User bounce Sell log, syslog from the user terminal, representing that the user is infected with a virus after executing a malicious file; (7) Termination: VPN disconnection log, vpnlog from the router, representing that the attacker disconnects the proxy to end this attack.

[0049] Similarly, based on multiple preset paths of attack methods, the construction of the "threat assessment" dataset is completed. Among them, taking three attack scenarios of FTP attack, mysql database attack, and intranet forum attack as examples, 100 rounds of FTP attacks are launched, generating more than 104,200 abnormal logs, 110 rounds of mysql database attacks are launched, generating more than 133,800 abnormal logs, and 700 rounds of intranet forum injection attacks are launched, generating more than 191,400 abnormal logs.

[0050] 2) Abnormal log collection

[0051] Based on the foregoing attack scenarios and paths, relevant logs need to be collected, and the abnormal log collection is completed by the log collection module. After extensive research and testing in the early stage, for abnormal log data, Elasticsearch in the Elastic Stack technology stack is used in this application to provide persistent services, Filebeat-Logstash is used for log collection and preprocessing, and Kibana is used for visual display. The overall process of log collection is as follows:

[0052] ① Download the filebeat installation package, edit the filebeat configuration file, which mainly includes the types and paths of logs collected on the local machine and the information pushed to Logstash, and run the. / filebeat script on the target host where filebeat is deployed;

[0053] ② Download the Logstash installation package, edit the Logstash configuration file, which mainly includes the log source port, log content filtering rules, and the configuration of pushing logs to ES-Kibana, run the. / Logstash script, configure different Logstash server clusters for different types of logs, and start the Logstash service on the server through the above configuration;

[0054] ③ Store various log entries as a search engine;

[0055] ④ Integrate various log data, count and visualize the logs;

[0056] ⑤ Log in to each host to view the log index and specific data.

[0057] Regarding the construction of the policy generation dataset:

[0058] Since the disposal strategy needs to be aligned with the original input log data in time to achieve real-time analysis and disposal. Therefore, consistent with the previous threat assessment dataset, this application sets a 10-second data window, aggregates the attack stage and problem node outputs of the previous threat assessment LLM, the network topology information provided by the knowledge base Agent, the device importance, the information of the area where the problem node is located, and the disposal node information associated with the problem node, designs the disposal strategy logic according to the actual online scenario, and conducts manual annotation of the dataset.

[0059] It should be noted that this application is mainly targeted at the internal / special network environment. Therefore, threat / anomaly scenarios covering the target network can be designed according to the target network environment and asset situation. Then, by launching multiple rounds of attacks, log data is obtained, combined with historical anomaly data, the attack stage is judged by the previous threat assessment LLM, the problem node is located, and combined with relevant information such as the disposal node provided by the knowledge base Agent, a defense strategy logic is established for the attack path and logic, and manual annotation is carried out.

[0060] Specifically, based on the attack scenarios and paths during the construction of the previous threat assessment dataset, a threat strategy dataset is constructed. The dataset fields are: {threat event, strategy}. To clearly describe the construction of the dataset, the attack on the internal network FTP server mentioned above is used as an example for elaboration. The dataset construction logic is as follows: (1) Threat event: An attack in the reconnaissance stage from 10.1.40.X is identified on ROUTER-1. Strategy: Disposal node IP: 192.168.1.1, alarm; (2) Threat event: An attack in the reconnaissance stage from 10.1.40.X is identified on SERVER-4. Strategy: Disposal node IP: 192.168.1.1, alarm; (3) Threat event: An attack in the reconnaissance stage from 10.1.40.X and an attack in the credential access stage from 10.1.40.X are identified on SERVER-4. Strategy 1: Disposal node IP: 192.168.1.1, alarm; (4) Threat event: An attack in the credential access stage from 10.1.40.X is identified on SERVER-4. Strategy: Disposal node IP: 192.168.1.1, delete the VPN access credentials of 10.1.40.X; (5) Threat event: An attack in the execution stage from 10.1.40.X is identified on SERVER-4. Strategy: Disposal node IP: 192.168.0.1, create a new security rule in the firewall to prohibit 10.1.40.X from accessing the trust area and the dmz area; (6) Threat event: An attack in the execution stage from 10.1.10 / 20.X is identified on SERVER-4. Strategy: Disposal node IP: 192.168.0.1, alarm; (7) Threat event: An attack in the command and control stage from 10.1.10 / 20.X is identified on TERMINAL-1. Strategy: Disposal node IP: 192.168.0.1, create a new security rule in the firewall to prohibit 10.1.10 / 20.X from accessing the trust area and the dmz area; (8) Threat event: An attack in the exfiltration stage from 10.1.10 / 20.X is identified on SERVER-4. Strategy: Disposal node: 192.168.0.1, block port 4444 of 10.1.10 / 20.X.

[0061] Similarly, based on the preset FTP attack, mysql database attack, and internal network forum attack scenarios, the construction of the "threat assessment" dataset is completed.

[0062] Regarding the fine-tuning of the threat assessment LLM and the strategy generation LLM:

[0063] Based on the existing research foundation and comprehensively considering the control of training and inference costs, this application adopts a technical solution of fine-tuning a relatively small-scale large model with a high-quality dataset, and compares the performance of the large model through actual measurements in the task scenario. Taking Llama3-8B as the base model, LoRA + reinforcement learning based on human feedback (RLHF) as the fine-tuning technology, and threat judgment dataset and policy generation dataset as the fine-tuning datasets, model fine-tuning is carried out, and finally the threat judgment LLM and policy generation LLM are obtained.

[0064] The embodiment of this application also includes the collection of abnormal traffic:

[0065] In this application, abnormal traffic needs to be collected and output to the simulated attack module, and the simulated attack module replays the abnormal traffic to achieve the simulation and reproduction of the attack. The collection of abnormal traffic is completed by the traffic collection module. For abnormal traffic data, this application uses Wireshark to collect network traffic data and analyze each data packet in the network. Wireshark is a common network data packet analysis tool that can intercept various network packets online, display the detailed information of network packets, and can also analyze existing message data, such as message data collected by tcpdump / Win Dump, Wireshark, etc. Wireshark provides a variety of filtering rules for message filtering. It has seven major functions: data packet capture, protocol analysis, traffic analysis, filtering function, data packet decoding, marking and annotation, and report generation.

[0066] Step 2: The threat judgment LLM conducts threat judgment on the abnormal logs output by the log collection module and outputs a threat information quadruple;

[0067] In this application, after receiving the input abnormal data, the threat judgment LLM outputs a quadruple information, defined as: {attack stage, threat source node, victim node, security event}, and this quadruple data is used as the input of the policy generation LLM to assist the policy generation LLM in generating device security policies under the current threat scenario.

[0068] Step 3: The policy generation LLM generates device security policies by calling the knowledge base Agent according to the threat information quadruple output by the threat judgment LLM;

[0069] The policy generation LLM receives the quadruple information generated by the threat assessment LLM. First, it performs preliminary data processing to repair some errors in the LLM generation process, such as output character errors, network information errors, or null values. Then, it pre-packages the target network topology information and device information into tools. The policy generation LLM realizes tool invocation through the langchain framework, invokes the knowledge base Agent for knowledge retrieval, and collects the responses at each step, and continues to judge whether it is necessary to use the tool again. Such an iterative loop continues until the final answer is given. The answer includes network topology knowledge, disposal device node information, and security event information, etc. It should be noted that the logic for presetting the search for disposable node devices in this application is as follows: if there is a threat source node, query the information of the nearest disposable node device that can dispose of the threat source node; if there is no threat source node, query the information of the nearest disposable node device that can dispose of the victim node. Finally, the policy generation LLM outputs the security policy binary information for the threat assessment result of the threat assessment LLM, and the output is defined as: {threat event, security policy}.

[0070] Specifically, this application constructs a knowledge base Agent according to the network topology of the target network. In the system of this application, the knowledge base Agent only involves the network topology, so a graph database is used to establish the knowledge base Agent in this application. The construction steps of the graph database are as follows: (1) Construct a node attribute table; (2) Construct a connection relationship table; (3) Import the node attribute table and the connection relationship table into the neo4j graph database to form a network topology knowledge graph. Among them, the fields of the node attribute table are: {node serial number, host name, node IP address, node port, node operating system, affiliated network segment, affiliated network area, log type, node grouping, device that can dispose of this node, node device type}; the fields of the connection relationship table are: {starting point serial number, ending point serial number, connection type}.

[0071] Step 4: The policy generation LLM invokes the evaluation Agent to evaluate the device security policy from three aspects: policy correctness evaluation, policy effectiveness evaluation, and policy adaptability evaluation.

[0072] Specifically, the evaluation Agent in this application consists of three parts: policy correctness evaluation, policy effectiveness evaluation, and policy adaptability evaluation. The evaluation results will be used as the benchmark for fine-tuning the framework LLM, and the fine-tuning method uses the proximal policy optimization (PPO) reinforcement learning algorithm method.

[0073] (1)Policy correctness evaluation: Policy correctness evaluation is completed based on a reference policy set. Since it is for an internal / private network, this application presets a reference policy set based on the target network environment and devices. First, the security policies output by the policy generation LLM are retrieved and compared with the reference policy set. By importing the pre-trained model "cross-encoder / stsb-distilroberta-base" in sentence_transformers, the confidence ranking calculation of semantic similarity is performed. If the output policy has a very high similarity with a certain reference policy, it can be determined that the policy generation of the model is correct. Then, the specific CLI command configurations of the corresponding network / security devices are accurately matched, and the policy is issued automatically / manually. If the correctness match fails, the result is fed back to the policy generation LLM. It should be noted that the reference policy set in this application refers to all possible security policies of the devices in this network in the face of threat scenarios for the target network.

[0074] (2)Policy effectiveness evaluation: Policy effectiveness evaluation is completed by the policy effectiveness module. The traffic collection module pushes abnormal traffic data to the simulated attack module, and the simulated attack module replays the attack in the isolation area to simulate the initiation of the attack. In the simulation network, the security policies output by the policy generation LLM are issued to the virtual network / security device terminals for policy execution, and the module continuously monitors the host logs to observe the policy implementation effect. If the attack fails after a period of policy execution, it is determined that the output network / security device security policy is "effective", and the time from policy issuance to defense effectiveness is recorded; otherwise, it is determined that the security policy is "invalid", and the result is fed back to the attack research and judgment model. Specifically, the policy effectiveness module uses network / security device images and terminal device images to construct a simulation test network, where one terminal acts as a virtual attacker and will receive abnormal traffic data from the traffic collection module, and a certain stage of attack behavior is implemented by the simulated attack module of the policy effectiveness module. After the policy generation model outputs the correct network / security device security policy, the instruction is issued to the virtual network / security device terminals for policy execution.

[0075] (3)Policy Adaptability Evaluation: The policy adaptability evaluation is completed by the policy adaptability module. The output network / security device security policy is issued to the virtual network / security device terminal in the simulation network for policy execution. The policy adaptability module continuously monitors the traffic fluctuations in the simulation network and the impact on the important business flows in the policy intranet. According to the preset network traffic fluctuation threshold, the policy adaptability result is output. It should be noted that considering that simple threshold settings are not sufficient to comprehensively reflect the real business situation of the network, this application introduces a human-computer interaction mechanism, sets up a network security expert interface, pushes the policy adaptability result to the network security expert, and the security expert can conduct a comprehensive judgment. If it fails, the result is fed back to the policy generation model.

[0076] In the above embodiments, the device security policy refers to a series of rules, configurations, and practices formulated for devices such as firewalls, routers, and switches to protect the device itself and the information assets it processes, aiming to ensure the confidentiality, integrity, and availability of the device, prevent unauthorized access, attacks, and other threats. These policies usually include mechanisms for blocking terminal devices (such as black and white lists), strong authentication mechanisms (such as multi-factor authentication), access control lists (ACLs) to restrict network traffic, enabling logging and monitoring to detect abnormal activities, and implementing encryption technologies to protect data transmission. In addition, physical security measures are also involved, such as restricting physical access to the hardware to comprehensively ensure the secure operation of the network / security device.

[0077] See Figure 2 , the embodiment of this application also provides a network / security device security policy generation and evaluation system based on a large language model, which is used to implement the above-mentioned network / security device security policy generation and evaluation method based on a large language model. It includes a log collection module, a traffic collection module, a threat judgment LLM, a policy generation LLM, a knowledge base Agent, and an evaluation Agent; the evaluation Agent includes a policy effectiveness module, a policy correctness module, and a policy adaptability module; the log collection module is connected to the policy generation LLM through the threat judgment LLM; the policy generation LLM is respectively connected to the knowledge base Agent and the evaluation Agent; the policy effectiveness module is connected to the threat judgment LLM, and the simulation attack module in the policy effectiveness module is connected to the traffic collection module; the policy correctness module and the policy adaptability module are respectively connected to the policy generation LLM; the policy effectiveness module is used for policy effectiveness evaluation; the policy adaptability module is used for policy adaptability evaluation; the policy correctness module is used for policy correctness evaluation.

[0078] Optionally, in the training phase, for the target network, the present application constructs a threat judgment dataset and a policy generation dataset, and completes the fine-tuning of the base model. In the testing phase, first, the abnormal data enters the threat judgment LLM after preprocessing, and the model outputs quadruple information related to attacks; then, the policy generation LLM obtains the information of the disposal node device by calling the knowledge base Agent according to the quadruple information, and generates the corresponding security response policy; next, in order to measure the quality of the policy generated by the LLM, the present application designs an evaluation Agent to evaluate the quality of the generated security policy, and feeds back the evaluation result to the corresponding LLM according to the preset feedback mechanism to optimize and iterate the policy until the preset conditions are met and the iteration stops to output the policy. The preset conditions can be set according to user requirements, such as the number of iterations reaching the required number, or the network fluctuation reaching a certain value.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present application can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present application shall be covered by the protection scope of the claims of the present application.

Claims

1. A network / security device security policy generation and evaluation method based on a large language model, characterized in that: include: Step 1: Use the threat analysis dataset and strategy generation dataset to fine-tune the base model to obtain the threat analysis LLM and strategy generation LLM; Step 2: Threat analysis LLM performs threat analysis on the abnormal logs output by the log collection module and outputs a threat information quadruple; Step 3: The policy generation LLM generates the device security policy by calling the knowledge base Agent based on the threat information quadruple output by the threat analysis LLM; Step 4: The policy generation LLM calls the evaluation agent to evaluate the device security policy from three aspects: policy correctness evaluation, policy effectiveness evaluation, and policy adaptability evaluation; The step 1 comprises: According to the application service scenarios of the target internal / private network, multiple attack paths under various application scenarios are designed to launch attacks on the target server or terminal. The relevant log data is collected through the abnormal log collection module, and different types of log data from different devices are sorted out. Finally, a threat analysis data set is constructed. The fields of the threat analysis data set are: {time, log, attacker IP, attacker port, victim IP, victim port, attack stage, threat level}; Design threat / abnormal scenarios covering the target network according to the target network environment and asset conditions. Obtain log data by launching multiple rounds of attacks. Combined with historical abnormal data, identify the attack stage through threat analysis LLM and locate the problem node. Combined with the relevant information provided by the knowledge base agent, establish defense strategy logic for the attack path and logic, perform annotation, and finally obtain the threat strategy data set. The fields of the threat strategy data set are: {threat event, strategy}; related information includes disposal nodes; A base model is selected, and LoRA+ based on human feedback reinforcement learning is used as the fine-tuning technology. The threat analysis dataset and strategy generation dataset are used as the fine-tuning datasets to fine-tune the base model and obtain the threat analysis LLM and strategy generation LLM.

2. The method for generating and evaluating network / security device security policies based on a large language model according to claim 1, characterized in that: The abnormal log collection module collects relevant log data and sorts out different types of log data from different devices, including: Download the filebeat installation package, edit the filebeat configuration file, and run the . / filebeat script on the target host where the filebeat configuration file is deployed; the filebeat configuration file includes the type and path of the logs collected locally and the information pushed to Logstash; Download the Logstash installation package, edit the Logstash configuration file, run the . / Logstash script, configure different Logstash server clusters for different types of logs on different devices, and start the Logstash service on the server through the Logstash configuration file; the Logstash configuration file includes the log source port, log content filtering rules, and the configuration of pushing logs to ES-Kibana; Store various log entries from different devices as a search engine; integrate various log data from different devices, count and visualize logs; log in to each host to view log indexes and specific data.

3. The network / security device security policy generation and evaluation method based on a large language model according to claim 1 is characterized in that: The step 2 comprises: The threat analysis LLM performs threat analysis on the abnormal logs received from the log collection module and outputs a threat information quadruple, which is: {attack stage, threat source node, victim node, security event}. The threat quadruple data is used as the input of the policy generation LLM, which assists the policy generation LLM in generating the device security policy under the current threat scenario.

4. The network / security device security policy generation and evaluation method based on a large language model according to claim 1 is characterized in that: The step 3 comprises: Fix the errors in the process of generating threat information quadruple, encapsulate the relevant information into tools in advance, and use the langchain framework to call the policy generation LLM to call the tool, call the knowledge base Agent for knowledge retrieval, and collect the response of each step, and continue to determine whether the tool needs to be used again. Repeat this iterative cycle until the answer is given. The answer includes network topology knowledge, disposal node device information and security event information; the relevant information includes the target network topology information and device information.

5. The method for generating and evaluating network / security device security policies based on a large language model according to claim 4, characterized in that: The method for searching the processing node device information is: if there is a threat source node, query the nearest disposable node device information that can process the threat source node; if there is no threat source node, query the nearest disposable node device information that can process the victim node. The policy generates LLM to output security policy binary information based on the threat analysis LLM analysis result. The security policy binary information is: {threat event, security policy}.

6. The network / security device security policy generation and evaluation method based on a large language model according to claim 4 is characterized in that: The construction process of the knowledge base includes: Construct a node attribute table and a connection relationship table, import the node attribute table and the connection relationship table into the graph database to form a network topology knowledge graph; the fields of the node attribute table are: {node serial number, host name, node IP address, node port, node operating system, network segment, network area, log type, node grouping, devices that can handle the node, node device type}; the fields of the connection relationship table are: {starting point serial number, end point serial number, connection type}.

7. The network / security device security policy generation and evaluation method based on a large language model according to claim 1 is characterized in that: The step 4 comprises: Policy correctness assessment: The device security policy output by the policy generation LLM is retrieved and compared with the reference policy set, and the confidence ranking calculation of the semantic similarity is performed. If the similarity between the output device security policy and any reference policy in the reference policy set exceeds the preset threshold, the device security policy output by the policy generation LLM is considered correct; then the command configuration of the corresponding network / security device is matched, and the policy is issued. If the command configuration of the corresponding network / security device is not matched, the result is fed back to the policy generation LLM; the reference policy set refers to the set of all possible security policies for devices in the target network when the target network faces a threat scenario.

8. The network / security device security policy generation and evaluation method based on a large language model according to claim 1 is characterized in that: The step 4 also includes: Policy effectiveness evaluation: The policy effectiveness evaluation is completed by the policy effectiveness module. The traffic collection module pushes the abnormal traffic data it collects to the simulated attack module in the evaluation agent. The simulated attack module replays the attack in the isolated area to simulate the launch of the attack. In the simulated network, the device security policy generated by the policy LLM output is sent to the virtual network / security device terminal for policy execution. The policy effectiveness module monitors the host log to observe the effect of policy implementation. If the attack fails after a period of execution of the device security policy, the output device security policy is judged to be valid, and the time from the device security policy being sent to the defense taking effect is recorded; otherwise, the device security policy is judged to be invalid, and the result is fed back to the threat assessment model.

9. The network / security device security policy generation and evaluation method based on a large language model according to claim 1 is characterized in that: The step 4 also includes: Policy adaptability assessment: Policy adaptability assessment is completed by the policy adaptability module in combination with the human-computer interaction mechanism. The device security policy is sent to the virtual network / security device terminal in the simulated network for policy execution. The policy adaptability module monitors the traffic fluctuations of the simulated network and the impact on the important business flows of the device security policy intranet. According to the preset network traffic fluctuation threshold, the policy adaptability results are output and pushed to network security experts for comprehensive analysis. If it fails, the policy adaptability results are fed back to the policy generation LLM.

10. A network / security device security policy generation and evaluation system based on a large language model, implementing the network / security device security policy generation and evaluation method based on a large language model according to any one of claims 1 to 9, characterized in that: It includes log collection module, traffic collection module, threat analysis LLM, strategy generation LLM, knowledge base agent and evaluation agent; the evaluation agent includes strategy validity module, strategy correctness module and strategy adaptability module; The log collection module is connected to the policy generation LLM through the threat analysis LLM; the policy generation LLM is connected to the knowledge base agent and the evaluation agent respectively; the policy effectiveness module is connected to the threat analysis LLM, the simulated attack module in the policy effectiveness module is connected to the traffic collection module, and the policy correctness module and the policy adaptability module are connected to the policy generation LLM respectively; the policy effectiveness module is used for policy effectiveness evaluation; the policy adaptability module is used for policy adaptability evaluation; and the policy correctness module is used for policy correctness evaluation.

Citation Information

Patent Citations

  • LLM-driven industrial network intrusion detection method and response system

    CN118381627A

  • Network security threat perception identification response method based on security knowledge graph

    CN119011251A