Network security decision-making method and device based on large language model
By leveraging the semantic understanding and contextual learning capabilities of Large Language Models (LLM), network state data is collected to generate natural language descriptions, semantic feature vectors are determined, target threat knowledge is retrieved, and protection strategies are generated. This solves the problem of response delay in existing network security defense systems when facing new threats, and enables efficient and accurate network security decision-making.
Patent Information
- Application Number
- CN202510935404.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-28
AI Technical Summary
Existing network security defense systems have limitations when facing new threats. Traditional signature-based intrusion detection systems and static firewall rules cannot cope with zero-day exploits and fileless attacks. Remote network unit decision-making methods have response delays in high-dimensional heterogeneous data and dynamic network environments.
By leveraging the semantic understanding and contextual learning capabilities of Large Language Models (LLM), natural language descriptions are generated by collecting network state data, semantic feature vectors are determined, target threat knowledge is retrieved, protection strategies are generated, and distributed to distributed execution units for execution and effect monitoring, thus realizing a closed-loop intelligent defense process from information perception to collaborative response.
It improves the real-time nature, accuracy, and scalability of security decisions, enabling rapid response to complex network threats and achieving intelligent network security defense.
Smart Images

Figure CN120856385A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of cybersecurity technology, and more specifically, to a cybersecurity decision-making method and apparatus based on a large language model. Background Technology
[0002] With the rapid development of the internet and information technology, the complexity and threat of cyberattacks have significantly increased. They have evolved from traditional malware and virus propagation to advanced persistent threats (APTs), zero-day exploits, and AI-driven attacks, posing unprecedented challenges to cybersecurity defenses. While current cybersecurity defense systems have achieved a certain level of protection through multi-layered technology stacking, their limitations are becoming increasingly apparent in the face of new threats. Traditional signature-based intrusion detection systems (IDS) and static firewall rules perform stably under known attack patterns but are unable to cope with emerging threats such as zero-day exploits and fileless attacks.
[0003] Existing remote network unit decision-making methods mostly rely on pre-defined rules or traditional machine learning models, which have significant limitations when dealing with high-dimensional heterogeneous data and dynamic network environments. For example, in edge computing scenarios, rule-based systems struggle to respond quickly to changes in network state, leading to task scheduling delays. Summary of the Invention
[0004] This disclosure provides at least one network security decision-making method and apparatus based on a large language model (LLM). It applies the semantic understanding, reasoning, and context learning capabilities of the LLM to remote network unit decision-making, enabling a closed-loop intelligent defense process from information perception to semantic understanding to collaborative response, thereby improving the real-time performance, accuracy, and scalability of security decisions.
[0005] This disclosure provides a network security decision-making method based on a large language model, including:
[0006] Collect network status data corresponding to the target network environment, and convert the network status data into a natural language description with contextual semantics;
[0007] Determine the semantic feature vector corresponding to the natural language description, and retrieve target threat knowledge that matches the semantic feature vector from a knowledge base that stores various basic network threat knowledge;
[0008] The target threat knowledge and the natural language description are combined to generate analysis prompt text, which is then input into the large language model. The network threat analysis steps indicated in the preset thinking chain are executed sequentially according to the analysis prompt text to determine the network anomaly attribution and the corresponding protection strategy.
[0009] The protection policy is sent to each distributed execution unit deployed in the target network environment, the protection policy is converted into a device configuration command, the device configuration command is executed and the changes in the network status data are monitored to determine the execution effect of the protection policy.
[0010] In one optional implementation, collecting network status data corresponding to the target network environment specifically includes:
[0011] The Syslog service deployed on each network device in the target network environment records network operation events in real time.
[0012] The network device's structure and performance attribute information is collected using a simple network management protocol, and abnormal traffic data during network operation is marked and captured using a network data collection and analysis tool.
[0013] Extract the expanded data corresponding to the preset key network indicators and the context information corresponding to the network operation events and the abnormal traffic data;
[0014] The network status data is generated by formatting the network operation events, the structure and performance attribute information, and the abnormal traffic data.
[0015] In one optional implementation, a semantic feature vector corresponding to the natural language description is determined, and target threat knowledge matching the semantic feature vector is retrieved from a knowledge base storing various basic network threat knowledge. Specifically, this includes:
[0016] Construct the knowledge base, in which threat feature descriptions and corresponding normal thresholds for various basic network threats are pre-stored in text form as the basic network threat knowledge;
[0017] Each piece of basic network threat knowledge is transformed into a high-dimensional vector embedding through a pre-trained embedding model, and an index corresponding to each high-dimensional vector embedding is constructed using a vector retrieval tool.
[0018] The natural language description is converted into the semantic feature vector through the embedding model, and the target threat knowledge that matches the semantic feature vector is retrieved from the index.
[0019] In one optional implementation, the analysis prompt text generated by combining the target threat knowledge with the natural language description is input into the large language model, specifically including:
[0020] The target threat knowledge is combined with the natural language description to form context-enhanced information;
[0021] Determine the threat type corresponding to the target threat knowledge, and select a prompt text template that matches the threat type from the preset prompt text generation templates;
[0022] The contextual enhancement information is embedded into the prompt text template to generate the analysis prompt text used to prompt the analysis steps and reasoning basis of the large language model.
[0023] In one optional implementation, the network threat analysis steps indicated in the preset thought chain are executed sequentially according to the analysis prompt text to determine the network anomaly attribution and corresponding protection strategy, specifically including:
[0024] Determine whether the network status data deviates from the normal threshold indicated in the target threat knowledge; if the network status data deviates from the normal threshold, then determine that the target network status is abnormal.
[0025] Based on the target threat knowledge, determine the network anomaly attribution corresponding to the target network state, and determine whether it is a potential threat;
[0026] Based on the analysis results of the network anomaly attribution, the protection strategy is generated under the prompt of the analysis prompt text.
[0027] In one optional implementation, the protection policy is distributed to each distributed execution unit deployed in the target network environment, the protection policy is converted into a device configuration command, and the device configuration command is executed, specifically including:
[0028] The protection policy is transmitted to the intelligent agent, which is the distributed execution unit, through a preset protocol, and the protection policy is converted into a policy instruction that can be recognized by the network unit in the target network environment;
[0029] The intelligent agent parses the policy instructions into device configuration commands that can be executed by the local network device.
[0030] Verify the integrity of the device configuration command and execute the device configuration command, and record the operation log corresponding to the device configuration command.
[0031] In one optional implementation, the network state data is converted into a natural language description with contextual semantics, specifically including:
[0032] Based on the Jinja2 template engine and Python preprocessing logic, the network state data is mapped to the natural language description;
[0033] The syntax and semantics of the natural language description are optimized using the T5-small model, and the natural language description is transformed into a structured input format that can be directly processed by the large language model.
[0034] This disclosure also provides a network security decision-making device based on a large language model, comprising:
[0035] The natural language description conversion module is used to collect network state data corresponding to the target network environment and convert the network state data into a natural language description with contextual semantics.
[0036] The retrieval enhancement module is used to determine the semantic feature vector corresponding to the natural language description, and to retrieve target threat knowledge that matches the semantic feature vector from a knowledge base that stores various basic network threat knowledge.
[0037] The protection strategy analysis module is used to combine the target threat knowledge with the natural language description to generate analysis prompt text, which is then input into the large language model. Based on the analysis prompt text, the module sequentially executes the network threat analysis steps indicated in the preset thought chain to determine the network anomaly attribution and the corresponding protection strategy.
[0038] The policy execution module is used to distribute the protection policy to each distributed execution unit deployed in the target network environment, convert the protection policy into device configuration commands, execute the device configuration commands, and monitor changes in the network status data to determine the execution effect of the protection policy.
[0039] This disclosure also provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the above-described network security decision-making method based on a large language model, or any possible implementation of the above-described network security decision-making method based on a large language model.
[0040] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the above-described network security decision-making method based on a large language model, or any possible implementation of the above-described network security decision-making method based on a large language model.
[0041] This disclosure also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the above-described network security decision-making method based on a large language model, or any possible implementation steps of the above-described network security decision-making method based on a large language model.
[0042] This disclosure provides a network security decision-making method and apparatus based on a large language model (LLM). The method involves collecting network state data corresponding to a target network environment and converting the network state data into a natural language description with contextual semantics. It then determines the semantic feature vector corresponding to the natural language description and retrieves target threat knowledge matching the semantic feature vector from a knowledge base storing various basic network threat knowledge. The method combines the target threat knowledge with the natural language description to generate analysis prompt text, which is then input into the LLM. Based on the analysis prompt text, the method sequentially executes network threat analysis steps indicated in a preset thought chain to determine the network anomaly attribution and corresponding protection strategy. The method distributes the protection strategy to each distributed execution unit deployed in the target network environment, converts the protection strategy into device configuration commands, executes the device configuration commands, and monitors changes in the network state data to determine the effectiveness of the protection strategy. Applying the semantic understanding, reasoning, and contextual learning capabilities of a large language model (LLM) to remote network unit decision-making enables a closed-loop intelligent defense process from information perception—semantic understanding—coordinated response, improving the real-time performance, accuracy, and scalability of security decisions.
[0043] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0044] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.
[0045] Figure 1 A flowchart illustrating a network security decision-making method based on a large language model provided in an embodiment of this disclosure is shown.
[0046] Figure 2 A flowchart illustrating another network security decision-making method based on a large language model provided in an embodiment of this disclosure is shown;
[0047] Figure 3 A schematic diagram of a network security decision-making device based on a large language model provided in an embodiment of this disclosure is shown;
[0048] Figure 4A schematic diagram of an electronic device provided in an embodiment of the present disclosure is shown. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0050] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0051] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0052] Research has revealed that while current cybersecurity defense systems achieve a certain level of protection through multi-layered technology stacking, their limitations are increasingly apparent in the face of emerging threats. Traditional signature-based intrusion detection systems (IDS) and static firewall rules perform stably under known attack patterns but are unable to cope with emerging threats such as zero-day exploits and fileless attacks. Existing remote network unit decision-making methods largely rely on pre-set rules or traditional machine learning models, which have significant limitations when handling high-dimensional heterogeneous data and dynamic network environments. For example, in edge computing scenarios, rule-based systems struggle to respond quickly to changes in network status, leading to task scheduling delays.
[0053] Based on the above research, this disclosure provides a network security decision-making method and apparatus based on a large language model (LLM). The method involves collecting network state data corresponding to a target network environment and converting the network state data into a natural language description with contextual semantics; determining the semantic feature vector corresponding to the natural language description and retrieving target threat knowledge matching the semantic feature vector from a knowledge base storing various basic network threat knowledge; combining the target threat knowledge with the natural language description to generate analysis prompt text, which is then input into the LLM. The method executes network threat analysis steps indicated in a preset thought chain according to the analysis prompt text to determine the network anomaly attribution and corresponding protection strategy; distributing the protection strategy to each distributed execution unit deployed in the target network environment; converting the protection strategy into device configuration commands; executing the device configuration commands; and monitoring changes in the network state data to determine the effectiveness of the protection strategy. Applying the semantic understanding, reasoning, and contextual learning capabilities of a large language model (LLM) to remote network unit decision-making enables a closed-loop intelligent defense process from information perception—semantic understanding—coordinated response, improving the real-time performance, accuracy, and scalability of security decisions.
[0054] To facilitate understanding of this embodiment, a detailed description of the network security decision-making method based on a large language model disclosed in this disclosure is provided first. The execution entity of the network security decision-making method based on a large language model provided in this disclosure is generally a computer device with a certain computing capability. This computer device may include, for example, a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. In some possible implementations, this network security decision-making method based on a large language model can be implemented by a processor calling computer-readable instructions stored in memory.
[0055] See Figure 1 The diagram shows a flowchart of a network security decision-making method based on a large language model provided in this disclosure. The method includes steps S101 to S104, wherein:
[0056] S101. Collect network status data corresponding to the target network environment, and convert the network status data into a natural language description with contextual semantics.
[0057] In practical implementation, in response to the security protection requirements of the target network environment, the first step is to collect and semantically convert network status data. This involves using Syslog services deployed on various network devices in the target network environment to record network operation events in real time; collecting the structural and performance attribute information of network devices using the Simple Network Management Protocol; and marking and capturing abnormal traffic data during network operation using network data collection and analysis tools; extracting extended data corresponding to preset key network indicators and contextual information corresponding to network operation events and abnormal traffic data; and finally, formatting the network operation events, structural and performance attribute information, and abnormal traffic data to generate network status data.
[0058] Here, in the target network environment, the System Log (Syslog) service deployed on each network device is used as a means of real-time event collection to record network operation events, including login events, configuration changes, connection establishment or disconnection, etc. In addition, the Simple Network Management Protocol (SNMP) is used to periodically retrieve the structural information (such as device topology and port information) and performance attributes (such as CPU utilization, memory utilization, and interface bandwidth) of each network device.
[0059] In order to capture potential abnormal behavior, network traffic is further captured using network data acquisition and analysis tools such as Wireshark, Zeek, or Suricata to extract abnormal traffic data (e.g., port scanning, abnormal connection frequency, packet distortion, etc.).
[0060] At the same time, in conjunction with the network security indicator system, expanded data of preset key network indicators (such as traffic surges, response latency, and packet loss rate) are extracted, and the above network operation events, structural and performance attribute information, and abnormal traffic data are uniformly timestamped to preserve context information.
[0061] As one possible implementation, streaming parsing and pattern matching techniques are applied to filter and denoise the collected network status data. For example, irrelevant traffic is filtered out using a multi-threaded parsing engine (implemented in Python). Key network metrics, including the number of abnormal connections, packet loss rate, and traffic rate, are calculated, and timestamps and contextual information of security events are extracted to optimize the quality of decision data input to the Large Language Model (LLM). A dynamic threshold detection mechanism is designed to automatically mark network traffic exceeding a preset threshold (e.g., exceeding 5Gbps) or the abnormal connection growth rate exceeding the normal range as a high-risk event and record detailed traffic context data.
[0062] It should be noted that the MQTT protocol is used for data transmission. This protocol is suitable for resource-constrained environments and uses a TLS encrypted channel to ensure the confidentiality and integrity of data during transmission. A high-performance memory caching system (such as Redis) is deployed on the local server to store data in JSON format, improving query and parsing efficiency. A low-latency queue mechanism is implemented, and priority scheduling and high-concurrency processing technologies are used to ensure the real-time performance of data transmission and avoid delays or data loss caused by traffic peaks.
[0063] Furthermore, to achieve a context-aware input format for large language models, the aforementioned network state data needs to be semantically converted into natural language descriptions. By transforming network states (such as traffic logs, network topology, and protocol data) into concise and semantically rich natural language descriptions, the language understanding capabilities of LLM can be fully utilized to directly analyze complex network states, achieving a seamless transition from technical language to decision-making language, thereby significantly improving the intelligence level of remote network unit decision-making.
[0064] In this embodiment, network state data is mapped to natural language descriptions based on the Jinja2 template engine and Python preprocessing logic; the syntax and semantics of the natural language descriptions are optimized through the T5-small model, and the natural language descriptions are transformed into a structured input format that can be directly processed by a large language model.
[0065] Here, Python scripts are used in conjunction with the Jinja2 template engine to construct templated transformation rules. For example, a network event data can be transformed using the following template: by filling the original structured data fields (such as device_name, timestamp, etc.) into the template, a preliminary natural language sentence can be generated. Simultaneously, data mapping rules are predefined for different types of devices (such as switches and servers), covering key indicators such as traffic status, CPU load, and abnormal event logs, ensuring that the transformation results conform to professional terminology and practical needs in the field of network security.
[0066] To further enhance the readability of the description and the understanding capability of the language model, the T5-small language generation model is introduced to optimize the syntax and semantics of the initial natural language description, improving syntactic accuracy and semantic clarity, making the description more fluent and easier to understand. During the optimization process, the model retains key entities, temporal logic, and contextual causality, ultimately transforming structured network data into natural language description input text with contextual semantics, which is then used for subsequent threat knowledge matching and analysis prompt generation.
[0067] It should be noted that a description length control mechanism can be implemented during this process. It is recommended that the output length be controlled between 50 and 100 characters to meet the transmission requirements in low-bandwidth environments while still fully conveying key network status information.
[0068] In this way, based on the collaborative processing of Jinja2 and Python, the raw network state data is mapped into a structured natural language description, ensuring the accuracy and completeness of the information. The syntax and semantics of the description are optimized through the T5-small model, improving the readability of the description and the understanding efficiency of LLM. At the same time, a dedicated parsing module is designed to transform the network state data into a structured input format that can be directly processed by LLM, thereby improving the analysis efficiency and decision accuracy.
[0069] S102. Determine the semantic feature vector corresponding to the natural language description, and retrieve the target threat knowledge that matches the semantic feature vector from a knowledge base that stores various basic network threat knowledge.
[0070] In practical implementation, in order for the large language model to efficiently match existing threat knowledge based on network state natural language descriptions to guide security decisions, the natural language descriptions need to be converted into semantic feature vectors and then retrieved in a vectorized manner within the constructed threat knowledge base.
[0071] Specifically, a knowledge base is constructed, which pre-stores threat feature descriptions and corresponding normal thresholds for various basic network threats in text form as basic network threat knowledge. Each piece of basic network threat knowledge is transformed into a high-dimensional vector embedding through a pre-trained embedding model, and an index corresponding to each high-dimensional vector embedding is constructed using a vector retrieval tool. The natural language description is converted into a semantic feature vector through the embedding model, and target threat knowledge that matches the semantic feature vector is retrieved from the index.
[0072] First, a pre-built network threat knowledge base is used to store textual descriptions of various basic network threats. Each basic network threat knowledge entry includes at least the following: threat type (e.g., ARP spoofing, DDoS attack, port scanning, etc.); corresponding typical characteristics (e.g., "frequent changes in destination IP" or "a surge in connection requests per unit time"); relevant normal / abnormal threshold ranges; and recommended response suggestions or historical protection schemes.
[0073] Here, through deep integration (Retrieval-Augmented Generation, RAG) technology, a cybersecurity knowledge base containing threat patterns and predefined normal traffic thresholds from the MITREATT&CK framework is combined with a locally deployed large language model, significantly enhancing the model's semantic understanding and professional reasoning capabilities for complex cybersecurity scenarios.
[0074] Knowledge is stored in structured text format to ensure consistency in subsequent semantic vector calculations and retrieval. This includes threat patterns described in the MITRE ATT&CK framework (such as "DDoS attack characteristics: a large number of SYN requests in a short period, accompanied by target server response delays or crashes") and normal traffic thresholds set based on actual enterprise network operation data (such as "peak traffic on the enterprise's internal network does not exceed 2Gbps; exceeding this may indicate an anomaly"). These thresholds can be dynamically expanded to incorporate the latest threat intelligence.
[0075] Furthermore, to support the vectorized representation of natural language semantics, a pre-trained semantic embedding model (e.g., BGE-Base, SBERT, MPNet, etc.) is selected to process each of the aforementioned textual threat knowledge items, encoding them into high-dimensional dense vectors. To accelerate retrieval, a vector indexing tool (such as FAISS) is then used to construct an embedded vector index, supporting fast and accurate similarity searches to meet the requirements of real-time response. The vector of each threat knowledge item is then saved as an index item in the database.
[0076] Here, the natural language network state description generated by the aforementioned modules is encoded using the same semantic embedding model to generate its corresponding semantic feature vector. Subsequently, this semantic feature vector is input into a vector retrieval system, where it is compared with the embedded vectors in the knowledge base based on the cosine similarity (or Euclidean distance) index between the vectors. Finally, the highest-scoring threat knowledge entries are selected as the target threat knowledge for the current network state.
[0077] For example, when the large language model receives a network state description (e.g., "vSwitch-1 detected 10Gbps of traffic at 14:30, with an average traffic of 2Gbps"), it first converts the description into a vector representation through an embedding model. Then, it retrieves the Top-K most relevant knowledge entries from the FAISS index, such as DDoS attack characteristics and normal traffic threshold information. These retrieval results, together with the original state description, constitute an enhanced context. This context is then input into the large language model through a carefully designed prompt template, guiding the model to perform multi-step reasoning and generate a decision strategy that meets network security requirements (e.g., "Abnormal traffic was detected; it is recommended to limit vSwitch-1 traffic to 2Gbps on firewall FW-1 and activate the defense mechanism").
[0078] It should be noted that, to ensure efficiency and robustness, the knowledge base is managed through a database or file system and supports regular updates; FAISS runs as an independent service or integrated module to optimize retrieval latency; and the large language model performs inference on a local server, balancing performance and security. This solution, by integrating external expertise with the generation capabilities of the large language model through RAG technology, not only significantly improves the accuracy and scenario adaptability of decision-making but also achieves low-latency real-time response, providing an innovative and highly scalable intelligent solution for addressing dynamic network threats.
[0079] S103. Combine the target threat knowledge with the natural language description to generate analysis prompt text, input it into the large language model, and execute the network threat analysis steps indicated in the preset thinking chain according to the analysis prompt text to determine the network anomaly attribution and the corresponding protection strategy.
[0080] In this step, after obtaining the natural language description of the target network environment and the target threat knowledge that matches its semantics, the anomaly attribution and protection strategy decision-making are further realized through the construction of the prompt text and the reasoning mechanism of the big language model.
[0081] In practice, the retrieved target threat knowledge is semantically fused with the natural language description to form enhanced contextual information. For example, if the natural language description is "a core switch received more than 1,000 ICMP requests within one minute," and the matching target threat knowledge is "short-term, high-frequency ICMP requests may indicate a Ping Flood attack," then the enhanced contextual information will reflect that "the current behavior is highly consistent with a Ping Flood attack, and further analysis is needed to determine if it is abnormal."
[0082] Specifically, target threat knowledge is combined with natural language descriptions to form context-enhanced information; the threat type corresponding to the target threat knowledge is determined, and a prompt text template matching the threat type is selected from the preset prompt text generation templates; the context-enhanced information is embedded into the prompt text template to generate analysis prompt text for prompting the analysis steps and reasoning basis of the large language model.
[0083] Here, multiple prompt text templates are pre-set, each corresponding to a specific type of network threat analysis task (such as denial-of-service analysis, permission anomaly analysis, configuration tampering analysis, etc.). Based on the threat type marked in the knowledge of the selected target threat, a matching prompt template is selected from the template library, and contextual enhancement information is embedded into the template to finally generate analysis prompt text used to guide the large language model in reasoning analysis.
[0084] For example, the prompt text could be as follows: "Based on the following context, please determine if there is a network anomaly and provide possible threat types, cause analysis, and protection recommendations. Context: A core switch received more than 1,000 ICMP requests within one minute. This behavior is uncommon under normal circumstances and may be related to a Ping Flood attack."
[0085] Furthermore, this solution introduces chain-thinking technology to guide the large language model to reason step by step according to structured logical steps, thereby improving the transparency and reliability of decision-making. The aforementioned analytical prompt text is input into the large language model (such as GPT, Claude, Baichuan, etc.), which performs a structured analysis of the current network state based on a pre-defined chain-thinking reasoning strategy.
[0086] For details on the logical reasoning process of the large language model, please refer to [link / reference]. Figure 2 The diagram shows another network security decision-making method based on a large language model provided in this disclosure. The method includes steps S1031 to S1033, wherein:
[0087] S1031. Determine whether the network status data deviates from the normal threshold indicated in the target threat knowledge. If the network status data deviates from the normal threshold, determine that the target network status is abnormal.
[0088] S1032. Determine the network anomaly attribution corresponding to the target network state based on the target threat knowledge, and determine whether it is a potential threat.
[0089] S1033. Based on the analysis results of the network anomaly attribution, generate the protection strategy under the prompt of the analysis prompt text.
[0090] In practice, the first step is to assess the network status to determine if it deviates from the normal threshold. For example, it checks whether "vSwitch-1 current traffic 10Gbps" exceeds the predefined normal threshold of 2Gbps. Next, anomaly attribution is performed, combining knowledge base information retrieved from RAG (such as DDoS attack characteristics from MITRE ATT&CK) to analyze the cause of the anomaly and determine if it is a potential threat (such as traffic flooding). Based on the analysis results, an executable protection policy is then generated, such as "limiting vSwitch-1 traffic to 2Gbps on firewall FW-1".
[0091] Here, during the reasoning process, the large language model records the intermediate outputs of each step, forming a complete thought chain log. For example, the model output might be: "The current traffic anomaly is caused by a Ping Flood attack. It is recommended to limit the number of ICMP packets through access control lists and enable firewall rate limiting mechanisms."
[0092] It should be noted that, to adapt to different network conditions and threat scenarios, the prompt template is adaptively adjusted to ensure that the output of the large language model meets the specific needs of the scenario. Based on the input network state type (such as traffic anomaly, protocol anomaly), the corresponding prompt template is automatically selected. For example, the template for traffic anomaly is: "Check if the traffic exceeds the threshold and determine if it is a DDoS attack." Knowledge base information retrieved by RAG (such as "normal traffic threshold 2Gbps") is embedded in the prompt text to enhance the reasoning basis of the large language model. For example: "Based on the knowledge base (normal threshold 2Gbps), analyze the abnormal reason for vSwitch-1 traffic of 10Gbps."
[0093] S104. The protection policy is sent to each distributed execution unit deployed in the target network environment, the protection policy is converted into a device configuration command, the device configuration command is executed and the changes in the network status data are monitored to determine the execution effect of the protection policy.
[0094] In this step, after the large language model outputs the protection policy, the policy is automatically transmitted to various distributed execution units in the target network environment, where it performs policy translation, command execution, and effect feedback monitoring. By designing an efficient decision-making distribution mechanism, intelligent agents, and a closed-loop feedback optimization process, the decision-making policy generated by the large language model is automatically deployed and continuously improved in remote network units.
[0095] In practical implementation, the protection policy is transmitted to the intelligent agent, which acts as a distributed execution unit, through a preset protocol, and the protection policy is converted into policy instructions that can be recognized by network units in the target network environment. The intelligent agent parses the policy instructions into device configuration commands that can be executed by local network devices. The integrity of the device configuration commands is verified and the device configuration commands are executed, and the operation logs corresponding to the device configuration commands are recorded.
[0096] Here, the generated protection policy is sent to distributed execution units deployed in the target network environment via a preset communication protocol (such as gRPC, MQTT, or HTTP API). Execution units can be intelligent agent nodes with certain local processing capabilities, such as edge computing devices, security middleware, or network management modules. A standard data packet format is defined, limiting the data packet size to within 1KB to ensure transmission efficiency in low-bandwidth environments; TLS encryption is applied before decision issuance to ensure data confidentiality and integrity during transmission.
[0097] Furthermore, after receiving the protection policy, the intelligent agent converts the policy into configuration commands supported by the device, based on the model, interface characteristics, and operating system of the local network device (such as Cisco IOS, Juniper Junos, Huawei VRP, etc.). For example, for the policy "Restrict external access to port 22", the agent verifies the integrity of the instruction before execution (such as checking whether JSON fields are missing) and records the operation log, including the execution timestamp, instruction content, and execution status (success / failure).
[0098] Here, the generated configuration commands, after passing authorization and integrity verification, are directly sent to the target network device for execution. The intelligent agent records execution logs and returns the execution status (success, failure, exception) and related feedback information (such as device response time, status code, etc.) to ensure the correct implementation of the protection policy.
[0099] As one possible implementation, to achieve collaborative execution in complex scenarios, multiple intelligent agents share execution status and resource information via the MQTT protocol and optimize policy deployment based on a distributed negotiation mechanism. When a policy involves multi-device linkage (e.g., "limit vSwitch-1 traffic and adjust vRouter routing"), the relevant intelligent agents negotiate a step-by-step execution plan to ensure configuration synchronization. In high-load scenarios, intelligent agents dynamically allocate execution tasks, for example, transferring complex instruction computation tasks to low-load intelligent agents. Intelligent agents use a heartbeat mechanism to detect the execution progress of neighboring intelligent agents; if inconsistencies are detected (e.g., an intelligent agent fails to execute), a collaborative retry or policy adjustment is triggered.
[0100] Here, to overcome the inefficiencies of centralized processing and insufficient global coordination in traditional network security systems, intelligent agents are introduced as distributed intelligent units, endowing them with innovative capabilities for autonomous collaborative optimization. Deployed at key nodes in the virtual network (such as virtual switches, virtual routers, or virtual firewalls), intelligent agents form a closed-loop decision-making system with the large language model on the local server. Through real-time collaboration among multiple agents, resource allocation, task scheduling, and policy execution consistency are optimized, providing intelligent support for security protection in dynamic network environments.
[0101] The intelligent agent employs a modular design and is deployed at key nodes in the virtual network to ensure distributed processing capabilities. Each intelligent agent contains a core collaboration optimization module, supplemented by necessary instruction parsing and state awareness functions, supporting efficient operation through a lightweight architecture. The intelligent agent communicates with the local server and other intelligent agents via the MQTT protocol, forming a distributed awareness and execution network that significantly improves overall performance.
[0102] Furthermore, after the strategy is executed, a real-time monitoring mechanism is activated to re-collect network status data, including the communication behavior of the target protected object, traffic changes of relevant interfaces, and changes in system logs. By comparing the status with that before protection, the actual effectiveness of the strategy can be evaluated. A time-series database is used to store execution logs, recording the correlation between decision instructions, execution results, and feedback data; the data format supports efficient querying and analysis, providing support for subsequent model optimization.
[0103] Here, if the protection effect is significant (such as abnormal traffic being blocked or attack behavior being interrupted), the policy will be marked as effective; if the effect is not obvious or side effects are introduced (such as legitimate access being blocked), the policy generation module will be notified through the feedback mechanism to re-analyze or generate an improved version of the protection policy.
[0104] As another possible implementation, the network security decision-making method based on a large language model provided in this application can be tested in advance in a virtual network environment. Specifically, PNETLab is used as the core virtualization platform to construct an experimental network environment with a multi-layered topology. Distributed data acquisition, autonomous anomaly detection, decision-making interaction with the local large language model, and collaborative optimization among multiple intelligent agents are achieved by deploying intelligent agents. This virtual environment supports network attack simulation and real-time data acquisition, providing a controllable and realistic experimental foundation for subsequent decision generation based on the large language model.
[0105] Here, a virtual network environment is built using the PNETLab platform. By configuring virtual network devices such as virtual switches (vSwitch), virtual routers (vRouter), and virtual firewalls (vFirewall), a real network architecture is simulated. Various predefined topologies are designed and implemented, including but not limited to star topologies, ring topologies, and hybrid topologies. Automated configuration templates are provided to support dynamic adjustment of network nodes, ensuring that the experimental environment can flexibly adapt to different test scenarios.
[0106] Specifically, each intelligent agent broadcasts and subscribes to network status information (such as traffic load and anomaly alarms) in real time via the MQTT protocol, building a distributed awareness network to ensure the rapid dissemination of global threat information. When an intelligent agent detects an anomaly (such as "vSwitch-1 traffic suddenly increased to 10Gbps"), its collaborative optimization module immediately broadcasts an alarm message via an MQTT topic (such as " / agent / status"), triggering joint analysis by neighboring intelligent agents. Information sharing uses a lightweight JSON format, containing key fields (such as node ID, status type, and timestamp), with the data packet size controlled within 512 bytes to adapt to low-bandwidth environments.
[0107] Optionally, a distributed task scheduling algorithm is employed to dynamically adjust the scheduling strategy based on task urgency and node load. For example, when a strategy requires "limiting vSwitch-1 traffic and synchronously adjusting vFirewall rules," the task is decomposed into subtasks (rate limiting, rule updates), and the execution order is allocated according to the real-time status of each Agent, prioritizing traffic control for critical nodes. Scheduling decisions are synchronized via the MQTT protocol to ensure consistency.
[0108] Optionally, a heartbeat synchronization mechanism (sending a status report via MQTT every 30 seconds) can be designed to monitor the running status of neighboring agents. If an agent fails to execute (e.g., vFirewall fails to update rules), the module triggers a coordinated retry, prioritizing low-load agents to re-execute the task. Abnormal states are logged for subsequent analysis.
[0109] Furthermore, using network testing tools such as Scapy, typical network threats such as distributed denial-of-service (DDoS) attacks, port scans, and packet forgery are generated in a virtual network environment to verify threat detection and response capabilities. Network security risks are introduced through manual configuration, such as opening unauthorized ports or setting overly lenient firewall rules, to simulate potential vulnerabilities caused by human error in real networks. The triggering conditions of attacks, network response behaviors, and abnormal traffic data are recorded to ensure that the data acquisition module can comprehensively capture key information and provide support for optimizing the input data quality of large language models.
[0110] This disclosure provides a network security decision-making method based on a large language model (LLM). The method involves collecting network state data corresponding to a target network environment and converting the network state data into a natural language description with contextual semantics. It then determines the semantic feature vector corresponding to the natural language description and retrieves target threat knowledge matching the semantic feature vector from a knowledge base storing various basic network threat knowledge. The method combines the target threat knowledge with the natural language description to generate analysis prompt text, which is then input into the LLM. Based on the analysis prompt text, the method sequentially executes network threat analysis steps indicated in a preset thought chain to determine the network anomaly attribution and corresponding protection strategy. The protection strategy is then distributed to each distributed execution unit deployed in the target network environment, converting the protection strategy into device configuration commands. These commands are executed, and changes in the network state data are monitored to determine the effectiveness of the protection strategy. Applying the semantic understanding, reasoning, and contextual learning capabilities of a large language model (LLM) to remote network unit decision-making enables a closed-loop intelligent defense process from information perception to semantic understanding to collaborative response, improving the real-time performance, accuracy, and scalability of security decisions.
[0111] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0112] Based on the same inventive concept, this disclosure also provides a network security decision-making device based on a large language model, which corresponds to the network security decision-making method based on a large language model. Since the principle of the device in this disclosure is similar to the network security decision-making method based on a large language model described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0113] Please see Figure 3 , Figure 3 This is a schematic diagram of a network security decision-making device based on a large language model, provided as an embodiment of this disclosure. Figure 3 As shown in the illustration, the network security decision-making device 300 based on a large language model provided in this embodiment includes:
[0114] The natural language description conversion module 310 is used to collect network state data corresponding to the target network environment and convert the network state data into a natural language description with contextual semantics.
[0115] The retrieval enhancement module 320 is used to determine the semantic feature vector corresponding to the natural language description, and to retrieve target threat knowledge that matches the semantic feature vector from a knowledge base storing various basic network threat knowledge.
[0116] The protection strategy analysis module 330 is used to combine the target threat knowledge with the natural language description to generate analysis prompt text, which is then input into the large language model. Based on the analysis prompt text, the module sequentially executes the network threat analysis steps indicated in the preset thought chain to determine the network anomaly attribution and the corresponding protection strategy.
[0117] The policy execution module 340 is used to send the protection policy to each distributed execution unit deployed in the target network environment, convert the protection policy into a device configuration command, execute the device configuration command and monitor the changes in the network status data to determine the execution effect of the protection policy.
[0118] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.
[0119] This disclosure provides a network security decision-making device based on a large language model (LLM). The device collects network status data corresponding to a target network environment and converts the network status data into a natural language description with contextual semantics. It determines the semantic feature vector corresponding to the natural language description and retrieves target threat knowledge matching the semantic feature vector from a knowledge base storing various basic network threat knowledge. It combines the target threat knowledge with the natural language description to generate analysis prompt text, which is then input into the LLM. Based on the analysis prompt text, it sequentially executes network threat analysis steps indicated in a preset thought chain to determine the network anomaly attribution and corresponding protection strategy. The protection strategy is then distributed to each distributed execution unit deployed in the target network environment, converting the protection strategy into device configuration commands. The device configuration commands are executed, and changes in the network status data are monitored to determine the effectiveness of the protection strategy. Applying the semantic understanding, reasoning, and contextual learning capabilities of a large language model (LLM) to remote network unit decision-making enables a closed-loop intelligent defense process from information perception—semantic understanding—coordinated response, improving the real-time performance, accuracy, and scalability of security decisions.
[0120] Corresponding to Figure 1 and Figure 2 The present disclosure also provides an electronic device 400, such as a network security decision-making method based on a large language model. Figure 4 The diagram shown is a structural schematic of an electronic device 400 provided in an embodiment of this disclosure, including:
[0121] Processor 41, memory 42, and bus 43; memory 42 is used to store execution instructions, including main memory 421 and external memory 422; the main memory 421, also called internal memory, is used to temporarily store the computational data in processor 41, as well as the data exchanged with external memory 422 such as hard disk. Processor 41 exchanges data with external memory 422 through main memory 421. When the electronic device 400 is running, processor 41 and memory 42 communicate through bus 43, enabling processor 41 to execute... Figure 1 and Figure 2 The steps of the network security decision-making method based on a large language model.
[0122] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the network security decision-making method based on a large language model as described in the above-described method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0123] This disclosure also provides a computer program product, which includes computer instructions. When the computer instructions are executed by a processor, they can perform the steps of the network security decision-making method based on a large language model as described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.
[0124] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0125] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0126] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0127] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0128] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0129] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. A network security decision-making method based on a large language model, characterized in that, include: Collect network status data corresponding to the target network environment, and convert the network status data into a natural language description with contextual semantics; Determine the semantic feature vector corresponding to the natural language description, and retrieve target threat knowledge that matches the semantic feature vector from a knowledge base that stores various basic network threat knowledge; The target threat knowledge and the natural language description are combined to generate analysis prompt text, which is then input into the large language model. The network threat analysis steps indicated in the preset thinking chain are executed sequentially according to the analysis prompt text to determine the network anomaly attribution and the corresponding protection strategy. The protection policy is sent to each distributed execution unit deployed in the target network environment, the protection policy is converted into a device configuration command, the device configuration command is executed and the changes in the network status data are monitored to determine the execution effect of the protection policy.
2. The method according to claim 1, characterized in that, Collect network status data corresponding to the target network environment, specifically including: The Syslog service deployed on each network device in the target network environment records network operation events in real time. The network device's structure and performance attribute information is collected using a simple network management protocol, and abnormal traffic data during network operation is marked and captured using a network data collection and analysis tool. Extract the expanded data corresponding to the preset key network indicators and the context information corresponding to the network operation events and the abnormal traffic data; The network status data is generated by formatting the network operation events, the structure and performance attribute information, and the abnormal traffic data.
3. The method according to claim 1, characterized in that, Determine the semantic feature vector corresponding to the natural language description, and retrieve target threat knowledge matching the semantic feature vector from a knowledge base storing various basic network threat knowledge, specifically including: Construct the knowledge base, in which threat feature descriptions and corresponding normal thresholds for various basic network threats are pre-stored in text form as the basic network threat knowledge; Each piece of basic network threat knowledge is transformed into a high-dimensional vector embedding through a pre-trained embedding model, and an index corresponding to each high-dimensional vector embedding is constructed using a vector retrieval tool. The natural language description is converted into the semantic feature vector through the embedding model, and the target threat knowledge that matches the semantic feature vector is retrieved from the index.
4. The method according to claim 1, characterized in that, The analysis prompt text generated by combining the target threat knowledge with the natural language description is input into the large language model, specifically including: The target threat knowledge is combined with the natural language description to form context-enhanced information; Determine the threat type corresponding to the target threat knowledge, and select a prompt text template that matches the threat type from the preset prompt text generation templates; The contextual enhancement information is embedded into the prompt text template to generate the analysis prompt text used to prompt the analysis steps and reasoning basis of the large language model.
5. The method according to claim 1, characterized in that, Based on the analysis prompt text, the network threat analysis steps indicated in the preset thought chain are executed sequentially to determine the network anomaly attribution and corresponding protection strategies, specifically including: Determine whether the network status data deviates from the normal threshold indicated in the target threat knowledge; if the network status data deviates from the normal threshold, then determine that the target network status is abnormal. Based on the target threat knowledge, determine the network anomaly attribution corresponding to the target network state, and determine whether it is a potential threat; Based on the analysis results of the network anomaly attribution, the protection strategy is generated under the prompt of the analysis prompt text.
6. The method according to claim 1, characterized in that, The protection policy is distributed to each distributed execution unit deployed in the target network environment, the protection policy is converted into a device configuration command, and the device configuration command is executed, specifically including: The protection policy is transmitted to the intelligent agent, which is the distributed execution unit, through a preset protocol, and the protection policy is converted into a policy instruction that can be recognized by the network unit in the target network environment; The intelligent agent parses the policy instructions into device configuration commands that can be executed by the local network device. Verify the integrity of the device configuration command and execute the device configuration command, and record the operation log corresponding to the device configuration command.
7. The method according to claim 1, characterized in that, Converting the network state data into a natural language description with contextual semantics specifically includes: Based on the Jinja2 template engine and Python preprocessing logic, the network state data is mapped to the natural language description; The syntax and semantics of the natural language description are optimized using the T5-small model, and the natural language description is transformed into a structured input format that can be directly processed by the large language model.
8. A network security decision-making device based on a large language model, characterized in that, include: The natural language description conversion module is used to collect network state data corresponding to the target network environment and convert the network state data into a natural language description with contextual semantics. The retrieval enhancement module is used to determine the semantic feature vector corresponding to the natural language description, and to retrieve target threat knowledge that matches the semantic feature vector from a knowledge base that stores various basic network threat knowledge. The protection strategy analysis module is used to combine the target threat knowledge with the natural language description to generate analysis prompt text, which is then input into the large language model. Based on the analysis prompt text, the module sequentially executes the network threat analysis steps indicated in the preset thought chain to determine the network anomaly attribution and the corresponding protection strategy. The policy execution module is used to distribute the protection policy to each distributed execution unit deployed in the target network environment, convert the protection policy into device configuration commands, execute the device configuration commands, and monitor changes in the network status data to determine the execution effect of the protection policy.
9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the network security decision-making method based on a large language model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the network security decision-making method based on a large language model as described in any one of claims 1 to 7.
Citation Information
Cited By
Mobile security protection method and system based on AI dialogue and context awareness
CN121093338A
An AI conversation and context-aware based mobile security protection method and system
CN121093338B
Network security event handling method and system based on knowledge consistency verification
CN121441642A
Systems and methods for automatic security rule generation
US20260122108A1