Network equipment vulnerability mining method based on large model and knowledge graph

By applying large models and knowledge graphs in network device vulnerability mining, the problems of traditional methods are solved, with low efficiency, limited coverage and strong knowledge dependence, and efficient, accurate and automated vulnerability mining is achieved to adapt to the ever-changing network security environment.

CN120151033APending Publication Date: 2025-06-13BEIJING INST OF COMP TECH & APPL
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510298006.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional network equipment vulnerability mining methods are inefficient, have limited coverage, and have strong knowledge dependence, making it difficult to achieve efficient, accurate and automated vulnerability mining.

Method used

Using methods based on big models and knowledge graphs, efficient, accurate and automated vulnerability mining is achieved through steps such as data collection, knowledge graph construction, data preprocessing, training of big models, designing RAG modules and large model analysis.

Benefits of technology

It realizes efficient, accurate and automated vulnerability mining of network equipment, significantly improves vulnerability mining efficiency, ensures the accuracy and comprehensiveness of vulnerability mining results, and adapts to the ever-changing network security environment.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention relates to a network equipment vulnerability mining method based on a large model and a knowledge graph, and belongs to the technical field of network security. According to the method, efficient, accurate and automatic vulnerability mining is realized by combining the semantic comprehension capability of a large model, the accurate information retrieval capability of the RAG and the structured knowledge representation capability of the knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security, and particularly relates to a method for network device vulnerability mining based on large models and knowledge graphs. Background Art

[0002] Network devices (such as routers, switches, firewalls, etc.) are core components of network infrastructure, and their security directly affects the security of the entire network. The following problems exist in traditional vulnerability mining methods:

[0003] 1. Low efficiency: relying on manual analysis, time-consuming and prone to missing vulnerabilities.

[0004] 2. Limited coverage: difficult to cope with complex network device systems and diverse attack scenarios.

[0005] 3. Strong knowledge dependence: requiring the experience of security experts and difficult to achieve automation.

[0006] The development of large models, RAG, and knowledge graph technologies provides new ideas for solving the above problems, but they have not been fully applied in network device vulnerability mining. Summary of the Invention

[0007] (I) Technical Problems to be Solved

[0008] The technical problem to be solved by the present invention is: to provide a method for network device vulnerability mining to achieve efficient, accurate, and automated vulnerability mining.

[0009] (II) Technical Solutions

[0010] To solve the above technical problems, the present invention provides a method for network device vulnerability mining based on large models and knowledge graphs, including the following steps:

[0011] 1. Knowledge Graph Construction

[0012] Data collection: Widely collect detailed information of various network devices from multiple channels, such as authoritative cybersecurity databases, official websites of device manufacturers, open-source vulnerability platforms, professional technical forums, and previous security audit reports. It covers the hardware specifications of network devices (such as chipset information, interface types and quantities, etc.), software versions (including operating system kernel versions, firmware versions, etc.), network protocol support (such as specific implementation details of the TCP / IP protocol stack, types of supported routing protocols, etc.), common configuration methods (such as default configuration parameters, typical enterprise-level configuration cases, etc.), known security vulnerabilities (including vulnerability numbers, vulnerability descriptions, affected device version ranges, vulnerability discovery times, and repair measures, etc.), and security patch release records. Collect vulnerability data related to network devices (such as CVE, NVD databases), device configuration information, protocol specifications (such as TCP / IP, HTTP), attack patterns (such as MITRE ATT&CK framework). Obtain security announcements, patch information, and device manuals released by device manufacturers.

[0013] The traditional static construction mode of knowledge graphs is difficult to adapt to the evolutionary characteristics of zero-day vulnerabilities. Three categories are added on the basis of traditional entity relationships: one is the time dimension attribute, including the active period of vulnerabilities and the effective time window of patches; the second is the spatial topology attribute, including the location characteristics of devices in the network architecture; the third is the dynamic propagation weight, including the attack path probability based on the graph attention mechanism.

[0014] Knowledge extraction and graph construction: Use natural language processing technology and knowledge extraction tools to process the collected data. Identify and extract entities therein, such as device names, software components, vulnerability names, protocol names, etc., and at the same time extract the relationships between them, such as the ownership relationship between devices and software versions, the association relationship between vulnerabilities and affected devices, the support relationship between protocols and devices, etc. Use graph database or semantic web technology to construct a network device knowledge graph with nodes representing entities and edges representing relationships, and add corresponding attributes to nodes and edges to enrich the semantic information of the knowledge graph. Use a graph database (such as Neo4j) to construct the knowledge graph, define entity types (such as devices, vulnerabilities, protocols, attack patterns) and their relationships (such as "devices have vulnerabilities", "vulnerabilities exploit protocols"). Extract entities and relationships from unstructured data (such as vulnerability reports) through natural language processing (NLP) technology and store them in the knowledge graph.

[0015] Design an incremental update mechanism for the knowledge graph to reduce repeated calculations.

[0016] 2. Data collection and preprocessing

[0017] Data source integration: In addition to the data sources collected during the above knowledge graph construction process, it also includes collecting data from the actual operating environment of network devices, such as obtaining real-time traffic data of network devices through network traffic monitoring tools, collecting system logs and event logs of devices using system log management tools, and obtaining information such as port opening status and service running status of devices through legal network scanning techniques. At the same time, comprehensively collect text materials such as user manuals, technical white papers, and online technical support documents of network devices to ensure the integrity and diversity of data.

[0018] Multimodal data processing:

[0019] Firmware binary analysis: Extract control flow graph features through a disassembler engine

[0020] Protocol traffic parsing: Generate a protocol state machine based on deep packet inspection (DPI)

[0021] Configuration semantic modeling: Convert CLI commands into an abstract syntax tree (AST)

[0022] Data cleaning and standardization: Perform cleaning and standardization operations on data with different formats and qualities. Remove noise information in text data, such as garbled characters, special characters, and redundant whitespace; perform structuring on unstructured text, for example, convert tabular data into a unified format; perform normalization on numerical data to make the same type of data in different data sources comparable; uniformly encode various data into a format suitable for subsequent processing, such as UTF-8 encoding.

[0023] Text tokenization and vectorization: Use professional text tokenization tools to tokenize the preprocessed text data, splitting the text into basic language units such as words or phrases. Then, use pre-trained word vector models (such as Word2Vec, GloVe, etc.) or deep learning-based language models (such as BERT, etc.) to convert the tokenized text into vector representations for efficient semantic calculation and similarity matching in subsequent RAG retrieval and large model analysis.

[0024] 3. Train a large model

[0025] Pre-train the large model using a dataset related to network device vulnerabilities (such as CVE descriptions, vulnerability reports, device logs).

[0026] Introduce domain-specific data (such as device configuration templates, protocol specifications) for fine-tuning to improve the model's semantic understanding ability in the field of network devices.

[0027] Use the Transformer architecture to train the large model and optimize its capabilities in vulnerability description generation, semantic parsing, and context understanding.

[0028] Introduce contrastive learning techniques to enhance the model's ability to distinguish between similar vulnerabilities.

[0029] 4. Design a Retrieval-Augmented Generation (RAG) module

[0030] Question generation and vector conversion: According to the goals and requirements of network device vulnerability mining, formulate a series of targeted questions, such as "Potential vulnerabilities of [device model X] when interacting with [a certain network protocol Z] under [specific software version Y]" "[Known security issues and unfixed vulnerabilities of [device name A]

[0031] under [specific configuration parameter B] settings". Input these questions into a pre-trained language model, and through the encoding layer of the model, convert the questions into vector representations so that they can perform semantic matching with text data in the vector space.

[0032] Vector indexing and retrieval: Utilize vector databases (such as Milvus, Faiss, etc.) to index and store the preprocessed text vector data, and construct an efficient vector index structure. When the question vectors are generated, perform similarity retrieval in the vector database, using measurement methods such as cosine similarity and Euclidean distance to find the text fragments corresponding to the text vectors most similar to the question vectors. These retrieved text fragments will serve as important reference materials for the large model analysis and may contain key information related to network device vulnerabilities, such as vulnerability cases of similar devices, discussions on security risks under specific configurations, analysis reports on network protocol vulnerabilities, etc. Using the knowledge graph as an external knowledge base, the RAG module retrieves entities and relationships related to the input information through a graph query language (such as Cypher). Example: Input "Router configuration vulnerability", the RAG module retrieves vulnerability patterns, attack paths, and repair suggestions related to routers in the knowledge graph.

[0033] Design a Dynamic Retrieval Strategy Optimizer (DRSO) based on reinforcement learning to achieve intelligent path planning for the RAG module;

[0034] Generation mechanism: The large model combines the retrieval results to generate vulnerability descriptions, exploitation suggestions, and repair solutions. Example: Generate "This router has a default configuration vulnerability. Attackers can exploit this vulnerability to obtain administrator privileges. It is recommended to modify the default password and enable access control."

[0035] 5. Large model analysis and reasoning

[0036] Multi-source data fusion input: Integrate the text snippets retrieved by RAG and the entities, relationships, and attribute information related to the problem queried from the knowledge graph as input data for the large model. The large model can be a language model based on the Transformer architecture (such as the GPT series, etc.). Through its powerful multi-head attention mechanism, it can focus on different parts of the input text at the same time, capture the semantic associations and logical relationships between texts, and thus conduct a comprehensive and in-depth analysis of the vulnerabilities of network devices.

[0037] Vulnerability pattern recognition and reasoning: The large model uses its pre-trained knowledge and reasoning ability on large-scale corpus to analyze the input fusion data. Identify possible vulnerability patterns in network equipment in terms of software code implementation, protocol processing logic, configuration parameter settings, etc. For example, by analyzing the text description of the network device configuration file and the vulnerability information caused by improper configuration of similar devices in the knowledge graph, it is possible to infer the permission bypass vulnerability that may exist in the current device under a specific configuration; or based on the specification text of the network protocol and the anomalies in the actual traffic data, the possibility of buffer overflow vulnerabilities in the protocol implementation process can be inferred. At the same time, the large model can also combine known attack methods and security vulnerability types to predict the vulnerability of network equipment in the face of new attacks and explore potential security risk points.

[0038] 6. Vulnerability mining process

[0039] Input stage:

[0040] Users provide configuration information, log data, or protocol descriptions of network devices.

[0041] Example: Enter a router's configuration file or a firewall's log file.

[0042] Semantic analysis:

[0043] The large model performs semantic analysis on the input data and extracts key information (such as device type, protocol version, and configuration parameters).

[0044] Example: Parse out "Device type: Cisco router, Protocol: SNMPv1, Configuration: Default community string is public".

[0045] Knowledge retrieval:

[0046] The RAG module retrieves vulnerability patterns, attack paths, and repair suggestions related to the input information from the knowledge graph.

[0047] Example: Retrieve "SNMPv1 default community string vulnerability, attackers can exploit this vulnerability to obtain device information."

[0048] Vulnerability Generation:

[0049] The large model combines the retrieval results to generate a potential vulnerability list, and gives vulnerability exploitation suggestions and repair solutions.

[0050] Example: Generate "Vulnerability: SNMPv1 default community string vulnerability, Risk level: High, Repair suggestion: Modify the default community string and upgrade to SNMPv3".

[0051] Output phase:

[0052] Output a vulnerability report, including vulnerability description, risk level, repair suggestions, etc.

[0053] Example: Output "Vulnerability report: The device has an SNMPv1 default community string vulnerability, Risk level: High, Repair suggestion: Modify the default community string and upgrade to SNMPv3".

[0054] 7. Iterative optimization

[0055] Feedback mechanism:

[0056] Dynamically update the knowledge graph and the large model according to the vulnerability mining results and user feedback.

[0057] Example: If a user feedbacks that a certain vulnerability is not recognized, the system adds the vulnerability information to the knowledge graph.

[0058] Model optimization:

[0059] Use incremental learning technology to update the large model and improve its performance in new vulnerability scenarios.

[0060] Example: Fine-tune the model for new vulnerabilities (such as zero-day vulnerabilities).

[0061] 8. Vulnerability verification and report generation

[0062] Vulnerability verification experiment design: According to the potential vulnerability situation analyzed by the large model, design targeted vulnerability verification experiments. This includes constructing specific network packets, simulating abnormal user operations or configuration scenarios, and making preliminary attack attempts using vulnerability exploitation tools. During the experiment, closely monitor the changes in indicators such as the running status of network devices, system logs, and network traffic to determine whether potential vulnerabilities can be triggered and observe the impacts they generate.

[0063] Vulnerability verification and assessment: If abnormal behaviors occur in network devices during the verification experiment, such as system crashes, service interruptions, illegal data access, etc., further analyze the causes, exploitation conditions, and impact scope of the vulnerabilities. Evaluate the severity of the vulnerabilities, refer to standards such as the Common Vulnerability Scoring System (CVSS), and quantitatively score the vulnerabilities from multiple dimensions such as exploitability, impact scope and degree, and repair difficulty to determine their risk levels.

[0064] Vulnerability report generation: Based on the results of vulnerability verification and assessment, a detailed vulnerability report is generated. The report content includes the discovery process of the vulnerability, the specific location (such as the software module name, the parameter location in the configuration file, the specific layer of the network protocol stack, etc.), the vulnerability type (such as buffer overflow, SQL injection, privilege escalation, etc.), the scope of network devices affected (including device models, software versions, network topology locations, etc.), the potential impact of the vulnerability (such as data leakage risk, system paralysis possibility, network intrusion consequences, etc.), and detailed repair suggestions (such as software update version numbers, configuration parameter modification methods, temporary protection measures, etc.), so that network security personnel can quickly and accurately understand the vulnerability situation and take effective repair and prevention measures.

[0065] The present invention also provides a system for implementing the above method.

[0066] The present invention also provides a network security analysis method implemented based on the above method.

[0067] (III) Beneficial effects

[0068] The present invention provides a network device vulnerability mining method based on the combined use of large models, RAG, and knowledge graphs. By combining the semantic understanding ability of large models, the accurate information retrieval ability of RAG, and the structured knowledge representation ability of knowledge graphs, efficient, accurate, and automated vulnerability mining is achieved. The method has the following advantages:

[0069] 1. High efficiency: Significantly improve the vulnerability mining efficiency through an automated process.

[0070] 2. Accuracy: The knowledge graph provides structured knowledge support to ensure the accuracy of the vulnerability mining results.

[0071] 3. Comprehensiveness: Cover a variety of network device types and vulnerability scenarios to reduce omissions.

[0072] 4. Scalability: Support dynamic updating of the knowledge graph and large models to adapt to the constantly changing network security environment. Specific implementation manners

[0073] To make the objectives, content, and advantages of the present invention clearer, the following further describes in detail the specific implementation manners of the present invention in conjunction with embodiments.

[0074] The present invention provides a method for network device vulnerability mining by jointly using a large language model (LLM), retrieval-augmented generation (RAG), and a knowledge graph (KG). By combining the semantic understanding ability of the large language model, the accurate information retrieval ability of RAG, and the structured knowledge representation ability of the knowledge graph, efficient, accurate, and automated vulnerability mining is achieved.

[0075] The present invention provides a method for network device vulnerability mining based on a large language model and a knowledge graph, which generally includes the following steps: constructing a network device vulnerability knowledge graph; training a large language model; designing a RAG module; implementing a vulnerability mining process; and iteratively optimizing the system. The knowledge graph includes entities such as vulnerabilities, devices, protocols, attack patterns, etc. and their relationships. The RAG module enhances the generation ability of the large language model by retrieving the knowledge graph. The vulnerability mining process includes semantic parsing, knowledge retrieval, vulnerability generation, and report output. This method is implemented through the following system:

[0076] 1. System architecture:

[0077] The system consists of a large language model module, a RAG module, a knowledge graph module, and a user interface module; the large language model module is responsible for semantic parsing and vulnerability generation; the RAG module is responsible for retrieving relevant information from the knowledge graph; the knowledge graph module stores structured knowledge related to network device vulnerabilities to form a knowledge graph; the user interface module provides input-output interaction functions.

[0078] 2. Vulnerability mining process:

[0079] The user inputs network device information.

[0080] The large language model module parses the input information and generates preliminary vulnerability hypotheses.

[0081] The RAG module retrieves relevant vulnerability patterns from the knowledge graph.

[0082] The large language model module generates a final vulnerability report in combination with the retrieval results.

[0083] 3. Iterative optimization:

[0084] According to the vulnerability report and user feedback, update the knowledge graph and the large language model to improve the system performance.

[0085] Example: Conduct vulnerability mining on an enterprise-level firewall of a certain brand

[0086] 1. Knowledge graph construction stage:

[0087] Obtain detailed hardware architecture drawings, firmware update instructions for different versions, and various network protocol specification documents supported by the device from the official technical documentation library of this firewall. For example, it is learned that this firewall uses a specific model of processor, its firmware version has gone through multiple iterations, and it supports various routing protocols such as OSPF and BGP, as well as common VPN protocols such as IPsec and SSL / TLS.

[0088] Search for all known vulnerability information related to this firewall in authoritative network security vulnerability databases (such as the CVE database), including detailed descriptions of vulnerabilities, the range of affected firmware versions, and corresponding repair measures. At the same time, collect vulnerability analysis reports and discussion posts of security researchers on this firewall from professional security research forums, and extract valuable information such as the causes and exploitation methods of vulnerabilities.

[0089] Use knowledge extraction tools to extract entities (such as firewall devices, firmware versions, network protocols, vulnerability numbers, security patches, etc.) and relationships (such as the firewall uses a specific firmware version, a specific vulnerability affects a specific version of the firewall, a security patch fixes a specific vulnerability, etc.) from the above information. Use the graph database Neo4j to build a knowledge graph, with the firewall device as a node, firmware versions, network protocols, vulnerabilities, etc. as related nodes, and represent their relationships through edges, and add detailed attribute information to each node and edge, such as the severity of the vulnerability, the version number of the protocol, the release time of the firmware, etc.

[0090] 2. Data collection and preprocessing stage:

[0091] In the enterprise's network environment, use professional network traffic monitoring devices (such as Snort, Suricata, etc.) to capture and store the in-and-out traffic of this firewall in real time for one week, ensuring that the network traffic data in different time periods and business scenarios is covered. At the same time, use the log export function of the firewall to collect various log files such as system logs, access control logs, and intrusion detection logs.

[0092] Obtain the initial configuration file of this firewall and the configuration backup files after multiple modifications from the enterprise's internal IT document management system. In addition, search the Internet for user manuals, technical white papers, and relevant discussion posts in online technical support forums of this firewall, and summarize these text materials.

[0093] Clean the collected network traffic data, remove invalid data packets (such as duplicate, abnormally long, and protocol - non - compliant data packets), and perform protocol parsing and feature extraction on it to convert the traffic data into structured traffic feature vectors. For log files, perform denoising processing, remove irrelevant routine log information such as system startup and shutdown, extract key log entries related to security events, network connections, access control, etc., and classify and organize them. For text materials, perform text cleaning to remove garbled characters, special characters, and redundant whitespace, then use natural language processing tools for pre - processing operations such as word segmentation, part - of - speech tagging, and named entity recognition. Finally, use a pre - trained word vector model (such as the GloVe model) to convert the text into a vector representation for subsequent retrieval and analysis.

[0094] 3. RAG - based relevant information retrieval stage:

[0095] o According to the goals of vulnerability mining, formulate a series of questions, such as "Potential vulnerabilities of [Firewall Model X] under [Current Running Firmware Version Y] for [Internal Network to External Network Access Control Policy Z]" and "Known security issues and unpatched vulnerabilities of [Firewall Name A] when handling [Large - scale Concurrent SSL / TLS Connections]".

[0096] Input these questions into a pre - trained language model (such as the GPT - 3 model), and convert the questions into vector representations through the encoding layer of the model. Use the vector database Faiss to index and store the pre - processed text vector data, and construct an efficient vector index structure. Conduct similarity retrieval in Faiss, using cosine similarity as the metric method, to find the text fragments corresponding to the text vectors most similar to the question vector. For example, retrieve relevant text fragments such as case analysis reports on access control vulnerabilities of similar firewalls under the same firmware version and discussions on security risks caused by resource exhaustion when handling high - concurrent encrypted connections. These fragments will serve as important reference materials for subsequent large - model analysis.

[0097] 4. Large - model analysis and reasoning stage:

[0098] Integrate the text fragments retrieved by RAG and the entity, relationship, and attribute information related to the question queried from the knowledge graph as the input data for the large - model (such as a network security analysis model fine - tuned based on GPT - 3). The large - model conducts in - depth analysis of the input integrated data through its powerful multi - head attention mechanism.

[0099] For example, when analyzing the access control policy vulnerabilities of a firewall, the model combines the access control rule settings of the firewall in the knowledge graph, similar vulnerability cases retrieved by RAG, and the network security knowledge pre-trained by itself to infer that there may be a privilege bypass vulnerability in the current firewall under specific access control rule combinations. Because in some complex network topology and rule configuration scenarios, internal users may bypass the access restrictions of the firewall by constructing special network requests, thereby accessing unauthorized external resources. At the same time, the model also predicts the potential SYN flood attack vulnerability risk when processing these abnormal traffic patterns (such as a large number of SYN half-connection requests from a specific IP segment) in the network traffic data and the TCP protocol processing logic of the firewall in the knowledge graph. Although there is no clear report of such vulnerabilities at present, the model discovers potential security risk points through comprehensive analysis.

[0100] 5. Vulnerability verification and report generation phase:

[0101] According to the potential vulnerability situation analyzed by the large model, design targeted vulnerability verification experiments. For example, for the speculated privilege bypass vulnerability, construct a series of data packets simulating internal user network requests, send these data packets to the firewall in the experimental environment, and at the same time monitor the access control logs and network connection status of the firewall. For the possible SYN flood attack vulnerability, use network attack tools to generate a large number of SYN half-connection requests and send them to specific ports of the firewall, and observe the resource usage of the firewall (such as CPU utilization, memory occupancy, connection table status, etc.) and the availability of network services.

[0102] If abnormal behaviors are found in the firewall during the verification experiment, such as the access control logs record unexpected allowed access entries, or the network services of the firewall are interrupted or the response latency is severely affected during a SYN flood attack, further analyze the causes, exploitation conditions, and scope of influence of the vulnerability. Refer to the CVSS standard to evaluate and score the severity of the vulnerability. For example, for the privilege bypass vulnerability, if it can directly obtain access to the administrator privilege and affects the security of the entire enterprise internal network, rate its severity as high (above 9.0 points); for the SYN flood attack vulnerability, if it can cause the firewall service to be unavailable under medium attack intensity and affect the normal operation of some network services, rate it as medium (6.0 - 8.9 points).

[0103] Generate a detailed vulnerability report based on the results of vulnerability verification and assessment. The report details the process of discovering the vulnerability, including the initially raised questions, the key information retrieved by RAG, and the analysis and reasoning process of the large model. Clearly indicate the specific location of the vulnerability, such as a privilege bypass vulnerability in a certain rule parsing function of the access control module of the firewall, and a SYN flood vulnerability risk in the connection request handling part of the TCP protocol processing module. Elaborate on the vulnerability type (privilege bypass, SYN flood attack risk, etc.), the scope of the affected firewalls (including specific models, currently running firmware versions, and deployment locations in the enterprise network), the potential impacts (such as leakage of sensitive information within the enterprise, business losses due to network service interruptions, etc.), and provide detailed repair suggestions, such as immediately upgrading the firewall firmware to the latest version (providing specific firmware download links and upgrade steps), modifying the access control policy to avoid specific rule combination vulnerabilities (giving specific rule modification examples), optimizing the firewall parameters to enhance the resistance to SYN flood attacks (such as adjusting the connection timeout, increasing the connection table capacity, etc., specific parameter setting suggestions), so that the enterprise's network security personnel can quickly take effective measures to repair and prevent the vulnerability, ensuring the safe and stable operation of the enterprise network.

[0104] Through the above specific steps, it is possible to more systematically and comprehensively utilize the advantages of large models, RAG, and knowledge graphs, improve the efficiency and accuracy of network device vulnerability mining, discover more hidden security vulnerabilities, and provide strong support for network security protection. In practical applications, the parameters and technical means of each step can be further optimized and adjusted according to the specific network device types, application scenarios, and security requirements to adapt to different vulnerability mining tasks. This method can be widely applied in the field of network security to enhance the security of network devices.

[0105] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the technical principles of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A network device vulnerability mining method based on a large model and knowledge graph, characterized in that: The following steps are involved: Step 1: Knowledge graph construction Data collection: Collect information about various network devices, including hardware specifications, software versions, network protocol support, configuration methods, and security patch release records of network devices; Knowledge extraction and graph construction: Use natural language processing technology and knowledge extraction tools to process the collected data, identify and extract entities. Entity information includes device name, software component, vulnerability name, network protocol name, and extract the relationship between them, including the relationship between network devices and software versions, the relationship between vulnerabilities and affected devices, and the support relationship between network protocols and devices; use graph databases or semantic web technology to build these entities and relationships into a network device knowledge graph, where nodes represent entities and edges represent relationships, and add corresponding attributes to nodes and edges to enrich the semantic information of the knowledge graph; use graph databases to build knowledge graphs and define entity types and their relationships; extract entities and relationships from unstructured data through natural language processing technology and store them in knowledge graphs; Step 2: Data collection and preprocessing Data source integration: Collect data from the actual operating environment of network devices, including obtaining real-time traffic data of network devices through network traffic monitoring tools, collecting system logs and event logs of devices using system log management tools, and obtaining port opening status and service operation status information of devices through legal network scanning technology; Data cleaning and standardization: Clean and standardize data of different formats and qualities; remove noise information in text data, including garbled characters, special characters, and redundant blanks; structure unstructured text and convert tabular data into a unified format; normalize numerical data to make the same type of data in different data sources comparable; uniformly encode various types of data into a format suitable for subsequent processing, i.e., preprocessed text data; Text segmentation and vectorization: Use text segmentation tools to segment the preprocessed text data and split the text into basic language units such as words or phrases. Then, use a pre-trained word vector model or a language model based on deep learning to convert the segmented text into a vector representation, i.e., the preprocessed text vector data. Step 3: Train a large model Use datasets related to network device vulnerabilities to pre-train large models; Use domain-specific data to fine-tune the big model and improve its semantic understanding ability in the field of network devices; Use the Transformer architecture to train large models and optimize their capabilities in vulnerability description generation, semantic parsing, and context understanding; Use contrastive learning technology to enhance the ability of large models to distinguish similar vulnerabilities; Step 4: Design the retrieval enhancement generation module, i.e., the RAG module The RAG module includes question generation and vector conversion module, vector indexing and retrieval module, and information generation module; The question generation and vector conversion module is designed to formulate a series of targeted questions according to the goals and requirements of network device vulnerability mining, input these questions into the pre-trained large model, and convert the questions into vector representations through the encoding layer of the large model, so that it can be semantically matched with text data in the vector space; The vector index and retrieval module is designed as follows: using the vector database to index and store the preprocessed text vector data and construct a vector index structure; when the question vector is generated, similarity retrieval is performed in the vector database, using cosine similarity and Euclidean distance measurement methods to find the text fragments corresponding to the text vector that is most similar to the question vector. These retrieved text fragments will be used as reference materials for large model analysis; using the knowledge graph as an external knowledge base, the entities and relationships related to the input information are retrieved through the graph query language; The information generation module is designed to combine the large model with the vector index and the search results of the search module to generate vulnerability descriptions, exploit suggestions and repair solutions; Step 5: Large model analysis and reasoning Multi-source data fusion input: Integrate the text snippets retrieved in step 4 and the entity, relationship, and attribute information related to the question queried from the knowledge graph as input data for the large model; Vulnerability pattern recognition and reasoning: The large model uses its pre-trained knowledge and reasoning ability on large-scale corpus to analyze the input data and identify possible vulnerability patterns in network equipment in terms of software code implementation, protocol processing logic, and configuration parameter settings; Step 6: Execute the vulnerability mining process Input stage: Configuration information, log data or protocol descriptions of network devices provided by users; Semantic analysis: The large model performs semantic analysis on the input data and extracts key information, including device type, protocol version, and configuration parameters; Knowledge retrieval: Use the RAG module to retrieve vulnerability patterns, attack paths, and repair suggestions related to the input information from the knowledge graph; Vulnerability Generation: Combine the large model with the search results to generate a list of potential vulnerabilities, and provide vulnerability exploitation suggestions and repair solutions; Output stage: Output vulnerability report, including vulnerability description, risk level, and repair suggestions.

2. The method according to claim 1, characterized in that The method further comprises step 7, iterative optimization: Feedback mechanism design: Dynamically update the knowledge graph and large model based on vulnerability mining results and user feedback; Model optimization: Use incremental learning techniques to update large models and improve their performance in new vulnerability scenarios.

3. The method according to claim 1, characterized in that The method also includes step 8, vulnerability verification and report generation: Vulnerability verification experiment design: Based on the potential vulnerabilities obtained from the large model analysis, design targeted vulnerability verification experiments, including constructing specific network data packets, simulating abnormal user operations or configuration scenarios, and using vulnerability exploitation tools to conduct preliminary attack attempts; During the experiment, monitor the operating status of network devices, system logs, and changes in network traffic indicators to determine whether potential vulnerabilities can be triggered and observe their impact; Vulnerability verification and assessment: If abnormal behavior of network devices is found in the vulnerability verification experiment, the cause, exploitation conditions and impact scope of the vulnerability will be further analyzed to assess the severity of the vulnerability. The vulnerability will be quantitatively scored from multiple dimensions such as exploitability, impact scope and degree, and difficulty of repair to determine its risk level. Vulnerability report generation: Generates a vulnerability report based on the results of vulnerability verification and assessment. The report content includes the vulnerability discovery process, specific location, vulnerability type, range of affected network devices, potential impact of the vulnerability, and repair suggestions.

4. The method according to claim 1, characterized in that When extracting knowledge and building graphs, we also design an incremental update mechanism for the knowledge graph to reduce repeated calculations.

5. The method according to claim 1, characterized in that In step 2, multimodal data processing is also performed before data cleaning and standardization: Firmware binary analysis: extract control flow graph features through the disassembly engine; Protocol traffic analysis: Generate protocol state machine based on deep packet inspection DPI; Configuration semantic modeling: Convert CLI commands into abstract syntax trees (ASTs).

6. The method according to claim 1, characterized in that In step 4, a dynamic retrieval strategy optimizer based on reinforcement learning is also designed to realize intelligent path planning of the RAG module.

7. The method according to claim 1, characterized in that The large model is a language model based on the Transformer architecture. Through the multi-head attention mechanism, it can focus on different parts of the input text at the same time, capture the semantic associations and logical relationships between texts, and thus analyze the vulnerabilities of network devices.

8. The method according to claim 1, characterized in that In the vulnerability pattern recognition and reasoning process of step 5, the large model analyzes the text description of the network device configuration file and the vulnerability information caused by improper configuration of similar devices in the knowledge graph, and infers the permission bypass vulnerability that may exist in the current device under a specific configuration; based on the specification text of the network protocol and the anomalies in the actual traffic data, the possibility of buffer overflow vulnerability in the protocol implementation process is inferred; at the same time, the large model also combines known attack methods and security vulnerability types to predict the vulnerability of network equipment in the face of new attacks and explore potential security risk points.

9. A system for implementing the method according to any one of claims 1 to 8.

10. A network security analysis method implemented based on the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Attack path planning method based on graph attention network and deep reinforcement learning

    CN121309217A

  • Industrial control equipment vulnerability detection method and system combining firmware analysis and network scanning

    CN121356818A

  • Knowledge-enhanced false positive judgment method and system for vulnerabilities

    CN122595335A