RAG-based multi-dimensional network penetration test vulnerability mining method
Through the multi-dimensional network penetration testing framework based on RAG and combined with the search enhancement generation technology of large language models, intelligent multi-dimensional network penetration testing is realized, solving the problem of high cost of existing penetration testing frameworks and improving testing efficiency and accuracy.
Patent Information
- Application Number
- CN202510358543.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-25
AI Technical Summary
Existing penetration testing frameworks such as Metasploit and DeepExploit (DE) require a large amount of human resources for data integration and vulnerability verification, and cannot detect vulnerability exploit links between multidimensional networks, resulting in high cost of frequent penetration testing.
The multidimensional network penetration testing framework based on RAG is adopted, combined with the search enhancement generation technology of large language models, and dynamic interaction of agents is realized by simulating attacker behavior, and multidimensional network penetration testing is carried out, including scanning modules, query and generation modules, vulnerability utilization modules and report generation modules. The large language model is used to collaborate with the local knowledge base to generate vulnerability utilization information and build a secure tunnel.
It reduces the work burden of security experts in data integration and vulnerability verification, reduces the cost of penetration testing, realizes automated penetration testing across networks and systems, and improves testing efficiency and accuracy.
Smart Images

Figure CN120378137A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cyberspace security and relates to a method for mining vulnerabilities in multi-dimensional network penetration testing based on retrieval enhancement of large language models. Background Art
[0002] And the way of entertainment and leisure is being profoundly changed by Web technology, forming a life mode of information interconnection. At the same time, Web technology is also constantly expanding our working mode, innovating business operation processes, and promoting the development of the education field towards digitalization and remoteness. Whether it is personal life or social development, the role of Web technology in promoting the modernization process is becoming more and more prominent. However, while providing convenience for life, Web technology has also brought new security challenges and threats. With the evolution of network attack means, malicious attackers have emerged with various attack means against Web sites, including but not limited to attack methods such as SQL injection, Remote Code Execution, and Distributed Denial of Service (DDoS). These attacks not only threaten personal privacy and property security, but also pose a serious potential threat to enterprise operations and national security.
[0003] Therefore, in the current rapid development of information technology, it is particularly crucial to promptly identify security vulnerabilities in information systems, evaluate their potential risks, and quickly take repair measures. Although existing penetration testing frameworks, such as Metasploit and DeepExploit (DE), can assist security experts in performing penetration testing tasks to a certain extent, these frameworks still require a large amount of human resources for data integration and vulnerability verification, and cannot detect the connection of vulnerability exploitation between multi-dimensional networks, resulting in a high cost of frequently performing penetration testing. Summary of the Invention
[0004] To address the limitations of existing penetration testing frameworks and achieve more intelligent and efficient penetration testing, the present invention proposes a multi-dimensional network penetration testing framework based on RAG. This framework deeply integrates the retrieval enhancement generation technology in the field of large language models during the testing process to achieve dynamic interaction of agents and perform penetration testing in a multi-dimensional network environment. This method aims to reduce the workload of security experts in data integration and vulnerability verification, reduce the cost of frequently performing penetration testing, and improve the efficiency and accuracy of testing. The framework proposed by the present invention is different from traditional frameworks such as Metasploit and DeepExploit (DE), and it achieves breakthroughs through the following key technologies:
[0005] (1) Agent Dynamic Interaction: The framework simulates the behavior of attackers to achieve deeper penetration testing of the network. This dynamic interaction ability enables the framework to face multi-dimensional network penetration testing tasks and respond to more variable network environments based on the knowledge information obtained;
[0006] (2) Retrieval-Augmented Generation (RAG) of Large Language Models: Utilize the inherent knowledge of large language models and the dynamic repository of external databases to synergistically integrate and improve the credibility and availability of large language models in generating exploits;
[0007] (3) Multi-Dimensional Network Penetration Testing: The framework is designed for penetration testing in a multi-dimensional network environment, which means it can handle the overall security testing from the external network to the internal network, rather than just a single website or system.
[0008] To achieve the above objectives, the present invention provides the following technical solutions:
[0009] A vulnerability mining method for multi-dimensional network penetration testing based on RAG designs a multi-dimensional network penetration testing framework and adopts an attack system using Agent technology. The attack system includes: a scanning module, a query and generation module, an exploit module, and a report generation module. The method includes the following steps:
[0010] (1) Identify the possible subdomains of the target website through the scanning module and sequentially expand the attack surface of the penetration testing; in the obtained subdomain list, perform IP and port scanning and vulnerability scanning on each subdomain to form "IP-port, vulnerability information" vulnerability information pairs, and the high-risk vulnerability information pairs will be put into the exploitation queue;
[0011] (2) The query and generation module adopts a retrieval-augmented generation design model to extract vulnerability information pairs from the exploitation queue, utilize the local knowledge base and network intelligent search for relevant knowledge items, and finally the large model generates exploit information based on the knowledge items;
[0012] (3) The exploit module constructs a command injector according to the exploit information obtained in step (2) to obtain the permissions of the target machine, establish a secure tunnel between the victim target machine and the local machine, use the victim target machine as a jump target machine, and further scan other internal network hosts accessed by the jump target machine;
[0013] (4) Store the successfully exploited exploit information in the local exploit library during the full-link mining process; the report generation module divides and fills the penetration testing report template according to the assets, outputs the penetration testing reports for different assets, and summarizes the information using a large language model.
[0014] Furthermore, the step (1) specifically includes the following sub-steps:
[0015] (1.1) The scanning module in the attack system calls open-source tools to enumerate subdomains that may exist on the target website;
[0016] (1.2) The scanning module calls open-source tools to perform port scanning on each subdomain obtained in step (1.1), and preliminarily filters out relevant information on live IPs and open ports;
[0017] (1.3) The vulnerability scanning sub-module in the scanning module calls open-source tools to perform vulnerability scanning on the screened assets, cleans the scanning results, and classifies the vulnerabilities into high-risk vulnerabilities and ordinary vulnerabilities according to the exploitability of the vulnerabilities, and temporarily stores the vulnerability information (including host IP, port, service information, vulnerability information, etc.). Among them, high-risk vulnerability entries are simultaneously put into the exploitable vulnerability queue.
[0018] Further, step (2) specifically includes the following sub-steps:
[0019] (2.1) The attack system monitors the exploitable vulnerability queue in real time, extracts the vulnerability entries, and the extracted vulnerability entries will be removed from the exploitable vulnerability queue;
[0020] (2.2) The query module in the attack system first extracts the vulnerability type, service version information, and vulnerability number (if any) from the vulnerability entries obtained in step (2.1) to form a query statement;
[0021] (2.3) The query module uses the query statement obtained in step (2.2) to retrieve vulnerability exploitation information from the local knowledge base. If the local retrieval fails, the query statement is passed to the network search module to search relevant Internet documents (web blogs, PDF files, etc. on major security websites) as data sources, and key information extraction is performed on the large model to eliminate irrelevant noise information to obtain a pure data source; the query module inputs the query statement into the embedding model to generate embeddings, and retrieves the pure data source through similarity measurement to obtain text blocks related to the vulnerability information;
[0022] (2.4) The generation module in the attack system reorders the retrieved text blocks using the long-context reordering algorithm, combines the POC exploitation information in the vulnerability entries to form context information, and combines the chain of thought and few-shot prompting techniques to construct prompt words and input them into the large language model to generate Exploit structured information, which is used as the final vulnerability exploitation information.
[0023] Further, step (3) specifically includes the following sub-steps:
[0024] (3.1) The vulnerability exploitation module in the attack system generates a command injector according to the vulnerability exploitation information generated in step (2.3), and uses the command injector and network proxy technology to build a secure tunnel between the jump target machine and the local host;
[0025] (3.2) The internal network scanning sub-module of the scanning module in the attack system uses the security tunnel constructed in step (3.1) to use the current exploitable target machine as a jump target machine to scan for vulnerabilities in the internal network, obtain vulnerability information pairs of internal network hosts, and put the high-risk vulnerability information pairs into the exploitation queue, and then jump to step (2).
[0026] Further, the step (4) specifically includes the following sub-steps:
[0027] (4.1) When the vulnerability exploitation queue is empty, the report generation module stores the successfully exploited vulnerability exploitation information in the local knowledge base during this full-link mining process;
[0028] (4.2) The report generation module divides and fills the penetration test report template according to assets, outputs the penetration test report, and summarizes the information using a large language model at the same time.
[0029] The beneficial effects of the present invention are as follows:
[0030] 1. This framework utilizes the current cutting-edge LLM (Large Language Model) technology and its powerful natural language processing capabilities, and combines the Retrieval-Augmented Generation (RAG) design pattern to form a more flexible and intelligent penetration test process. By introducing RAG agents into the penetration test process, the human resource consumption in penetration testing can be significantly reduced, and the testing cost can be significantly lowered.
[0031] 2. While this framework can perform penetration testing from the external network and exploit the vulnerabilities scanned from the external network, it can thereby achieve the scanning of internal network vulnerabilities, achieving the effect of multi-dimensional network automated penetration testing. And initially establish the connection between internal and external network vulnerabilities. It realizes comprehensive penetration testing across networks and systems, helping security experts quickly identify, verify, and repair vulnerabilities in complex and ever-changing network environments, thereby achieving efficient and low-cost network security maintenance. Description of the Drawings
[0032] Figure 1 It is a schematic diagram of the framework in the present invention.
[0033] Figure 2 It is an execution flow chart of the framework in the present invention.
[0034] Figure 3 It is the vulnerability information pair obtained by data cleaning in the present invention. Detailed Embodiments
[0035] The following will detail the technical solutions provided by the present invention in combination with specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0036] As Figure 2 shown, the present invention provides a method for mining vulnerabilities in multi-dimensional network penetration testing based on RAG, and designs a multi-dimensional network penetration testing framework (such as Figure 1 ), an attack system using Agent technology (composed of a scanning module, a query and generation module, a vulnerability exploitation module, and a report generation module), including the following steps:
[0037] (1) The scanning module identifies possible sub-domains of the target website, and sequentially expands the attack surface of penetration testing; in the obtained sub-domain name list, IP and port scanning and vulnerability scanning are performed on each sub-domain name to form a vulnerability information pair of "(IP-port, vulnerability information)", and the high-risk vulnerability information pairs will be put into the exploitation queue at the same time;
[0038] (2) The query and generation module adopts a design model based on retrieval-augmented generation, extracts vulnerability information pairs from the exploitation queue, retrieves relevant knowledge items using the local knowledge base and network intelligent search, and finally the large model generates vulnerability exploitation information according to the knowledge items;
[0039] (3) The vulnerability exploitation module constructs a command injector according to the vulnerability exploitation information obtained in step (2) to obtain the target machine's permission, builds a secure tunnel between the victim target machine and the local machine, uses the victim target machine as a jump target machine, and further scans other internal network hosts accessed by the jump target machine;
[0040] (4) Store the successfully exploited vulnerability exploitation information in the local vulnerability exploitation library during this full-link mining process; the report generation module divides and fills the penetration testing report template according to the assets, outputs penetration testing reports for different assets, and summarizes the information using a large language model.
[0041] In this embodiment, step (1) specifically includes the following steps:
[0042] (1.1) Use open-source tools such as Sublist3r and Amass to identify possible sub-domain names of the target website;
[0043] (1.2) Perform data cleaning through an automated script, filter out newly discovered or unrecognized sub-domain names, and dynamically update the sub-domain name list to form a complete initial attack surface data set.
[0044] (1.3) Use the Nuclei open-source vulnerability scanning tool to comprehensively scan the sub-domain name services obtained, and extract the vulnerability information pairs of "(IP-port, vulnerability information)", as Figure 2 shown;
[0045] (1.4) Classify and process the vulnerability information pairs. Classify the vulnerabilities into ordinary vulnerabilities and high-risk vulnerabilities according to exploitability. Record the ordinary penetration testing vulnerabilities, and adopt the producer-consumer model to generate an Exploit_Queue vulnerability exploitation queue, focusing on the rapid processing of high-risk vulnerabilities.
[0046] In this embodiment, the step (2) specifically includes the following sub-steps:
[0047] (2.1) The attack system monitors the exploitable vulnerability queue in real time, extracts the vulnerability entries, and the extracted vulnerability entries will be removed from the exploitable vulnerability queue;
[0048] (2.2) For the extracted vulnerability information pairs, the query module in the attack system first extracts the vulnerability type, service version information, and vulnerability number from the vulnerability entries obtained in step (2.1) to form a query statement;
[0049] (2.3) The query module uses the query statement obtained in step (2.2) to retrieve vulnerability exploitation information from the local knowledge base. If the local retrieval fails, the query statement is passed to the network search module to search for relevant Internet documents (webpage blogs, PDF files, etc. of major security websites) as data sources, and input into the large model for key information extraction to eliminate irrelevant noise information to obtain a pure data source; the query module inputs the query statement into the embedding model to generate an embedding, and retrieves the pure data source through similarity measurement to obtain the context related to the vulnerability information.
[0050] (2.4) The generation module in the attack system reorders the retrieved text blocks using the long context reordering algorithm, combines the POC exploitation information in the vulnerability entries to form context information, and combines the chain of thought and few-shot prompting techniques to construct a prompt word and input it into the large language model to generate Exploit structured information, which is used as the final vulnerability exploitation information.
[0051] In this embodiment, the step (3) specifically includes the following sub-steps:
[0052] (3.1) The vulnerability exploitation module in the attack system generates a command injector according to the vulnerability exploitation information generated in step (2.3), tests whether the command injector can be successfully executed. If it fails, regenerate the Exploit structured information and pass it into the command injector until the set upper limit is reached; after successful detection, build a secure tunnel between the jump target machine and the local host with the help of the command injector and network proxy technology;
[0053] (3.2) The internal network scanning sub-module of the scanning module in the attack system scans the internal network of the jump target machine with the jump target machine as a jump server through the secure tunnel constructed in step (3.1), marks the jump target machine as exploited to prevent waste of system resources by reusing the jump target machine, obtains the vulnerability information pairs of internal network hosts, and puts the high-risk vulnerability information into the Exploit_Queue vulnerability exploitation queue for further exploitation.
[0054] In this embodiment, step (4) specifically includes the following sub-steps:
[0055] (4.1) When the Exploit_Queue vulnerability exploitation queue is empty, store the successfully exploited vulnerability exploitation information in the local knowledge base during this full-link mining process;
[0056] (4.2) The report generation module fills the report template, outputs penetration test reports for different assets, and at the same time summarizes the information with a large language model and outputs the report.
[0057] The technical means disclosed in the solution of the present invention are not limited to the technical means disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A vulnerability mining method for multi-dimensional network penetration testing based on RAG, characterized in that, An attack system based on a multi-dimensional network penetration testing method and adopting Agent technology, the attack system comprising: a scanning module, a query and generation module, a vulnerability exploitation module, and a report generation module, the method comprising the following steps: (1) Identify the possible subdomains of the target website through the scanning module, and sequentially expand the attack surface of the penetration testing; in the obtained subdomain list, perform IP and port scanning and vulnerability scanning on each subdomain to form "IP-port, vulnerability information" vulnerability information pairs, and the high-risk vulnerability information pairs will be put into the exploitation queue; (2) The query and generation module adopts a design model based on Retrieval-Augmented Generation (RAG), extracts the vulnerability information pairs from the exploitation queue, utilizes the local knowledge base and network intelligent search for relevant knowledge items, and finally the large model generates vulnerability exploitation information according to the knowledge items; (3) The vulnerability exploitation module constructs a command injector according to the vulnerability exploitation information obtained in step (2) to obtain the target machine permission, builds a secure tunnel between the victim target machine and the local machine, uses the victim target machine as a jump target machine, and further scans other internal network hosts accessed by the jump target machine; (4) Store the successfully exploited vulnerability exploitation information in the local vulnerability exploitation library during this full-link mining process; the report generation module divides and fills the penetration testing report template according to the assets, outputs the penetration testing reports for different assets, and summarizes the information using a large language model.
2. The method for mining vulnerabilities in multi-dimensional network penetration testing based on RAG according to claim 1, wherein The specific steps of step (1) are as follows: (1.1) The scanning module in the attack system calls open-source tools to enumerate the possible subdomains of the target website; (1.2) The scanning module calls open-source tools to perform port scanning on each subdomain obtained in step (1.1), and initially filters out the relevant information of the live IPs and open ports; (1.3) The vulnerability scanning sub-module in the scanning module calls open-source tools to perform vulnerability scanning on the filtered assets, cleans the scanning results, classifies the vulnerabilities into high-risk vulnerabilities and ordinary vulnerabilities according to the exploitable degree of the vulnerabilities, and temporarily stores the vulnerability information, and the high-risk vulnerability entries are simultaneously put into the exploitable vulnerability queue.
3. The method for mining vulnerabilities in multi-dimensional network penetration testing based on RAG according to claim 2, characterized in that, The specific steps of step (2) are as follows: (2.1) The attack system monitors the exploitable vulnerability queue in real time, extracts the vulnerability entries, and the extracted vulnerability entries will be removed from the exploitable vulnerability queue; (2.2) The query module in the attack system first extracts the vulnerability type, service version information, and vulnerability number from the vulnerability entries obtained in step (2.1) to form a query statement; (2.3) The query module uses the query statement obtained in step (2.2) to retrieve vulnerability exploitation information from the local knowledge base. If the local retrieval fails, the query statement is passed to the network search module to search for relevant Internet documents as the data source, and key information extraction is performed on the input to the large model to eliminate irrelevant noise information to obtain a pure data source; The query module inputs the query statement into the embedding model to generate an embedding, retrieves the pure data source through similarity measurement, and obtains the text blocks related to the vulnerability information; (2.4) The generation module in the attack system reorders the retrieved text blocks using the long-context reordering algorithm, combines the POC exploitation information in the vulnerability entries to form context information, and constructs prompts by combining the chain of thought and few-shot prompting techniques to input into the large language model to generate Exploit structured information, which serves as the information for the final vulnerability exploitation.
4. The method for mining vulnerabilities in multi-dimensional network penetration testing based on RAG according to claim 3, wherein, (3) The specific steps include the following sub-steps: (3.1) The vulnerability exploitation module in the attack system generates a command injector based on the vulnerability exploitation information generated in step (2.3), and uses the command injector and network proxy technology to build a secure tunnel between the jump target machine and the local host. (3.2) The internal network scanning sub-module of the scanning module in the attack system uses the secure tunnel constructed in step (3.1) to perform vulnerability scanning on the internal network with the current exploitable target machine as the jump target machine, obtains the vulnerability information pairs of the internal network hosts, puts the high-risk vulnerability information pairs into the queue to be exploited, and then jumps to step (2).
5. The method for mining vulnerabilities in multi-dimensional network penetration testing based on RAG according to claim 3, wherein (4) The specific steps include the following sub-steps: (4.1) When the vulnerability exploitation queue is empty, the report generation module stores the successfully exploited vulnerability exploitation information in the local knowledge base during this full-link mining process. (4.2) The report generation module divides and fills the penetration test report template according to the assets, outputs the penetration test report, and summarizes the information using the large language model.
Citation Information
Cited By
Graph-driven self-attention compressible memory management method
CN120994822A
Education network penetration vulnerability auxiliary analysis method based on retrieval enhancement generation
CN121037122A
An education network penetration vulnerability auxiliary analysis method based on retrieval enhancement generation
CN121037122B