An LLM driven method comprising asset identification and topology construction
Patent Information
- Application Number
- CN202611034884.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-25
AI Technical Summary
[0007]现有的测试框架对于系统资产的推断较为薄弱,并且多数框架是针对Web应用、服务器等目标,受到资产类别的影响较小,因此亟须探索能够应用于工业系统的资产识别与拓扑构建方法
(1)本发明提出了一种针对进攻性安全中环境信息获取阶段的LLM驱动方法,该框架将MIA推广为目标系统的资产推断,并基于工业信息物理系统中不同资产的特征,构建了涵盖隐蔽性资产推断与核心资产识别的资产评分模型,一方面判断某个资产是否属于目标系统,剔除非资产,另一方面将系统中的资产按照先验知识划分为隐蔽性资产、一般显性资产与核心资产,并根据所得到的资产信息构建目标系统网络拓扑,将获取到的全部环境信息融入拓扑图。
Smart Images

Figure CN122817484A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of asset identification and topology construction in industrial cyber-physical systems, and specifically relates to an LLM-driven method that includes asset identification and topology construction. Background Technology
[0002] Environmental information acquisition, also known as information gathering, is the first stage in realizing offensive security technologies. This stage involves probing and collecting information about the target system's hosts, ports, and services to reconnoiter the system's situation, thereby enabling more targeted penetration missions.
[0003] Currently, a common information gathering method is to integrate scanning tools such as Nmap into a penetration testing framework, remotely invoked to probe the target system. Compared to writing scanning commands manually, this method is simpler and more efficient. However, the environmental information obtained solely from scanning tools is relatively discrete and disorganized, and when multiple scanning tools are invoked, different tools collect a lot of duplicate content, making the environmental information redundant. An effective solution is to introduce Retrieval Augmentation (RAG) technology, utilizing the semantic reasoning capabilities of RAG and LLM to deduplicate and classify the environmental information output by various scanning tools, facilitating subsequent invocation. However, tool-based information scanning can only obtain relatively superficial environmental information and cannot achieve in-depth detection and reconnaissance of the target system.
[0004] For offensive security testers, simply using scanning tools is insufficient to gather as much information as possible about a target system. In reality, typical industrial systems contain a vast array of assets, including servers, gateways, routers, and circuit breakers. Due to the sheer number of these assets, it's impractical for either testers or system security personnel to fully penetrate or protect every single one.
[0005] Therefore, classifying these assets involves two aspects: firstly, determining whether a certain asset exists in the target system; and secondly, identifying the core and hidden assets within the system. Determining the existence of a particular asset is essentially similar to a Membership Inference (MIA) attack. MIA is a privacy attack targeting machine learning models, where attackers use the model's behavior differences between training and non-training data to determine whether a specific piece of data was used to train the model.
[0006] Extending this concept to asset inference in industrial systems, we can arrive at the following definition: Asset inference is the process of using prior knowledge and publicly available rules to determine whether a specific asset exists in a target system by comparing its characteristics in terms of equipment type, response time, etc. Taking ECPS as an example, this definition helps to exclude information equipment or assets in other Industrial Cyber-Physical Systems (ICPS), thus distinguishing ECPS from information equipment, ICPS, etc.
[0007] Existing testing frameworks are relatively weak in inferring system assets, and most frameworks are designed for targets such as web applications and servers, which are less affected by asset categories. Therefore, there is an urgent need to explore asset identification and topology construction methods that can be applied to industrial systems. Summary of the Invention
[0008] To overcome the shortcomings of existing technologies, this invention provides an LLM-driven method that includes asset identification and topology construction. By introducing LLM and ECPS-related knowledge into the environment information acquisition stage of penetration testing, and based on LLM assistance, the guiding framework constructs an asset scoring model considering ECPS characteristics, thereby classifying assets in ECPS and taking into account the category of each asset when drawing the system network topology, thus achieving a more refined and complete acquisition of ECPS environment information.
[0009] The specific technical solution is as follows: S1, invoke multiple scanning tools to scan the target industrial cyber-physical system, and identify the assets, asset attribute information and connection relationships between assets in the target industrial cyber-physical system based on the scanning results; S2: Using all assets as nodes and the connections between assets as edges, a preliminary network topology is constructed. S3, based on the asset's attribute information, combined with LLM calculations, yields the evidence index, violation index, and core index for each asset; S3 specifically involves: First, using LLM to determine whether the asset meets the judgment conditions for each type of evidence, and calculating the asset's evidence index based on the judgment results; the evidence types include publish or subscribe relationships, industrial control port services, industrial control protocols, business response time, and production control region affiliation. Secondly, LLM is used to determine whether the asset has protocol semantic violations, topology logic violations, behavior pattern violations, or business flow paradox violations, and the asset's violation index is calculated based on the determination results; Finally, based on the list of core assets given by the target industrial cyber-physical system, the core index of the asset is calculated. S4 combines the evidence index, violation index, and core index to calculate the asset score for each asset. Based on the asset score, the assets in the target industrial cyber-physical system are divided into four types: non-assets, hidden assets, general explicit assets, and core assets. S5. Based on the asset type classification results in S4, the nodes corresponding to different asset types in the preliminary network topology diagram and the edges connected to nodes of different asset types are adjusted respectively. A label is established next to each node to describe the asset corresponding to that node, thereby obtaining a network topology diagram to describe the asset distribution of the target industrial cyber-physical system.
[0010] Furthermore, in S1, the scanning tools include Nmap, Masscan, and Zmap.
[0011] Furthermore, in S1, the assets include the hardware and software of the target industrial cyber-physical system; the hardware refers to the physical devices in the target industrial cyber-physical system, including switches and servers; the software refers to the programs and logical functions running on the hardware, including gateway programs, operating systems, and application software.
[0012] Furthermore, in S3, the judgment conditions for each type of evidence are specifically as follows: The condition for determining the publish or subscribe relationship is whether the asset has a publish or subscribe relationship with a known control device in the target physical information system; The determination condition for the industrial control specific port service is whether the asset has a preset industrial control protocol port open and whether the banner information of the port contains the dedicated equipment characteristics of the asset. The determination criterion for the industrial control protocol is whether the asset uses a publicly available industrial control protocol. The condition for determining the business response time is whether the response time of the asset is within a specified range; The criteria for determining the ownership of the production control area are whether the asset directly participates in the production execution or process control of the target industrial cyber-physical system. The evidence index for each asset is calculated as follows: set the initial evidence weight of the asset to 0, give weights for different evidence types, and sum the evidence weights of all evidence types that the asset satisfies based on the LLM judgment result to obtain the asset's evidence index.
[0013] Furthermore, in S3, the specific content of each violation is as follows: The protocol semantic violation is that the interval between hardware asset reports of status changes is less than the mechanical limit of the hardware asset. The aforementioned topology violation is that the asset directly establishes a connection with the upper-level management network of the target industrial cyber-physical system; The violation of the behavioral pattern is that the asset fails to communicate with other assets at fixed time intervals; The aforementioned business flow paradox is that the asset, as a data subscriber, did not receive the data; The violation index for each asset is calculated as follows: based on the LLM's judgment results, the violation index of an asset with no violations or with 1 violation is recorded as 1, the violation index of an asset with 2 violations is recorded as 2, the violation index of an asset with 3 violations is recorded as 3, and the violation index of an asset with 4 violations is recorded as 4. Furthermore, in S3, the calculation formula for the core index of each asset is as follows: ; in, The core index for asset e. The list of core assets is provided by the target industrial cyber-physical system; i represents the core asset. In the preliminary network topology diagram, with assets The number of edges that are connected.
[0014] Furthermore, in S4, the asset score calculation formula for each asset is as follows: ; in, For assets; Assess asset rating; and For parameters, ; , and These are the evidence index, violation index, and core index for asset e, respectively. For industrial cyber-physical systems.
[0015] Furthermore, in S4, the criteria for classifying assets based on asset rating are as follows: ; Where e represents assets, To score assets.
[0016] Furthermore, in S5, the process of adjusting the nodes and edges of the preliminary network topology graph based on the asset partitioning results is specifically as follows: Remove non-asset nodes from the initial network topology diagram, and represent hidden asset nodes with gray solid circles, core asset nodes with red solid circles, and general visible asset nodes with black solid circles. Further, represent edges connected to non-asset nodes with dashed lines, edges connected to hidden asset nodes with black solid lines, and other edges with blue solid lines.
[0017] Furthermore, if the adjustment results in a node or edge that is not connected to the other nodes, then that node or edge is deleted directly.
[0018] The beneficial effects of this invention are: (1) This invention proposes an LLM-driven method for the environmental information acquisition stage in offensive security. This framework extends MIA to asset inference of the target system and constructs an asset scoring model covering the inference of concealed assets and the identification of core assets based on the characteristics of different assets in the industrial cyber-physical system. On the one hand, it determines whether an asset belongs to the target system and eliminates non-assets. On the other hand, it divides the assets in the system into concealed assets, general explicit assets and core assets according to prior knowledge, and constructs the network topology of the target system based on the obtained asset information, and integrates all the obtained environmental information into the topology graph.
[0019] (2) By combining relevant industrial domain knowledge with the method of the present invention, a method for information collection in industrial cyber-physical systems is constructed. In a specific embodiment, the method of the present invention is integrated into a penetration testing method, verifying the performance improvement of the present invention on the corresponding method, enabling more penetration testing methods to be adapted to power scenarios.
[0020] (3) Compared with the prior art, the present invention takes into account the problem that the existing penetration testing methods pay less attention to the acquisition of environmental information. First, it constructs an environmental information acquisition framework and adds two major functions, asset identification and topology construction, on the basis of the preliminary scanning of the existing penetration testing methods. On the other hand, it connects to multiple scanning tools and integrates ECPS domain knowledge into the framework. While ensuring that the preliminary scanning obtains richer and more detailed information, it enables the framework to match ECPS, forming an environmental information acquisition framework for ECPS. It can also connect to other penetration testing methods and is ultimately used for offensive security and active defense of ECPS. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating the method of the present invention; Figure 2 It is the actual topology of the target system used for testing; Figure 3 This is the topology of the target system drawn by the method of this invention; Figure 4 This is a bar chart showing the performance improvement of the method of this invention for DeepAttacker; Figure 5 This is a bar chart showing the performance improvement of the method of the present invention for a four-agent offensive security method. Detailed Implementation
[0022] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in the various embodiments of the present invention can be combined accordingly without mutual conflict.
[0023] Before introducing this invention, the technical terms used in this invention will be explained first: Cyber-physical systems (ECPS): refers to systems that apply cyber-physical systems to the power system field, such as power plants, power grids, electric vehicle charging stations, and other power systems.
[0024] Environmental information refers to the collection and identification of various data sets of the target system before the penetration test begins. It mainly includes basic information such as system configuration, network topology, service version, user accounts, security policies, and log configuration.
[0025] Assets: Generally refers to all tangible and intangible assets owned or managed by the target system that are necessary to ensure the system's production, transportation, distribution and supporting facilities. It includes not only equipment, but also related software, data and supporting facilities.
[0026] Evidence: observable facts or available information used to support or prove a certain judgment or conclusion.
[0027] Network topology: refers to the use of graph theory to abstract assets in a target system into nodes and edges, describing the connections and connectivity between assets.
[0028] Due to the powerful semantic reasoning and enhanced generation capabilities of Large Language Models (LLM), and the unique characteristics of industrial domain knowledge compared to other ICPS, this invention considers introducing LLM and domain-related knowledge into the environmental information acquisition stage of penetration testing. Based on LLM assistance, the guiding framework constructs an asset scoring model considering domain characteristics, thereby classifying assets in the system and considering the category of each asset when drawing the system network topology, thus achieving a more refined and complete acquisition of environmental information.
[0029] like Figure 1As shown, this invention proposes an LLM-driven method that includes asset identification and topology construction. For all industrial control systems (industrial cyber-physical systems), the method first takes the target system's IP address or domain name as input. Through the embedded LLM, a scanning tool is invoked to perform a preliminary scan of the target system and acquire information. Then, a pre-designed asset scoring model is used to identify and classify all assets in the system. Finally, based on the results of the preliminary scan and asset identification, a network topology diagram of the target system is drawn, including descriptions of all asset categories and corresponding environmental information, forming a complete LLM-driven framework that includes asset identification and topology construction, achieving the goal of detailed and complete acquisition of environmental information.
[0030] The specific steps are as follows: Step 1: Use multiple scanning tools to perform a preliminary scan of the target system.
[0031] Step 1 specifically includes: Step 1.1: Determine the IP address or domain name of the target system. Based on the preset LLM prompts, automatically generate instructions to call the scanning tools. Then, according to the instructions, call the corresponding scanning tools in sequence to perform a preliminary scan of the target system. The scanning tools include Nmap, Masscan, Zmap, etc. Step 1.2: Use Retrieval Enhancement Generation (RAG) technology to remove redundancy from the environmental information scanned by the scanning tool, and classify it according to device type, operating system, live host, open ports, software version, user agreement, and service information to obtain the target system assets.
[0032] The asset includes the hardware and software of the target industrial cyber-physical system. The hardware refers to all physically existing physical devices in the target industrial cyber-physical system, including but not limited to switches and servers. The software refers to various programs and logical functions running on the aforementioned hardware devices, including but not limited to gateway programs, operating systems, and application software.
[0033] In this embodiment, Nmap, Masscan and Zmap are used as examples. The comparison of the discovery of the main environmental information by the three tools is shown in Table 1. "-" indicates that the corresponding scanning tool did not collect the corresponding information.
[0034] Table 1 Comparison of environmental information for the three scanning tools Step 2: Establish an asset scoring model and use the model to classify the assets in the target system, including non-assets, hidden assets, general explicit assets, and core assets.
[0035] Step 2 specifically includes: Step 2.1: Construct an asset scoring model: in, For a specific asset in the target system; To rate the assets; and For the parameters of the model, ; , and These are the asset evidence index, violation index, and core index. For the target system.
[0036] The specific calculation method for each index is as follows: Evidence Index for: LLM is used to determine which of the following criteria a device meets: publish / subscribe relationship (i.e., whether the asset has a publish or subscribe relationship with known control devices in the target physical information system), industrial control port service (whether the asset opens typical industrial control protocol ports, such as port 102 of the MMS protocol, port 502 of the Modbus protocol, or port 20000 of the GOOSE protocol, and the banner information of the port shows non-default or dedicated device characteristics), industrial control protocol (using publicly available standard industrial control protocols, including but not limited to Modbus protocol, IEC 61850 protocol, IEC 60870-5-104 protocol, DNP3 protocol), service response time (meeting the preset response time interval), and production control area affiliation (deployed in the core area of the industrial production control network, directly participating in production execution or process control). The weight corresponding to each criterion is preset by the user. The initial criterion weight is 0. For each criterion met by the device, the corresponding weight score is accumulated, and finally, the sum of the weights of all satisfied criterions is obtained.
[0037] The selection of the above five types of evidence is based on a comprehensive consideration of the technical difficulties in identifying industrial control system assets and the sources of false alarms. In industrial control systems, publish / subscribe relationships are strongly correlated with the corresponding industry's industrial control systems. Therefore, once a relevant publish / subscribe relationship exists, the corresponding asset can be considered to likely belong to that industrial control system. An active and reachable asset has frequent and monitorable subscription / publishing. This evidence aims to exclude false alarm assets that only have general IP connectivity (such as ICMP reachability) but no industrial control semantic interaction. Industrial control port services are used to ensure that industrial control assets open typical ports and support specific services. Unlike the widespread openness of IT system ports, key equipment in industrial control systems typically opens specific standard ports. This evidence aims to distinguish real industrial control equipment from simulators, honeypots, or non-industrial control equipment. Industrial control protocols are used to exclude some IT equipment in industrial control systems. The essential difference between industrial control systems and general IT systems is... Regarding the communication protocol stack, if the relevant assets use non-industrial control protocols such as HTTP and FTP, this evidence is not satisfied, thus anchoring the asset identification scope to devices with industrial control communication capabilities, eliminating general IT devices that have no business connection with the industrial control system from the communication semantic level; the business response time is related to the mechanical cycle of equipment operation, and this evidence ensures that industrial control assets are physical devices, not purely information devices; the attribution of the production control area is determined by the location of the equipment. According to the "Regulations on Security Protection of Power Monitoring Systems" and other industrial control security industry standards, the production control area (including the control area and the non-control area) is the core area where equipment directly involved in production execution or process control is deployed. This evidence aims to exclude non-industrial control equipment deployed in the management information area or external networks through network area boundary constraints.
[0038] When applied to different industries, the specific judgment criteria for each piece of evidence vary. Taking ECPS as an example, the method for calculating the weight of evidence is shown in Table 2.
[0039] Table 2. Types of Evidence and Their Weights Violation Index for: The LLM (Local Management Model) is used to determine any violations found in the device. These violations include four categories: protocol semantic violations (i.e., the interval between hardware asset status change reports is lower than the physical limit for that type of asset), topology logic violations (i.e., the device directly establishes a connection with the upper-level management network, which includes the manufacturing execution system server, enterprise resource planning server, and office management terminal in the target system), behavioral pattern violations (i.e., the device does not communicate with other devices at fixed time intervals, and its communication intervals vary randomly), and business flow paradoxes (i.e., the device, as a data subscriber, never receives data packets from the publisher, forming a communication island).
[0040] Similarly, the above four types of violations are selected based on the three essential constraints of industrial control cyber-physical systems, and the selection criteria are as follows: Protocol semantic violations take into account that industrial control equipment is subject to the physical response limits of the actuators. There must be a lower bound threshold for judging its status changes or reporting intervals. If the status reporting interval of a certain asset is detected to be significantly lower than the physical limit of this type of equipment, it indicates that the communication behavior cannot be driven by physical processes and is most likely caused by forged messages, replay attacks or emulator injection.
[0041] Topology violation: Considering that industrial control networks adopt a strict hierarchical and partitioned architecture according to industrial control system architecture standards such as IEC 62264 and IEC 62351, if devices directly connect to the upper-level management network, it violates the basic design principles of this hierarchical architecture and is considered a high-risk topology violation.
[0042] Behavioral pattern violations: Considering that periodicity is one of the essential characteristics that distinguishes industrial control systems from general IT systems, if the communication interval between devices changes randomly and does not have stable periodicity or quasi-periodicity, it indicates that the communication behavior does not conform to the driving mode of normal industrial control tasks, and may be malicious scanning, data leakage or abnormal external connection behavior.
[0043] The business flow paradox addresses the issue that data transmission in industrial control systems follows a directed publish / subscribe pattern. A normal publish / subscribe relationship is bidirectional and verifiable, with the publisher continuously sending data and the subscriber continuously receiving it. If the device cannot obtain control commands or acquire data, it indicates that the subscription relationship has failed or been maliciously disrupted. Taking ECPS as an example, the specific details of each violation are shown in Table 3.
[0044] Table 3. Penalties and Weights for Violations Core Index for: in, The list of core assets is provided by the management personnel of the target system. Is with assets The number of connected edges (connections). The larger the value, the more assets and resources one possesses. Connected.
[0045] Step 2.2: Calculate the score for each asset in the target system. Assets are classified into non-assets, hidden assets, general explicit assets, and core assets according to the following discrimination formula: The discrimination formula is: Step 3: Construct a network topology diagram of the target system and mark the environmental information obtained from the preliminary scan on the topology diagram.
[0046] Step 3 specifically includes: Step 3.1: Treat all assets as topology nodes, and connect the nodes with undirected edges using the asset connection relationships detected by the initial scan to form a preliminary network topology.
[0047] Step 3.2: Based on the asset classification, directly delete the nodes corresponding to non-assets in the network topology, and change the edges connected to non-assets to dashed lines, indicating that the connection involves networks outside the target system such as office network and information system. If, after deletion, a node or edge not connected to other nodes is generated in the topology, then delete the node and edge together.
[0048] Step 3.3: Mark the nodes corresponding to hidden assets in the topology as gray, the nodes corresponding to core assets as red, and the remaining nodes, representing general visible assets, as black; change the edges connected to the nodes corresponding to hidden assets to black, and the remaining edges to blue.
[0049] Step 3.4: Create a label for each node in the network topology and write the environmental information collected in Step 1 into the label, including the IP address of the corresponding asset, device name and type, operating system, software version, port status, power protocol, etc., and finally complete the construction of the target system network topology.
[0050] To verify the effectiveness of the method proposed in this invention, a GOAD testbed deployed on a Linux host machine was used as the experimental platform to conduct related experiments. The specific experimental parameters are designed as follows: In this embodiment, considering the efficiency, accuracy, and cost of LLM, DeepSeek-v4 is used as the LLM framework for acquiring target ECPS environment information in the embedded experiment. Nmap, Masscan, and Zmap are selected as the scanning tools used in the initial scan of this experiment. The scan reveals the assets contained in the target ECPS in this experiment: one switch, three domain controller servers, and two member servers.
[0051] In this experiment, under the target environment of ECPS, the embedded LLM is first used to generate instructions to call the scanning tool, perform a preliminary scan of the environmental information, and use RAG technology to remove redundancy. Then, the constructed asset scoring model is used to identify and classify the assets in ECPS, dividing all assets into non-assets, hidden assets, general visible assets, and core assets. Finally, using the environmental information obtained from the preliminary scan and the asset information obtained from asset identification, a network topology diagram for the target ECPS is drawn. For each node corresponding to each asset in the diagram, the relevant information is input into the node's label. The overall environmental information acquisition framework is integrated with two methods in the ECPS field, and the performance of the methods before and after the integration of the framework of this invention is compared, thereby demonstrating the superiority of this invention.
[0052] To demonstrate the effectiveness of the method of the present invention, Table 4 shows different... , and Impact on asset identification results. Assets are scored in descending order as core assets (++), general explicit assets (+), implicit assets (+-), and non-assets (-). It is evident that... That is, without referring to the core asset list, The larger the value, the lower the asset score. Secondly, The value does not change the assets and non-assets in the system, but for assets, The larger the value, the higher the asset score. (Based on 6 different groups) , Value, as can be seen in At that time, the accuracy of asset identification was the highest.
[0053] Table 4 , and Impact on asset identification Figure 2 and Figure 3 The actual topology of the target system used in the experiment and the topology drawn using the framework of the present invention are shown respectively. For ease of presentation, Figure 3 Only the DC01 label is retained; the labels for all other assets have been hidden. It's important to note that due to the inaccessibility of domain controller DC02, this asset... Figure 2 The actual topology is not shown; it is only marked "DC02 down" on the diagram.
[0054] The results show that the network topology diagram drawn by this invention is more complete and detailed. On the one hand, it takes into account DC02 as a concealed asset and incorporates it into the topology; on the other hand, it attaches relatively complete information labels next to each node, with the information derived from the initial scan. Comparing the topology constructed by the framework of this invention with the actual topology, the actual topology distinguishes different asset devices through different shapes and colors, while the topology of the framework of this invention classifies asset categories according to color, making it easier to highlight the importance of assets. Therefore, the topology diagram of the framework of this invention is more detailed and complete.
[0055] This specific embodiment further verifies the impact of the method of the present invention on the performance of the two existing ECPS penetration testing methods, DeepAttacker and the Four Agents Offensive Security Approach (FAF).
[0056] Three types of experiments were conducted for each penetration testing method: using the penetration testing method itself, combining the penetration testing method with Nmap, and combining the penetration testing method with the framework of this invention. The effectiveness was evaluated primarily based on the accuracy and quantity of the test links generated by the method. The experimental results are as follows: Figure 4 and Figure 5 As shown, where Figure 4 In this context, DeepAttacker represents the pure method itself, DeepAttacker+Nmap represents the method combined with Nmap, and DeepAttacker+HINT represents the method combined with the framework of this invention. Figure 5 In this context, FAF represents the pure method itself, FAF + Nmap represents the method combined with Nmap, and FAF + HINT represents the method combined with the framework of this invention.
[0057] It can be observed that, for DeepAttacker, the framework of this invention significantly improves both the accuracy and quantity of the generated test links. This is because the framework, through reconnaissance of the target ECPS, obtains more complete environmental information and a wider variety of system vulnerabilities, thereby helping DeepAttacker to build test links more effectively. Furthermore, since DeepAttacker itself collects environmental information via RAG without relying on scanning tools, the assistance of Nmap also results in a slight performance improvement.
[0058] Compared to DeepAttacker, FAF utilizes a wider range of methods, tools, and databases, resulting in significantly superior performance. However, replacing the scanning portion of FAF with Nmap actually reduces the number of tools available, leading to a decrease in the accuracy of the generated test links. The reduced number of scanning tools broadens the environmental information acquired by FAF, potentially generating more useless test links and increasing the overall number of links. Nevertheless, after incorporating the framework of this invention, the quantity and quality of FAF test links are significantly improved, maintaining approximately 90% accuracy while generating over 250 test links. This further demonstrates that the environmental information provided by the framework described in this invention can help offensive security frameworks generate more systematic and targeted test links, helping to expose more ECPS risks.
[0059] In summary, the innovation and technical contributions of the framework constructed in this invention are mainly reflected in the following aspects: (1) An LLM-driven framework that includes asset identification and topology construction is proposed. This framework extends MIA to asset inference of the target system and constructs the network topology of the target system based on the obtained asset information, integrating all the obtained environmental information into the topology graph; (2) An asset scoring model that covers hidden asset inference and core asset identification is constructed. This model can determine whether an asset belongs to the target system and eliminate non-assets. On the other hand, it divides the assets in the system into hidden assets, general explicit assets and core assets according to prior knowledge and marks them when drawing the network topology; (3) The HINT is combined with ECPS domain knowledge to construct a method that can be used for ECPS information collection. The HINT is then integrated into the penetration testing framework, verifying the performance improvement of the corresponding framework by the HINT, so that more penetration testing frameworks can be adapted to the power scenario.
[0060] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.
Claims
1. An LLM-driven method comprising asset identification and topology construction, characterized in that, include: S1, invoke multiple scanning tools to scan the target industrial cyber-physical system, and identify the assets, asset attribute information and connection relationships between assets in the target industrial cyber-physical system based on the scanning results; S2: Using all assets as nodes and the connections between assets as edges, a preliminary network topology is constructed. S3, based on the asset's attribute information, combined with LLM calculations, yields the evidence index, violation index, and core index for each asset; S3 specifically involves: First, using LLM to determine whether the asset meets the judgment conditions for each type of evidence, and calculating the asset's evidence index based on the judgment results; the evidence types include publish or subscribe relationships, industrial control port services, industrial control protocols, business response time, and production control region affiliation. Secondly, LLM is used to determine whether the asset has protocol semantic violations, topology logic violations, behavior pattern violations, or business flow paradox violations, and the asset's violation index is calculated based on the determination results; Finally, based on the list of core assets given by the target industrial cyber-physical system, the core index of the asset is calculated. S4 combines the evidence index, violation index, and core index to calculate the asset score for each asset. Based on the asset score, the assets in the target industrial cyber-physical system are divided into four types: non-assets, hidden assets, general explicit assets, and core assets. S5. Based on the asset type classification results in S4, the nodes corresponding to different asset types in the preliminary network topology diagram and the edges connected to nodes of different asset types are adjusted respectively. A label is established next to each node to describe the asset corresponding to that node, thereby obtaining a network topology diagram to describe the asset distribution of the target industrial cyber-physical system.
2. The LLM-driven method comprising asset identification and topology construction according to claim 1, characterized in that, In S1, the scanning tools include Nmap, Masscan, and Zmap.
3. The LLM-driven method according to claim 1, comprising asset identification and topology construction, is characterized in that, In S1, the assets include the hardware and software of the target industrial cyber-physical system; the hardware refers to the physical devices in the target industrial cyber-physical system, including switches and servers; the software refers to the programs and logical functions running on the hardware, including gateway programs, operating systems, and application software.
4. The LLM-driven method including asset identification and topology construction according to claim 1, characterized in that, In S3, the specific judgment conditions for each type of evidence are as follows: The condition for determining the publish or subscribe relationship is whether the asset has a publish or subscribe relationship with a known control device in the target physical information system; The determination condition for the industrial control specific port service is whether the asset has a preset industrial control protocol port open and whether the banner information of the port contains the dedicated equipment characteristics of the asset. The determination criterion for the industrial control protocol is whether the asset uses a publicly available industrial control protocol. The condition for determining the business response time is whether the response time of the asset is within a specified range; The criteria for determining the ownership of the production control area are whether the asset directly participates in the production execution or process control of the target industrial cyber-physical system. The evidence index for each asset is calculated as follows: set the initial evidence weight of the asset to 0, give weights for different evidence types, and sum the evidence weights of all evidence types that the asset satisfies based on the LLM judgment result to obtain the asset's evidence index.
5. The LLM-driven method according to claim 1, comprising asset identification and topology construction, characterized in that, In S3, the specific content of each violation is as follows: The protocol semantic violation is that the interval between hardware asset reports of status changes is less than the mechanical limit of the hardware asset. The aforementioned topology violation is that the asset directly establishes a connection with the upper-level management network of the target industrial cyber-physical system; The violation of the behavioral pattern is that the asset fails to communicate with other assets at fixed time intervals; The aforementioned business flow paradox is that the asset, as a data subscriber, did not receive the data; The violation index for each asset is calculated as follows: based on the LLM's judgment results, the violation index of an asset with no violations or one violation is recorded as 1, the violation index of an asset with two violations is recorded as 2, the violation index of an asset with three violations is recorded as 3, and the violation index of an asset with four violations is recorded as 4.
6. The LLM-driven method including asset identification and topology construction according to claim 1, characterized in that, In S3, the formula for calculating the core index of each asset is as follows: ; in, The core index for asset e. The list of core assets is provided by the target industrial cyber-physical system; i represents the core asset. In the preliminary network topology diagram, with assets The number of edges that are connected.
7. The LLM-driven method comprising asset identification and topology construction according to claim 1, characterized in that, In S4, the asset score calculation formula for each asset is as follows: ; in, For assets; Assess asset rating; and For parameters, ; , and These are the evidence index, violation index, and core index for asset e, respectively. For industrial cyber-physical systems.
8. The LLM-driven method comprising asset identification and topology construction according to claim 1, characterized in that, In S4, the criteria for classifying assets based on asset rating are as follows: ; Where e represents assets, To score assets.
9. The LLM-driven method comprising asset identification and topology construction according to claim 1, characterized in that, In step S5, the process of adjusting the nodes and edges of the preliminary network topology graph based on the asset partitioning results is as follows: Remove non-asset nodes from the initial network topology diagram, and represent hidden asset nodes with gray solid circles, core asset nodes with red solid circles, and general visible asset nodes with black solid circles. Further, represent edges connected to non-asset nodes with dashed lines, edges connected to hidden asset nodes with black solid lines, and other edges with blue solid lines.
10. An LLM-driven method comprising asset identification and topology construction according to claim 9, characterized in that, If the adjustment results in a node or edge that is not connected to the other nodes, then that node or edge should be deleted directly.