A multi-stage vulnerability mining and threat tracing method and system

CN122640245BActive Publication Date: 2026-09-29HUAYI DIGITAL TECHNOLOGY (JILIN PROVINCE) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202611122166.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-29
Estimated Expiration
2046-07-28

AI Technical Summary

Technical Problem

[0008]本发明解决了攻防演练中现有技术使用的工具之间数据孤立、全量日志扫描资源浪费、测试引导缺乏战术视角的问题

Benefits of technology

1、攻击原语第一阶段记录、攻击原语第二阶段记录和攻击原语第三阶段记录的串行执行,省略了需要人工手动导入数据、转格式的过程,节省了人工转译环节,提高了效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122640245B_ABST
    Figure CN122640245B_ABST
Patent Text Reader

Abstract

The application relates to a multi-stage vulnerability mining and threat tracing method and system, and relates to the technical field of network security. The method solves the problems of data isolation between tools used by the prior art, waste of full-quantity log scanning resources, and lack of tactical perspective in test guidance. The method comprises the following steps: acquiring system call logs and permission change logs in an isolated sandbox, performing semantic analysis to obtain test cases; executing the test cases; extracting system call sequences to generate attack primitive first-stage records; retrieving a full-quantity log subset from full-quantity logs; performing two classifications through a graph neural network to obtain an attack path candidate set and generate attack primitive second-stage records; extracting key entities from the attack primitive second-stage records, retrieving intelligence, splicing the attack path, the retrieved intelligence and a MITRE ATT&CK mapping result into reasoning context; generating an attacker portrait, marking an uncovered test blind area, generating attack primitive third-stage records, and writing the test blind area into a blind area list.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, and more specifically to the assessment of security vulnerabilities in attack and defense exercises. Background Technology

[0002] Attack and defense drills (red team / blue team exercises) are a routine method for enterprises and governments to test their defense systems. These drills generate a large amount of security data, including system call logs, network traffic, and permission change records. Analysts need to perform three tasks from this data: discovering vulnerabilities, reconstructing attack paths, and identifying attackers. Currently, these three tasks are accomplished using different tool systems: fuzzing tools such as AFL and LibFuzzer are used to perform random mutation tests on target programs to discover vulnerabilities; log analysis platforms such as Splunk and ELK use manually written query rules to search for attack traces in logs and reconstruct attack paths; and online threat intelligence platforms such as Microstep are used to manually compare threats against a threat intelligence database to identify attackers.

[0003] In existing technologies, the tools used for the three tasks mentioned above run independently, and their output formats are incompatible. Security analysts must manually convert the output of the vulnerability discovery tools into a format that the platform for reconstructing attack paths can parse, and then manually enter the results of the attack path reconstruction into the intelligence platform. This manual translation process takes 3 to 6 hours, and critical contextual information (such as the temporal relationship of system call sequences) is easily lost during the conversion.

[0004] The attack path reconstruction system currently performs indiscriminate scanning of all logs. In attack and defense drill environments, a single server generates millions of logs per hour. Logs actually related to the attack typically only account for 5% to 10% of the total, with the remaining 90% or more of the computing resources consumed on irrelevant logs, resulting in tracing taking 2 to 6 hours.

[0005] Existing fuzzing tools use code coverage as the sole test guidance signal, prioritizing the generation of test cases that can execute more code branches. This guidance method ignores coverage at the attack tactical level. The MITRE ATT&CK framework defines 14 attack tactical phases (reconnaissance, initial access, privilege escalation, lateral movement, etc.), and high code coverage does not mean that all tactical phases have been tested. Some tactical phases may have no corresponding test cases at all, creating tactical blind spots.

[0006] Chinese patent document CN120389903A discloses a method and system for predicting inter-state cyberattacks based on the fusion of GNN and LLM. It constructs a directed multi-view dynamic graph by fusing multimodal data such as news, satellite imagery, and public opinion. Then, it uses GNN (Graph Neural Network) to generate nodes, embeds them, and injects them into various layers of LLM (Large Language Model) for joint fine-tuning, ultimately outputting a predicted attack probability. However, this approach suffers from high model complexity and poor interpretability due to the deep fusion of GNN and LLM.

[0007] Chinese patent document CN120579194A discloses a closed-loop vulnerability management method and system based on intelligent collaboration. It constructs a multi-dimensional vulnerability semantic model and a state-aware graph, achieving closed-loop vulnerability management by modifying model parameters and edge weights. However, continuously modifying model parameters leads to unstable model behavior, historical learning results are easily overwritten by new data, and it is difficult to distinguish between short-term fluctuations and long-term trends. Furthermore, closed-loop modifications at the parameter level lack interpretability, making it impossible for operations and maintenance personnel to trace the specific reasons for model decision changes, which is detrimental to security auditing. Summary of the Invention

[0008] This invention solves the problems of data isolation between tools used in existing technologies during attack and defense exercises, resource waste from full log scanning, and lack of tactical perspective in test guidance.

[0009] The technical solution provided by this invention: Option 1: A multi-stage vulnerability discovery and threat attribution method, comprising the following steps: S1: Obtain system call logs and permission change logs from the isolation sandbox, and perform semantic analysis on code snippets, known vulnerability patterns, and permission change logs of the target system to obtain test cases; Execute test cases and detect abnormal behavior based on system call logs; Extract the system call sequence at the moment of abnormal behavior and generate the first phase record of attack primitives; S2: Based on the first-stage record of the attack primitives, retrieve a subset of the full logs that satisfy the time constraints, entity constraints and fingerprint matching from the full logs; The full log subset is converted into a heterogeneous tracing graph that includes three node types and four edge types; The three node types are classified into two categories using a graph neural network to obtain attack behavior nodes; The attack behavior nodes are concatenated according to the time sequence to obtain the attack path candidate set, and the second phase record of the attack primitive is generated. S3: Map the attack behavior nodes to the tactical stages corresponding to the MITRE ATT&CK matrix to obtain the MITRE ATT&CK mapping results; Key entities are extracted from the second-stage record of the attack primitives, and intelligence is retrieved from the closed attack and defense exercise knowledge base. The attack path, retrieved intelligence, and MITRE ATT&CK mapping results are then combined into a reasoning context. The inference context is input into a large language model to generate an attacker profile. The MITRE ATT&CK mapping result is compared with all tactical phases in the MITRE ATT&CK matrix. Uncovered test blind spots are marked, third-phase records of attack primitives are generated, and the test blind spots are written into the blind spot list.

[0010] In a further preferred embodiment, the method further includes feeding back the blind zone list in step S3 to step S1, and retrieving the corresponding test case template from a preset tactical test case mapping library according to the tactical phase in the blind zone list. The test case template is semantically mutated using a large language model and then injected into a seed library.

[0011] In a further preferred embodiment, in step S1, the first-stage record of the attack primitive includes triggering conditions, call sequence fingerprint, scope of influence, and confidence score; In step S2, the second phase record of the attack primitive adds attack path, time window and topology context to the first phase record of the attack primitive. In step S3, the third-stage record of the attack primitives adds attacker profile, tactical mapping and blind spot list to the second-stage record of the attack primitives.

[0012] In a further preferred embodiment, the abnormal behavior in step S1 includes program crash, unauthorized access, and abnormal call chain.

[0013] In a further preferred embodiment, in step S2, the time constraint is to extract only the full logs for 10 minutes before and after the abnormal behavior; The entity constraint is to extract only entries related to processes, files, and ports within the scope of influence; The fingerprint matching method prioritizes extracting full logs that are similar to the system call sequence pattern.

[0014] In a further preferred embodiment, in step S2, the three node types include process nodes, file nodes, and network connection nodes; The four edge types include call edge, read / write edge, communication edge, and timing edge.

[0015] In a further preferred embodiment, in step S3, the key entities include the name of the attack tool, the type of the target service, and the characteristics of the attack method.

[0016] In a further preferred embodiment, in step S3, the attacker profile includes any one or more combinations of technical skill assessment, toolchain used, attack motive speculation, or possible organizational affiliation.

[0017] Option 2: A multi-stage vulnerability discovery and threat tracing system, wherein the system is used to implement any of the methods described in this invention, and the system includes: Vulnerability discovery module: Used to obtain system call logs and permission change logs in the isolation sandbox, and perform semantic analysis on code snippets, known vulnerability patterns and permission change logs of the target system to obtain test cases; Execute test cases and detect abnormal behavior based on system call logs; Extract the system call sequence at the moment of abnormal behavior and generate the first phase record of attack primitives; The attack tracing module retrieves a subset of the full logs that meet the time constraints, entity constraints, and fingerprint matching based on the first-stage record of the attack primitives. The full log subset is converted into a heterogeneous tracing graph that includes three node types and four edge types; The three node types are classified into two categories using a graph neural network to obtain attack behavior nodes; The attack behavior nodes are concatenated according to the time sequence to obtain the attack path candidate set, and the second phase record of the attack primitive is generated. Threat reasoning module: used to map attack behavior nodes to the tactical stages corresponding to the MITRE ATT&CK matrix, and obtain the MITRE ATT&CK mapping results; Key entities are extracted from the second-stage record of the attack primitives, and intelligence is retrieved from the closed attack and defense exercise knowledge base. The attack path, retrieved intelligence, and MITRE ATT&CK mapping results are then combined into a reasoning context. The inference context is input into a large language model to generate an attacker profile. The MITRE ATT&CK mapping result is compared with all tactical phases in the MITRE ATT&CK matrix. Uncovered test blind spots are marked, third-phase records of attack primitives are generated, and the test blind spots are written into the blind spot list.

[0018] Option 3: A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of any of the methods described in this invention.

[0019] Option 4: A computer-readable storage medium storing a computer program thereon, characterized in that the computer program, when executed by a processor, implements the steps of any of the methods described in this invention.

[0020] Option 5: A computer program product, comprising computer instructions, characterized in that, when executed by a processor, the computer instructions implement the steps of any of the methods described in this invention.

[0021] The above-mentioned technical solution of the present invention solves the problems of data silos between tools, resource waste in full log scanning, and lack of tactical perspective in test guidance used in the prior art, and achieves the following beneficial effects: 1. The sequential execution of the first, second, and third phase records of the attack primitive eliminates the need for manual data import and format conversion, saving on manual translation steps and improving efficiency.

[0022] 2. The second-stage recording process of attack primitives is based on the call sequence fingerprint and scope of influence recorded in the first stage of attack primitives. It is no longer necessary to scan the entire log. Only 5% to 10% of the relevant entries need to be processed, which saves a lot of computing resources and improves work efficiency.

[0023] 3. After each completion of the third phase of attack primitive recording, the blind zone list is returned to step S1. Based on the tactical phase in the blind zone list, the corresponding test case template is retrieved from the preset tactical test case mapping library, which solves the problem of lack of tactical perspective in test guidance. Attached Figure Description

[0024] Figure 1 This is a flowchart illustrating a multi-stage vulnerability discovery and threat tracing method and system as described in Implementation Method 1.

[0025] Figure 2 This is a schematic diagram illustrating the hierarchical construction process of the attack primitives of this invention. Detailed Implementation

[0026] To facilitate understanding of the technical solutions claimed in this invention, the following specific embodiments are provided in conjunction with the accompanying drawings. These embodiments are merely used to explain the technical principles of the embodiments of this application and are not intended to limit the scope of protection of this application. Those skilled in the art can make reasonable adjustments to the embodiments given in the following specific embodiments based on their general knowledge in the art to adapt them to specific application scenarios.

[0027] Implementation Method 1, see Figure 1 and Figure 2 This embodiment describes a multi-stage vulnerability discovery and threat tracing method, which includes the following steps: S1: Obtain system call logs and permission change logs from the isolation sandbox, and perform semantic analysis on code snippets, known vulnerability patterns, and permission change logs of the target system to obtain test cases; Execute test cases and detect abnormal behavior based on system call logs; Extract the system call sequence at the moment of abnormal behavior and generate the first phase record of attack primitives; S2: Based on the first-stage record of the attack primitives, retrieve a subset of the full logs that satisfy the time constraints, entity constraints and fingerprint matching from the full logs; The full log subset is converted into a heterogeneous tracing graph that includes three node types and four edge types; The three node types are classified into two categories using a graph neural network to obtain attack behavior nodes; The attack behavior nodes are concatenated according to the time sequence to obtain the attack path candidate set, and the second phase record of the attack primitive is generated.

[0028] S3: Map the attack behavior nodes to the tactical stages corresponding to the MITRE ATT&CK matrix to obtain the MITRE ATT&CK mapping results; Key entities are extracted from the second-stage record of the attack primitives, and intelligence is retrieved from the closed attack and defense exercise knowledge base. The attack path, retrieved intelligence, and MITRE ATT&CK mapping results are then combined into a reasoning context. The inference context is input into the large language model to generate an attacker profile. The MITRE ATT&CK mapping result is compared with all tactical phases in the MITRE ATT&CK matrix. Uncovered test blind spots are marked, third-phase records of attack primitives are generated, and the test blind spots are written into the blind spot list. The retrieval of a subset of full logs that meets the time constraints, entity constraints, and fingerprint matching is performed using the call sequence fingerprint as the keyword and the scope of influence as the retrieval boundary.

[0029] The attack and defense exercise knowledge base is an independently built and proprietary knowledge base of this invention. It is not purchased from external sources, nor is it a commercially available or open-source knowledge base. It has no fixed external source. This knowledge base is built by the company's internal red team and blue team over a long period of time. All materials come from the team's practical work output over the years. It includes three core scenarios: various attack and defense exercise competitions, experience summaries of network protection special operations at all levels, and routine daily attack and defense training and security handling work. It summarizes real attack and defense samples, attack behavior records and incident handling review materials from multiple cycles and scenarios.

[0030] The database contains structured storage of multi-dimensional attack and defense related data, specifically including: (1) Historical complete attack chain samples: simulated attacks of each APT, complete attack paths of red team penetration, and full process logs of vulnerability exploitation; (2) ATT&CK technical and tactical related materials: attack tools, execution commands, traffic characteristics, permission tampering behavior, and persistence methods corresponding to each tactical stage; (3) Attacker profile: the toolchains, attack preferences, typical behavior patterns, and technical level stratification of different attack groups and red team players; (4) Vulnerability and Exploitation Database: Triggering conditions, system call fingerprints, and scope of impact for various general vulnerabilities, local privilege escalation vulnerabilities, and privilege tampering vulnerabilities; (5) Historical event review intelligence: high-frequency attack methods, interference behaviors, defense false alarm samples, and tactical coverage blind spot cases that have appeared in the network protection and attack and defense competitions over the years.

[0031] The technical meaning of each field in this embodiment and its usage in subsequent methods are shown in Table 1.

[0032] Table 1

[0033] The system call sequence log includes open, read, write, and exec; the permission change log includes sudo, chmod, and setuid.

[0034] The two categories include nodes that are operating normally and nodes that are engaging in attack behavior.

[0035] MITRE ATT&CK, an abbreviation for Adversarial Tactics, Techniques, and Common Knowledge, is a cybersecurity knowledge base created by the non-profit organization MITRE. Based on real-world attack observation data, it systematically categorizes attacker behavior. The matrix comprises 14 standardized attack tactical phases: Reconnaissance (TA0043), Resource Development (TA0042), Initial Access (TA0001), Execution (TA0002), Persistence (TA0003), Privilege Escalation (TA0004), Defense and Avoidance (TA0005), Credential Access (TA0006), Discovery (TA0007), Lateral Movement (TA0008), Data Collection (TA0009), Command and Control (TA0011), Data Infiltration (TA0010), and Impact (TA0040). Combined with the technical solution of this invention, this matrix plays a core role in the entire three-phase system chain. The third-stage threat inference module automatically maps the attack paths identified by the second-stage graph neural network to the corresponding tactic numbers of MITRE ATT&CK node by node through the tactic_mapping field, realizing the standardized classification of attack behaviors; based on the process, file, network, and permission change events corresponding to the attack behavior nodes output by the graph neural network binary classification, it matches the subdivided attack techniques under each MITRE ATT&CK tactic, and completely restores the attacker's complete kill chain. After mapping is completed, the system compares the covered tactics with all 14 MITRE ATT&CK tactics. The tactical phases that are not hit are stored in the attack primitive blind zone list, which serves as a directional test guidance signal for the first-stage vulnerability discovery module. Unlike traditional fuzz testing that is only guided by code coverage, this invention uses the MITRE ATT&CK tactical blind zone to drive semantic mutation of the large language model to generate targeted test cases, supplement the seed library, and fill in the missing attack tactic test samples in attack and defense exercise scenarios, thus solving the problem of incomplete tactical coverage in traditional tools. The invention has built a closed attack and defense exercise knowledge base, which fully stores the association mapping relationship between MITRE ATT&CK tactics, techniques and vulnerability exploitation, permission tampering and system call fingerprints. In the third stage of retrieval, the MITRE ATT&CK tactical tags are used as the retrieval dimension to match historical attack cases of the same tactics and attacker toolchain characteristics, supporting the large language model to complete attacker profile reasoning. The final structured analysis report uses the MITRE ATT&CK matrix as a unified standard, associating vulnerability discovery results, attack chains reconstructed by graph neural networks, and attacker profiles with corresponding MITRE ATT&CK tactical identifiers, and outputting standardized and auditable red team / blue team adversarial analysis conclusions.

[0036] Unlike the random mutation of traditional fuzzing tools, large language models, based on an understanding of code logic, can generate test cases that are more likely to hit boundary conditions and abnormal paths, and each generated test case comes with a trigger condition label.

[0037] In this embodiment, the method further includes feeding back the blind zone list in step S3 to step S1, and retrieving the corresponding test case template from the preset tactical test case mapping library according to the tactical phase in the blind zone list. The test case template is semantically mutated using a large language model and injected into a seed library. This is used to supplement test coverage for blind spots in testing. This step only expands the content of the seed library and does not modify any model parameters.

[0038] In this embodiment, in step S1, the first-stage record of the attack primitive includes the triggering condition, the call sequence fingerprint, the scope of influence, and the confidence score. In step S2, the second phase record of the attack primitive adds attack path, time window and topology context to the first phase record of the attack primitive. In step S3, the third-stage record of the attack primitives adds attacker profile, tactical mapping and blind spot list to the second-stage record of the attack primitives.

[0039] In this embodiment, the abnormal behavior in step S1 includes program crash, unauthorized access, and abnormal call chain.

[0040] In this embodiment, the time constraint in step S2 is to extract only the full logs of the 10 minutes before and after the abnormal behavior. The entity constraint is to extract only entries related to processes, files, and ports within the scope of influence; The fingerprint matching method prioritizes extracting full logs that are similar to the system call sequence pattern.

[0041] The heterogeneous source graph is input into the graph neural network. Through multiple rounds of message iteration, the graph neural network enables each node to aggregate the feature information of its neighborhood, and then performs binary classification on the node type.

[0042] By employing triple constraints of time constraints, entity constraints, and fingerprint matching, the search scope can be reduced to 5% to 10% of the entire log.

[0043] In this embodiment, the three node types in step S2 include process nodes, file nodes, and network connection nodes. The four edge types include call edge, read / write edge, communication edge, and timing edge.

[0044] Call edges represent the parent-child relationship between process nodes, read-write edges represent the operation of a process node on a file node, communication edges represent data transmission between network connection nodes, and timing edges represent the order of events.

[0045] After obtaining the second-stage record of the attack primitives, receive the complete second-stage record of the attack primitives containing the attack path, and complete tactical mapping, attacker profile generation, and blind spot marking.

[0046] In the tactical phase, for example: port scanning is mapped to TA0043 (reconnaissance), vulnerability exploitation is mapped to TA0002 (execution), and privilege escalation is mapped to TA0004 (privilege escalation).

[0047] Key entities are extracted from the second-stage record of the attack primitives. Intelligence is retrieved from a closed attack and defense exercise knowledge base using the key entities as query terms. The attack path, retrieved intelligence, and MITRE ATT&CK mapping results are concatenated to form the reasoning context. A closed knowledge base is used here instead of publicly available Internet data to ensure the reliability and confidentiality of the reasoning.

[0048] In this embodiment, the key entities in step S3 include the name of the attack tool, the type of the target service, and the characteristics of the attack method.

[0049] In this embodiment, the attacker profile in step S3 includes at least one or more combinations of technical level assessment, toolchain used, attack motive speculation, or possible organizational affiliation.

[0050] The large language model and graph neural network used in this invention process data separately. The large language model runs in the first and third stages, while the graph neural network runs only in the second stage. They are not integrated in the same computational process, but rather cooperate indirectly by attacking the structured fields of the primitives.

[0051] In this invention, after the third phase of the attack primitive records the blind spot list feedback step S3, what is changed is the content of the seed library (adding test cases for blind spot tactics), rather than modifying the model parameters of the large language model or graph neural network.

[0052] Implementation Method Two: This implementation method describes the experimental environment for the multi-stage vulnerability discovery and threat tracing method described in Implementation Method Two.

[0053] The network environment required for the experiment consisted of 12 isolated sandboxes, divided into three zones: 2 sandboxes in the DMZ for running web services and an email gateway, 6 sandboxes in the office terminal zone (running Windows 10 and Ubuntu 20.04 desktop versions), and 4 sandboxes in the intranet server zone (two Windows Server 2019 and two CentOS 7). This scale was chosen because a single experimental cycle would be too long, and a smaller scale would not be sufficient to simulate lateral movement scenarios.

[0054] The full log collection took 72 hours, collecting approximately 820,000 records from sources including Sysmon event logs, Windows security audit logs, network traffic logs generated by Zeek, and EDR terminal logs. Scanning this volume of full logs to build a heterogeneous source map would take about 4 minutes (on the test server of this invention, configured with a Xeon E5-2680v4 and 64GB of memory). After applying time constraints, entity constraints, and fingerprint matching retrieval, this time was reduced to about 1 minute and 10 seconds, and CPU usage decreased from nearly 100% to around 30%.

[0055] Regarding the samples of attack behavior nodes, this invention references ATT&CK v13, injecting 23 technical points corresponding to 9 tactical stages. The attack paths mainly refer to the publicly available APT29 simulation scheme, which includes 3 real attack paths plus 2 interference paths (the interference paths are used to measure the false positive rate of the graph neural network).

[0056] The first round of testing covered about 10 of the 14 tactical phases (approximately 73.9%). After feeding back the blind spot list from step S3 to step S1, the second round of testing covered TA0003 (initial access) and TA0010 (data leakage). The third round covered the remaining sub-techniques within the initial access phase, bringing the overall success rate to 92.9%. However, this data is highly dependent on the sample type of the injected attack nodes; changing the sample may result in fluctuations of a few points.

[0057] By having three SOC analysts with over five years of experience independently reconstruct the attack path for the same scenario, and averaging their responses, a human benchmark for comparison was established. The fastest analyst took 1 hour and 12 minutes, the slowest nearly 3 hours, and the average time was 2 hours. The 38-second time on the system side is the end-to-end time from when the vulnerability alert is triggered to when the attack path field is output in the second stage, excluding the attacker profiling generation in the third stage (which requires an additional ten seconds).

[0058] Implementation Method 3 is an example illustrating the multi-stage vulnerability discovery and threat tracing method described in Implementation Method 3.

[0059] A company conducted its annual attack and defense drill. The red team (attackers) gained initial access via SSH (Secure Shell) brute-force attack, then exploited a local privilege escalation vulnerability to obtain superuser privileges. They subsequently moved laterally to other servers on the internal network and stole sensitive data. The blue team (defenders) deployed the system of this invention to automate the analysis of the drill process.

[0060] The environment deployment for this embodiment is shown in Table 2.

[0061] Table 2

[0062] Taking an SSH brute-force attack scenario as an example, the complete data flow process of the system is as follows: S1: After analyzing the SSH service code of the target server using a large language model, test cases containing numerous combinations of abnormal behaviors were generated. After execution in the isolation sandbox, an SSH login exception was captured (test case number TC-0087). The system call sequence fingerprint sha256 was extracted: b4e97289d0fc45ae2187bc96e370581cc297466da83bf1590ce72edc9f8f301, affecting process pid: 2341, file / etc / passwd, and port 22. The confidence score was 0.92. Based on this, the first phase of the attack primitives was generated.

[0063] S2: Read the call sequence fingerprint and impact range from the first stage of the attack primitives to construct retrieval constraints. Extract 62,000 relevant logs (approximately 6.2%) from 1 million full logs to construct a heterogeneous source graph. Graph neural network identifies the attack path: SSH brute-force login → superuser privilege escalation → / etc / shadow reading → lateral movement to the 192.168.1.0 / 24 intranet segment. Attack time window: 14:23:05~14:47:32. After appending 3 fields, the second stage of the attack primitives record is formed.

[0064] Phase 3: Each attack behavior node in the attack path is mapped to the MITRE ATT&CK matrix. Covered tactical phases include TA0001 (reconnaissance), TA0006 (credential access), TA0004 (privilege escalation), and TA0008 (lateral movement). Uncovered phases are TA0003 (initial access), TA0005 (persistence), and TA0010 (data exfiltration), marked as blind spots. After RAG retrieval, a large language model generates an attacker profile: intermediate technical level, using the Hydra+LinPEAS toolchain, with the objective of data theft. Adding 3 fields creates a Phase 3 attack primitive record (10 fields). The blind_spots list is fed back to the Phase 1 seed database to supplement initial access, persistence, and data exfiltration test cases.

[0065] The key parameters used in this embodiment are shown in Table 3.

[0066] Table 3

[0067] Based on the data in Table 3, it can be determined that when the anchored retrieval time window is set to 10 minutes, the log extraction volume is approximately 9.8% of the total logs. Combined with triple constraints, the overall retrieval scope can be reduced to 5%~10% of the total logs. When the confidence threshold is set to 0.85, approximately 8.2% of false alarms can be filtered while retaining 91.8% of real vulnerabilities. When the GNN message passing rounds are set to 3, the path identification accuracy can reach 98%. When the RAG retrieval Top-K is set to 5, the inference accuracy can reach 97.5%. When each blind zone is expanded by 10 seeds, approximately 8 new tactical phases can be added to the coverage per iteration, and the tactical coverage rate increases from 73.9% to 92.9% after three iterations. When the LLM mutation temperature is set to 0.6, the optimal balance between test case diversity and effectiveness can be achieved. Under the above optimal parameter combination, the end-to-end attack path reconstruction time is approximately 38 seconds, and the computational resource consumption is reduced by approximately 70% compared to the full analysis, verifying the synergistic effectiveness between the various technical means of this invention.

Claims

1. A multi-stage vulnerability discovery and threat attribution method, characterized in that, The method includes the following steps: S1: Obtain system call logs and permission change logs from the isolation sandbox, and perform semantic analysis on code snippets, known vulnerability patterns, and permission change logs of the target system to obtain test cases; Execute test cases and detect abnormal behavior based on system call logs; Extract the system call sequence at the moment of abnormal behavior and generate the first phase record of attack primitives; S2: Based on the first-stage record of the attack primitives, retrieve a subset of the full logs that satisfy the time constraints, entity constraints and fingerprint matching from the full logs; The full log subset is converted into a heterogeneous tracing graph that includes three node types and four edge types; The three node types are classified into two categories using a graph neural network to obtain attack behavior nodes; The attack behavior nodes are concatenated according to the time sequence to obtain the attack path candidate set, and the second phase record of the attack primitive is generated. S3: Map the attack behavior nodes to the tactical stages corresponding to the MITRE ATT&CK matrix to obtain the MITRE ATT&CK mapping results; Key entities are extracted from the second-stage record of the attack primitives, and intelligence is retrieved from the closed attack and defense exercise knowledge base. The attack path, retrieved intelligence, and MITRE ATT&CK mapping results are then combined into a reasoning context. The inference context is input into the large language model to generate an attacker profile. The MITRE ATT&CK mapping result is compared with all tactical phases in the MITRE ATT&CK matrix. Uncovered test blind spots are marked, third-phase records of attack primitives are generated, and the test blind spots are written into the blind spot list. The method further includes feeding back the blind zone list in step S3 to step S1, and retrieving the corresponding test case template from the preset tactical test case mapping library according to the tactical phase in the blind zone list. The test case template is semantically mutated using a large language model and then injected into a seed library. In step S3, the key entities include the name of the attack tool, the type of the target service, and the characteristics of the attack method.

2. The multi-stage vulnerability discovery and threat tracing method according to claim 1, characterized in that, In step S1, the first-stage record of the attack primitive includes the triggering condition, the call sequence fingerprint, the scope of influence, and the confidence score. In step S2, the second phase record of the attack primitive adds attack path, time window and topology context to the first phase record of the attack primitive. In step S3, the third-stage record of the attack primitives adds attacker profile, tactical mapping and blind spot list to the second-stage record of the attack primitives.

3. The multi-stage vulnerability discovery and threat tracing method according to claim 1, characterized in that, In step S1, the abnormal behaviors include program crashes, unauthorized access, and abnormal call chains.

4. The multi-stage vulnerability discovery and threat tracing method according to claim 1, characterized in that, In step S2, the time constraint is to extract the full logs for only 10 minutes before and after the abnormal behavior. The entity constraint is to extract only entries related to processes, files, and ports within the scope of influence; The fingerprint matching method prioritizes extracting full logs that are similar to the system call sequence pattern.

5. The multi-stage vulnerability discovery and threat tracing method according to claim 1, characterized in that, In step S2, the three node types include process nodes, file nodes, and network connection nodes. The four edge types include call edge, read / write edge, communication edge, and timing edge.

6. The multi-stage vulnerability discovery and threat tracing method according to claim 1, characterized in that, In step S3, the attacker profile includes any one or more combinations of technical skill assessment, toolchain used, attack motive speculation, or organizational affiliation.

7. A multi-stage vulnerability discovery and threat tracing system, characterized in that, The system is used to implement the method according to any one of claims 1 to 6, the system comprising: Vulnerability discovery module: Used to obtain system call logs and permission change logs in the isolation sandbox, and perform semantic analysis on code snippets, known vulnerability patterns and permission change logs of the target system to obtain test cases; Execute test cases and detect abnormal behavior based on system call logs; Extract the system call sequence at the moment of abnormal behavior and generate the first phase record of attack primitives; The attack tracing module retrieves a subset of the full logs that meet the time constraints, entity constraints, and fingerprint matching based on the first-stage record of the attack primitives. The full log subset is converted into a heterogeneous tracing graph that includes three node types and four edge types; The three node types are classified into two categories using a graph neural network to obtain attack behavior nodes; The attack behavior nodes are concatenated according to the time sequence to obtain the attack path candidate set, and the second phase record of the attack primitive is generated. Threat reasoning module: used to map attack behavior nodes to the tactical stages corresponding to the MITRE ATT&CK matrix, and obtain the MITRE ATT&CK mapping results; Key entities are extracted from the second-stage record of the attack primitives, and intelligence is retrieved from the closed attack and defense exercise knowledge base. The attack path, retrieved intelligence, and MITRE ATT&CK mapping results are then combined into a reasoning context. The inference context is input into a large language model to generate an attacker profile. The MITRE ATT&CK mapping result is compared with all tactical phases in the MITRE ATT&CK matrix. Uncovered test blind spots are marked, third-phase records of attack primitives are generated, and the test blind spots are written into the blind spot list.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes a computer program to implement the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • International network attack prediction method and system based on fusion of GNN and LLM

    CN120389903A

  • Vulnerability closed-loop processing method and system based on intelligent collaboration

    CN120579194A

  • Algorithm test method and device based on KG enhanced large language model

    CN120492571A

  • Data security event real-time monitoring method and system

    CN120528657A