A verifiable apt attack chain extraction method based on cognitive distillation and gating technology
By combining cognitive distillation with gating technology, the problems of stage misjudgment and knowledge base lag in APT attack chain reconstruction are solved, enabling accurate extraction and real-time defense of APT attack chains, and improving the robustness and accuracy of the defense system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU UNIVERSITY
- Filing Date
- 2026-01-26
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies suffer from stage misjudgment, outdated static knowledge base, lack of semantic interference defense, and fragmented semantic understanding when extracting APT attack chains, resulting in inaccurate attack chain reconstruction and delayed defense.
By employing cognitive distillation and gating techniques, a semantic understanding engine is used to achieve deep alignment between unstructured text and graph nodes. The attack phase transition patterns are dynamically modeled, and an incremental learning mechanism is used to integrate novel attack features. Combined with multi-head attention mechanism and long short-term memory network analysis of historical patterns, gating weights are generated to perform hierarchical decision-making and ensure adversarial robustness.
It achieves accurate mapping and real-time response to APT attack chains, improves attack chain reconstruction efficiency, reduces false positive rate, and enhances the dynamic adaptability of the defense system.
Smart Images

Figure CN121567486B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network security technology, and in particular to a verifiable APT attack chain extraction method based on cognitive distillation and gating technology. Background Technology
[0002] Cybersecurity defense systems are facing a severe challenge from Advanced Persistent Threats (APTs). These attacks, launched by highly specialized organizations, aim to steal sensitive data, conduct espionage, or disrupt critical infrastructure. They are characterized by strong organization, clear targeting, high technical complexity, and long-term, covert infiltration. The essence of APT attacks lies in their highly structured, multi-stage technical depth: attack behavior typically follows a progressive framework similar to "initial penetration → lateral movement → target achievement," with close logical connections and tactical dependencies between the technical means used at each stage (such as spear phishing, credential theft, and command and control). This systematic nature requires defenders to build protective capabilities from a global attack evolution perspective, rather than responding only to isolated attack events.
[0003] Threat intelligence, as a core resource in combating APT attacks, describes the tactics, techniques, and procedures (TPs) of attack organizations through unstructured text (such as security incident reports and vulnerability announcements). This intelligence is crucial for understanding attacker intent, detecting anomalous behavior, and tracing the origins of attack organizations. However, with the continuous evolution of APT attack techniques and the increasing sophistication of threat intelligence, a fundamental technical bottleneck has become increasingly apparent: how to automatically reconstruct the attack chain and TTP chain.
[0004] Existing technologies have yielded some results in the research on attack chains or tactical chains that extract common features of responses from threat intelligence, and have formed different technical routes, such as system log behavior chain analysis technology, semantic parsing and ontology mapping technology, multimodal fusion reasoning technology, and lightweight behavior feature extraction technology.
[0005] However, existing techniques for extracting attack chains still have the following limitations:
[0006] 1. Discontinuity in the attack chain timing leads to misjudgment of stages.
[0007] Existing technologies generally lack explicit modeling of the causal logic between attack phases, failing to quantify temporal constraints such as "macro activation must precede script triggering." Their positional encoding merely marks text order rather than tactical dependencies, leading to isolated phase segmentation. When dealing with complex actions such as "enabling macros to trigger PowerShell download," traditional character-level masking strategies decompose technical entities (such as "PowerShell") into discrete characters, disrupting the complete semantic understanding of the "payload delivery pipeline." This fragmented parsing causes phase classification errors in the Conti case (such as confusing "initial access" with "execution"), stemming from the lack of a mathematically structured model of phase relationships.
[0008] 2. Static knowledge base lags behind in identifying new types of attacks.
[0009] Traditional solutions rely on predefined rule bases and manual annotation mechanisms, resulting in slow responses to novel attacks such as "cloud storage API leaks." Their knowledge update cycles are typically long, and tactical mappings require manual verification, leading to significant delays in MITRE adoption. Existing models cannot dynamically distill the essence of attacks (e.g., stripping away the "API" legitimacy facade to focus on the "data leak" core), only mechanically matching known features. When attackers abuse legitimate services (such as OneDrive), the rule engine misses detections due to a lack of real-time knowledge expansion capabilities, stemming from the absence of a self-evolving cognitive distillation system.
[0010] 3. Lack of semantic interference defenses led to system crashes.
[0011] Existing methods lack robustness guarantees against semantic tampering (such as "PowerShell → Command Line"). Their rule filtering cannot distinguish the tactical homology of technical variants (such as memory-resident vs. disk write), resulting in a false positive rate as high as 38%. More seriously, the system lacks probability distribution stability monitoring, leading to uncontrolled decision shifts after tampering. The black-box model lacks decision traceability capabilities (such as attention weight visualization), making it impossible for analysts to verify whether "command line download" matches ransomware behavior patterns. The root cause lies in the defense design's failure to integrate noise simulation with mathematical stability proofs.
[0012] 4. Fragmented semantic understanding weakens entity association.
[0013] Character-level masking strategies fragment the semantic integrity of complex technical entities (e.g., "PowerShell" is broken down into character fragments), disrupting the coherent expression of key tactical features such as "pipeline operation chains." Existing positional encoding only marks text positions and lacks phase constraints for action logic such as "macro activation takes precedence over script triggering," resulting in the inability to quantify the strength of stage dependencies. Multi-head attention mechanisms fail to separate the technical carrier dimension (e.g., the physical characteristics of Excel attachments) from the malicious behavior dimension (e.g., the aggressiveness of download actions), leading to semantic confusion under ambiguous descriptions. The root cause lies in the lack of a collaborative architecture between full-word masking and phase-aware encoding. Summary of the Invention
[0014] The purpose of this invention is to provide a verifiable APT attack chain extraction method based on cognitive distillation and gating technology, and to construct a multi-engine collaborative intelligent analysis framework: The semantic understanding engine achieves deep alignment between unstructured text and graph nodes through a domain-adaptive pre-trained model, accurately mapping "PowerShell pipeline operations" to ATT&CK tactical T1059.001 (command interpreter), significantly improving semantic disambiguation accuracy. The dynamic coupling modeling engine captures attack phase transition patterns (such as the causal constraint of "macro activation → script triggering"), solving the problem of attack chain breakage. The incremental learning mechanism integrates new attack features (such as "cloud storage API leakage" → dynamic encoding T1537) under the protection of KL divergence anchoring, realizing real-time evolution of the knowledge base. This technology bridges the technical gap between dynamic cognition and real-time response of threat intelligence: the semantic upscaling stripping technology focuses on the attack kernel (such as downweighting the "API" legal attribute to 0.31 and upweighting "leakage" to 0.69), temporal modeling enforces causal constraints between phases, and incremental evolution breaks through the bottleneck of static knowledge base lag. In the SolarWinds incident, the improved efficiency of attack chain reconstruction contributed to the shift of APT defense from "feature matching" to "cognitive immunity."
[0015] The present invention provides a verifiable APT attack chain extraction method based on cognitive distillation and gating technology, which adopts the following technical solution:
[0016] A verifiable APT attack chain extraction method based on cognitive distillation and gating techniques, specifically including:
[0017] S1. Input processing and extraction, including semantic atomization aggregation of textual threat intelligence, fusion of fragmented technical descriptions into tactical units, locking the causal chain between stages, initiating multi-dimensional feature distillation, extracting four types of tactical paths: attack medium, execution fingerprint, payload pattern and infiltration path, constructing an anti-obfuscation topology network, disambiguating fuzzy semantics based on context distance weight, and generating a tactical vector map with core attack nodes as the hub.
[0018] S2. Dynamic gating weight generation includes positional encoding of the input attack sequence, capturing the temporal relationship and dependency logic between different attack actions, using multi-head attention mechanism to focus on the technical implementation of identifying the carrier and malicious behavior characteristics, identifying key elements in the attack, and using long short-term memory network to analyze historical attack patterns, calculate the correlation strength between the current action and historical behavior, and generate gating weight values by fusing historical dependencies and semantic similarity.
[0019] S3. Hierarchical Prototype Decision Making: Based on the generated gating weights and tactical vector graph, hierarchical decision making is performed. This includes combining the input vector with information from the knowledge base, activating the few-shot reasoning ability of the large language model, determining the current attack mode, and if it is a novel attack mode, then the essential features of the attack are distilled out through a self-attention mechanism to generate a new tactical vector. This vector is then compared with known attack modes for similarity, and mapped to the existing tactical coding system or marked as an unknown attack requiring further research based on the similarity result. If it is a conventional attack, a standard matching process is used. The standard matching process refers to the attack mode matching process based on a predefined knowledge base and a similarity threshold. Specifically, it includes: retrieving the known attack mode most similar to the current tactical vector from the ATT&CK framework knowledge base, calculating the cosine similarity, and directly mapping to the corresponding TTP code when the similarity exceeds the preset threshold of 0.85. When the similarity is insufficient, it is marked as requiring manual review.
[0020] S4. To ensure robustness against attacks and output attack chains, the TTP encoding output from the hierarchical prototype decision in step 3 (i.e., the attack tactical feature encoding based on the MITRE ATT&CK framework) is injected with noise of a specific intensity. This TTP encoding is a standard attack feature identifier generated by performing hierarchical decision-making on the tactical vector output in step 2 using the LLaMA-3 model. Injecting noise of a specific intensity simulates possible semantic tampering by attackers. The model output results of the original encoding and the noisy encoding are calculated separately. The system stability is evaluated by comparing the differences in the model output results. This evaluation includes quantifying the interference level by combining the distance metric of the output vector and the relative entropy of the probability distribution. When interference exceeding a threshold is detected, a semantic reconstruction mechanism is activated to correct and restore the encoding based on attention focus and historical tactical logic. For cases where the threshold is not exceeded, a gating weight value is generated to update the signal, forming a defensive closed loop. The TTP encoding is a standard attack feature identifier generated by performing hierarchical decision-making on the tactical vector output in step S1 using the LLaMA-3 model.
[0021] Specifically, the process of performing semantic atomization aggregation on textual threat intelligence and integrating fragmented technical descriptions into tactical units in S1 includes performing semantic standardization on the original descriptions through a text preprocessing engine, aggregating scattered technical elements, binding malicious Excel attachments containing macro code as delivery carriers for tactical units, and retaining the initial access attributes of email distribution.
[0022] Furthermore, the causal chain between stages in S1 includes the causal chain of "enabling macros → triggering PowerShell → downloading ransomware".
[0023] Furthermore, in S2, the input attack sequence is positionally encoded and then understood using a full-word masking strategy, wherein the masking loss function is:
[0024] ;
[0025] in This represents the loss value of the masked language model, indicating the degree of error in the model's prediction of the masked words; This is the mask position index, representing the position number of the masked token in the sequence; It is a mask set, containing the set of all masked token positions in the current sequence; The true value of the target token, representing the value at the 1st... The original token before the masked position; This is non-mask context information, representing information other than the mask set. The set of all visible tokens outside; This is the conditional probability value, representing the probability that the model predicts the target token based on the context;
[0026] The position code is:
[0027] ;
[0028] in The location-coded value represents the location. The position encoding at dimension index 2i; The position index represents the sequential position (integer, starting from 0) of the token or attack phase in the sequence. For dimension index, it represents the dimension number (integer, from 0 to d / 2-1) in the location encoding vector. 10000 is the model dimension, representing the total dimension size of the location coding vector (usually 768); 10000 is the frequency cardinality, a constant value that controls the frequency distribution of location coding.
[0029] In the multi-head attention mechanism described in S2, the third attention head focuses on the carrier dimension, capturing the physical characteristics of the Excel attachment as the initial access carrier; the seventh attention head locks onto the malicious behavior dimension, reinforcing the aggressive nature of the download action; and the eleventh attention head analyzes the tactical intent through a gating weight formula, which is:
[0030] ;
[0031] in The gating weight value represents the quantification strength of the dependency relationship between attack phases, and its value ranges from (0,1).
[0032] 0.6 is the sigmoid activation function, used to compress the linear combination result into the probability interval; 0.6 is the history dependency weight coefficient, reflecting the strong temporal prior characteristic of the attack chain; 0.4 represents the historical dependency probability, which is the output of the temporal modeling of historical patterns of attack sequences based on the LSTM network; 0.4 is the semantic similarity weight coefficient, which balances the consistency requirements of the current semantic context. For semantic similarity, the cosine similarity is used to calculate the semantic coherence between the current action and the context.
[0033] If text is altered to vaguely describe system operations, a semantic weighting mechanism is activated. This semantic weighting mechanism is as follows:
[0034] ;
[0035] in, This is the final gating weight value. The output of the gate weight value after semantic security verification; The model predicts the probability of a keyword term based on the context, and then predicts the conditional probability of the correct term. The semantic credibility threshold. A value of 0.5 is used to determine the credibility of the semantic description.
[0036] The algorithm flow for S3 is as follows:
[0037] S31. Obtain the input tactical vector, and concatenate the tactical vector with a preset prototype feature set according to a preset format to construct a prompt word template sequence;
[0038] S32. Input the prompt word template sequence into the LLaMA-3-70B model for inference calculation and obtain the attack behavior analysis results output by the model;
[0039] S33. Analyze the attack behavior analysis results and determine whether the currently input tactical vector belongs to a new type of attack through preset discrimination logic;
[0040] S34. If it is determined to be a new type of attack, the attack attributes in the attack behavior analysis results are encoded and calculated to generate a new prototype vector, the new prototype vector is updated to the prototype feature set, and the corresponding tactical and technical process code is generated based on the new prototype vector.
[0041] S35. If the attack type is determined to be known, the feature distance between the tactical vector and each existing vector in the prototype feature set is calculated, the target prototype vector with the smallest feature distance is selected, and the corresponding tactical technical process code is generated based on the target prototype vector.
[0042] S36. Output the finalized tactical and technical process code;
[0043] In S32, the hierarchical prototype decision-making stage achieves precise mapping of tactical intent through the LLaMA-3-70B model, when processing the gating weight vector of the dynamic gating output. At this time, the model first performs basic tactical classification:
[0044]
[0045] This prompt template activates the model's task-aware state via the [INST] instruction header, and sets the gating weights. =0.89 is used as the historical dependency prior injection decision system. During the decoding process, an attention head is selected as the attention head with the highest focus weight on PowerShell.
[0046] S33 includes the sub-prototype fine matching stage's enhanced technical identification capability:
[0047] ;
[0048] Model construction feature comparison matrix The differences in row vector quantization techniques are as follows:
[0049] ;
[0050] Among them, the first column, which features PowerShell-specific Get-Content | Invoke-Expression pipeline chain operations, has a weight of 0.85, and the second column, which exposes batch installation defects in command lines, has a weight of 0.12.
[0051] Dynamic expansion mechanisms trigger semantic distillation:
[0052] ;
[0053] In the 12-layer self-attention mechanism, the model performs hierarchical feature decoupling, reducing the weight of legitimate "API" attributes (stripping away technical appearances) and increasing the weight of "leakage" attack kernels (focusing on tactical essence). Generate vectors. The mathematical expression is:
[0054] ;
[0055] in The newly generated feature vector represents the attack kernel features extracted after hierarchical feature decoupling; It is an encoder function that maps text descriptions to a high-dimensional feature space. It is an average pooling operation that reduces the dimensionality and aggregates the encoded features;
[0056] Similarity verification formula:
[0057] ;
[0058] in It is the cosine similarity value, ranging from [-1, 1], which represents the degree of directional similarity between two vectors; It is the feature vector of the T1537 attack, representing the T1537 attack pattern features extracted from the data; It is the feature vector of Conti ransomware, representing the feature representation of the Conti attack family; It is the formula for calculating cosine similarity, which measures the similarity of two vectors in a direction.
[0059] The algorithm flow of S4 includes:
[0060] S41: Obtain the tactical and technical process code, and superimpose a preset Gaussian noise perturbation on the tactical and technical process code to generate the corresponding adversarial sample code;
[0061] S42: Input the tactical technical process code and the adversarial sample code into the preset verification model for inference, and obtain the original output features and adversarial output features respectively;
[0062] S43: Based on the original output features and the adversarial output features, calculate the robustness loss value that characterizes the stability of the model output. The robustness loss value includes at least the Euclidean distance and KL divergence between the original output features and the adversarial output features.
[0063] S44: Determine whether the robustness loss value is greater than a preset threshold;
[0064] S45: If the robustness loss value is determined to be greater than the threshold, then construct an error correction prompt word containing attack type information, input it into a preset correction model to generate and output the corrected tactical and technical process code;
[0065] S46: If the robustness loss value is determined to be less than or equal to the threshold, the preset gating weight parameters are updated based on the KL divergence calculation parameter increment, and the updated weight parameters are output.
[0066] In Phase Two, when the key text is maliciously altered to trigger the command-line tool to download the encryption module after macros are enabled, the system activates a triple collaborative protection mechanism upon detecting semantic tampering: First, noise injection training simulates attacker behavior; second, KL divergence distribution constraints and system stability monitoring; third, a semantic reconstruction correction engine restores the correct semantics. These three mechanisms form a complete defense-cognition closed loop, ensuring the robustness and accuracy of attack chain analysis. The noise injection training method is as follows:
[0067] ;
[0068] in For adversarial examples, For the original input, The noise intensity coefficient, It is Gaussian noise;
[0069] The core defense response mechanism is implemented through a KL divergence distribution system, which is constructed as follows:
[0070] ;
[0071] Full power start;
[0072] in For robust loss, The Kullback-Leibler divergence;
[0073] The semantic reconstruction and correction engine construction instructions are as follows:
[0074] ;
[0075] in To correct the output, the instruction triggers a four-layer cognitive analysis. The first layer, an attention mechanism, focuses on the essence of the download action with a certain weight, removing the interference of the tool attributes of the command line. The second layer analyzes the semantic relationship of the context. The third layer associates the phishing attack characteristics of the historical stage T1566.001 to confirm that the subsequent action should be an execution-type action. Finally, the output is accurately mapped to T1059.005 and maintains a high confidence level. The fourth layer performs confidence-weighted decision fusion.
[0076] In S4, the defense feedback closed loop is updated using an adaptive formula, which is:
[0077] ;
[0078] in ;in This is the updated defense strength parameter, used to dynamically adjust the strength level of the system's adversarial defense. The value range is [0,1], with a larger value indicating a stronger defense.
[0079] In step S4, noise is added by adding a noise intensity parameter. Implementation, including:
[0080] ;
[0081] in It is the noise intensity parameter of the t-th iteration, the noise injection coefficient at the current time, which controls the noise amplitude when the adversarial example is generated; It is the noise intensity parameter of the (t-1)th iteration, and the noise injection coefficient of the previous moment, which serves as the starting point for the current optimization. It is the gradient descent learning rate, a hyperparameter that controls the step size of parameter updates, and determines the convergence speed and stability of noise intensity optimization. A value that is too large will cause oscillations, while a value that is too small will cause slow convergence. It is the partial derivative of robustness loss with respect to noise intensity, representing the gradient of the effect of changes in noise intensity on the defense effect.
[0082] The beneficial effects of the verifiable APT attack chain extraction method based on cognitive distillation and gating technology provided by this invention are as follows:
[0083] 1. A semantic continuity modeling and dynamic temporal gating mechanism is created, achieving a fundamental breakthrough through a full-word masking strategy and phase-aware positional encoding: The full-word masking loss function forces the model to learn composite technical entities such as "PowerShell" as complete semantic units, improving the accuracy of term recognition in the Conti case and completely solving the semantic fragmentation problem caused by traditional character masking; Positional encoding explicitly injects the temporal logic of attack actions, and its unique frequency decay design accurately models the causal constraint that "macro activation must precede script triggering"; The multi-head attention gating formula innovatively integrates LSTM historical dependency probability and semantic similarity, triggering a hard weighting mechanism when encountering ambiguous descriptions, thus blocking semantic noise pollution from a mathematical perspective.
[0084] 2. A dynamic knowledge evolution system driven by cognitive distillation is constructed, employing a prototyping matching technique based on prompts and a decoupling technique for self-attention features: structured prompt templates activate the model's few-sample reasoning ability, and cosine distance is used to quantify decision determinism, enabling precise matching of threat behaviors with tactical prototypes; the self-attention mechanism performs tactical essence distillation, intelligently stripping away technical appearances and focusing on the attack core when analyzing new attacks; dynamically expanding verification formulas establish quantitative standards for tactical homology, driving real-time evolution of the knowledge base and achieving a cognitive leap from specific cases to general tactics. This dynamic knowledge evolution mechanism improves the response speed to new attacks such as "cloud storage API leaks".
[0085] 3. Create a verifiable adversarial immunity system by constructing a mathematical-level defense system through noise parameterization injection and distribution constraint closed loop: Gaussian noise injection achieves accurate simulation of attacker behavior, and parameters are optimized through massive adversarial training to form an industrial-grade defense benchmark; KL divergence double constraint innovatively integrates output space anchoring and probability distribution deformation monitoring, providing mathematical proof for decision stability; the semantic reconstruction engine reconstructs the logic chain through attention focus and historical dependency, and its adaptive feedback forms the world's first defense-cognition closed loop, suppressing the false positive rate of tampering attacks to a high accuracy level in the Conti ransomware test. Attached Figure Description
[0086] Figure 1 This is a schematic diagram of the process of the present invention;
[0087] Figure 2 This is a flowchart illustrating step S3 in this invention;
[0088] Figure 3 This is a flowchart illustrating step S4 in this invention. Detailed Implementation
[0089] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Unless otherwise defined, the technical or scientific terms used herein should have the ordinary meaning understood by those skilled in the art. The terms "comprising" and similar expressions used herein mean that the element or object preceding the word covers the element or object listed after the word and its equivalents, but does not exclude other elements or objects.
[0090] See Figures 1 to 3 As shown, this embodiment of the invention provides a verifiable APT attack chain extraction method based on cognitive distillation and gating technology.
[0091] This method achieves highly robust TTP chain extraction through a three-order collaborative mechanism: First, it parses threat intelligence text based on the RoBERTa-wwm-ext model, enhances technical entity understanding through a full-word masking strategy, and combines positional encoding to model attack temporal logic, dynamically generating gating weights to quantify inter-stage dependency strength. Second, it utilizes LLaMA-3-70B's cue engineering to achieve few-sample prototype matching, dynamically expanding the new attack knowledge base. Finally, it constructs a dual defense mechanism through Phi-3-mini, injecting noise to simulate tampering attacks and constraining prototype distribution stability to achieve semantic interference immunity. The entire process forms a cognitive closed loop of "semantic parsing → knowledge evolution → defense hardening," achieving high accuracy and improved speed of new attack identification in industrial scenarios.
[0092] A verifiable APT attack chain extraction method based on cognitive distillation and gating techniques, specifically including:
[0093] S1. Input processing and extraction, including semantic atomization aggregation of textual threat intelligence, fusion of fragmented technical descriptions into tactical units, locking the causal chain between stages, initiating multi-dimensional feature distillation, extracting four types of tactical paths: attack medium, execution fingerprint, payload pattern and infiltration path, constructing an anti-obfuscation topology network, disambiguating fuzzy semantics based on context distance weight, and generating a tactical vector map with core attack nodes as the hub.
[0094] S2. Dynamic gating weight generation includes positional encoding of the input attack sequence, capturing the temporal relationship and dependency logic between different attack actions, using multi-head attention mechanism to focus on the technical implementation of identifying the carrier and malicious behavior characteristics, identifying key elements in the attack, and using long short-term memory network to analyze historical attack patterns, calculate the correlation strength between the current action and historical behavior, and generate gating weight values by fusing historical dependencies and semantic similarity.
[0095] S3. Hierarchical Prototype Decision Making: Based on the generated gating weights and tactical vector graph, hierarchical decision making is performed. This includes combining the input vector with information from the knowledge base, activating the few-shot reasoning ability of the large language model, determining the current attack mode, and if it is a novel attack mode, then the essential features of the attack are distilled out through a self-attention mechanism to generate a new tactical vector. This vector is then compared with known attack modes for similarity, and mapped to the existing tactical coding system or marked as an unknown attack requiring further research based on the similarity result. If it is a conventional attack, a standard matching process is used. The standard matching process refers to the attack mode matching process based on a predefined knowledge base and a similarity threshold. Specifically, it includes: retrieving the known attack mode most similar to the current tactical vector from the ATT&CK framework knowledge base, calculating the cosine similarity, and directly mapping to the corresponding TTP code when the similarity exceeds the preset threshold of 0.85. When the similarity is insufficient, it is marked as requiring manual review.
[0096] S4. To ensure robustness against attacks and output attack chains, the TTP encoding output from the hierarchical prototype decision in step 3 (i.e., the attack tactical feature encoding based on the MITRE ATT&CK framework) is injected with noise of a specific intensity. This TTP encoding is a standard attack feature identifier generated by performing hierarchical decision-making on the tactical vector output in step 2 using the LLaMA-3 model. Injecting noise of a specific intensity simulates possible semantic tampering by attackers. The model output results of the original encoding and the noisy encoding are calculated separately. The system stability is evaluated by comparing the differences in the model output results. This evaluation includes quantifying the interference level by combining the distance metric of the output vector and the relative entropy of the probability distribution. When interference exceeding a threshold is detected, a semantic reconstruction mechanism is activated to correct and restore the encoding based on attention focus and historical tactical logic. For cases where the threshold is not exceeded, a gating weight value is generated to update the signal, forming a defensive closed loop. The TTP encoding is a standard attack feature identifier generated by performing hierarchical decision-making on the tactical vector output in step S1 using the LLaMA-3 model.
[0097] Specifically, the process of performing semantic atomization aggregation on textual threat intelligence and integrating fragmented technical descriptions into tactical units in S1 includes performing semantic standardization on the original descriptions through a text preprocessing engine, aggregating scattered technical elements, binding malicious Excel attachments containing macro code as delivery carriers for tactical units, and retaining the initial access attributes of email distribution.
[0098] Furthermore, the causal chain between stages in S1 includes the causal chain of "enabling macros → triggering PowerShell → downloading ransomware".
[0099] Furthermore, in S2, the input attack sequence is positionally encoded and then understood using a full-word masking strategy, wherein the masking loss function is:
[0100] ;
[0101] in, This represents the loss value of the masked language model, indicating the degree of error in the model's prediction of the masked words; This is the mask position index, representing the position number of the masked token in the sequence; It is a mask set, containing the set of all masked token positions in the current sequence; The true value of the target token, representing the value at the 1st... The original token before the masked position; This is non-mask context information, representing information other than the mask set. The set of all visible tokens outside; This is the conditional probability value, representing the probability that the model predicts the target token based on the context;
[0102] The core value of the above process lies in its mask set. Prioritizing the coverage of complete technical entities rather than fragmented characters, this mechanism manifests in the instance as a complete mask of "PowerShell" rather than character-level segmentation. When the model is faced with the input "Enable macro-triggered [MASK] download", it must reconstruct the complete semantics based on the context. This training forces the model to deeply grasp the tactical essence of "PowerShell" as an attack execution vehicle—its pipeline operation characteristics and memory-resident capabilities constitute the typical payload delivery pattern of ransomware, rather than a superficial understanding of isolated words.
[0103] The position code is:
[0104] ;
[0105] in The location-coded value represents the location. The position encoding at dimension index 2i; The position index represents the sequential position (integer, starting from 0) of the token or attack phase in the sequence. For dimension index, it represents the dimension number (integer, from 0 to d / 2-1) in the location encoding vector. 10000 is the model dimension, representing the total dimension size of the location coding vector (usually 768); 10000 is the frequency cardinality, a constant value that controls the frequency distribution of location coding.
[0106] The mathematical expression described above precisely models the causal logic of the attack process: macro activation must precede script triggering, reflecting the inherent sequence constraints of the attacker's actions.
[0107] In the multi-head attention mechanism described in S2, the third attention head focuses on the carrier dimension, capturing the physical characteristics of the Excel attachment as the initial access carrier; the seventh attention head locks onto the malicious behavior dimension, reinforcing the aggressive nature of the download action; and the eleventh attention head analyzes the tactical intent through a gating weight formula, which is:
[0108] ;
[0109] in The gating weight value represents the quantification strength of the dependency relationship between attack phases, and its value ranges from (0,1).
[0110] 0.6 is the sigmoid activation function, used to compress the linear combination result into the probability interval; 0.6 is the history dependency weight coefficient, reflecting the strong temporal prior characteristic of the attack chain; 0.4 represents the historical dependency probability, which is the output of the temporal modeling of historical patterns of attack sequences based on the LSTM network; 0.4 is the semantic similarity weight coefficient, which balances the consistency requirements of the current semantic context. For semantic similarity, the cosine similarity is used to calculate the semantic coherence between the current action and the context.
[0111] If text is altered to vaguely describe system operations, a semantic weighting mechanism is activated. This semantic weighting mechanism is as follows:
[0112] ;
[0113] in This is the final gating weight value. The output of the gate weight value after semantic security verification; The model predicts the probability of a keyword term based on the context, and then predicts the conditional probability of the correct term. The semantic credibility threshold. The value is 0.5, which is used to determine the credibility of the semantic description.
[0114] like Figure 2 As shown, the algorithm flow of S3 is as follows:
[0115] S31. Obtain the input tactical vector, and concatenate the tactical vector with a preset prototype feature set according to a preset format to construct a prompt word template sequence;
[0116] S32. Input the prompt word template sequence into the LLaMA-3-70B model for inference calculation and obtain the attack behavior analysis results output by the model;
[0117] S33. Analyze the attack behavior analysis results and determine whether the currently input tactical vector belongs to a new type of attack through preset discrimination logic;
[0118] S34. If it is determined to be a new type of attack, the attack attributes in the attack behavior analysis results are encoded and calculated to generate a new prototype vector, the new prototype vector is updated to the prototype feature set, and the corresponding tactical and technical process code is generated based on the new prototype vector.
[0119] S35. If the attack type is determined to be known, the feature distance between the tactical vector and each existing vector in the prototype feature set is calculated, the target prototype vector with the smallest feature distance is selected, and the corresponding tactical technical process code is generated based on the target prototype vector.
[0120] S36. Output the finalized tactical and technical process code;
[0121] In S32, the hierarchical prototype decision-making stage achieves precise mapping of tactical intent through the LLaMA-3-70B model, when processing the gating weight vector of the dynamic gating output. At this time, the model first performs basic tactical classification:
[0122]
[0123] This prompt template activates the model's task-aware state via the [INST] instruction header, and sets the gating weights. As a historically dependent prior injection decision system, a priority head with a value of 0.89 is selected during the decoding process as the attention head with the highest focus weight on PowerShell, capturing its memory-resident characteristics—distinct from the disk write mode of traditional command-line batch processing. This design enables the model to output a probability distribution p=[0.92,0.08] within 3ms.
[0124] S33 includes the sub-prototype fine matching stage's enhanced technical identification capability:
[0125] ;
[0126] Model construction feature comparison matrix The differences in row vector quantization techniques are as follows:
[0127] ;
[0128] Among them, the first column, which features PowerShell-specific Get-Content | Invoke-Expression pipeline chain operations, has a weight of 0.85, and the second column, which exposes batch installation defects in command lines, has a weight of 0.12.
[0129] Dynamic expansion mechanisms trigger semantic distillation:
[0130] ;
[0131] In the 12-layer self-attention mechanism, the model performs hierarchical feature decoupling, reduces the weight of "API" legitimate attributes (stripping away technical appearances), increases the weight of "leakage" attack kernels (focusing on tactical essence), and generates vectors. The mathematical expression is:
[0132] ;
[0133] in The newly generated feature vector represents the attack kernel features extracted after hierarchical feature decoupling; It is an encoder function that maps text descriptions to a high-dimensional feature space. It is an average pooling operation that reduces the dimensionality and aggregates the encoded features;
[0134] Similarity verification formula:
[0135] ;
[0136] in It is the cosine similarity value, ranging from [-1, 1], which represents the degree of directional similarity between two vectors; It is the feature vector of the T1537 attack, representing the T1537 attack pattern features extracted from the data; It is the feature vector of Conti ransomware, representing the feature representation of the Conti attack family; It is the formula for calculating cosine similarity, which measures the similarity of two vectors in a direction.
[0137] The numerator dot product captures the common "cloud service cover" strategy of OneDrive API abuse in the Conti ransomware case, while denominator normalization eliminates technical differences such as protocol type. High similarity reflects the model's ability to abstract general attack patterns from specific implementations, enabling threats without MITRE mappings to obtain high-confidence decision support.
[0138] like Figure 3 As shown, the algorithm flow of S4 includes:
[0139] S41: Obtain the output attack tactics code, and superimpose a preset Gaussian noise perturbation on the tactical process code to generate the corresponding adversarial sample code;
[0140] S42: Input the tactical technical process code and the adversarial sample code into the preset verification model for inference, and obtain the original output features and adversarial output features respectively;
[0141] S43: Based on the original output features and the adversarial output features, calculate the robustness loss value that characterizes the stability of the model output. The robustness loss value includes at least the Euclidean distance and KL divergence between the original output features and the adversarial output features.
[0142] S44: Determine whether the robustness loss value is greater than a preset threshold;
[0143] S45: If the robustness loss value is determined to be greater than the threshold, then construct an error correction prompt word containing attack type information, input it into a preset correction model to generate and output the corrected tactical and technical process code;
[0144] S46: If the robustness loss value is determined to be less than or equal to the threshold, the preset gating weight parameters are updated based on the KL divergence calculation parameter increment, and the updated weight parameters are output.
[0145] In Phase Two, when the key text is maliciously altered to trigger the command-line tool to download the encryption module after macros are enabled, the system activates a triple collaborative protection mechanism upon detecting semantic tampering: First, noise injection training simulates attacker behavior; second, KL divergence distribution system constrains and monitors system stability; third, the semantic reconstruction correction engine restores the correct semantics. These three mechanisms form a complete defense-cognition closed loop, ensuring the robustness and accuracy of attack chain analysis. The noise injection training method is as follows:
[0146] ;
[0147] in For adversarial examples, For the original input, The noise intensity coefficient, It is Gaussian noise;
[0148] The core defense response mechanism is implemented through a KL divergence distribution system, which is constructed as follows:
[0149] ;
[0150] Full power start;
[0151] in For robust loss, The Kullback-Leibler divergence;
[0152] The semantic reconstruction and correction engine construction instructions are as follows:
[0153] ;
[0154] in To correct the output, the instruction triggers a four-layer cognitive analysis. The first layer, an attention mechanism, focuses on the essence of the download action with a certain weight, removing the interference of the tool attributes of the command line. The second layer analyzes the semantic relationship of the context. The third layer associates the phishing attack characteristics of the historical stage T1566.001 to confirm that the subsequent action should be an execution-type action. Finally, the output is accurately mapped to T1059.005 and maintains a high confidence level. The fourth layer performs confidence-weighted decision fusion.
[0155] In S4, the defense feedback closed loop is updated using an adaptive formula, which is:
[0156] ;
[0157] in ;in This is the updated defense strength parameter, used to dynamically adjust the strength level of the system's adversarial defense. The value range is [0,1], with a larger value indicating a stronger defense.
[0158] In step S4, noise is added by adding a noise intensity parameter. Implementation, including:
[0159] ;
[0160] in It is the noise intensity parameter of the t-th iteration, the noise injection coefficient at the current time, which controls the noise amplitude when the adversarial example is generated; It is the noise intensity parameter of the (t-1)th iteration, and the noise injection coefficient of the previous moment, which serves as the starting point for the current optimization. The learning rate is a hyperparameter controlling the step size of parameter updates during gradient descent. It determines the convergence speed and stability of noise intensity optimization; a value that is too large will cause oscillations, while a value that is too small will lead to slow convergence. It is the partial derivative of robustness loss with respect to noise intensity, representing the gradient of the impact of noise intensity changes on the defense effect. The entire defense mechanism and attack chain extraction system form a cognitive closed loop. Noise injection simulates attacker behavior patterns, KL divergence anchors knowledge representation stability, semantic reconstruction inherits historical tactical logic, and finally outputs a complete attack chain with defense audit logs.
[0161] While embodiments of the present invention have been described in detail above, it will be apparent to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations fall within the scope and spirit of the invention as set forth in the claims. Furthermore, the invention described herein may have other embodiments and can be implemented or carried out in various ways.
Claims
1. A verifiable APT attack chain extraction method based on cognitive distillation and gating technology, characterized in that, Includes the following steps: S1. Input processing and extraction, including semantic atomization aggregation of textual threat intelligence, fusion of fragmented technical descriptions into tactical units, locking the causal chain between stages, initiating multi-dimensional feature distillation, extracting four types of tactical paths: attack medium, execution fingerprint, payload pattern and infiltration path, constructing an anti-obfuscation topology network, disambiguating fuzzy semantics based on context distance weight, and generating a tactical vector map with core attack nodes as the hub. S2. Dynamic gating weight generation includes positional encoding of the input attack sequence, capturing the temporal relationship and dependency logic between different attack actions, using multi-head attention mechanism to focus on identifying carrier and malicious behavior characteristics, identifying key elements in the attack, and using long short-term memory network to analyze historical attack patterns, calculate the correlation strength between the current action and historical behavior, and generate gating weight values by fusing historical dependencies and semantic similarity. S3. Hierarchical Prototype Decision Making: Based on the generated gating weights and tactical vector graph, hierarchical decision making is performed. This includes combining the tactical vector graph with information from the knowledge base, activating the few-shot reasoning ability of the large language model, determining the current attack mode, and if it is a novel attack mode, then the essential features of the attack are distilled through a self-attention mechanism to generate a new tactical vector. This vector is then compared with known attack modes for similarity. Based on the similarity results, the new tactical vector is mapped to the existing tactical coding system or marked as an unknown attack that needs further research. If it is a conventional attack, a standard matching process is used. The standard matching process refers to the attack mode matching process based on a predefined knowledge base and a similarity threshold. Specifically, it includes: retrieving the known attack mode most similar to the current tactical vector from the ATT&CK framework knowledge base, calculating the cosine similarity, and directly mapping it to the corresponding TTP code when the similarity exceeds the preset threshold of 0.
85. When the similarity is insufficient, it is marked as requiring manual review. S4. To ensure robustness against attacks and output attack chains, the TTP encoding of the hierarchical prototype decision output from S3 (i.e., the attack tactical feature encoding based on the MITREATT&CK framework) is injected with noise of a specific intensity to simulate possible semantic tampering by attackers. The model output results of the original encoding and the noisy encoding are calculated separately. By comparing the differences in the model output results, the system stability is evaluated. This evaluation includes quantifying the degree of interference by combining the distance metric of the output vector and the relative entropy of the probability distribution. When interference exceeding a threshold is detected, a semantic reconstruction mechanism is initiated to correct and restore the encoding based on attention focus and historical tactical logic. For cases where the threshold is not exceeded, a gating weight value is generated to update the new signal, forming a defensive closed loop. The TTP encoding is a standard attack feature identifier generated by performing hierarchical decision-making on the tactical vector output from S1 using the LLaMA-3 model.
2. The verifiable APT attack chain extraction method based on cognitive distillation and gating technology according to claim 1, characterized in that, The process of performing semantic atomization aggregation on textual threat intelligence and integrating fragmented technical descriptions into tactical units in S1 includes performing semantic standardization on the original descriptions through a text preprocessing engine, aggregating scattered technical elements, binding malicious Excel attachments containing macro code as delivery carriers for tactical units, and retaining the initial access attributes of email distribution.
3. The verifiable APT attack chain extraction method based on cognitive distillation and gating technology according to claim 2, characterized in that, The inter-stage causal chain mentioned in S1 includes the causal chain of "enabling macro → triggering PowerShell → downloading ransomware".
4. The verifiable APT attack chain extraction method based on cognitive distillation and gating technology according to claim 3, characterized in that, In step S2, the input attack sequence is positionally encoded, and understanding is first achieved through a full-word masking strategy. The masking loss function is: ; in, This represents the loss value of the masked language model, indicating the degree of error in the model's prediction of the masked words; This is the mask position index, representing the position number of the masked token in the sequence; It is a mask set, containing the set of all masked token positions in the current sequence; The true value of the target token, representing the value at the 1st... The original token before the masked position; This is non-mask context information, representing information other than the mask set. The set of all visible tokens outside; This is the conditional probability value, representing the probability that the model predicts the target token based on the context; The position code is: ; in The location-coded value represents the location. The position encoding at dimension index 2i; The position index represents the sequential position of the token or attack phase in the sequence, where the sequential position is an integer, starting from 0; is the dimension index, representing the dimension number in the location encoding vector, where the dimension number is an integer from 0 to d / 2-1; is the model dimension, representing the total dimension size of the location encoding vector, where the total dimension size is 768; 10000 is the frequency cardinality, a constant value that controls the frequency distribution of location encoding.
5. The verifiable APT attack chain extraction method based on cognitive distillation and gating technology according to claim 1, characterized in that, In the multi-head attention mechanism described in S2, the third attention head focuses on the carrier dimension, capturing the physical characteristics of the Excel attachment as the initial access carrier; the seventh attention head locks onto the malicious behavior dimension, reinforcing the aggressive nature of the download action; and the eleventh attention head analyzes the tactical intent through a gating weight formula, which is: ; in The gating weight value represents the quantification strength of the dependency relationship between attack phases, and its value ranges from (0,1). 0.6 is the sigmoid activation function, used to compress the linear combination result into the probability interval; 0.6 is the history dependency weight coefficient, reflecting the strong temporal prior characteristic of the attack chain; 0.4 represents the historical dependency probability, which is the output of the temporal modeling of historical patterns of attack sequences based on the LSTM network; 0.4 is the semantic similarity weight coefficient, which balances the consistency requirements of the current semantic context. For semantic similarity, the semantic coherence between the current action and the context is calculated based on cosine similarity. If text is altered to vaguely describe system operations, a semantic weighting mechanism is activated. This semantic weighting mechanism is as follows: ; in, This is the final gating weight value. The output of the gate weight value after semantic security verification; The model predicts the probability of a keyword term based on the context, and then predicts the conditional probability of the correct term. The semantic credibility threshold. The value is 0.5, which is used to determine the credibility of the semantic description.
6. The verifiable APT attack chain extraction method based on cognitive distillation and gating technology according to claim 1, characterized in that, The algorithm flow for S3 is as follows: S31. Obtain the input tactical vector, and concatenate the tactical vector with a preset prototype feature set according to a preset format to construct a prompt word template sequence; S32. Input the prompt word template sequence into the LLaMA-3-70B model for inference calculation and obtain the attack behavior analysis results output by the model; S33. Analyze the attack behavior analysis results and determine whether the currently input tactical vector belongs to a new type of attack through preset discrimination logic; S34. If it is determined to be a new type of attack, the attack attributes in the attack behavior analysis results are encoded and calculated to generate a new prototype vector, the new prototype vector is updated to the prototype feature set, and the corresponding tactical and technical process code is generated based on the new prototype vector. S35. If the attack type is determined to be known, the feature distance between the tactical vector and each existing vector in the prototype feature set is calculated, the target prototype vector with the smallest feature distance is selected, and the corresponding tactical technical process code is generated based on the target prototype vector. S36. Output the finalized tactical and technical process code; In S32, the hierarchical prototype decision-making stage achieves precise mapping of tactical intent through the LLaMA-3-70B model, and the gating weights of the dynamic gating output are processed. At this time, the model first performs basic tactical classification: ; The prompt template activates the model's task-aware state via the [INST] instruction header, and sets the gating weights. =0.89 is used as the historical dependency prior injection decision system. During the decoding process, an attention head is selected as the attention head with the highest focus weight on PowerShell. S33 includes the sub-prototype fine matching stage's enhanced technical identification capability: ; Model construction feature comparison matrix : ; Among them, the first column, which features PowerShell-specific Get-Content | Invoke-Expression pipeline chain operations, has a weight of 0.85, and the second column, which exposes batch installation defects in command lines, has a weight of 0.
12. Dynamic expansion mechanisms trigger semantic distillation: ; In the 12-layer self-attention mechanism, the model performs hierarchical feature decoupling; generating vectors. The mathematical expression is: ; in The newly generated feature vector represents the attack kernel features extracted after hierarchical feature decoupling; It is an encoder function that maps text descriptions to a high-dimensional feature space; It is an average pooling operation that reduces the dimensionality and aggregates the encoded features; Similarity verification formula: ; in It is the cosine similarity value, ranging from [-1, 1], which represents the degree of directional similarity between two vectors; It is the feature vector of the T1537 attack, representing the T1537 attack pattern features extracted from the data; It is the feature vector of Conti ransomware, representing the feature representation of the Conti attack family; It is the formula for calculating cosine similarity, which measures the similarity of two vectors in a direction.
7. The verifiable APT attack chain extraction method based on cognitive distillation and gating technology according to claim 6, characterized in that, The algorithm flow of S4 includes: S41: Obtain the tactical and technical process code, and superimpose a preset Gaussian noise perturbation on the tactical and technical process code to generate the corresponding adversarial sample code; S42: Input the tactical technical process code and the adversarial sample code into the preset verification model for inference, and obtain the original output features and adversarial output features respectively; S43: Based on the original output features and the adversarial output features, calculate the robustness loss value that characterizes the stability of the model output. The robustness loss value includes at least the Euclidean distance and KL divergence between the original output features and the adversarial output features. S44: Determine whether the robustness loss value is greater than a preset threshold; S45: If the robustness loss value is determined to be greater than the threshold, then construct an error correction prompt word containing attack type information, input it into a preset correction model to generate and output the corrected tactical and technical process code; S46: If the robustness loss value is determined to be less than or equal to the threshold, the preset gating weight parameters are updated based on the KL divergence calculation parameter increment, and the updated weight parameters are output. When the key text is maliciously altered to trigger the command-line tool to download the encryption module after macros are enabled, the system activates a triple collaborative protection mechanism upon detecting semantic tampering: First, noise injection training simulates attacker behavior; second, KL divergence distribution system constrains and monitors system stability; third, the semantic reconstruction correction engine restores the correct semantics. These three mechanisms form a complete defense-cognition closed loop, ensuring the robustness and accuracy of attack chain analysis. The noise injection training method is as follows: ; in For adversarial examples, For the original input, For noise intensity parameters, It is Gaussian noise; The core defense response mechanism is implemented through a KL divergence distribution system, which is constructed as follows: ; ; in For robust loss, The Kullback-Leibler divergence; The semantic reconstruction and correction engine construction instructions are as follows: ; in To correct the output, the instruction triggers a four-layer cognitive analysis. The first layer, an attention mechanism, focuses on the essence of the download action with a certain weight, removing the interference of the tool attributes of the command line. The second layer analyzes the semantic relationship of the context. The third layer associates the phishing attack characteristics of the historical stage T1566.001 to confirm that the subsequent action should be an execution-type action. The fourth layer performs confidence-weighted decision fusion, and the final output is accurately mapped to T1059.005 while maintaining a high confidence level.
8. The verifiable APT attack chain extraction method based on cognitive distillation and gating technology according to claim 1, characterized in that, The defense feedback closed loop in S4 is updated using an adaptive formula, which is: ; in ;in This is the gating weight value. This is the updated defense strength parameter, used to dynamically adjust the strength level of the system's adversarial defense. The value range is [0,1], with a larger value indicating a stronger defense.
9. The verifiable APT attack chain extraction method based on cognitive distillation and gating technology according to claim 1, characterized in that, In step S4, noise of a specific intensity is injected by adding a noise intensity parameter. Implementation, in which: ; in It is the noise intensity parameter in the t-th iteration; It is the noise intensity parameter for the (t-1)th iteration; It is the gradient descent learning rate, a hyperparameter that controls the step size of parameter updates, and determines the convergence speed and stability of noise intensity optimization; It is the partial derivative of robustness loss with respect to noise intensity, representing the gradient of the effect of changes in noise intensity on the defense effect.
Citation Information
Patent Citations
APT killing chain reconstruction and prediction method and system based on causal reasoning
CN119598455A
APT attack traceability and path restoration method and system
CN121151119A