APT attack source determination method and device, and electronic equipment
By acquiring multi-source data to determine the target confidence of APT attacks, a set of high-confidence attack events is selected. By adopting a multi-detector parallel architecture and joint training mechanism, the problem of inaccurate identification of the source of APT attacks is solved, and higher accuracy and interpretability of source tracing are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- STATE GRID BEIJING ELECTRIC POWER CO
- Filing Date
- 2026-05-25
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies often fail to accurately identify the source of APT attacks. Attack behavior-based tracing methods have poor adaptability, while machine learning and deep learning-based methods are costly and lack interpretability.
By acquiring multi-source data corresponding to multiple attack events in the target network, the target confidence of multiple attack events is determined, and the attack source of APT attacks is determined based on the high-confidence attack event set. A multi-detector parallel architecture and a joint training mechanism of weak supervision + remote supervision are adopted to achieve multi-modal feature differentiation modeling and collaborative decision-making.
It improves the accuracy and interpretability of identifying the source of APT attacks, reduces false alarms and false negatives, and enhances the credibility of emergency response and attribution accountability.
Smart Images

Figure CN122457355A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network security technology, and more specifically, to a method, apparatus, and electronic device for determining the source of an APT attack. Background Technology
[0002] Attribution analysis of Advanced Persistent Threat (APT) attacks is a core and crucial component of cybersecurity defense systems, and a vital support for ensuring the security of critical information infrastructure and core data. Its core objective is to accurately pinpoint the source of an APT attack by analyzing attack chains, uncovering attack characteristics, and correlating threat intelligence, thus providing essential evidence for emergency response, attack containment, and attribution and accountability.
[0003] Related technologies employ attribution methods based on attack behavior characteristics, machine learning, and deep learning to determine the source of APT attacks. Attribution methods based on attack behavior characteristics have significant limitations and poor adaptability to new attack and evasion techniques; machine learning and deep learning-based attribution methods have high deployment costs and insufficient model interpretability. Therefore, these technologies suffer from the technical problem of inaccurate attribution determination of APT attack sources.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This application provides a method, apparatus, and electronic device for determining the source of an APT attack, in order to at least solve the technical problem of inaccurate determination of the source of an APT attack in related technologies.
[0006] According to one aspect of the embodiments of this application, a method for determining the source of an APT attack is provided, comprising: acquiring multi-source data corresponding to multiple attack events of a target network in the current period, wherein an attack event refers to a unit that supports independent detection in APT attack detection; determining the target confidence level corresponding to each of the multiple attack events based on the multi-source data corresponding to the multiple attack events, wherein the target confidence level is used to indicate the credibility of the corresponding attack event as an APT attack; filtering out a set of high-confidence attack events from the multiple attack events based on the target confidence levels corresponding to the multiple attack events; and determining the source of the APT attack based on the target confidence levels corresponding to the multiple high-confidence attack events included in the set of high-confidence attack events.
[0007] According to another aspect of the embodiments of this application, an APT attack source determination device is provided, comprising: a data acquisition module, configured to acquire multi-source data corresponding to multiple attack events of a target network in the current period, wherein an attack event refers to a unit that supports independent detection in APT attack detection; a first determination module, configured to determine the target confidence level corresponding to each of the multiple attack events based on the multi-source data corresponding to each of the multiple attack events, wherein the target confidence level is used to indicate the credibility of the corresponding attack event as an APT attack; a filtering module, configured to filter out a set of high-confidence attack events from the multiple attack events based on the target confidence level corresponding to each of the multiple attack events; and a second determination module, configured to determine the attack source of the APT attack based on the target confidence level corresponding to each of the multiple high-confidence attack events included in the high-confidence attack event set.
[0008] According to another aspect of the embodiments of this application, a non-volatile storage medium is provided, which stores multiple instructions, any one of which is adapted to be loaded and executed by a processor for determining the source of an APT attack.
[0009] According to another aspect of the embodiments of this application, an electronic device is provided, including: one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the methods for determining the source of an APT attack.
[0010] According to another aspect of the embodiments of this application, a computer program product is provided, which, when executed on a data processing device, is suitable for performing the steps of a method for determining the source of an APT attack.
[0011] In this embodiment, multi-source data corresponding to multiple attack events of the target network within the current period is obtained. An attack event refers to a unit that supports independent detection in APT attack detection. Based on the multi-source data corresponding to each attack event, the target confidence level corresponding to each attack event is determined. The target confidence level indicates the credibility of the corresponding attack event as an APT attack. Based on the target confidence levels corresponding to each attack event, a high-confidence attack event set is selected from the multiple attack events. Based on the target confidence levels corresponding to the multiple high-confidence attack events included in the high-confidence attack event set, the attack source of the APT attack is determined. This achieves the goal of obtaining multi-source data corresponding to multiple attack events, selecting a high-confidence attack event set, and determining the attack source of the APT attack based on the target confidence levels corresponding to the multiple high-confidence attack events included in the high-confidence attack event set. This improves the accuracy of the APT attack source determination result and solves the technical problem of inaccurate APT attack source determination results in related technologies. Attached Figure Description
[0012] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0013] Figure 1 This is a flowchart of a method for determining the source of an APT attack according to an embodiment of this application;
[0014] Figure 2 This is a flowchart of an optional method for determining the source of an APT attack, according to an embodiment of this application.
[0015] Figure 3 This is a structural diagram of an optional APT organization tracing architecture provided according to an embodiment of this application;
[0016] Figure 4 This is a schematic diagram of an APT attack source determination device according to an embodiment of this application;
[0017] Figure 5 This is a structural diagram of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0018] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0019] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0020] It should be noted that the information collected in this application (including but not limited to historical confidence levels, historical network layer attack probabilities, historical host attack probabilities, historical sample attack probabilities, and addresses of high-confidence attack events) and data (including but not limited to multi-source data and source tracing report text data) are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of this data all comply with relevant laws, regulations, and standards, and necessary confidentiality measures have been taken. This process does not violate public order and good morals, and corresponding operation entry points are provided for users to choose to authorize or refuse. For example, interfaces are set up between this system and relevant users or organizations, providing users with corresponding operation entry points for them to choose to agree to or refuse automated decision results; if the user chooses to refuse, the process proceeds to the expert decision-making process.
[0021] For ease of description, the following explains some of the nouns or terms used in the embodiments of this application:
[0022] APT attacks refer to long-term, covert, multi-stage coordinated attacks launched by highly organized, well-resourced, and long-term lurking attackers against specific targets (such as critical information infrastructure, energy, and other core systems).
[0023] The network layer refers to the logical functional layer in computer network architecture that is responsible for routing, forwarding, and transmitting data packets across networks. Its core task is to enable end-to-end communication between different hosts or network nodes. In the context of network security analysis, the network layer specifically refers to the raw traffic data and related metadata layer that can be collected, monitored, and analyzed during network communication.
[0024] Indicators of Compromise (IOCs) are digital traces or characteristics left unintentionally or intentionally by attackers in target systems, networks, or data streams during cybersecurity incidents. These traces are objectively detectable and quantifiable and used to characterize the existence, stage, or behavioral pattern of a specific attack activity. Essentially, they are the "technical fingerprints" of attack behavior in cyberspace, which can be collected, stored, compared, and correlated using automated tools to support attack detection, threat assessment, and attribution response.
[0025] According to an embodiment of this application, a method embodiment for determining the source of an APT attack is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0026] Figure 1 This is a flowchart of a method for determining the source of an APT attack according to an embodiment of this application, as shown below. Figure 1 As shown, the method includes the following steps:
[0027] Step S102: Obtain multi-source data corresponding to multiple attack events of the target network in the current period, wherein an attack event refers to a unit that supports independent detection in APT attack detection.
[0028] It is understandable that acquiring multi-source data corresponding to multiple attack events on the target network within the current period is necessary. This includes network layer data, host data, and APT sample data for each attack event. An attack event refers to the smallest independently detectable unit of behavior in APT attack detection, such as a C2 (Command and Control) heartbeat communication, the creation of a malicious process, or the loading and execution of a suspicious binary file. By defining attack events as the smallest independently detectable unit of behavior and simultaneously collecting their multi-source data, a refined and structured characterization of the entire APT attack chain can be achieved. This improves the completeness of APT attack identification and the accuracy of attribution, avoiding missed detections, false alarms, and broken attribution chains caused by fragmented and coarse-grained data.
[0029] Optionally, the data acquisition phase focuses on constructing a cross-domain APT multi-source dataset. This addresses the issues of data disconnect from real-world scenarios and insufficient coverage of new attacks, while also providing support for improved accuracy and scalability. First, it aggregates core data from multiple sources, integrating source tracing reports and IOC data from publicly available threat intelligence databases, attack logs, and new, unknown APT attack data captured by honeypot systems. This achieves complementarity between publicly available intelligence and real-world scenario data, broadening the sample boundaries for model learning. Second, it employs hierarchical labeling and sample expansion, performing refined hierarchical labeling of data by industry, attack stage, and APT organization type. Simultaneously, it introduces data augmentation strategies, such as randomly supplementing attack behavior chains with nodes to simulate code obfuscation and behavioral camouflage by APT attackers. This effectively improves sample diversity, alleviates the scarcity of new attack samples, and allows the model to learn more comprehensive features within the source tracing report data space (including multi-source data corresponding to multiple attack events and source tracing report text data), laying a solid foundation for improved accuracy.
[0030] Step S104: Based on the multi-source data corresponding to multiple attack events, determine the target confidence level corresponding to each of the multiple attack events, wherein the target confidence level is used to indicate the credibility of the corresponding attack event as an APT attack;
[0031] It is understandable that by fusing multi-source data from attack events, it is possible to quantitatively characterize the credibility of APT attacks, improve the accuracy and anti-interference capability of malicious behavior identification, reduce false positives and false negatives caused by single-modal noise, attack camouflage, or tool reuse, and lay a reliable confidence foundation for the subsequent identification of highly credible attack sources.
[0032] In an optional embodiment, when the multi-source data includes network layer data, host data, and APT sample data corresponding to the attack events, the target confidence level corresponding to each of the multiple attack events is determined based on the multi-source data corresponding to each attack event. This includes: obtaining the source tracing report text data in the current period; preprocessing the multi-source data and source tracing report text data corresponding to each of the multiple attack events to obtain standard data and source tracing report text standard data corresponding to each of the multiple attack events, wherein the standard data includes network layer standard data, host standard data, and APT sample standard data corresponding to the attack events; and extracting features from the network layer standard data corresponding to each of the multiple attack events to obtain the network layer standard data corresponding to each of the multiple attack events. Layer features: Feature extraction is performed on the host standard data corresponding to multiple attack events to obtain host features for each attack event; feature extraction is performed on the APT sample standard data corresponding to multiple attack events to obtain APT sample features for each attack event; feature extraction is performed on the source tracing report text standard data to obtain source tracing report text features; based on the source tracing report text features, network layer weights, host weights, and sample weights are determined; based on the network layer features, host features, APT sample features, network layer weights, host weights, and sample weights corresponding to multiple attack events, the target confidence level corresponding to each attack event is determined.
[0033] It is understandable that, given multi-source data including network layer data, host data, and APT sample data corresponding to the attack event, the target confidence of the attack event is determined as follows: First, the source tracing report text data for the current period is obtained. This source tracing report text data describes the tactics, techniques, and procedures (TTPs) used by known APT groups in their attack activities, including the initial intrusion vector, lateral movement methods, malicious tools and C2 infrastructure used, vulnerabilities exploited, data infiltration paths, attack trace characteristics, and organizational attribution basis. Second, the multi-source data and source tracing report text data corresponding to multiple attack events are preprocessed to obtain standard data and source tracing report text standard data corresponding to each attack event. The standard data includes network layer standard data, host standard data, and APT sample standard data corresponding to the attack event. Next, feature extraction is performed on the standard network layer data, host data, APT sample data, and source tracing report text data corresponding to multiple attack events, respectively, to obtain the network layer features, host features, APT sample features, and source tracing report text features corresponding to each attack event. Then, based on the source tracing report text features, the network layer weight, host weight, and sample weight are determined. Finally, based on the network layer features, host features, APT sample features, network layer weight, host weight, and sample weight corresponding to each attack event, the target confidence level for each attack event is determined. This approach enables a shift in the source tracing decision-making paradigm for APT attacks from being driven by human experience to being driven by data semantics, improving the accuracy and anti-interference capabilities in scenarios involving attack tool reuse, behavioral camouflage, and cross-organizational obfuscation, and enhancing the accuracy of attack source identification.
[0034] Optionally, network layer data is used to describe characteristics related to network communication behavior in the attack event, including C2 heartbeat communication, DNS (Domain Name System) tunneling, abnormal outbound traffic, and protocol abuse behavior; host data is used to describe abnormal behavior at the terminal system level in the attack event, including malicious process creation, privilege escalation operations, registry modification, file encryption or hiding, and persistent persistence actions; APT sample data is used to describe the static structural characteristics (such as PE (Portable Executable) header, import table, strings, and entropy values) and dynamic execution behavior characteristics (such as API (Application Programming Interface) call sequences, network connections, process injection, and memory self-modification behavior in the sandbox) of the malicious binary files associated with the attack event, in order to characterize the uniqueness of the attack tools and the organizational affiliation.
[0035] Optionally, preprocessing includes three steps. First, basic data cleaning: deleting duplicate records, filling in missing fields, and removing invalid data to reduce noise interference with model training. Second, adaptive attack noise filtering: constructing a dedicated filter based on anomaly detection and attention mechanisms. First, the isolated forest algorithm identifies obvious anomalies in the logs, initially distinguishing normal business data from potential attack behaviors. Then, the attention mechanism assigns weights to the filtered data, weakening interference from false alarms and normal business fluctuations, strengthening real attack traces, and accurately improving data quality, directly contributing to improved detection accuracy and robustness. Third, industry-adaptive feature masking and data standardization: adding an industry feature masking layer to mask normal high-frequency features specific to different industry business characteristics, avoiding interference from normal business features on attack features. Simultaneously, numerical features are normalized, and textual features are encoded and converted. Cross-modal data alignment is achieved based on key identifiers, generating a unified multi-source data matrix. This ensures both the accuracy of feature extraction and the flexibility of industry adaptation, laying the foundation for scalability.
[0036] Optionally, for network layer features, core statistical features are extracted based on NetFlow (Network Traffic Acquisition Protocol) data to accurately capture network communication anomalies in APT attacks, providing fine-grained support for improving accuracy. For host features, process behavior features, resource operation features, and API call features are extracted to depict the behavioral trajectory of a host after it has been compromised, enhancing the accurate identification of attack behaviors. For sample features, static and dynamic features of binary samples are extracted to resist feature distortion caused by code obfuscation and improve the model's accuracy in identifying complex attacks. For source tracing report text features, semantic features such as attack tool descriptions, vulnerability exploitation methods, and attack tactic summaries are extracted from the source tracing report text data, providing support for weakly supervised training and collaborative decision-making between text (source tracing report text features) and behavioral features (including network layer features, host features, and sample features). At the same time, this type of feature has a general extraction logic and can adapt to the analysis needs of newly added APT organization source tracing reports.
[0037] In one optional embodiment, determining network layer weights, host weights, and sample weights based on the text features of the source tracing report includes: obtaining the historical confidence scores, network layer historical attack probabilities, host historical attack probabilities, and sample historical attack probabilities corresponding to multiple historical attack events in historical periods prior to the current period for the target network; filtering out a fuzzy confidence attack event set from multiple historical attack events based on the historical confidence scores corresponding to multiple historical attack events, and preset first and second thresholds, wherein the preset second threshold is greater than the preset first threshold; determining the network layer F1 value of the network layer detector based on the network layer historical attack probabilities corresponding to multiple historical attack events; determining the host F1 value of the host detector based on the host historical attack probabilities corresponding to multiple historical attack events; determining the sample F1 value of the sample detector based on the sample historical attack probabilities corresponding to multiple historical attack events; and determining the sample F1 value based on the network layer F1 value, host F1 value, and sample F1 value. The following steps are taken: First, network layer performance drivers, host performance drivers, and sample performance drivers are determined. The network layer performance drivers indicate the relative stability and generalization ability of the network layer detector in response to APT attacks. The host performance drivers indicate the discrimination reliability of the host detector under complex behavioral noise. The sample performance drivers indicate the robustness of the sample detector in identifying APT attacks. Second, network layer text drivers, host text drivers, and sample text drivers are determined based on the features of the source tracing report text. The network layer text drivers indicate the degree of matching between the source tracing report text data and network layer features. The host text drivers indicate the degree of matching between the source tracing report text data and host features. The sample text drivers indicate the degree of matching between the source tracing report text data and APT sample features. Third, network layer weights, host weights, and sample weights are determined based on the network layer performance drivers and network layer text drivers.
[0038] The network layer weights, host weights, and sample weights are determined as follows: First, the historical confidence levels, network layer historical attack probabilities, host historical attack probabilities, and sample historical attack probabilities for multiple historical attack events within a historical period are obtained for the target network. Second, based on the historical confidence levels, a preset first threshold, and a preset second threshold, a fuzzy confidence attack event set is selected from the multiple historical attack events. The historical confidence levels of the historical attack events in the fuzzy confidence attack event set are greater than the preset first threshold and less than or equal to the preset second threshold. This set is used to characterize potential threat events that possess some APT attack characteristics but are not yet sufficient for high-confidence judgment, providing an interpretable intermediate decision range for subsequent manual review and model feedback optimization. Then, based on the network layer historical attack probabilities, host historical attack probabilities, and sample historical attack probabilities corresponding to multiple historical attack events, the network layer F1 score of the network layer detector, the host F1 score of the host detector, and the sample F1 score of the sample detector are determined. The network layer detector is used to identify abnormal communication behaviors in APT attacks, the host detector is used to detect lateral movement and privilege escalation operations, and the sample detector is used to identify malware families and attack tools. The F1 score is used to comprehensively measure the balance between precision and recall of each detector, serving as an indicator of its stability and reliability in identifying APT attacks over a historical period. Next, based on the network layer F1 score, host F1 score, and sample F1 score, network layer performance driving factors, host performance driving factors, and sample performance driving factors are determined. Furthermore, based on the characteristics of the source tracing report text, network layer text driving factors, host text driving factors, and sample text driving factors are determined. Finally, network layer weights are determined based on network layer performance driving factors and network layer text driving factors; host weights are determined based on host performance driving factors and host text driving factors; and sample weights are determined based on sample performance driving factors and sample text driving factors. By dynamically calculating network layer weights, host weights, and sample weights, APT attack attribution decision-making based on the dual-dimensional collaborative drive of performance reliability and tactical semantics can be achieved. This improves the accuracy and anti-interference capability of attribution in scenarios involving tool reuse, behavioral camouflage, and cross-organizational obfuscation, and enhances the accuracy of attack source identification results.
[0039] Optionally, dedicated modeling and collaborative decision-making design can be used to achieve a dual improvement in detection accuracy and robustness, while relying on a generalized model architecture to support scalability. A multi-detector parallel architecture is adopted, and dedicated detectors (i.e., network layer detector, host detector, and sample detector) are designed for the network layer, host, and sample, respectively, to address the multimodal feature differences of APT attacks. Targeted modeling is conducted based on the APT origination report data space, focusing only on learning the multimodal behavioral features in the origination report, avoiding interference from irrelevant data, and significantly improving detection accuracy. At the same time, a weakly supervised + remotely supervised joint training mechanism is introduced to adapt to the training needs of newly added APT organization samples, enhancing scalability.
[0040] Optionally, the network layer detector can employ a Temporal Attention Gradient Boosting (TAGB) detector, trained based on the network layer feature set. It introduces a temporal attention mechanism into the gradient boosting framework to strengthen the feature weights of key time windows, while incorporating industry-adaptive features to accurately capture temporal communication anomalies in APT attacks and obtain the network layer attack probabilities corresponding to each attack event. The host detector can employ a Graph-enhanced Hierarchical Gradient Boosting (GHGB) detector, constructing processes, APIs, and files as behavioral graphs. It extracts topological features through graph convolutional networks and incorporates a hierarchical gradient boosting framework and behavioral constraint regularization terms to accurately characterize the host intrusion trajectory. The sample detector can employ a Multi-View Random Forest Ensemble (MRFE) detector, fusing static and dynamic dual-view features to resist code obfuscation interference.
[0041] Optionally, the network layer historical attack probability, host historical attack probability, and sample historical attack probability of historical attack events in the fuzzy confidence attack event set are manually verified. Based on the verification results, the precision and recall of the corresponding detector in the historical period are determined. Based on the precision and recall of the corresponding detector, the F1 score of the corresponding detector is calculated.
[0042] Alternatively, the performance driving factors for each detector type can be determined in the following manner:
[0043]
[0044] in, This is the performance driving factor corresponding to the j-th detector type. Detector types include network layer detectors, host detectors, and sample detectors. The F1 score corresponding to the j-th detector type. It is the minimum F1 value among the multiple detectors. This represents the maximum F1 value among the multiple detectors.
[0045] Optionally, the text-driving factor corresponding to each detector can be determined as follows: Construct a corpus of APT source tracing reports and extract semantic content such as attack tactics, tool names, IOC information, and behavioral descriptions; construct three semantic keyword dictionaries for network layer, host, and sample based on the type of each detector, with each dictionary containing terms strongly related to the corresponding modality and their semantic weights; perform semantic parsing on the text features of the source tracing reports and extract keywords that match each dictionary; calculate the sum of the semantic weights of the keywords matched by each detector as its semantic matching score; normalize the three scores to generate network layer text-driving factors, thus obtaining the text-driving factor corresponding to each detector, with a value range of [0,1].
[0046] Alternatively, the weights corresponding to the j-th detector type can be determined in the following manner. :
[0047]
[0048] in, These are the weighting coefficients of the text-driven factors. This is the text-driving factor corresponding to the j-th detector type.
[0049] In one optional embodiment, the target confidence level corresponding to each of the multiple attack events is determined based on the network layer features, host features, APT sample features, network layer weights, host weights, and sample weights corresponding to the multiple attack events. This includes: determining the network layer attack probability corresponding to each of the multiple attack events using a network layer detector based on the network layer features, wherein the network layer attack probability is used to indicate the confidence level of the network attack behavior initiated by the corresponding attack event as an APT attack; and determining the host attack probability corresponding to each of the multiple attack events using a host detector based on the host features. The attack probability is calculated as follows: host attack probability indicates the credibility of a host attack initiated by an APT attack event; based on the APT sample features corresponding to multiple attack events, a sample detector is used to determine the sample attack probability corresponding to each of the multiple attack events, where the sample attack probability indicates the credibility of the APT sample data associated with the corresponding attack event as an attack sample used by an APT group; based on the network layer attack probability, host attack probability, sample attack probability, network layer weight, host weight, and sample weight corresponding to multiple attack events, the target confidence level corresponding to each of the multiple attack events is determined.
[0050] It can be understood that the network layer features corresponding to multiple attack events are input into a network layer detector to obtain the network layer attack probabilities corresponding to each attack event. These network layer attack probabilities indicate the credibility of the corresponding attack event as an APT-initiated network attack. Similarly, the host features corresponding to multiple attack events are input into a host detector to obtain the host attack probabilities, which indicate the credibility of the corresponding attack event as an APT-initiated host attack. Finally, the APT sample features corresponding to multiple attack events are input into a sample detector to obtain the sample attack probabilities, which indicate the credibility of the APT sample data associated with the corresponding attack event as an attack sample used by an APT group. Based on the network layer attack probabilities, host attack probabilities, sample attack probabilities, network layer weights, host weights, and sample weights corresponding to multiple attack events, the target confidence level corresponding to each attack event is determined. By inputting multimodal features into the corresponding detectors, corresponding attack probabilities are generated. The network layer weights, host weights, and sample weights, which are dynamically calculated by the F1 performance driving factor and the semantic driving factor of the source tracing report text, are weighted and fused to achieve APT attack target confidence quantification from isolated modality discrimination to cross-dimensional semantic-performance co-weighting. This improves the accuracy of target confidence determination results and thus improves the accuracy of attack source determination results.
[0051] Optionally, the network layer attack probability of the i-th attack event The following method is used to determine:
[0052]
[0053] in, Here, m is the m-th gradient boosting tree, and M is the number of gradient boosting trees. For the m-th gradient boosting tree, Let be the weight coefficients of the m-th gradient boosting tree. Let be the network layer feature of the i-th attack event.
[0054] Optionally, the objective function of the host detector It is constructed in the following way:
[0055]
[0056]
[0057] in, This is the true label of the host data corresponding to the i-th attack event. For the host characteristics of the i-th attack event, Let be the predicted score based on the host features of the i-th attack event after the (m-1)-th round of gradient boosting. This is the complexity regularization coefficient. For behavioral constraint regularization coefficients, For complexity regularization, For behavioral constraint regularization terms, For learning rate, To enhance the gradient boosting tree for the m-th graph, The gradient operator for the complexity regularization term, For the gradient operator of the behavioral constraint regularization term, This is the cumulative prediction function updated after the m-th round of gradient boosting.
[0058] The probability of host attack for the i-th attack event The following method is used to determine:
[0059]
[0060] Optionally, the sample attack probability of the i-th attack event The following method is used to determine:
[0061]
[0062] in, Let be the static features in the sample features of the i-th attack event. Let be the dynamic feature among the sample features of the i-th attack event. These are the weighting coefficients for static features. These are the weighting coefficients for dynamic features. This represents the total number of trees in the random forest in the static view. This represents the total number of trees in the random forest in the dynamic view. For the t-th static view decision tree, Let t be the dynamic view decision tree.
[0063] Optionally, the three types of dedicated detectors mentioned above focus solely on feature learning from source tracing report data and mine attack patterns from multiple dimensions, significantly improving detection accuracy. Simultaneously, a joint training model combining weak supervision and remote supervision is introduced. This maps publicly available threat intelligence text to feature labels to generate weakly labeled samples and generates pseudo-labels for unlabeled data, expanding the training sample while adapting the model to new APT attacks, thus balancing accuracy and scalability requirements.
[0064] In one optional embodiment, the target confidence level corresponding to each of the multiple attack events is determined based on the network layer attack probability, host attack probability, sample attack probability, network layer weight, host weight, and sample weight of each attack event. This includes: for any attack event among the multiple attack events, determining the network layer confidence level of any attack event based on the network layer attack probability and network layer weight of any attack event; determining the host confidence level of any attack event based on the host attack probability and host weight of any attack event; determining the sample confidence level of any attack event based on the sample attack probability and sample weight of any attack event; determining the maximum value among the network layer confidence level, host confidence level, and sample confidence level as the target confidence level of any attack event; and determining the target confidence level corresponding to each of the multiple attack events using the same method as determining the target confidence level of any attack event.
[0065] It is understandable that for any attack event among multiple attack events, the network layer confidence level of any attack event is obtained by multiplying its network layer attack probability and network layer weight; the host confidence level is obtained by multiplying its host attack probability and host weight; and the sample confidence level is obtained by multiplying its sample attack probability and sample weight. The network layer confidence level, host confidence level, and sample confidence level of any attack event are compared, and the maximum value is determined as the target confidence level for that attack event. By using the method of determining the target confidence level for any attack event, the target confidence levels corresponding to multiple attack events are determined. By calculating the weighted confidence levels of attack events under each modality and taking the maximum value as the target confidence level, an "optimal modality priority" attribution mechanism can be implemented, which prioritizes the most reliable modality in decision-making. This effectively avoids the interference of multimodal noise and the misjudgment of low-confidence detectors, improving the rationality and accuracy of the target confidence level determination results.
[0066] Optionally, precise source tracing and flexible adaptation of ATP attacks can be achieved through dynamic fusion and closed-loop optimization. The outputs of the three dedicated detectors are combined through a dynamic weighted fusion mechanism, and the detector weights (including network layer weights, host weights, and sample weights) are automatically updated hourly based on detection accuracy and recall, ensuring a higher proportion of high-confidence results and further improving accuracy and robustness. The target confidence of the i-th attack event... The following method is used to determine:
[0067]
[0068] in, The weights corresponding to the j-th detector type include network layer weights, host weights, and sample weights. Detector types include network layer detectors, host detectors, and sample detectors. Let be the attack probability corresponding to the j-th detector type for the i-th attack event. This indicates that the output should be the one that maximizes the weighted score.
[0069] Step S106: Based on the target confidence corresponding to each of the multiple attack events, select a set of high-confidence attack events from the multiple attack events;
[0070] It is understandable that classifying and filtering attack events based on target confidence levels to construct a high-confidence attack event set can automatically identify and isolate highly reliable APT attack evidence, reduce the burden of manual review and the risk of erroneous responses, and provide reliable and actionable decision input for emergency response and source tracing.
[0071] In one optional embodiment, based on the target confidence levels corresponding to multiple attack events, a set of high-confidence attack events is selected from the multiple attack events, including: identifying attack events with target confidence levels greater than a preset second threshold as high-confidence attack events in the set of high-confidence attack events.
[0072] It is understandable that the target confidence scores corresponding to multiple attack events are compared with a preset second threshold, and attack events with target confidence scores greater than the preset second threshold are identified as high-confidence attack events, thereby filtering out a set of high-confidence attack events from multiple attack events. By dynamically comparing the target confidence scores with the preset second threshold, a highly automated screening of high-confidence attack events based on quantifiable confidence boundaries can be achieved, effectively filtering low-confidence noise and false positives, ensuring that the set of high-confidence attack events only contains APT attack events supported by sufficient evidence, and improving the rationality and accuracy of attack source tracing results.
[0073] Step S108: Based on the target confidence levels corresponding to the multiple high-confidence attack events included in the high-confidence attack event set, determine the attack source of the APT attack.
[0074] It is understandable that source attribution aggregation can be performed based on the target confidence of high-confidence attack events in a high-confidence attack event set, thereby achieving accurate source tracing of APT attacks driven by evidence strength and improving the credibility and interpretability of the tracing results.
[0075] In one optional embodiment, determining the source of an APT attack based on the target confidence levels corresponding to multiple high-confidence attack events included in a high-confidence attack event set includes: determining the addresses corresponding to the multiple high-confidence attack events; dividing the multiple high-confidence attack events into multiple attack event subsets based on the addresses corresponding to the multiple high-confidence attack events; for any attack event subset in the multiple attack event subsets, determining a risk score for any attack event subset based on the target confidence levels corresponding to the multiple first high-confidence attack events included in the any attack event subset, a first matching coefficient of the any attack event subset, and a second matching coefficient of the any attack event subset; determining the risk scores corresponding to the multiple attack event subsets using the same method as determining the risk scores of the any attack event subsets; determining the target attack event subset in the multiple attack event subsets based on the risk scores corresponding to the multiple attack event subsets and a preset risk threshold; and determining the target attack event subset as the attack source based on the address corresponding to the target attack event subset.
[0076] It is understandable that the following method is used to determine the source of an APT attack. First, the addresses corresponding to multiple high-confidence attack events are identified. These addresses refer to the source IP address (Internet Protocol Address), C2 server domain name, or network identifier associated with the malicious sample of the corresponding high-confidence attack event. Second, based on the addresses corresponding to the multiple high-confidence attack events, the multiple high-confidence attack events are divided into multiple attack event subsets. Among them, the high-confidence attack events in each attack event subset have the same address or belong to the same attack infrastructure (i.e., network assets actively deployed, controlled, or reused by the attacker to achieve communication, payload distribution, persistence, or covert channels in the attack chain). Next, based on the target confidence levels corresponding to the multiple first high-confidence attack events included in any subset of attack events, and the first and second matching coefficients of any subset of attack events, a risk score for any subset of attack events is determined. The risk score for each subset of attack events is then determined using the same method. The first matching coefficient describes the matching degree between the address corresponding to any subset of attack events and the known IOC of an APT organization, while the second matching coefficient describes the matching degree between the network entity to which the address corresponding to any subset of attack events belongs (i.e., the entity to which the network assets belong) and the industry label of the target network. Then, attack event subsets with risk scores greater than a preset risk threshold are identified as target attack event subsets. These target attack event subsets are characterized as systemic APT attack vectors (i.e., the entity resources carrying the attack behavior (such as C2 servers, malicious samples, and domain names), possessing high-confidence attack behavior, strong IP correlation, and industry targeting, and are core components of the attack infrastructure. Finally, the addresses corresponding to the subset of target attack events are identified as the source of the APT attack, meaning these addresses are the core infrastructure nodes that initiated and led this APT attack chain. By clustering high-confidence attack events into subsets based on their addresses, we achieve precise source tracing of APT attacks using attacker-controlled infrastructure as the anchor, evidence strength as the driver, and industry context as the constraint. This improves the credibility and interpretability of the tracing results, overcoming the fundamental shortcomings of traditional fuzzy attribution methods based on isolated IOC matching or statistical frequency, which suffer from high false positive rates, inability to correlate attack chains, and neglect of industry semantics.
[0077] Optionally, a dual threshold (including a preset first threshold and a preset second threshold) can be set to divide the decision interval. This allows high-confidence attack events in the high-confidence interval to automatically trigger a source tracing response, while attack events in the ambiguous interval are subject to manual review, further reducing the false positive rate. The source tracing response process is based on address information such as the malicious source IP, resulting in multiple attack event subsets. A risk score is then calculated by combining the industry characteristics corresponding to these subsets. The risk score for the k-th attack event subset is then calculated. The following methods can be used to determine this:
[0078]
[0079] in, This represents the number of high-confidence attack events included in the k-th subset of attack events. The address of the subset of the k-th attack event. Let be the target confidence level for the i-th attack event. Let be the multimodal feature vector of the i-th attack event, including network layer features, host features, and sample features. The weighting coefficient of the first matching coefficient. Let be the first matching coefficient of the k-th subset of attack events. The weighting coefficient for the second matching coefficient. It is the second matching coefficient of the k-th subset of attack events.
[0080] Optionally, the first matching coefficient of the attack event subset can be determined as follows: For any address associated with a subset of attack events, the matching degree is calculated by combining exact and fuzzy matching scores with historical indicators in the known APT group threat intelligence database (IOC database). For exact matches, i.e., matches in the attack event subset address that completely match a known C2 IP, domain name, or malicious sample hash of a certain APT group in the IOC database, a match type indicator variable is assigned. For fuzzy matches, i.e., matches in a subset of attack event addresses that are partially similar to items in the IOC library (e.g., domain name spelling variations, IPs located in the same C segment, file hash similarity > 90%), a matching score is calculated based on the similarity function. .
[0081]
[0082] in, Here, N is the index of the matching item, and N is the total number of successfully matched items in the address subset of attack events. This is a variable indicating the type of the match; if it is an exact match, If it is a fuzzy match, , For the first The base weight of each matching item The address of the subset of attack events The feature values of each matching item are used for similarity calculation. For similarity function, For the first in the IOC library Reference feature quantities for each matching item.
[0083] First matching coefficient for:
[0084]
[0085] in, This is the theoretical matching score if all N matching items are exact matches.
[0086] Optionally, when there are multiple subsets of target attack events, the address corresponding to the subset of target attack events with the highest risk score is determined as the main source of this APT attack, so as to achieve single-point accurate tracing based on multi-dimensional evidence strength ranking, and ensure that a unique, interpretable and highly confident tracing result can still be output in complex attack environments.
[0087] Optionally, the second matching coefficient of the attack event subset can be determined as follows: If the network entity to which the attack event subset addresses belong has launched APT attacks against the target network's industry (such as the energy industry) multiple times, then a weight is assigned to the industry's historical attack preference. If the network entity to which the subset of attack event addresses belongs primarily serves the industry of the target network, then a service matching weight is assigned. If the network entity to which the address subset of the attack event belongs has a communication association with industry equipment (such as SCADA (Supervisory Control and Data Acquisition) systems or industrial control protocol ports) of the target network in historical attack events, then a behavioral association weight is assigned. Calculate the second matching coefficient comprehensively. :
[0088]
[0089] in, Score the industry's historical attack preferences. To match service ratings, Scoring of behavioral associations.
[0090] Through the above steps S102 to S108, the goal is to obtain multi-source data corresponding to multiple attack events, filter out high-confidence attack event sets, and determine the attack source of APT attacks based on the target confidence of the multiple high-confidence attack events included in the high-confidence attack event sets. This achieves the technical effect of improving the accuracy of the attack source determination results of APT attacks, thereby solving the technical problem of inaccurate attack source determination results of APT attacks in related technologies.
[0091] Based on the above embodiments and optional embodiments, this application proposes an implementation method for an optional method of determining the source of an APT attack. This optional implementation method can be understood as a multi-layered, lightweight method for tracing the source of APT organizations.
[0092] With the deepening of digital transformation, APT attacks have shown an evolutionary trend towards greater concealment, long-term nature, and multi-carrier capabilities. Attack targets cover key sectors such as government affairs, energy, and finance. Attack methods integrate various techniques, including zero-day exploitation, social engineering, code obfuscation, and covert communication, posing significant challenges to attribution investigations. The heterogeneity of modern APT attacks is mainly reflected in three dimensions: 1) multi-modal attack carriers, encompassing various types of data such as network traffic, malicious samples, host behavior, and threat intelligence; 2) multi-level attack chains, forming a complete closed loop from initial intrusion, lateral movement, privilege escalation to data infiltration; and 3) diverse evasion methods, such as dynamically adjusting C2 servers, forging attack traces, and reusing attack tools to evade attribution tracking. The complexity and concealment of these attacks pose a severe challenge to the accuracy, robustness, and generalization ability of attribution investigation technologies. Attribution methods optimized for single data modalities or attack phases often struggle to form a complete chain of evidence when facing cross-modal, full-chain APT attacks, and may even lead to attribution biases.
[0093] Existing APT attack attribution techniques include threat intelligence matching-based methods, attack behavior feature-based methods, machine learning-based methods, and deep learning-based methods. Threat intelligence matching-based methods are currently the mainstream approach. They quickly associate suspected attack sources by comparing Indicators of Computation (IOCs) captured during the attack with APT organization characteristics in public or private threat intelligence databases. This type of method has the advantages of fast response time and low deployment cost, enabling rapid attribution of attacks from known APT organizations. However, its attribution effectiveness heavily relies on the completeness and timeliness of the intelligence database. It is extremely unsuitable for new attack tools and dynamic C2 communication evasion techniques used by APT organizations, and it is difficult to distinguish between attacks using reused attack tools and attacks directly launched by the organization, easily leading to misattribution. Attack behavior feature-based methods extract attack behavior patterns at the host and network layers to construct a behavioral profile of the APT organization, achieving precise location of the attack source. This type of method can, to some extent, reduce its dependence on static IOCs and has good results in identifying new attack variants of known organizations. However, this method heavily relies on manual extraction of behavioral features, making it difficult to cover the entire chain of APT attacks. Furthermore, it lacks robustness against attack behavior camouflage and trace erasure techniques, and its granularity for attribution is insufficient. Machine learning-based attribution methods utilize algorithms such as ensemble learning and graph neural networks to automatically learn the characteristic patterns of APT organizations from multi-source attack data, constructing end-to-end attribution models with strong generalization capabilities. Although these methods can achieve high attribution accuracy on laboratory datasets, in real-world complex network environments, due to the heterogeneity of multi-source data, noise interference, and the dynamic evolution of attack features, a single model struggles to simultaneously adapt to the feature distributions of different modalities. This leads to a significant performance drop in cross-scenario and cross-industry attribution, and the models have weak interpretability, making it difficult to provide a clear chain of attribution evidence. Deep learning-based attribution methods further enhance the ability to uncover complex attack patterns, including using convolutional neural networks to extract deep features from APT samples, recurrent neural networks to model the temporal dependencies of attack behavior, and graph attention networks to capture the correlations between multi-source data. These models can automatically learn high-level feature representations and have superior potential for tracing the origins of highly covert APT attacks. However, deep learning models generally suffer from high training data labeling costs, slow inference speeds, and opaque model decision-making processes. They are difficult to deploy in high-throughput, real-time network environments and have poor adaptability to resource-constrained edge devices, making it difficult to meet the efficiency requirements of practical tracing work.
[0094] A review of existing technologies reveals that the evolution of APT organization attribution techniques has been a continuous process of balancing "attribution efficiency" and "attribution accuracy," as well as "generalization ability" and "interpretability." Threat intelligence matching and attack behavior characteristic-based attribution methods are lightweight and efficient but have significant limitations and poor adaptability to new attack and evasion techniques. Machine learning and deep learning-based attribution methods are flexible and accurate but complex and costly to deploy, with insufficient model interpretability. These methods all aim to solve the end-to-end attribution problem using a single model or single-modal data, neglecting the inherent heterogeneity of APT attacks—multimodal and across the entire chain. This makes it difficult to find the optimal balance between the macroscopic communication characteristics of network traffic, the microscopic structural characteristics of APT samples, and the temporal evolution characteristics of host behavior. Consequently, when facing hybrid and covert real-world APT attacks, it is difficult to construct a complete and credible chain of attribution evidence, severely limiting overall attribution effectiveness. Therefore, a multi-level lightweight APT organization tracing method is proposed to overcome the inherent limitations of single-model and single-modal tracing methods, and to achieve accurate, efficient and reliable tracing of the source of APT attacks.
[0095] Figure 2 This is a flowchart of an optional method for determining the source of an APT attack, provided according to an embodiment of this application. Figure 2 As shown, the steps of the multi-level, lightweight APT organization attribution method include:
[0096] Step S1: Data acquisition and preprocessing.
[0097] To address issues such as limited training data samples, noise interference, and poor cross-domain adaptability during APT source tracing, a cross-source data aggregation design is adopted to provide high-quality data support for subsequent specialized modeling and feature extraction, while also considering feature universality to adapt to extended needs. This specifically includes two main stages: data acquisition and preprocessing optimization.
[0098] The data acquisition phase focuses on building a cross-domain APT multi-source dataset, addressing the issues of data disconnect from real-world scenarios and insufficient coverage of new attacks, while also providing support for improved accuracy and scalability. First, it aggregates core data from multiple sources, integrating source tracing reports and IOC data from publicly available threat intelligence databases, attack logs, and new, unknown APT attack data captured by honeypot systems. This achieves complementarity between publicly available intelligence and real-world scenario data, broadening the sample boundaries for model learning. Second, it employs hierarchical labeling and sample expansion, performing refined hierarchical labeling of data by industry, attack stage, and APT organization type. Simultaneously, it introduces data augmentation strategies, such as randomly supplementing attack behavior chains with nodes to simulate APT attacker code obfuscation and behavioral camouflage, effectively improving sample diversity and alleviating the scarcity of new attack samples. This allows the model to learn more comprehensive features within the source tracing report data space (including multi-source data corresponding to multiple attack events and source tracing report text data), laying a solid foundation for improved accuracy.
[0099] The preprocessing optimization process simultaneously improves accuracy and robustness, and its core consists of three steps. First, basic data cleaning: removing duplicate records, filling in missing fields, and eliminating invalid data to reduce noise interference with model training. Second, adaptive attack noise filtering: building a dedicated filter based on anomaly detection and attention mechanisms. First, the isolated forest algorithm identifies obvious anomalies in logs, initially distinguishing normal business data from potential attack behaviors. Then, the attention mechanism assigns weights to the filtered data, weakening interference from false alarms and normal business fluctuations, strengthening real attack traces, and accurately improving data quality, directly contributing to improved detection accuracy and robustness. Third, industry-adaptive feature masking and data standardization: adding an industry feature masking layer to mask normal high-frequency features specific to different industry business characteristics, avoiding interference from normal business features on attack features. Simultaneously, numerical features are normalized, and textual features are encoded and converted. Cross-modal data alignment is achieved based on key identifiers, generating a unified multi-source data matrix. This ensures both the accuracy of feature extraction and the flexibility of industry adaptation, laying the foundation for scalability.
[0100] Step S2, feature extraction.
[0101] To ensure scalability, a feature system is constructed. For the three data modalities (network layer data, host data, and APT sample data) and the characteristics of APT source tracing report text data, a unique, shared, and universal feature set is designed. This ensures the specificity of features for each modality to enhance detection accuracy, while also supporting the system's scalability through generalized and easily extractable feature design.
[0102] For network layer features, core statistical features are extracted based on NetFlow (Network Traffic Acquisition Protocol) data to accurately capture network communication anomalies in APT attacks, providing fine-grained support for improving accuracy. For host features, process behavior features, resource operation features, and API call features are extracted to depict the behavioral trajectory of a host after it has been compromised, enhancing the accurate identification of attack behaviors. For sample features, static and dynamic features of binary samples are extracted to resist feature distortion caused by code obfuscation and improve the model's accuracy in identifying complex attacks. For source tracing report text features, semantic features such as attack tool descriptions, vulnerability exploitation methods, and attack tactic summaries are extracted from the source tracing report text data, providing support for weakly supervised training and collaborative decision-making between text (source tracing report text features) and behavioral features (including network layer features, host features, and sample features). At the same time, this type of feature has a general extraction logic and can be adapted to the analysis needs of newly added APT organization source tracing reports.
[0103] Table 1 shows an example of APT attribution report text features, which cover the tactics, tools, traces, and organizational associations of the entire attack chain. These features can be directly extracted from the attribution report text data without complex calculations, balancing feature extraction efficiency and specificity. Moreover, the feature definitions are universal and can be quickly adapted to the analysis needs of attribution reports from different APT organizations, providing core support for subsequent weakly supervised training and cross-modal collaborative decision-making.
[0104] Table 1. Textual Characteristics of APT Origin Reports
[0105]
[0106] Step S3: Determine the attack probability based on multiple detectors.
[0107] Through dedicated modeling and collaborative decision-making design, both detection accuracy and robustness are improved, while scalability is supported by a generalized model architecture. A multi-detector parallel architecture is adopted, and dedicated detectors (i.e., network layer detector, host detector, and sample detector) are designed for the network layer, host, and sample, respectively, to address the multimodal feature differences of APT attacks. Targeted modeling is carried out based on the APT origination report data space, focusing only on the multimodal behavioral features learned from the origination report to avoid interference from irrelevant data and significantly improve detection accuracy. At the same time, a weakly supervised + remotely supervised joint training mechanism is introduced to adapt to the training needs of newly added APT organization samples, enhancing scalability.
[0108] The improved detection accuracy is achieved through specialized modeling and precise feature adaptation.
[0109] The network layer detector employs a Temporal Attention Gradient Boosting (TAGB) detector. Trained on the network layer feature set, it introduces a temporal attention mechanism into the gradient boosting framework to strengthen the feature weights of key time windows. Simultaneously, it incorporates industry-adaptive features to accurately capture temporal communication anomalies in APT attacks, obtaining the network layer attack probability corresponding to each attack event. The network layer attack probability for the i-th attack event is given. The following method is used to determine:
[0110]
[0111] in, Here, m is the m-th gradient boosting tree, and M is the number of gradient boosting trees. For the m-th gradient boosting tree, Let be the weight coefficients of the m-th gradient boosting tree. Let be the network layer feature of the i-th attack event.
[0112] The host detector employs a graph-enhanced hierarchical gradient boosting (GHGB) detector, constructing a behavioral graph from processes, APIs, and files. It extracts topological features through a graph convolutional network, incorporating a hierarchical gradient boosting framework and behavioral constraint regularization terms to accurately characterize the host intrusion trajectory. The objective function of the host detector is... It is constructed in the following way:
[0113]
[0114]
[0115] in, This is the true label of the host data corresponding to the i-th attack event. For the host characteristics of the i-th attack event, Let be the predicted score based on the host features of the i-th attack event after the (m-1)-th round of gradient boosting. This is the complexity regularization coefficient. For behavioral constraint regularization coefficients, For complexity regularization, For behavioral constraint regularization terms, For learning rate, To enhance the gradient boosting tree for the m-th graph, The gradient operator for the complexity regularization term, For the gradient operator of the behavioral constraint regularization term, This is the cumulative prediction function updated after the m-th round of gradient boosting.
[0116] The probability of host attack for the i-th attack event The following method is used to determine:
[0117]
[0118] The sample detector employs a Multi-View Random Forest Ensemble (MRFE) detector, fusing static and dynamic dual-view features to resist code obfuscation interference. The sample attack probability for the i-th attack event is... The following method is used to determine:
[0119]
[0120] in, Let be the static features in the sample features of the i-th attack event. Let be the dynamic feature among the sample features of the i-th attack event. These are the weighting coefficients for static features. These are the weighting coefficients for dynamic features. This represents the total number of trees in the random forest in the static view. This represents the total number of trees in the random forest in the dynamic view. For the t-th static view decision tree, Let t be the dynamic view decision tree.
[0121] The three types of dedicated detectors focus solely on feature learning from source tracing report data and mine attack patterns from multiple dimensions, significantly improving detection accuracy. Simultaneously, they introduce joint training with weak supervision and remote supervision, mapping publicly available threat intelligence texts to feature labels to generate weakly labeled samples and generating pseudo-labels for unlabeled data. This expands the training sample while adapting the model to new APT attacks, balancing accuracy and scalability requirements.
[0122] The improvement in detection robustness relies on a multimodal, multi-level collaborative decision-making mechanism, achieved through text-behavioral feature linkage and dynamic weight adjustment.
[0123] A collaborative decision-making system combining "source tracing report text features + multimodal behavioral features" is constructed. Drawing on the concept of "multimodal information complementarity to enhance scenario representation," this system deeply integrates source tracing report text features with network, host, and sample behavioral features. Based on the recent attack performance of different APT groups, the system adaptively adjusts the decision weights of each detector and feature dimension—giving higher weights to feature dimensions and detectors with high recent attack pattern matching and high identification accuracy, highlighting more reliable source tracing results and effectively suppressing false positives. Simultaneously, high-confidence feature dimensions are prioritized for computational resource allocation, strengthening the feature capture of real attack traces and further reducing the false positive rate caused by noise interference. For example, if an APT group has recently frequently used a specific vulnerability + attack tool combination, the system will automatically increase the weights of the corresponding text features and behavioral features, relying on the cross-modal collaborative mechanism to enhance the identification reliability of this type of attack.
[0124] The scalability support is achieved by relying on a generalized model architecture and easily extracted features.
[0125] Each detector is designed based on a general framework, with unified and easily implementable feature extraction logic, eliminating the need to reconstruct the model for new APT organizations. When it is necessary to expand to more APT organizations, it is only necessary to preprocess the source report data of the new organizations according to the existing format, extract the general features, expand the category labels on the basis of the existing classification, and carry out incremental training in conjunction with the original samples, without the need for full retraining, which greatly improves training efficiency and reduces time costs.
[0126] Step S4: The source of the attack is identified.
[0127] Through dynamic fusion and closed-loop optimization, precise source tracing and flexible adaptation of ATP attacks are achieved. The outputs of three dedicated detectors are combined through a dynamic weighted fusion mechanism. Detector weights (including network layer weights, host weights, and sample weights) are automatically updated hourly based on detection accuracy and recall, ensuring a higher proportion of high-confidence results and further improving accuracy and robustness. The target confidence level of the i-th attack event is... The following method is used to determine:
[0128]
[0129] in, The weights corresponding to the j-th detector type include network layer weights, host weights, and sample weights. Detector types include network layer detectors, host detectors, and sample detectors. Let be the attack probability corresponding to the j-th detector type for the i-th attack event. This indicates that the output should be the one that maximizes the weighted score.
[0130] A dual threshold system (including a preset first threshold and a preset second threshold) is used to divide the decision interval. High-confidence attack events in the high-confidence interval automatically trigger a source tracing response, while attack events in the ambiguous interval undergo manual review to further reduce the false positive rate. The source tracing response process is based on malicious source IP and other address information to divide the event into multiple attack event subsets. A risk score is calculated by combining the industry characteristics corresponding to these subsets. The risk score for the k-th attack event subset is shown below. The following method is used to determine:
[0131]
[0132] in, This represents the number of high-confidence attack events included in the k-th subset of attack events. The address of the subset of the k-th attack event. Let be the target confidence level for the i-th attack event. Let be the multimodal feature vector of the i-th attack event, including network layer features, host features, and sample features. The weighting coefficient of the first matching coefficient. Let be the first matching coefficient of the k-th subset of attack events. The weighting coefficient for the second matching coefficient. It is the second matching coefficient of the k-th subset of attack events.
[0133] The accuracy of source tracing is improved by precise risk scoring. At the same time, a closed-loop feedback mechanism is introduced to dynamically adjust the threshold, detector weight and industry feature mask parameters, and update the weak supervision label mapping rules and collaborative decision weights in sync, so as to continuously optimize the accuracy and robustness of source tracing.
[0134] Based on the above-mentioned multi-level lightweight APT organization tracing method, a multi-level lightweight APT organization tracing architecture is proposed. Figure 3 This is a structural diagram of an optional APT organization tracing architecture provided according to an embodiment of this application, such as... Figure 3As shown, the APT organization attribution architecture includes: a data acquisition and preprocessing unit, a multimodal feature extraction unit, a multi-detector attack identification unit, and a decision fusion and attribution response unit. The data acquisition and preprocessing unit is responsible for aggregating multi-source data from networks, hosts, and samples, and completing cleaning and standardization. The multimodal feature extraction unit extracts specific feature sets for different data types, balancing fine granularity and anti-obfuscation. The multi-detector attack identification unit runs three dedicated detectors in parallel, each adapted to network layer features, host features, and APT sample features. The decision fusion and attribution response unit dynamically weights and fuses the outputs of each detector to generate a final risk score, simultaneously completing attack source attribution and response rule generation. This design ensures accurate capture of heterogeneous data features while achieving a balance between detection efficiency and attribution completeness.
[0135] An optional APT organization tracing method and architecture serves as the APT tracing analysis engine. Employing a distributed collaborative deployment model, it balances data security and analytical efficiency, providing precise support for APT attack tracing and prevention. Data acquisition, preprocessing, and basic feature extraction units are deployed on distributed edge nodes. Embedded security probes collect full data in real time, covering network traffic, host logs, APT sample traces, and cached APT tracing report data, strictly adhering to local data preprocessing requirements. Edge nodes perform lightweight processing, cleaning and removing invalid data. An attack noise adaptive filtering module distinguishes normal data from potential attacks, weakens interference items, and extracts core feature vectors, reducing transmission volume while retaining key information to meet system real-time requirements. Simultaneously, the platform identifies APT attacks in parallel using multiple detectors, generates weakly labeled samples based on tracing reports, optimizes the adaptation capability for new attacks, and outputs highly reliable tracing results through dynamic weighted fusion. The platform generates response rules and distributes them to edge nodes for rapid handling. Closed-loop feedback optimizes model parameters, continuously improving detection accuracy and robustness, constructing a fully closed-loop APT protection link.
[0136] The above optional implementation methods achieve at least the following effects: by constructing a cross-domain heterogeneous APT hybrid dataset, the dataset of APT organizations is expanded, thereby improving sample diversity; based on text features and attack behavior features, collaborative decision-making is carried out to adaptively adjust the decision weights according to the performance of different APT organizations in recent attacks, highlighting more reliable tracing results and effectively reducing the false positive rate.
[0137] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0138] This embodiment also provides an APT attack source identification device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated for details already described. As used below, the terms "module" and "device" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0139] According to an embodiment of this application, an apparatus embodiment for a method of determining the source of an APT attack is also provided. Figure 4 This is a schematic diagram of an APT attack source determination device according to an embodiment of this application, as shown in the figure. Figure 4 As shown, the APT attack source identification device includes a data acquisition module 402, a first identification module 404, a filtering module 406, and a second identification module 408. The device will be described below.
[0140] The data acquisition module 402 is used to acquire multi-source data corresponding to multiple attack events of the target network in the current period. Here, an attack event refers to a unit that supports independent detection in APT attack detection.
[0141] The first determining module 404, connected to the data acquisition module 402, is used to determine the target confidence level corresponding to each of the multiple attack events based on the multi-source data corresponding to the multiple attack events. The target confidence level is used to indicate the credibility of the corresponding attack event as an APT attack.
[0142] The filtering module 406, connected to the first determining module 404, is used to filter out a set of high-confidence attack events from multiple attack events based on the target confidence levels corresponding to the multiple attack events respectively.
[0143] The second determining module 408, connected to the filtering module 406, is used to determine the source of an APT attack based on the target confidence corresponding to each of the multiple high-confidence attack events included in the high-confidence attack event set.
[0144] This application provides an APT attack source determination device. By setting a data acquisition module 402, a first determination module 404, a filtering module 406, and a second determination module 408, it achieves the purpose of acquiring multi-source data corresponding to multiple attack events, filtering out a set of high-confidence attack events, and determining the attack source of the APT attack based on the target confidence levels corresponding to the multiple high-confidence attack events included in the high-confidence attack event set. This achieves the technical effect of improving the accuracy of the APT attack source determination result, thereby solving the technical problem of inaccurate APT attack source determination results in related technologies.
[0145] It should be noted that the above modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following ways: the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0146] It should be noted that the data acquisition module 402, the first determination module 404, the filtering module 406, and the second determination module 408 mentioned above correspond to steps S102 to S108 in the embodiments. The instances and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as part of the device, can run in a computer terminal.
[0147] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant descriptions in the embodiments, and will not be repeated here.
[0148] The aforementioned APT attack source identification device may also include a processor and a memory. The data acquisition module 402, the first determination module 404, the filtering module 406, the second determination module 408, etc., are all stored in the memory as program units, and the processor executes the aforementioned program units stored in the memory to realize the corresponding functions.
[0149] The processor contains a core that retrieves the corresponding program unit from memory. One or more cores may be configured. Memory may include non-persistent memory in computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory includes at least one memory chip.
[0150] This application provides a non-volatile storage medium storing a program that, when executed by a processor, implements a method for determining the source of an APT attack.
[0151] This application provides an electronic device. Figure 5 This is a structural diagram of an electronic device provided according to an embodiment of this application. For example... Figure 5 As shown, the electronic device may include: one or more ( Figure 5(Only one is shown in the diagram) Processor 502, memory 504, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module, and display. The electronic device includes a processor, memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it performs the following steps: acquiring multi-source data corresponding to multiple attack events of the target network within the current period, wherein an attack event refers to a unit that supports independent detection in APT attack detection; determining the target confidence level corresponding to each of the multiple attack events based on the multi-source data corresponding to the multiple attack events, wherein the target confidence level indicates the credibility of the corresponding attack event as an APT attack; filtering out a high-confidence attack event set from the multiple attack events based on the target confidence levels of the multiple high-confidence attack events included in the high-confidence attack event set; and determining the attack source of the APT attack based on the target confidence levels corresponding to the multiple high-confidence attack events included in the high-confidence attack event set. The device in this document can be a server, PC, etc.
[0152] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing an initialization program with the following method steps: acquiring multi-source data corresponding to multiple attack events of a target network within the current period, wherein an attack event refers to a unit that supports independent detection in APT attack detection; determining the target confidence level corresponding to each of the multiple attack events based on the multi-source data corresponding to the multiple attack events, wherein the target confidence level is used to indicate the credibility of the corresponding attack event as an APT attack; filtering out a set of high-confidence attack events from the multiple attack events based on the target confidence levels corresponding to the multiple attack events included in the set of high-confidence attack events; and determining the attack source of the APT attack based on the target confidence levels corresponding to the multiple high-confidence attack events included in the set of high-confidence attack events.
[0153] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0154] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0155] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0156] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0157] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0158] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0159] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0160] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0161] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0162] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for determining the source of an APT attack, characterized in that, include: Acquire multi-source data corresponding to multiple attack events of the target network in the current period, wherein the attack event refers to a unit that supports independent detection in APT attack detection; Based on the multi-source data corresponding to the multiple attack events, the target confidence level corresponding to each of the multiple attack events is determined, wherein the target confidence level is used to indicate the credibility of the corresponding attack event as an APT attack; Based on the target confidence levels corresponding to the multiple attack events, a set of high-confidence attack events is selected from the multiple attack events; Based on the target confidence levels corresponding to the multiple high-confidence attack events included in the high-confidence attack event set, the attack source of the APT attack is determined.
2. The method according to claim 1, characterized in that, When the multi-source data includes network layer data, host data, and APT sample data corresponding to the attack events, determining the target confidence level corresponding to each of the multiple attack events based on the multi-source data corresponding to each of the multiple attack events includes: Obtain the traceability report text data for the current period; The multi-source data and the source tracing report text data corresponding to the multiple attack events are preprocessed respectively to obtain standard data and source tracing report text standard data corresponding to the multiple attack events respectively. The standard data includes network layer standard data, host standard data and APT sample standard data corresponding to the attack events. Feature extraction is performed on the standard network layer data corresponding to each of the multiple attack events to obtain the network layer features corresponding to each of the multiple attack events. Feature extraction is performed on the host standard data corresponding to each of the multiple attack events to obtain the host features corresponding to each of the multiple attack events; Feature extraction is performed on the APT sample standard data corresponding to the multiple attack events respectively to obtain the APT sample features corresponding to the multiple attack events respectively; Feature extraction is performed on the standard data of the traceability report text to obtain the traceability report text features; Based on the text features of the source tracing report, the network layer weight, host weight, and sample weight are determined. Based on the network layer features corresponding to the multiple attack events, the host features corresponding to the multiple attack events, the APT sample features corresponding to the multiple attack events, the network layer weights, the host weights, and the sample weights, the target confidence levels corresponding to the multiple attack events are determined.
3. The method according to claim 2, characterized in that, The determination of network layer weights, host weights, and sample weights based on the text features of the source tracing report includes: The target network is obtained as follows: in the historical periods prior to the current period, the historical confidence of multiple historical attack events, the network layer historical attack probability of the multiple historical attack events, the host historical attack probability of the multiple historical attack events, and the sample historical attack probability of the multiple historical attack events. Based on the historical confidence levels corresponding to the multiple historical attack events, and a preset first threshold and a preset second threshold, a set of fuzzy confidence attack events is selected from the multiple historical attack events, wherein the preset second threshold is greater than the preset first threshold; Based on the network layer historical attack probabilities corresponding to the multiple historical attack events, the network layer F1 value of the network layer detector is determined. Based on the host historical attack probabilities corresponding to the multiple historical attack events, the host F1 value of the host detector is determined. Based on the sample historical attack probabilities corresponding to the multiple historical attack events, the sample F1 value of the sample detector is determined. Based on the network layer F1 value, the host F1 value, and the sample F1 value, a network layer performance driving factor, a host performance driving factor, and a sample performance driving factor are determined. The network layer performance driving factor is used to indicate the relative stability and generalization ability of the network layer detector in response to the APT attack. The host performance driving factor is used to indicate the discrimination reliability of the host detector under complex behavioral noise. The sample performance driving factor is used to indicate the robustness of the sample detector in identifying the APT attack. Based on the source tracing report text features, network layer text driving factors, host text driving factors, and sample text driving factors are determined. The network layer text driving factor is used to indicate the degree of matching between the source tracing report text data and the network layer features. The host text driving factor is used to indicate the degree of matching between the source tracing report text data and the host features. The sample text driving factor is used to indicate the degree of matching between the source tracing report text data and the APT sample features. The network layer weights are determined based on the network layer performance driving factor and the network layer text driving factor. The host weight is determined based on the host performance driving factor and the host text driving factor; The sample weights are determined based on the sample performance driving factors and the sample text driving factors.
4. The method according to claim 2, characterized in that, The determination of the target confidence level corresponding to each of the multiple attack events based on the network layer features, host features, APT sample features, network layer weights, host weights, and sample weights includes: Based on the network layer features corresponding to the multiple attack events, a network layer detector is used to determine the network layer attack probability corresponding to the multiple attack events, wherein the network layer attack probability is used to indicate the credibility of the corresponding attack event as a network attack behavior initiated by the APT attack. Based on the host characteristics corresponding to the multiple attack events, a host detector is used to determine the host attack probability corresponding to the multiple attack events, wherein the host attack probability is used to indicate the credibility of the corresponding attack event as a host attack behavior initiated by the APT attack. Based on the APT sample features corresponding to the multiple attack events, a sample detector is used to determine the sample attack probability corresponding to the multiple attack events. The sample attack probability is used to indicate the credibility of the APT sample data associated with the corresponding attack event as an attack sample used by an APT organization. Based on the network layer attack probability, host attack probability, sample attack probability, network layer weight, host weight, and sample weight corresponding to the multiple attack events, the target confidence level corresponding to each of the multiple attack events is determined.
5. The method according to claim 4, characterized in that, The determination of the target confidence level corresponding to each of the multiple attack events, based on the network layer attack probability, host attack probability, and sample attack probability corresponding to each of the multiple attack events, the network layer weight, the host weight, and the sample weight, includes: For any one of the plurality of attack events, the network layer confidence level of any one attack event is determined based on the network layer attack probability of any one attack event and the network layer weight; Based on the host attack probability of any attack event and the host weight, determine the host confidence level of any attack event; Based on the sample attack probability of any attack event and the sample weight, determine the sample confidence level of any attack event; The maximum value among the network layer confidence, the host confidence, and the sample confidence is determined as the target confidence for any attack event. The target confidence level corresponding to each of the multiple attack events is determined by using the method of determining the target confidence level of any one of the attack events.
6. The method according to claim 1, characterized in that, The step of filtering out a set of high-confidence attack events from the multiple attack events based on the target confidence levels corresponding to each of the multiple attack events includes: Among the target confidence scores corresponding to the multiple attack events, the attack events with target confidence scores greater than a preset second threshold are identified as high-confidence attack events in the high-confidence attack event set.
7. The method according to any one of claims 1 to 6, characterized in that, The step of determining the source of the APT attack based on the target confidence levels corresponding to multiple high-confidence attack events included in the high-confidence attack event set includes: Determine the addresses corresponding to the multiple high-confidence attack events; Based on the addresses corresponding to the multiple high-confidence attack events, the multiple high-confidence attack events are divided into multiple attack event subsets; For any subset of attack events in the plurality of attack event subsets, a risk score for the subset of attack events is determined based on the target confidence corresponding to the plurality of first high confidence attack events included in the subset of attack events, the first matching coefficient of the subset of attack events, and the second matching coefficient of the subset of attack events. The risk scores corresponding to the multiple attack event subsets are determined by using a method that determines the risk score of any subset of the attack events. Based on the risk scores corresponding to the multiple attack event subsets and the preset risk thresholds, the target attack event subsets in the multiple attack event subsets are determined; The source of the attack is determined based on the address corresponding to the subset of target attack events.
8. A device for determining the source of an APT attack, characterized in that, include: The data acquisition module is used to acquire multi-source data corresponding to multiple attack events of the target network in the current period, wherein the attack event refers to a unit that supports independent detection in APT attack detection. The first determining module is used to determine the target confidence level corresponding to each of the multiple attack events based on the multi-source data corresponding to each of the multiple attack events, wherein the target confidence level is used to indicate the credibility of the corresponding attack event as an APT attack; The filtering module is used to filter out a set of high-confidence attack events from the multiple attack events based on the target confidence levels corresponding to the multiple attack events respectively; The second determining module is used to determine the source of the APT attack based on the target confidence levels corresponding to the multiple high-confidence attack events included in the high-confidence attack event set.
9. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores multiple instructions, which are adapted to be loaded by a processor and executed by the method for determining the source of an APT attack as described in any one of claims 1 to 7.
10. An electronic device, characterized in that, include: One or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method for determining the source of an APT attack as described in any one of claims 1 to 7.