Multi-source threat intelligence data fusion and decision optimization method
By constructing a semantic fragment merging and comparison set and a behavioral consistency expression set, the behavioral chain progression path is reconstructed, target levels are divided, and the threat response sequence is optimized. This solves the problems of scattered expression and ambiguous target identification in the processing of multi-source threat intelligence data, and improves the uniformity of information processing and response efficiency.
Patent Information
- Application Number
- CN202511522987.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2026-01-09
AI Technical Summary
Existing methods, when processing multi-source threat intelligence data, lack precise mapping between actions and subject structures, resulting in scattered expressions, difficulties in merging, and an inability to adapt to changes in behavioral chains within the context. This affects the complete reconstruction of paths. Furthermore, the lack of a unified standard for the distribution hierarchy of target entities leads to ambiguity in the identification of key nodes, weakening the targeting of strategy deployment. The response strategy lacks a sorting mechanism for path terminals and cross-targets, which can easily lead to redundant resource scheduling or delays in the response to key targets, reducing the coordination and coverage accuracy of the defense process.
By constructing a semantic guidance index, clearing ambiguous descriptions, generating a semantic fragment merging and comparison set, focusing on a set of consistent behavioral expressions, reconstructing the behavioral chain progression path, dividing target levels, optimizing the threat response order, and generating priority threat response results.
It enhances the ability to focus on semantic direction, improves the uniformity and discriminability of information structure, strengthens the structural alignment and semantic cleanup of behavioral descriptions, clarifies the coherent path of action chains, optimizes the order of resource allocation in the response process, and improves the efficiency and accuracy of handling.
Smart Images

Figure CN121302273A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent decision-making and information processing technology, and in particular to a method for decision optimization based on multi-source threat intelligence data fusion. Background Technology
[0002] The field of intelligent decision-making and information processing technology encompasses the entire process of collecting, analyzing, fusing, and supporting decisions from multi-source data. Its core lies in forming a comprehensive information foundation that can be used to assist judgment and decision-making by uniformly modeling and extracting features from data from different sources. Its overall technical system includes data acquisition and cleaning, semantic understanding and feature construction, information fusion and correlation analysis, and rule-based or model-based decision generation mechanisms. It typically combines statistical analysis, knowledge reasoning, pattern recognition, and data semantic matching to achieve structured associations and logical inferences between data, providing information support for decision-making in complex environments.
[0003] Among them, the multi-source threat intelligence data fusion and decision optimization method refers to a method in cybersecurity and intelligence analysis scenarios that unifies and fuses threat intelligence data from different channels to form a decision-making basis to support security policy formulation. The technical aspects covered include threat data source identification, information semantic standardization, feature correlation matching, and multi-dimensional data fusion reasoning. It mainly establishes an attribute mapping system for multi-source intelligence data to uniformly express the structural characteristics of different data sources, utilizes semantic association rules of information content to complete data fusion processing, and then constructs a decision parameter set based on the fusion results, thereby forming an intelligence analysis framework that can be used for decision optimization.
[0004] Existing methods, when handling heterogeneous intelligence expressions, lack precise mapping between actions and subject structures, often resulting in scattered expressions and difficulties in merging. Sequential logic judgments rely on static templates, making it difficult to adapt to changes in behavioral chains within the context, affecting the complete reconstruction of paths. The lack of a unified standard for the distribution hierarchy of target entities leads to ambiguity in the identification of key nodes across multiple paths, weakening the targeting of strategy deployment. Response strategies lack a sorting mechanism for path endpoints and intersecting targets, easily causing redundant resource scheduling or delays in the response to key targets, reducing the coordination and coverage accuracy of the defense process. Summary of the Invention
[0005] To address the technical problems existing in the prior art, this invention provides a method for multi-source threat intelligence data fusion and decision optimization. The technical solution is as follows:
[0006] A method for decision optimization based on multi-source threat intelligence data fusion includes the following steps: S1: Extract statements involving targets, behaviors, and contexts from multi-source intelligence, construct a semantic guidance index, merge semantically similar expressions, remove ambiguous descriptions or duplicate content, and generate a semantic fragment merge comparison set; S2: Based on the semantic fragment merging and comparison set, focus on the behavioral content pointing to the same object, filter logical jump items, retain the main phrases with clear expression structure, restore the unified semantics in combination with the context, and generate a set of consistent behavioral expressions; S3: Based on the event progression clues located in the set of consistent behavioral expressions, mark the starting point and node connection relationship of the action, connect the behavioral order according to the context identifier, reconstruct the behavioral chain progression path, and generate the behavioral order triggering path; S4: Based on the behavioral sequence triggering path, extract the first and last associated positions of each target in the path, divide the key nodes into upper and lower levels according to the action succession relationship, unify and organize the hierarchical structure, and generate the target hierarchical division content. S5: Combine the target hierarchy classification to analyze the target concentration trend and path convergence, check the location distribution of high-frequency response targets, reorder each disposal object according to the endpoint path weight, and generate threat response priority results.
[0007] As a further aspect of the present invention, The semantic fragment merging and comparison set includes a unified item for action phrases, a set of subject-verb pairing structures, and semantic direction indicator markers; The set of behaviorally consistent expressions includes target-consistent instruction items, expression convergent phrase groups, and semantic cleanup key components; The behavioral sequence triggering path includes action ordering relationships, logical connection point annotations, and behavioral chain progression sequences. The target hierarchy classification includes the first occurrence location record, the belonging location identifier, and the structural hierarchy mapping label; The threat response priority results include a response exit order table, cross-path target ranking items, and a list of end-point objects.
[0008] As a further aspect of the present invention, the step of obtaining the semantic fragment merging comparison set is as follows: S101: After acquiring multi-source intelligence data, relevant segments involving attack behaviors are selected, the attack behaviors in each segment are labeled, target entities are identified and associated, and an index relationship between each segment and the corresponding target entity is established; the result is the preliminary association data between attack behaviors and target entities, forming the association data between attack behaviors and target entities. S102: Based on the data associated with the attack behavior and the target entity, analyze the pairing relationship between the behavior and the subject in the fragment, identify and match similar action phrases, and obtain a simplified set of action phrases by merging repeated action expressions and retaining key parts; the processing result is the merged action phrase data, which is called the merged action phrase data. S103: Further filter the merged data of the action phrases, extract content with clear semantic direction and relevance, and conduct comparative analysis to form a corresponding semantic fragment merging set; this set contains the merged semantic fragments, and the result is a semantic fragment merging comparison set.
[0009] As a further aspect of the present invention, the step of obtaining the behavior consistency expression set is as follows: S201: After obtaining the semantic fragment merging and comparison set, identify the behavioral descriptions pointing to the consistent target, mark the association between each behavior and the target, filter out phrases with repetitive structures or similar expressions, and remove components that cause ambiguity or meaning jumps; obtain behavioral description data pointing to the consistent target, called target consistent behavioral description data; S202: Based on the target consistency behavior description data, and combined with the context semantic continuity, behavior descriptions with similar semantics are merged, redundant information is removed and core behavior content is retained; the result after processing is simplified behavior description data, referred to as simplified behavior description data. S203: Based on the simplified behavior description data, mark the behavior instructions that can be merged, verify them in combination with the semantic continuity of the context, integrate them into a unified set of behavior descriptions, and obtain a set of consistent behavior expressions.
[0010] As a further aspect of the present invention, the step of obtaining the behavior sequence triggering path is as follows: S301: After obtaining the behavioral statements in the behavioral consistency expression set, analyze the connection between the preceding and following structures of each statement, extract the sequence relationship clues, identify the order of actions in multiple segments, and mark the logical connection points; obtain the sequence relationship clue data; S302: Based on the sequence relationship clue data, sort multiple action segments, clarify the logical connection position of each action segment, determine the start and end positions of the action in the behavior chain by combining the context semantics, and determine the progression order of the action sequence; generate action sorting judgment data; S303: Based on the action sorting judgment data, construct a combined path according to the action progression order, integrate the successive relationship of action segments, and generate a behavior sequence triggering path.
[0011] As a further aspect of the present invention, the process of analyzing the structural connection of each statement, extracting sequential relationship clues, identifying the order of actions in multiple segments, and marking logical connection points specifically involves: parsing the behavioral statements through dependency parsing, extracting time adverbs, logical conjunctions, or verb states in the statements as sequential relationship clues, and defining the preconditions and postconditions of related actions as logical connection points based on the sequential relationship clues. The process of determining the start and end positions of an action in a behavior chain by combining contextual semantics is as follows: the logical connection points of the action segment are analyzed. If a logical connection point of an action segment only has a subsequent result and no preceding condition, it is determined as the start position of the behavior chain. If a logical connection point of an action segment only has a preceding condition and no subsequent result, it is determined as the end position of the behavior chain.
[0012] As a further aspect of the present invention, the step of obtaining the target hierarchical division content is as follows: S401: Obtain each behavior node in the behavior sequence triggering path, identify the target entity corresponding to each node, and determine the first and first appearance positions of the target entity. Use these positions to determine the distribution level of the target in the path; obtain the target entity distribution level data. S402: Based on the target entity distribution hierarchy data, identify key objects that appear repeatedly or are in transit positions in the path, and combine the positional relationship of these objects in the path to perform hierarchical division and clarify their belonging in the structure; the result of this step is the key object hierarchy division data. S403: Based on the key object hierarchical classification data, organize the hierarchical relationship of the target entity in the path and generate a complete target hierarchical classification result; the result is the target hierarchical classification content.
[0013] As a further aspect of the present invention, the process of determining the distribution level of the target in the path through these positions specifically involves: calculating the node span between the first occurrence of the target entity and the location where it appears; if the node span exceeds a preset proportion threshold of the total number of nodes in the behavior sequence triggering path, then the distribution level of the target entity is determined as the global level; otherwise, it is determined as the local level. The process of hierarchically classifying key objects that repeatedly appear or are in transit positions in the identification path, combined with the positional relationship of these objects in the path, is as follows: The frequency of occurrence of each target entity in the behavior sequence triggering path is counted, and the target entities with a frequency greater than a preset frequency threshold are identified as key objects; the connection relationship between the key objects and other target entities is analyzed. If a key object is pointed to by multiple target entities from different sources, the key object is defined as a parent-level entity, and the multiple target entities from different sources pointing to the key object are defined as child-level entities.
[0014] As a further aspect of the present invention, the step of obtaining the priority result of the threat response is as follows: S501: Obtain the object list in the target hierarchy division content, check whether there are target items located at the end of the path or frequently associated across paths, identify and mark the location and association of these target items; the result of this step is target item aggregation and association data. S502: Based on the target item aggregation and association data, sort the target objects located at the end position and with cross features, adjust their priority in the response process, and clarify their response order boundaries; obtain target object sorting adjustment data; S503: Based on the target object sorting and adjustment data, organize the order of appearance of each target in the response process, and combine the target sorting information to generate a priority result for threat response.
[0015] As a further aspect of the present invention, the process of identifying and marking the position and association of these target items specifically comprises: calculating the ratio of the position index value of a target item in the behavior sequence triggering path to the total number of nodes in the behavior sequence triggering path; when the ratio is greater than a preset end position threshold, marking the target item as an end; counting the number of behavior sequence triggering paths associated with each target item; when the number exceeds a preset cross-path association threshold, marking the target item as an association. The process of sorting target objects located at the end position and having cross characteristics, and adjusting their priority in the response flow, specifically involves: calculating a threat response priority score for each target item that simultaneously has the end marker and the association marker. The threat response priority score is a weighted sum based on the number of behavioral sequence triggering paths associated with the target item and the position index value of the target item within the path; and prioritizing all target items that simultaneously have the end marker and the association marker according to the threat response priority score from high to low.
[0016] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: This invention strengthens the semantic focus by merging and representing attack behaviors and target elements, enhancing the uniformity and discriminability of the information structure. Structural alignment and semantic cleanup of behavioral descriptions improve instruction uniformity and expression consistency, avoiding misunderstandings caused by word and sentence differences. Sequential clue extraction, based on context, constructs a coherent path of action chains, clarifying dependencies. The first occurrence and position of target entities in the path are extracted to form a hierarchical index, revealing the structural status of transit nodes and high-frequency targets. Tail-end aggregation of target identification and priority ranking optimizes the order of resource allocation in the response process, improving efficiency and accuracy. Attached Figure Description
[0017] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a flowchart illustrating the process of obtaining the semantic fragment merging comparison set in this invention. Figure 3 This is a flowchart illustrating the process of obtaining the behavior-consistent expression set of the present invention. Figure 4 This is a flowchart illustrating the process of obtaining the behavior sequence triggering path in this invention. Figure 5 This is a flowchart illustrating the process of obtaining the target hierarchical division content of this invention. Figure 6 This is a flowchart illustrating the process of obtaining priority results for threat response in this invention. Detailed Implementation
[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0019] refer to Figures 1 to 6 A method for decision optimization based on multi-source threat intelligence data fusion includes the following steps: S1: Obtain expression fragments involving attack behavior, target entities and related environments from multi-source intelligence, establish sentence index relationships based on source information, identify the behavior and subject pairing method in the description order, merge action phrases with similar expressions, retain parts with clear semantic direction, and generate a semantic fragment merging comparison set; S2: Call the semantic fragment merging comparison set, identify behavioral descriptions pointing to the same target in the same source information, extract phrases with repetitive structures or similar expressions, clean up components that are prone to ambiguity or meaning jumps, sort them out in combination with the semantic continuity of the context, mark the behavioral instructions that can be merged, and generate a set of behaviorally consistent expressions. S3: Call the behavioral statements in the behavioral consistency expression set, analyze the connection between the preceding and following structures in the expression, extract the sequential relationship clues from the context, sort and judge multiple action segments, clarify the logical connection points and the start and end positions of the behavioral chain, construct the combination path according to the order of behavior progression, and generate the behavioral order trigger path; S4: Call the behavior sequence to trigger the path, identify the target entity corresponding to each behavior node in the path, determine its distribution level based on the first and second appearance of the target in the path, sort out the key objects that appear repeatedly or are in the transit position under multiple paths, classify their hierarchical affiliation in the structure, and generate target hierarchical division content. S5: Invoke the object list in the target hierarchy, check whether there are target items clustered at the end of the path or frequently associated across paths, sort and adjust the objects at the end and with cross characteristics, define the response order boundary, adjust the appearance order of each layer of targets in the response process, and generate priority results for threat response.
[0020] The semantic fragment merging and comparison set includes action phrase unification items, subject-predicate pairing structure set, and semantic direction indicator. The behavioral consistency expression set includes target consistency instruction items, expression convergence phrase groups, and semantic cleanup key components. The behavioral order triggering path includes action ordering relationship, logical connection point annotation, and behavioral chain advancement sequence. The target level division content includes first appearance location record, attribution location identifier, and structural level mapping label. The threat response priority results include response appearance order table, cross-path target ordering items, and closing position object list.
[0021] Please see Figure 2 The steps for obtaining the semantic fragment merging comparison set are as follows: S101: After acquiring multi-source intelligence data, relevant segments involving attack behaviors are selected, the attack behaviors in each segment are labeled, target entities are identified and associated, and an index relationship between each segment and the corresponding target entity is established; the result is the preliminary association data between attack behaviors and target entities, forming the association data between attack behaviors and target entities. When processing three network threat intelligence reports from different sources, identified as A, B, and C, the text is read line by line, and a keyword list containing offensive verbs such as "execute," "connect," "penetrate," and "read" is used for matching. Once a keyword appears in the text, the relevant context is automatically extracted to form an initial attack fragment and assigned a unique identifier. For example, from intelligence from source A, fragment F001 can be obtained: "The malicious code GoldDigger successfully executed a PowerShell script that modified the system registry." The fragment is then analyzed using pre-defined regular expressions and an entity dictionary to identify specific targets such as IP addresses, domain names, filenames, and organizational structures. For instance, in fragment F001... In the message 03, “…the phishing email successfully tricked employees of Trinity Bank into clicking, leading to GoldDigger's infiltration of the internal network…”, the target entity “Trinity Bank” can be identified. The fragment ID, source, original text, attack behavior, and target entity ID are structured and stored to build an index association between fragments and entities. For example, fragment F001 is parsed as the attack behavior “execution” and the target entity “PowerShell script”, while fragment F003 is associated with the behavior “penetration” and the target “Trinity Bank internal network”. After processing all the intelligence data, 245 attack fragments were identified and associated with 180 independent target entities, forming attack behavior and target entity association data.
[0022] S102: Based on the data on the association between attack behavior and target entity, analyze the pairing relationship between behavior and subject in the fragment, identify and match similar action phrases, and obtain a simplified set of action phrases by merging repeated action expressions and retaining key parts; the processing result is the merged action phrase data, called the merged action phrase data. After initial association, all attack behavior phrases, such as "successfully executed," "attempted to backtrack," and "caused...to penetrate," are extracted from the data. The core verbs "execute," "backtrack," and "penetrate" are then extracted to form a set of core actions to be processed. To merge semantically similar expressions, the system pairs phrases within the set and calculates their semantic similarity. The similarity determination does not rely on existing models but is achieved through a vector space comparison method. This method converts each phrase into a numerical vector based on a cybersecurity lexicon before calculating its score. Similarity merging threshold The threshold is set to 0.85. This value is the optimal result achieved after testing and verification with a large amount of sample data, ensuring a balance between high accuracy and low false alarm rate. During the calculation, if the scores of two phrases are greater than or equal to this threshold, they are considered similar expressions. For example, the similarity score between "establish a connection" and "create a session" is 0.89, which exceeds the threshold of 0.85. Therefore, the system classifies them into one category and selects "establish a connection" as the standard representative, which appears more frequently. All fragment indices of "create a session" are uniformly pointed to "establish a connection". Through this iterative processing, the original phrase set is simplified to 82 standardized action phrases, forming merged action phrase data.
[0023] S103: Further filtering is performed on the merged data of action phrases to extract content with clear semantic direction and relevance, and comparative analysis is conducted to form a corresponding semantic fragment merging set; this set contains the merged semantic fragments, and the result is a semantic fragment merging comparison set.
[0024] To ensure that the selected behavioral descriptions are clearly targeted, a semantic orientation score is introduced. A quantitative assessment is conducted, and the score calculation incorporates the specificity of the action. With goal certainty Two dimensions, specifically calculated as follows The weighting coefficients of 0.6 and 0.4 are derived from the statistical evaluation results of domain experts regarding the importance of actions and targets in attack events, and the specificity of the actions. Based on a predefined tiered list of values, high-specificity actions (such as "encryption") are assigned a value of 1.0, medium-specificity actions are assigned 0.7, low-specificity actions are assigned 0.3, and target certainty is assigned a value of 0.3. The threshold is determined by the entity type: 1.0 for specific entities (e.g., IP addresses), 0.8 for categorical entities (e.g., user accounts), and 0.4 for abstract concepts (e.g., system performance). The system sets the filtering threshold. The threshold is 0.70. Only phrases with a score not lower than this threshold are retained. For example, the action "data return" and its target "IP address 10.0.0.5" have a calculated score of 1.0 and are retained, while the action "impact" and its target "system performance" have a score of only 0.34 and are filtered out. After this round of evaluation, 65 action phrases with clear semantic directions remain. These filtered phrases, along with their associated subject-verb pairing structures and semantic direction identifiers (here, attack behaviors are uniformly marked as negative), are integrated to form a structured comparison set. For example, one record is "Unified action item: Execute; Associated subject-verb pairing structure: ['malicious code', 'execute', 'PowerShell script']; Semantic direction identifier: negative". This is the semantic fragment merged comparison set.
[0025] Please see Figure 3The steps to obtain the set of behavior-consistent expressions are as follows: S201: After obtaining the semantic fragment merging and comparison set, identify the behavioral descriptions pointing to the consistent target, mark the association between each behavior and the target, filter out phrases with repetitive structures or similar expressions, and remove components that cause ambiguity or meaning jumps; the behavioral description data pointing to the consistent target is obtained, which is called the target consistent behavioral description data. After obtaining the semantic fragment merged reference set, the processing flow uses the target entity as the primary key to aggregate all behavioral descriptions pointing to the same target. For example, for "C2 server 192.168.1.100", the two behavioral descriptions associated with it, "implant establishing a connection" and "backdoor program uploading data", are grouped together. After classification, each behavioral description undergoes structural review to identify and break down complex expressions containing multiple independent actions. This process relies on a "behavioral complexity score". The score is determined by subtracting 1 from the total number of standardized actions in the phrase, and a complexity threshold is set. If a description is 0, If the score is greater than 0, it is considered a compound expression and is split. For example, "malicious code downloaded and executed payload A" contains two actions: "download" and "execute". The value is 1, so it is split into two independent descriptions: "Malicious code download payload A" and "Malicious code execution payload A". At the same time, the modifiers in the expression, such as "random attempt", are also filtered out, and the core "script returns user credentials" is retained. After this round of review and cleanup, a list is formed, which lists the cleaned single behavior description corresponding to each target entity. This is the target consistent behavior description data.
[0026] S202: Based on the target consistency behavior description data and combined with the contextual semantic continuity, behavior descriptions with similar semantics are merged, redundant information is removed and the core behavior content is retained; the result after processing is the simplified behavior description data, which is called simplified behavior description data. To merge redundant information that is substantially the same but differs slightly in origin or expression, a contextual analysis will be performed on multiple behavioral descriptions under each target entity name. The core of this analysis lies in calculating the "semantic continuity score" between any two behavioral descriptions. This score is based on time proximity. Relevance to source The weighted composition is calculated as follows: The weighting ratio is set based on the analysis results of historical attack chain data, including time proximity. The source correlation is calculated based on the difference between the timestamps of the two actions within a 3600-second window. The value is assigned based on whether the intelligence sources are the same or belong to the same attack campaign, and a continuity merging threshold is set. The value is 0.80 when two behaviors are described. If the score is not lower than the threshold, merging is performed. For example, for two behaviors under the target "registry key HKCU": "virus writes to startup items" and "malware modifies the registry for persistence", because their timestamps are close and the intelligence sources are related to the same attack activity, the score is calculated as follows: The value is 0.9379, which is greater than 0.80. Therefore, the system merges the two behaviors and selects the core action "write" as the representative, retaining the core content of "virus write startup item". Through this process, each target entity will correspond to one or more simplified and deduplicated core behavior sequences, thereby forming simplified behavior description data.
[0027] S203: Based on the simplified behavior description data, mark the behavior instructions that can be merged, verify them in combination with the semantic continuity of the context, integrate them into a unified set of behavior descriptions, and obtain a set of consistent behavior expressions.
[0028] The core behavioral sequences in the simplified behavioral description data are marked and verified as merging instructions. This step requires context consistency verification for each instruction to be merged. During verification, the system backtracks to the original intelligence fragments that constitute the instruction and calculates the text similarity between the context environments of these fragments. The similarity is calculated using the Jaccard similarity coefficient, which is the number of words in the intersection of two context environments divided by the number of words in the union. The system sets a context verification threshold. The value is 0.65. This value was determined after conducting discrimination experiments on a large number of samples to achieve the optimal segmentation point. It is only possible when the average context similarity of the multiple original segments constituting the merging instruction is not lower than a certain threshold. Only when the merge instruction is valid is it confirmed to be valid. For example, the aforementioned merged "virus writes to startup item" instruction, after tracing back its source fragment, calculates the context's Jaccard similarity coefficient to be 0.72. Since it is greater than 0.65, the merge validity of the instruction is confirmed. After verifying all simplified behavioral descriptions, these verified merge instructions are integrated into a structured set, where each item defines a unified behavioral expression. For example, all operations related to writing to startup items are uniformly described as "writing to persistent registry key". The set of all these unified descriptions constitutes the set of consistent behavioral expressions.
[0029] Please see Figure 4 The steps to obtain the behavior sequence trigger path are as follows: S301: After obtaining the behavioral statements in the behavioral consistency expression set, analyze the connection between the preceding and following structures of each statement, extract the sequence relationship clues, identify the order of actions in multiple segments, and mark the logical connection points; obtain the sequence relationship clue data; After acquiring the set of consistent behavioral expressions, the system traces back to the corresponding original intelligence fragment for each consistent behavioral statement and parses these fragments using dependency parsing. This parsing process breaks down sentences into core components such as subject, verb, object, attributive, adverbial, and complement, and identifies the grammatical relationships connecting these components. The system focuses on extracting three types of sequence relationship clues: the first type is time adverbs or conjunctions, such as "after"; the second type is logical conjunctions, such as "therefore" and "following up"; and the third type is verb tense or state changes, such as changing from perfect tense to progressive tense. The system presets a basic weight value for each type of clue. This weight value is based on the statistical analysis results of a CII (Cyber Threat Intelligence) report database containing 5,000 labeled attack sequences. The analysis shows that explicit time conjunctions have an accuracy rate of 92% in indicating sequence, logical conjunctions 85%, and verb state changes 71%. Based on this, the basic weights of the three are assigned. The values are set to 0.9, 0.85, and 0.7 respectively. The system adjusts the weights based on the position of the clues in the syntactic structure. If the clues connect two main verb phrases, the weight adjustment coefficients are adjusted accordingly. The coefficient is 1.0; if the connection is between a main clause and a subordinate clause, the coefficient is 0.8, indicating the strength of the clue. pass and The product is calculated; for example, for the original fragment "After the malware downloaded the payload, it immediately executed the script," dependency parsing identifies "after" as a time-related word, connecting the two actions of "downloading" and "executing." It is 0.9. The value is 1.0, therefore the clue strength is... With a strength of 0.9, the system records this relationship as a directed connection: (download payload, execute script). The system sets a logical join point filtering threshold of 0.9. The threshold is 0.75. Only clues with a strength greater than or equal to this threshold are accepted. This threshold was determined by ROC curve analysis on the aforementioned CII report database. At a value of 0.75, the difference between the true positive rate and the false positive rate reaches its maximum, resulting in the best identification effect. By traversing all behavioral statements, the system filters out all directed connection pairs with sufficient strength and defines them as logical connection points, forming sequential relationship clue data.
[0030] S302: Based on the sequential relationship clue data, sort multiple action segments, clarify the logical connection position of each action segment, combine the context semantics to determine the start and end positions of the action in the behavior chain, determine the progression order of the action sequence, and generate action sorting judgment data; Based on sequential relational clue data—a set of directed connection pairs of the form (action A, action B)—the system begins by globally sorting all action fragments. The system initializes all action fragments as independent nodes and establishes directed edges between nodes according to the clue data, forming a preliminary behavioral relational network structure. The system calculates the in-degree and out-degree of each node in this network. The in-degree is the number of edges pointing to that node, representing how many preconditions it has; the out-degree is the number of edges pointing outwards from that node, representing how many post-results it can trigger. The system determines the start and end positions of actions in the behavioral chain according to the following rules: if a node has an in-degree of 0 and an out-degree greater than 0, then that node is determined to be the starting position of the behavioral chain because it has no preconditions but is the beginning of other actions. For example, in the analysis, the action fragment "downloading malicious payloads" is found to be related to... In all logical connection points, if a node appears only as a precondition (out-degree of 2) and not as a post-action result (in-degree of 0), the system marks it as one of the starting actions. Conversely, if a node has an out-degree of 0 and an in-degree greater than 0, it is determined to be the termination position of the action chain because it has a precondition but no longer triggers any subsequent known actions. For example, the action "returning stolen data" has an in-degree of 1 (precondition is "establishing a C2 connection"), but its out-degree is 0 in all threads, so it is marked as a termination action. For nodes with both in-degree and out-degree not equal to 0, they are determined to be intermediate links in the chain. By analyzing and calculating the in-degree and out-degree of all 65 action segments, the system identified 4 starting actions, 3 termination actions, and 58 intermediate actions, and based on this, determined the progression order of multiple action sequences and generated action sorting judgment data.
[0031] S303: Based on the action sorting, determine the data, construct a combined path according to the order of behavior progression, integrate the successive relationships of action segments, and generate a behavior sequence trigger path.
[0032] Based on the action sequencing data, which identifies the start and end positions and progression order of each action segment, the system constructs a combined path. This process begins with each action node marked as the start position and proceeds through a depth-first traversal along the directed connections established in the sequential relationship clue data, sequentially linking subsequent action segments until a node marked as the end position is reached, thus forming a complete behavioral chain. For example, starting with the initial action "download malicious payload," the system finds a logical connection pointing to "execute PowerShell script," and then moves to the next node. Starting from the "execute PowerShell script" node, the system finds a connection to "establish C2 connection," and then from "establish C2 connection" to "return stolen data." Since "return stolen data" is marked as the end position, the construction of this path is complete. The system integrates this sequence, which consists of four action segments linked together sequentially, and assigns a confidence level to each connection. This confidence level is the strength of all the original clues constituting the connection. The system summarizes all paths that are constructed from different starting points and have a length of two or more action nodes. Each path is stored in a structured sequence, clearly showing the complete process from one or more initial attack behaviors to the target. These structured sequences that integrate the successive relationships of all action segments are the generated behavior sequence trigger paths.
[0033] Please see Figure 5 The steps for obtaining the target hierarchy content are as follows: S401: Obtain each behavior node in the behavior sequence triggering path, identify the target entity corresponding to each node, and determine the first and second occurrence positions of the target entity. Use these positions to determine the distribution level of the target in the path; obtain the target entity distribution level data. The first step in processing the behavior sequence triggering path is to traverse all behavior nodes, extract the ordered sequence of target entities, and record the first and second occurrence positions of each unique entity. The distribution hierarchy is defined by calculating the node span between the first and second occurrence positions of an entity. The criterion here is a predefined proportional threshold. Its value is 0.50. When the node span of an entity exceeds 50% of the total path length, it is defined as a global level; otherwise, it is a local level. For example, in a 10-node path, the activity of the target "C2 server X" extends from node 5 to node 10, with a span of 6, accounting for 60%, which exceeds the threshold, so it is classified as a global level; while "PowerShell.exe" appears only once, with a span of 1, accounting for 10%, so it is classified as a local level. After this classification of all target entities, the target entity distribution hierarchy data is formed.
[0034] S402: Based on the target entity distribution hierarchy data, identify key objects that appear repeatedly or are in transit positions in the path, and combine the positional relationship of these objects in the path to perform hierarchical division and clarify their belonging in the structure; the result of this step is the key object hierarchy division data. Once the hierarchical data is available, key objects are identified by statistically analyzing the frequency of each entity's appearance in the behavioral sequence triggering path. A frequency threshold is set here. A value of 2 indicates that any entity appearing at least twice is considered a key object. For example, if "C2 server X" appears three times and "terminal A" appears twice in the path, both are identified as key objects. Hierarchical classification is determined by analyzing the connections between these key objects: if a key object becomes the common target of multiple entities from different sources, then that key object is defined as a parent-level entity, and the multiple source entities pointing to it become child-level entities. For example, if "registry key B," "terminal A," and "database D" all point to "C2 server X," then "C2 server X" is established as the parent level, and the other three are child levels. The output of this analysis process is the key object hierarchy classification data.
[0035] S403: Based on the key object hierarchy data, organize the hierarchical relationship of the target entity in the path and generate a complete target hierarchy division result; the result is the target hierarchy division content.
[0036] Based on the clearly defined parent-child relationships in the key object hierarchy data, a complete hierarchical structure is constructed. This process uses all "parent-level entities" as top-level nodes, such as "C2 server X" mentioned earlier, and then mounts all identified "child-level entities" (such as "registry key B", "terminal A", etc.) to their corresponding parent nodes. For entities in the path that have logical connections but whose hierarchical roles are not yet clearly defined, the system determines their affiliation based on their contextual connections. For example, if the behavior of entity "file C" directly originates from "terminal A", then "file C" is classified as a child node of "terminal A". In this way, all target entities are incorporated into a unified hierarchy, and a structured list that clearly shows the primary and secondary relationships of each asset is output. This is the result of the target hierarchy division.
[0037] Please see Figure 6 The steps for obtaining priority results in threat response are as follows: S501: Obtain the list of objects in the target hierarchy, check whether there are target items located at the end of the path or frequently associated across paths, identify and mark the location and association of these target items; the result of this step is target item aggregation and association data. Retrieve the object list from the target hierarchy. This list records the distribution of each target entity in the action sequence trigger path. Specific data includes: target entity "Internal File Server - FS02" associated with path P1, location index 11, total number of nodes 18; target entity "Database Master Server - DB01" associated with path P2, location index 18, total number of nodes 20; target entity "Gateway Firewall - FW01" associated with paths P1, P3, and P4, with their locations and node data as follows; target entity "Operation and Maintenance Management Terminal - OMT0"... 5” Associated paths P2 and P3, whose data is as follows: For each target item in the list, determine whether it is located at the end of each associated path. The determination method is to calculate the ratio of its position index value to the total number of nodes in the path and compare it with the preset end position threshold of 0.75. This threshold is set based on the statistical analysis of 1000 historical attack cases, in which 91% of the cases the key target was located in the last 25% of the path. Therefore, 0.75, which is the 95% confidence interval, is taken as the threshold. This value showed an accuracy of 94% in the backtesting of 500 samples and was lower than 3. The false alarm rate of % is sufficiently reasonable. For example, for "Database Master Server - DB01", its position ratio in path P2 is 18 divided by 20, which equals 0.90. This value is greater than 0.75, so an end marker is added to it. However, the position ratio of "Internal File Server - FS02" in path P1 is 11 divided by 18, which is approximately 0.61, less than 0.75, so no marker is added. After determining the end position, the cross-path association of each target item is further identified. This is done by counting the total number of paths associated with each target item and comparing it with the preset cross-path association threshold of 3. The threshold is based on research on known APT attack activities, in which 78% of attack campaigns utilize the same critical asset as the intersection of at least 3 paths. This setting was verified in a simulation environment containing 50 APT samples, with a recall rate of 88% and a precision rate of 92%. Taking "Gateway Firewall-FW01" as an example, its number of associated paths is 3, reaching the threshold, so an association tag is added. However, "Database Master Server-DB01" is only associated with 1 path, so no tag is added. The end tags and association tags of each target item are summarized to form target item aggregation and association data.
[0038] S502: Based on the target item aggregation and association data, sort the target objects that are located at the end position and have cross characteristics, adjust their priority in the response process, and clarify their response order boundaries; obtain the target object sorting adjustment data; Based on the aggregation and association data of target items, target items that simultaneously possess both end-point and association markers are selected for priority ranking. In this embodiment, only "Gateway Firewall-FW01" simultaneously meets the following conditions: its position ratios on paths P1, P3, and P4 are 0.833, 0.90, and 0.87, respectively, all greater than the threshold of 0.75, and the number of associated paths is equal to 3, satisfying the cross-path association condition. Therefore, the threat response priority score is calculated only for this target item. The score is obtained by weighted summation of the number of associated paths and the average position ratio within each path, where the association weight and position weight are... The weights were set to 0.65 and 0.35 respectively. This ratio was determined based on a survey of 50 senior cybersecurity analysts. The survey showed that experts believed the risk contribution of horizontal correlation breadth was 1.8 times that of single path location depth. Based on this, normalization was performed to obtain the above weights. This weighting ratio achieved a 96% consistency rate with the experts' manual ranking in 100 simulated attack simulations. For "Gateway Firewall-FW01," its threat response priority score was calculated as follows: the product of the number of correlation paths (3) and the weight (0.65) is 1.95, and its average position ratio among the three paths is (0.833). (+0.900 + 0.867) divided by 3 is approximately 0.867. The product of this value and the weight 0.35 is 0.30345. The sum of the two is 2.25345. After calculating the scores of all eligible target items, they are sorted from highest to lowest score, and response levels are divided according to preset score boundaries. For example, a score greater than 2.0 is defined as "Level 1 Emergency Response", and a score between 1.5 and 2.0 is "Level 2 Priority Response". Since the score of "Gateway Firewall-FW01" is 2.25345, which is higher than 2.0, it is classified into the "Level 1 Emergency Response" category. This sorting and level division results together constitute the target object sorting adjustment data.
[0039] S503: Adjust the data according to the target object sorting, organize the order of appearance of each target in the response process, and generate the threat response priority result by combining the target sorting information.
[0040] Based on the target object sorting and data adjustment, a response priority table is constructed. The target item with the highest ranking in the previous sub-step is placed at the top of the priority table, i.e., "Gateway Firewall-FW01" is priority 1, with a response level of "Level 1 Emergency Response". For the remaining target items in the list that are not double-marked, their subsequent priority is determined according to a set of secondary sorting rules. The dimensions of this rule are: structural hierarchy, presence or absence of end marker, and path position ratio. During execution, the remaining "Internal File Server-FS02", "Database Master Server-DB01", and "Operations Management Terminal-OMT05" are first judged by their hierarchy. "Database Master Server-DB01" and "Internal File Server-FS02" belong to the "Data Layer", while "Operations Management Terminal-OMT05" belongs to the "Access Layer". Therefore, the former two have priority over the latter. Among the two targets belonging to the "Data Layer", "Database Master Server-DB01" has an end marker (path P) The position ratio of "FW01" is 0.90, which is higher than "FW02" (path P1 position ratio is 0.61) which has no end marker. By integrating the primary and secondary sorting results, a complete response priority list is formed. This list clarifies the order of appearance and response level of each target, as follows: Rank 1 is "Gateway Firewall - FW01", response level "Level 1 Emergency Response", based on its priority score of 2.25345; Rank 2 is "Database Master Server - DB01", response level "Level 2 Priority Response", based on its "Data Layer" structural level and the presence of an end marker; Rank 3 is "Internal File Server - FS02", response level "Level 3 Normal Response", based on its "Data Layer" structural level; Rank 4 is "Operations and Maintenance Management Terminal - OMT05", response level "Level 4 Attention Response", based on its "Access Layer" structural level. This ordered list is the generated threat response priority result.
[0041] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for decision optimization based on multi-source threat intelligence data fusion, characterized in that, Includes the following steps: S1: Extract statements involving targets, behaviors, and contexts from multi-source intelligence, construct a semantic guidance index, merge semantically similar expressions, remove ambiguous descriptions or duplicate content, and generate a semantic fragment merge comparison set; S2: Based on the semantic fragment merging and comparison set, focus on the behavioral content pointing to the same object, filter logical jump items, retain the main phrases with clear expression structure, restore the unified semantics in combination with the context, and generate a set of consistent behavioral expressions; S3: Based on the event progression clues located in the set of consistent behavioral expressions, mark the starting point and node connection relationship of the action, connect the behavioral order according to the context identifier, reconstruct the behavioral chain progression path, and generate the behavioral order triggering path; S4: Based on the behavioral sequence triggering path, extract the first and last associated positions of each target in the path, divide the key nodes into upper and lower levels according to the action succession relationship, unify and organize the hierarchical structure, and generate the target hierarchical division content. S5: Combine the target hierarchy classification to analyze the target concentration trend and path convergence, check the location distribution of high-frequency response targets, reorder each disposal object according to the endpoint path weight, and generate threat response priority results.
2. The method for multi-source threat intelligence data fusion and decision optimization according to claim 1, characterized in that: The semantic fragment merging and comparison set includes a unified item for action phrases, a set of subject-verb pairing structures, and semantic direction indicator markers; The set of behaviorally consistent expressions includes target-consistent instruction items, expression convergent phrase groups, and semantic cleanup key components; The behavioral sequence triggering path includes action ordering relationships, logical connection point annotations, and behavioral chain progression sequences. The target hierarchy classification includes the first occurrence location record, the belonging location identifier, and the structural hierarchy mapping label; The threat response priority results include a response exit order table, cross-path target ranking items, and a list of end-point objects.
3. The method for multi-source threat intelligence data fusion and decision optimization according to claim 1, characterized in that: The steps for obtaining the semantic fragment merging comparison set are as follows: S101: After acquiring multi-source intelligence data, relevant segments involving attack behaviors are selected, the attack behaviors in each segment are labeled, target entities are identified and associated, and an index relationship between each segment and the corresponding target entity is established; the result is the preliminary association data between attack behaviors and target entities, forming the association data between attack behaviors and target entities. S102: Based on the data associated with the attack behavior and the target entity, analyze the pairing relationship between the behavior and the subject in the fragment, identify and match similar action phrases, and obtain a simplified set of action phrases by merging repeated action expressions and retaining key parts; the processing result is the merged action phrase data, which is called the merged action phrase data. S103: Further filter the merged data of the action phrases, extract content with clear semantic direction and relevance, and perform comparative analysis to form a corresponding semantic fragment merging set; the semantic fragment merging set includes the merged semantic fragments, and the result is a semantic fragment merging comparison set.
4. The method for multi-source threat intelligence data fusion and decision optimization according to claim 1, characterized in that: The steps for obtaining the set of consistent behavior expressions are as follows: S201: After obtaining the semantic fragment merging and comparison set, identify the behavioral descriptions that point to the same target, mark the association between each behavior and the target, filter out phrases with repetitive structures or similar expressions, and remove components that cause ambiguity or meaning jumps. Obtain target consistency behavior description data; S202: Based on the target consistency behavior description data, and in combination with the context semantic continuity, merge behavior descriptions with similar semantics, remove redundant information and retain the core behavior content to obtain simplified behavior description data; S203: Based on the simplified behavior description data, mark the behavior instructions that can be merged, and verify them in combination with the semantic continuity of the context to obtain a set of consistent behavior expressions.
5. The method for multi-source threat intelligence data fusion and decision optimization according to claim 1, characterized in that: The steps for obtaining the behavior sequence trigger path are as follows: S301: After obtaining the behavioral statements in the behavioral consistency expression set, analyze the connection between the preceding and following structures of each statement, extract the sequence relationship clues, identify the order of actions in multiple segments, and mark the logical connection points; Obtain sequence relationship clue data; S302: Based on the sequence relationship clue data, sort multiple action segments, clarify the logical connection position of each action segment, combine the context semantics to determine the start and end positions of the action in the behavior chain, determine the progression order of the action sequence, and generate action sorting judgment data; S303: Based on the action sorting judgment data, construct a combined path according to the action progression order, integrate the successive relationship of action segments, and generate a behavior sequence triggering path.
6. The method for multi-source threat intelligence data fusion and decision optimization according to claim 5, characterized in that: The analysis examines the structural connections between each statement, extracts clues about sequence relationships, and identifies the order of actions in multiple segments. The process of determining the start and end positions of an action in a behavior chain by combining contextual semantics is as follows: the logical connection points of the action segment are analyzed. If a logical connection point of an action segment only has a subsequent result but no preceding condition, it is determined as the start position of the behavior chain. If a logical connection point of an action segment only has a preceding condition but no subsequent result, it is determined as the end position of the behavior chain.
7. The method for multi-source threat intelligence data fusion and decision optimization according to claim 1, characterized in that: The steps for obtaining the target hierarchical content are as follows: S401: Obtain each behavior node in the behavior sequence triggering path, identify the target entity corresponding to each node, determine the first and first appearance position of the target entity, and determine the distribution level of the target in the path by the position; Obtain hierarchical data on the distribution of the target entity; S402: Based on the target entity distribution hierarchy data, identify key objects that appear repeatedly or are in transit positions in the path, and combine the positional relationship of these objects in the path to perform hierarchical division and clarify their belonging in the structure; the result is key object hierarchy division data. S403: Based on the key object hierarchical classification data, organize the hierarchical relationship of the target entity in the path and generate a complete target hierarchical classification result; The result is the content divided into target levels.
8. The method for multi-source threat intelligence data fusion and decision optimization according to claim 7, characterized in that: The process of determining the distribution hierarchy of the target in the path through these locations specifically involves: calculating the first and second relationships between the target entity and... If the node span between the locations where the target entity appears exceeds a preset proportion threshold of the total number of nodes in the behavior sequence triggering path, then the distribution level of the target entity is determined as the global level; otherwise, it is determined as the local level. The process of hierarchically classifying key objects that repeatedly appear or are in transit positions in the identification path, combined with the positional relationship of these objects in the path, is as follows: count the frequency of each target entity in the behavior sequence triggering path, and determine the target entities whose frequency of occurrence is greater than a preset frequency threshold as the key objects; Analyze the connection relationships between the key object and other target entities. If a key object is pointed to by multiple target entities from different sources, then the key object is defined as a parent entity, and the multiple target entities from different sources pointing to the key object are defined as child entities.
9. The method for multi-source threat intelligence data fusion and decision optimization according to claim 1, characterized in that: The steps for obtaining priority results in threat response are as follows: S501: Obtain the list of objects in the target hierarchy, check whether there are target items located at the end of the path or frequently associated across paths, identify and mark the location and association of these target items; Obtain aggregated and correlated data for the target items; S502: Based on the target item aggregation and association data, sort the target objects that are located at the end position and have cross characteristics, adjust their priority in the response process, and clarify their response order boundaries; obtain the target object sorting adjustment data; S503: Adjust the data according to the target object sorting, organize the order of appearance of each target in the response process, and generate the threat response priority result by combining the target sorting information.
10. The method for multi-source threat intelligence data fusion and decision optimization according to claim 9, characterized in that: The process of identifying and marking the location and association of these target items specifically involves: calculating the ratio of the position index value of a target item in the behavior sequence triggering path to the total number of nodes in the behavior sequence triggering path; when the ratio is greater than a preset end position threshold, the target item is marked as an end. The number of behavioral sequence triggering paths associated with each target item is counted. When the number exceeds a preset cross-path association threshold, the target item is marked as associated. The process of sorting target objects located at the end position and having cross features, and adjusting their priority in the response flow, specifically involves: calculating a threat response priority score for each target item that simultaneously has the end marker and the associated marker. The threat response priority score is a weighted sum based on the number of behavioral sequence triggering paths associated with the target item and the position index value of the target item within the path. Based on the threat response priority score, all target items that simultaneously possess both the terminal tag and the association tag are prioritized from high to low.