Threat Intelligence Entity Relationship Extraction Method for Unstructured Data

Through the threat intelligence entity relationship extraction method based on the STIX standard and the BERT-BiLSTM-CRF model, the problem of lack of methodology and sharing difficulties in threat intelligence is solved, and efficient identification and relationship extraction of unstructured threat intelligence is achieved, which improves the diversity and accuracy of the data set.

CN116450844BActive Publication Date: 2025-07-11JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310323400.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-29
Publication Date
2025-07-11
Estimated Expiration
2043-03-29

AI Technical Summary

Technical Problem

The lack of methodology of threat intelligence in the existing technology, difficulty in sharing unstructured threat intelligence, and difficulty in identifying and attribution of entities, making it difficult for enterprises to analyze the correlation between a large number of fallout indicators and specific threat environments.

Method used

The STIX threat intelligence standard is used to define entity types and relationships, build a vocabulary knowledge base in the threat intelligence field, and perform entity extraction and relationship extraction through the BERT-BiLSTM-CRF model, and build a threat intelligence knowledge graph with the Neo4j database.

Benefits of technology

It improves the accuracy and semantic understanding of the extraction of unstructured threat intelligence entity relationships, enhances the diversity of data sets, and realizes efficient identification and fusion of entities and relationships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116450844B_ABST
    Figure CN116450844B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of threat intelligence named entity recognition, and specifically relates to a method for extracting threat intelligence entity relationships for unstructured data, an approach for threat intelligence named entity recognition based on data augmentation and BERT, and a method for extracting threat intelligence entity relationships that integrates multi-source entity information to accurately extract the relationships of network threat intelligence entities in unstructured text. The present invention increases the number of entities of vulnerabilities, domain names, and IPs, enhances the sample diversity of attack organizations and malware entities, searches for sentences containing entities of the type to be augmented as template sentences, fills in entities of the same type in the knowledge base into the template sentences to generate new sentences containing specific types of entities, and adds the newly generated sentences to the training set to achieve data augmentation, thereby improving semantic accuracy. The present invention integrates entity semantic information and entity boundary information, and adds entity type information to the BERT sentence vector to help the model better perform relationship classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of threat intelligence named entity recognition, and particularly relates to a method for extracting threat intelligence entity relationships for unstructured data. Background Art

[0002] In recent years, with the rapid development of computer network and communication technologies, the connection between the Internet and people's daily lives has become increasingly close. The Internet has continuously facilitated the digital transformation of small and medium-sized enterprises in China, promoted the development of the digital economy, and enabled the digital dividend to benefit the general public. The Internet has been widely applied in fields such as intelligent manufacturing, intelligent transportation, and e-government, and its influence on all industries has been increasing year by year.

[0003] In addition, as a common and difficult-to-prevent attack method, distributed denial of service (DDoS) attacks have shown an obvious upward trend in frequency in recent years. In the current network environment, although traditional security technologies such as firewalls, antivirus software, and intrusion detection systems have been widely applied and achieved certain results, they also show some deficiencies when facing "zero-day" vulnerability attacks and advanced persistent threat (APT) attacks. People have started to turn their attention to threat intelligence to seek new ideas for solving network security problems. Threat intelligence is a kind of evidence-based knowledge, including context, mechanism, indication, meaning, and actionable advice. This knowledge is related to existing or emerging threats or hazards faced by assets and can be used to provide information support for asset-related entities to make decisions on threat or hazard response or handling.

[0004] In recent years, the threat intelligence-driven network security active defense model has attracted much attention from the academic and industrial circles as a new development direction. Many security organizations and security vendors have started to research and apply threat intelligence and provide related services accordingly. By actively collecting, refining, and analyzing data, threat intelligence can integrate information on network security incidents that have occurred or are occurring, detect network threats early, and then take measures to achieve the role of prior prevention or reduction of losses and hazards.

[0005] However, the development of threat intelligence still faces many challenges. Kris Ossthoek, a senior cyber threat intelligence analyst at the Dutch government, pointed out in the Journal of International Intelligence and CounterIntelligence the following difficulties in the development and utilization of threat intelligence: First, threat intelligence lacks a methodology. Most threat intelligence analyses are input-driven by alerts and log data rather than pre-determined methods or hypotheses. The lack of methodology makes it difficult for enterprises to analyze the relevance of the large number of Indicators of Compromise (IOCs) generated every day to a specific threat environment. Second, the sharing of threat intelligence is only verbal. For reasons such as trust and interests, a small part of structured and semi-structured threat intelligence is only shared within enterprises and organizations. In contrast, most threat intelligence is still shared on the Internet in an unstructured manner. Finally, there is no common naming convention in the field of threat intelligence, which brings great difficulties to entity recognition and attribution. For marketing purposes, security vendors and threat intelligence providers are keen to give various names to each threat organization. This leads to an increase in the number of entities to be identified and an increase in the difficulty of identification. Therefore, it is of practical significance to develop and utilize unstructured threat intelligence using computer-related technologies. Summary of the Invention

[0006] To address the deficiencies in the prior art, the present invention proposes a method for extracting threat intelligence entity relationships for unstructured data, which accurately extracts the network threat intelligence entity relationships in unstructured text based on a threat intelligence named entity recognition method based on data augmentation and BERT and a threat intelligence entity relationship extraction method that integrates multiple entity information.

[0007] To achieve the above-mentioned invention purposes, the technical solutions adopted by the present invention are as follows:

[0008] A method for extracting threat intelligence entity relationships for unstructured data, including the following three parts:

[0009] 1) Threat intelligence entity extraction, including the following steps:

[0010] S1: Define threat entity types and relationships between threat intelligence entities based on the STIX threat intelligence standard;

[0011] S2: Construct an NER original annotation data set and a threat intelligence domain vocabulary knowledge base;

[0012] S3: Find sentences containing entities of the type to be augmented in the original data set as template sentences, fill in entities of the same type in the threat intelligence domain vocabulary knowledge base into the template sentences to generate new sentences containing specific type entities, and add the newly generated sentences to the NER original data set;

[0013] S4: Fill the template sentence: Convert the template sentence into the BIO annotation mode, and use the annotation result and the vocabulary knowledge base in the threat intelligence field as inputs. Generate and output the sentence after template filling through the template sentence filling algorithm. The output sentences form the enhanced dataset;

[0014] S5: Use the BERT-BiLSTM-CRF model to extract entities from the sentences in the enhanced dataset output in step S4. Among them, the BERT layer is responsible for dynamically generating word vectors for each input word according to its context. The generated word vector sequence will be used as the input of the BiLSTM layer; the BiLSTM layer is responsible for encoding the temporal relationship of the input sequence and outputting the hidden state sequence; the CRF layer decodes the hidden state sequence to obtain the label sequence corresponding to the sentence, and the obtained label sequence is the entity type;

[0015] 2) After entity extraction, perform threat intelligence relationship extraction, including the following steps:

[0016] P1: Extract the sentences containing entity relationships from the original annotated dataset;

[0017] P2: Extract threat intelligence entity relationships: Use the threat intelligence text strings in the original annotated dataset, the entity list obtained in the entity extraction step, and the relationship list defined in S1 as inputs for sentence extraction. Then output the start index, end index, head entity information, and tail entity information of each sentence, and use the output information for threat intelligence entity relationship extraction;

[0018] 3) After all entity extraction and relationship extraction are completed, input the extracted entity and relationship information into the graph database to construct the threat intelligence knowledge graph. The graph database uses the Neo4j database.

[0019] Furthermore, the entity types in the above step S1 include 13 categories, namely Threat Actor, Campaign, Malware, Technique, Tool, Identity, Location, Industry, Vulnerability, Course of Action, URL, Domain, IP; there are 7 types of relationships between threat intelligence entities, namely use, attack, originate from, similar, same, own, respond to.

[0020] Furthermore, the original dataset in the above step S2 is derived from unstructured APT reports. The APT report text is manually annotated to obtain the original annotated dataset; the original annotated dataset is enhanced through the template sentence filling algorithm to obtain the enhanced dataset.

[0021] Further, the template sentence filling algorithm in step S4 specifically includes the following steps:

[0022] S4.1 Convert the sentences in the training set into BIO annotation mode;

[0023] S4.2 Use the BIO annotation results of the template sentences and the threat intelligence domain vocabulary knowledge base as inputs. For each line in the BIO annotation results of the template sentences, obtain its words and labels respectively;

[0024] S4.3 If the label is O, concatenate the word and the label and store them in a list; if the label is not O, obtain a domain vocabulary from the knowledge base and determine how many words the domain vocabulary consists of;

[0025] S4.4 If the domain vocabulary consists of one word, concatenate the domain vocabulary and the corresponding label and store them in a list; if the domain vocabulary consists of multiple words, concatenate the first word of the domain vocabulary with the B-label and store it in a list, and concatenate the second and subsequent words of the domain vocabulary with the I-label and store them in a list; finally, return a list of sentences generated by filling the template and output the list.

[0026] Further, step S5 specifically includes the following content:

[0027] S5.1 Input the sentences generated in S4 into BERT, and BERT dynamically generates word vectors for each input word according to its context, and the word vectors are used to represent the semantic information of the words;

[0028] S5.2 Input the word vectors generated in S5.1 into a bidirectional LSTM encoder, and the LSTM encoder encodes the temporal relationship of the input sequence from two directions (forward and backward) and outputs a hidden state sequence;

[0029] S5.3 Use a CRF conditional random field to decode the hidden state sequence output in S5.2 and output the label sequence corresponding to the sentence. For the context encoding module of the output sequence R = {r1, r2,..., rn}, the calculation method of the matching score given the input and output is expressed as:

[0030]

[0031] where P represents the feature matrix of the BiLSTM layer, A represents the state transition score matrix of the CRF layer, represents the conversion score from the y i label to the y i+1 label;

[0032] S5.4 Performs softmax on all possible tag sequences of the input sequence R to obtain the predicted tag y. The obtained predicted tag is the corresponding entity type, and its definition is expressed as:

[0033]

[0034] S5.5 In model training, the maximum log-likelihood estimation is used to obtain the loss function. The specific process is expressed as:

[0035]

[0036] Finally, the predicted tag obtained in S5.4 is mapped to the entity type and output.

[0037] Furthermore, the three gate structures of the computational unit of the above LSTM encoder: input gate, forget gate, and output gate;

[0038] The specific calculation process is as follows:

[0039] f t = σ(W f * [c t-1 , h t-1 , x t + b f )

[0040] i t = σ(W i * [c t-1 , h t-1 , x t + b i )

[0041] o t = σ(W o * [c t-1 , h t-1 , x t + b o )

[0042]

[0043] Among them, (f t , i t , o t , C t ) represent the forget gate, input gate, output gate, and cell state respectively. (W f , W i , W o , W c ) and (b f , b i , b o , b cdenote the weight matrices and bias vectors of the forget gate, input gate, output gate, and memory unit respectively; x t and h t represent the input vector and the hidden layer vector at time t; σ and tanh are activation functions.

[0044] Furthermore, the above step P1 includes the following specific contents:

[0045] P1.1 Input the full text string, entity list, and relationship list;

[0046] P1.2 For each record in the relationship list, obtain the two entities involved in the relationship from the entity list;

[0047] P1.3 Set the left pointer to the start index of the entity that comes first among the two entities, and set the right pointer to the end index of the entity that comes later among the two entities;

[0048] P1.4 Enter a loop until breaking out using the keyword break. If the left pointer is greater than zero and the left pointer points to a full stop, determine whether the part from the left pointer to the end of the string contains the characteristics of an English sentence ending. If it does, the left pointer index is the starting index of the target sentence; similarly, the ending index of the target sentence can be obtained;

[0049] P1.5 Return and output the starting index, ending index of the target sentence, head entity information, and tail entity information.

[0050] Furthermore, the above step P2 includes the following specific contents:

[0051] P2.1 Integrate the entity semantic information and entity boundary information output by P1.5, so that the boundary information and type information of the entity are reflected in the tags on both sides of the entity;

[0052] P2.2 Given a sentence s containing entities e1 and e2, input it into BERT and output a vector H;

[0053] P2.3 Calculate the average of the word vectors of each word constituting the entity to obtain the representation vector of the entity. Send the representation vectors of the two entities into the tanh activation function and a fully connected layer, and the outputs are denoted as H1 ′ and H2 ′ , and the calculation process is expressed as:

[0054]

[0055] where the vector H i to H j represent the word vectors of entity e1, and the vector H x to H y represent the word vectors of entity e2,

[0056] b1 = b2 is the deviation vector, and d is the size of the BERT hidden state vector;

[0057] P2.4 Sends the sentence vector generated by CLS into the tanh activation function and a fully connected layer, which is expressed as:

[0058] H0 ′ = W0[tanh(H0)] + b0

[0059] Among them, b0 is the deviation vector, and d is the size of the BERT hidden state vector;

[0060] The CLS represents classification and is a token used by the BERT model for text classification tasks; it obtains the sentence-level information representation through the self-attention mechanism, and its corresponding vector is the sentence vector containing the sentence semantic information;

[0061] P2.5 Sends H0 ′ , H1 ′ and H2 ′ After concatenation, it is sent into a fully connected layer and a SoftMax layer to obtain the output vector P. This process is expressed as:

[0062] H ″ = W3[concat(H0 ′ , H1 ′ , H2 ′ )] + b3

[0063] p = softmax(H ″ )

[0064] Among them, L is the number of relationship categories, b3 is the deviation vector, and the output vector P ∈ R L ; R L is the relationship set;

[0065] P2.6 Selects the relationship corresponding to the item with the largest value in P as the output relationship.

[0066] Advantages of the present invention:

[0067] 1. Compared with the existing data augmentation methods, this method increases the entity quantity of three types of entities, namely vulnerabilities, domain names, and IPs, increases the sample diversity of two types of entities, namely attack organizations and malware, searches for sentences containing entities of the type to be augmented as template sentences, fills in entities of the same type in the knowledge base into the template sentences to generate new sentences containing specific type entities, and adds the newly generated sentences to the training set to achieve data augmentation, thereby improving semantic accuracy.

[0068] 2. In the threat intelligence entity relationship extraction task, this method adds entity type information to the BERT sentence vector on the basis of fusing entity semantic information and entity boundary information to help the model better perform relationship classification. Brief Description of the Drawings

[0069] Figure 1 This is the full flow chart of the threat intelligence entity relationship extraction of the present invention;

[0070] Figure 2 This is the flow chart of the threat intelligence named entity recognition of the present invention;

[0071] Figure 3 This is the model diagram of the threat intelligence named entity recognition of the present invention;

[0072] Figure 4 This is the ontology diagram of the threat intelligence of the present invention;

[0073] Figure 5 This is the flow chart of the threat intelligence relationship extraction of the present invention;

[0074] Figure 6 This is the model diagram of the threat intelligence relationship extraction of the present invention. Detailed Embodiment

[0075] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0076] As Figure 1 shown, the present invention is a threat intelligence entity relationship extraction method for unstructured data, including three parts:

[0077] 1) Threat intelligence entity recognition. The overall threat intelligence entity recognition process is as Figure 2 shown, and the threat intelligence recognition model is as Figure 3 shown, including the following steps:

[0078] S1: Define threat entity types and relationships between threat intelligence entities based on the STIX threat intelligence standard. The defined entities and relationships are as Figure 4 shown;

[0079] S2: Construct an NER original annotation dataset and a threat intelligence domain vocabulary knowledge base;

[0080] S3: Find sentences containing entities of the type to be enhanced in the original annotation dataset as template sentences, fill in entities of the same type in the threat intelligence domain vocabulary knowledge base into the template sentences to generate new sentences containing specific type entities, and add the newly generated sentences to the NER original annotation dataset;

[0081] As a preferred embodiment of the present invention, although the above template sentences can be reused, it is still necessary to control the number of times each template sentence appears in the augmented data to prevent overfitting. The number of times N that each template sentence appears in the augmented data is expressed as:

[0082]

[0083] S4: Fill the template sentence: Convert the template sentence into the BIO annotation mode, and use the annotation result and the vocabulary knowledge base in the threat intelligence field as inputs. Generate and output the sentence after template filling through the template sentence filling algorithm. The output sentences constitute the augmented data set;

[0084] S5: Use the BERT - BiLSTM - CRF model to perform entity extraction on the sentences in the augmented data set output in step S4. Among them, the BERT layer is responsible for dynamically generating word vectors for each input word according to its context, and the generated word vector sequence will be used as the input of the BiLSTM layer; the BiLSTM layer is responsible for encoding the temporal relationship of the input sequence and outputting the hidden state sequence; the CRF layer decodes the hidden state sequence to obtain the label sequence corresponding to the sentence, and the obtained label sequence is the entity type;

[0085] 2) After entity recognition, perform threat intelligence relationship extraction. The overall threat intelligence relationship extraction process is as Figure 5 shown, and the threat intelligence extraction model is as Figure 6 shown, including the following steps:

[0086] P1: Extract the sentences containing entity relationships from the original threat intelligence text;

[0087] P2: Extract threat intelligence entity relationships: Use the threat intelligence text strings in the original annotation data set, the entity list obtained in the entity extraction step, and the relationship list defined in S1 as inputs for sentence extraction, and then output the start index, end index, head entity information, and tail entity information of each sentence. Then use the output information to perform threat intelligence entity relationship extraction;

[0088] 3) After all entity extraction and relationship extraction are completed, input the extracted entity and relationship information into the graph database to construct a threat intelligence knowledge graph. The graph database uses the Neo4j database.

[0089] When applying the threat intelligence ontology in the field of knowledge graphs, some threat entities in STIX2.1 may not be applicable. It is necessary to make a choice on whether the 18 types of threat entities in STIX2.1 are applicable to the threat intelligence ontology proposed in the present invention.

[0090] For Indicator, Infrastructure, and ObservedData. Indicator represents the Indicator of Compromise (IOC) during an attack; ObservedData represents the data that can be observed during an attack, with IOCs being the main focus in APT reports; Infrastructure refers to the physical or virtual resources involved in an attack activity, and in APT reports, IPs or domain names are usually used to describe the resources used by APT groups, such as C2 servers. There is semantic overlap among the three, all containing content related to Indicators of Compromise. Therefore, the three are refined into URLs, domain names, and IPs and classified into the Infrastructure category. Next are ThreatActor, Campaign, and Intrusion Set. Campaign is part of Intrusion Set, and APT reports usually describe a certain Campaign or some stages of a Campaign. Therefore, the two are merged into Campaign. Threat Actor, as the protagonist of the Campaign, has a participatory relationship with the Campaign and is thus retained.

[0091] Furthermore, consider Attack Pattern, Malware, Malware Analysis, and Tool. For Malware and Malware Analysis, there are specialized malware analysis reports in the threat intelligence field, but APT reports usually only contain basic descriptions such as the name and functions of the malware, without including the analysis of the malware. Therefore, Malware and MalwareAnalysis are merged into Malware. And Attack Pattern, Malware, and Tool all fall within the scope of TTP. In APT reports, Attack Pattern will be refined into descriptions of the techniques used by APT groups and the corresponding attack processes. Therefore, Technique is used instead of Attack Pattern, and Malware and Tool are retained.

[0092] Furthermore, consider Course of Action, Vulnerability, Identity, and Location. Course of Action describes the methods and measures for dealing with TTPs and certain security vulnerabilities; Location is common information in APT reports and is also important information constituting the threat intelligence knowledge graph, including the source locations of entities such as APT groups and malware, and the locations under attack; Identity can appear in APT reports with different identities, such as APT groups, attacked institutions, etc. Therefore, the above four types of threat entities are retained. In addition, Industry, as the target attacked by APT groups, often appears together with locations, so Industry is added as a threat entity.

[0093] Furthermore, consider Grouping, Report, Note, and Opinion. The relationships between Grouping, Note, and Opinion and other STIX objects are not clearly defined in STIX 2.1, and the three are more applicable to scenarios such as threat intelligence exchange or collaborative threat analysis. For constructing the threat intelligence knowledge graph, these three types of threat entities are not necessary. Report is used as a data source in the form of APT report examples in the present invention, so it is not listed separately as a type of threat entity. Therefore, the above four types of threat entities are not retained.

[0094] As a preferred embodiment of the present invention, the entity types in step S1 include 13 types, namely Threat Actor, Campaign, Malware, Technique, Tool, Identity, Location, Industry, Vulnerability, Course of Action, URL, Domain, IP; there are 7 types of relationships between threat intelligence entities, namely use, attack, originate from, similar, identical, own, and respond to.

[0095] As a preferred embodiment of the present invention, the original data set in step S2 is derived from unstructured APT reports, and the APT report text is manually annotated to obtain the original annotated data set; the original annotated data set is enhanced through a template sentence filling algorithm to obtain the enhanced data set.

[0096] As a preferred embodiment of the present invention, the threat intelligence domain vocabulary knowledge base consists of threat intelligence entity vocabularies, and a total of 5 knowledge bases are constructed, covering the entity types to be enhanced, namely attack organizations, malware, vulnerabilities, domain names, and IPs.

[0097] As a preferred embodiment of the present invention, the template sentence filling algorithm in step S4 specifically includes the following steps:

[0098] S4.1 Convert the sentences in the training set into the BIO annotation mode;

[0099] S4.2 Use the BIO annotation results of the template sentences and the vocabulary knowledge base in the threat intelligence field as inputs. For each line in the BIO annotation results of the template sentences, obtain its words and labels respectively;

[0100] S4.3 If the label is O, concatenate the word and the label and store them in a list; if the label is not O, obtain a domain vocabulary from the knowledge base and determine how many words the domain vocabulary consists of;

[0101] S4.4 If the domain vocabulary consists of one word, concatenate the domain vocabulary and the corresponding label and store them in a list; if the domain vocabulary consists of multiple words, concatenate the first word of the domain vocabulary and the B-label and store them in a list, and concatenate the second and subsequent words of the domain vocabulary and the I-label and store them in a list; finally, return a list of sentences generated by filling the template and output this list.

[0102] As a preferred embodiment of the present invention, step S5 specifically includes the following contents:

[0103] S5.1 Input the sentences generated in S4 into BERT, and BERT dynamically generates word vectors for each input word according to its context, and the word vectors are used to represent the semantic information of the words;

[0104] Figure 3 And Figure 6 The [CLS] in it represents classification and is a marker used by the BERT model for text classification tasks. [CLS] obtains the sentence-level information representation through the self-attention mechanism, and its corresponding vector is the sentence vector containing the sentence semantic information. Adding [CLS] in the present invention is required by the BERT input format and has nothing to do with the named entity recognition task.

[0105] S5.2 Input the word vectors generated in S5.1 into a bidirectional LSTM encoder, and the LSTM encoder encodes the temporal relationship of the input sequence from two directions (forward and backward) and outputs a hidden state sequence;

[0106] S5.3 Use a CRF conditional random field to decode the hidden state sequence output in S5.2 and output the label sequence corresponding to the sentence. For the context encoding module of the output sequence R = {r1, r2,..., rn}, the calculation method of the matching score given the input and output is expressed as:

[0107]

[0108] Among them, P represents the feature matrix of the BiLSTM layer, and A represents the state transition score matrix of the CRF layer. represents the transition score from y i label to y i+1 label;

[0109] S5.4 performs softmax on all possible label sequences of the input sequence R to obtain the predicted label y, and the obtained predicted label is the corresponding entity type, and its definition is expressed as:

[0110]

[0111] S5.5 In model training, the maximum log-likelihood estimation is used to obtain the loss function, and the specific process is expressed as:

[0112]

[0113] Finally, the predicted label obtained in S5.4 is mapped to the entity type and output.

[0114] As a preferred embodiment of the present invention, the calculation unit of the above LSTM encoder has three gate structures: an input gate, a forget gate, and an output gate;

[0115] The specific calculation process is as follows:

[0116] f t = σ(W f * [c t-1 , h t-1 , x t + b f )

[0117] i t = σ(W i * [c t-1 , h t-1 , x t + b i )

[0118] o t = σ(W o * [c t-1 , h t-1 , x t + b o )

[0119]

[0120] Among them, (f t , i t , o t , C t ) represent the forget gate, the input gate, the output gate, and the cell state respectively, (W f , Wi , W o , W c ), and (b f , b i , b o , b c ), respectively represent the weight matrices and bias vectors of the forget gate, input gate, output gate, and memory unit; x t and h t represent the input vector and the hidden layer vector at time t; σ and tanh are activation functions.

[0121] As a preferred embodiment of the present invention, step P1 includes the following specific contents:

[0122] P1.1 Input the full text string, entity list, and relationship list;

[0123] P1.2 For each record in the relationship list, obtain the two entities involved in the relationship from the entity list;

[0124] P1.3 Set the left pointer to the start index of the earlier entity among the two entities, and set the right pointer to the end index of the later entity among the two entities;

[0125] P1.4 Enter a loop until breaking out using the keyword break. If the left pointer is greater than zero and the left pointer points to a full stop in English, determine whether the part from the left pointer to the end of the string contains the English sentence ending feature. If it contains, the left pointer index is the start index of the target sentence; similarly, the end index of the target sentence can be obtained;

[0126] P1.5 Return and output the start index, end index, head entity information, and tail entity information of the target sentence.

[0127] As a preferred embodiment of the present invention, step P2 includes the following specific contents:

[0128] P2.1 Integrate the entity semantic information and entity boundary information output by P1.5, so that the boundary information and type information of the entity are reflected in the tags on both sides of the entity;

[0129] Examples of the tags are as follows: For example, [E11:att] represents the left boundary of entity 1, and this entity belongs to the attacker type; [E12:att] represents the right boundary of entity 1, and this entity belongs to the attacker type;

[0130] P2.2 Given a sentence s containing entities e1 and e2, input it into BERT and output the vector H;

[0131] The average of the word vectors of each word in the constituent entity is calculated by P2.3 to obtain the representation vector of the entity. The representation vectors of the two entities are fed into the tanh activation function and a fully connected layer, and the outputs are denoted as H1 ′ and H2 ′ , and the calculation process is expressed as:

[0132]

[0133] where the vector H i to H j represents the word vector of entity e1, and the vector H x to H y represents the word vector of entity e2,

[0134] b1 = b2 is the bias vector, and d is the size of the BERT hidden state vector;

[0135] P2.4 feeds the sentence vector generated by CLS into the tanh activation function and a fully connected layer, which is expressed as:

[0136] H0 ′ = W0[tanh(H0)] + b0

[0137] where, b0 is the bias vector, and d is the size of the BERT hidden state vector;

[0138] The CLS represents classification, which is a token used by the BERT model for text classification tasks. It obtains the sentence-level information representation through the self-attention mechanism, and its corresponding vector is the sentence vector containing the sentence semantic information;

[0139] P2.5 concatenates H0 ′ , H1 ′ and H2 ′ and feeds them into a fully connected layer and a SoftMax layer to obtain the output vector P. This process is expressed as:

[0140] H ″ = W3[concat(H0 ′ , H1 ′ , H2 ′ )] + b3

[0141] p = softmax(H ″ )

[0142] where, L is the number of relationship categories, b3 is the bias vector, and the output vector P ∈ R L ; R L is the relationship set;

[0143] The relationship corresponding to the item with the largest value in P for P2.6 is the output relationship.

[0144] The above embodiments are only used to illustrate the design concept and characteristics of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design ideas disclosed by the present invention are within the protection scope of the present invention.

Claims

1. A threat intelligence entity relationship extraction method for unstructured data, characterized in that, It includes the following three parts: 1) Threat intelligence entity extraction, including the following steps: S1: Define threat entity types and relationships between threat intelligence entities based on the STIX threat intelligence standard; S2: Construct the NER original annotation dataset and the vocabulary knowledge base in the threat intelligence field; S3: Find the sentences containing entities of the type to be enhanced in the original annotation dataset as template sentences, fill in the entities of the same type in the vocabulary knowledge base in the threat intelligence field into the template sentences to generate new sentences containing specific type entities, and add the newly generated sentences to the NER original annotation dataset; S4: Fill the template sentences: Convert the template sentences into the BIO annotation mode, and use the annotation results and the vocabulary knowledge base in the threat intelligence field as inputs. Through the template sentence filling algorithm, generate and output the sentences after template filling. The output sentences constitute the enhanced dataset; The template sentence filling algorithm in step S4 specifically includes the following steps: S4.1 Convert the sentences in the training set into the BIO annotation mode; S4.2 Use the BIO annotation results of the template sentences and the vocabulary knowledge base in the threat intelligence field as inputs. For each line in the BIO annotation results of the template sentences, obtain its word and label respectively; S4.3 If the label is O, splice the word and the label and store them in a list; if the label is not O, obtain a domain vocabulary from the knowledge base and judge how many words the domain vocabulary consists of; S4.4 If the domain vocabulary consists of one word, splice the domain vocabulary and the corresponding label and store them in a list; if the domain vocabulary consists of multiple words, splice the first word of the domain vocabulary and the B-label and store them in a list, splice the second and subsequent words of the domain vocabulary and the I-label and store them in a list; finally, return a list of sentences composed of the sentences generated by filling the template and output the list; S5: Use the BERT+BiLSTM+CRF model to perform entity extraction on the sentences in the enhanced dataset output in step S4. Among them, the BERT layer is responsible for dynamically generating word vectors for each input word according to its context, and the generated word vector sequence will be used as the input of the BiLSTM layer; the BiLSTM layer is responsible for encoding the temporal relationship of the input sequence and outputting the hidden state sequence; the CRF layer decodes the hidden state sequence to obtain the label sequence corresponding to the sentence, and the obtained label sequence is the entity type; 2) After entity extraction, perform threat intelligence relationship extraction, including the following steps: P1: Extract the sentences containing entity relationships from the original annotation dataset; P2: Extract threat intelligence entity relationships: Use the threat intelligence text strings in the original annotation dataset, the entity list obtained in the entity extraction step and the relationship list defined in S1 as inputs for sentence extraction, and then output the start index, end index, head entity information and tail entity information of each sentence, and then use the output information to perform threat intelligence entity relationship extraction; 3) After all entity extraction and relationship extraction are completed, input the extracted entity and relationship information into the graph database to construct a threat intelligence knowledge graph, and the graph database uses the Neo4j database.

2. The method for extracting threat intelligence entity relationships for unstructured data according to claim 1, wherein The entity types in step S1 include 13 categories, namely Threat Actor, Campaign, Malware, Technique, Tool, Identity, Location, Industry, Vulnerability, Course of Action, URL, Domain, IP; there are 7 types of relationships between threat intelligence entities, namely use, attack, originate from, similar, identical, possess, and respond to.

3. The method for extracting threat intelligence entity relationships for unstructured data according to claim 1, characterized in that, The original dataset in step S2 is sourced from unstructured APT reports, and the original labeled dataset is obtained by manually annotating the APT report text; The original labeled dataset is enhanced by the template sentence filling algorithm to obtain the enhanced dataset.

4. The threat intelligence entity relationship extraction method for unstructured data according to claim 1, characterized in that Step S5 specifically includes the following content: S5.1 Input the sentences generated in S4 into BERT, and BERT dynamically generates word vectors for each input word according to the context, and the word vectors are used to represent the semantic information of the words; S5.2 Input the word vectors generated in S5.1 into a bidirectional LSTM encoder, and the LSTM encoder encodes the temporal relationship of the input sequence from two directions (forward and backward) and outputs a hidden state sequence; S5.3 Use a CRF conditional random field to decode the hidden state sequence output in S5.2 and output the label sequence corresponding to the sentence. For the context encoding module of the output sequence R = {r1, r2,..., rn}, the calculation method of the matching score given the input and output is expressed as: Where P represents the feature matrix of the BiLSTM layer, and A represents the state transition score matrix of the CRF layer, represents the transition score from the y i label to the y i+1 label; S5.4 Perform softmax on all possible label sequences of the input sequence R to obtain the predicted label y, and the obtained predicted label is the corresponding entity type, and its definition is expressed as: S5.5 In model training, use the maximum log-likelihood estimation to obtain the loss function, and the specific process is expressed as: Finally, map the predicted label obtained in S5.4 to the entity type and output.

5. The method for extracting threat intelligence entity relationships for unstructured data according to claim 4, characterized in that, The computing unit of the LSTM encoder has three gate structures: input gate, forget gate, and output gate; The specific calculation process is as follows: f t = σ(W f * [c t-1 , h t-1 , x t + b f ) i t = σ(W i * [c t-1 , h t-1 , x t + b i ) o t = σ(W o * [c t-1 , h t-1 , x t + b o ) where (f t , i t , o t , C t ) represent the forget gate, input gate, output gate, and cell state respectively, and (W f , W i , W o , W c ) and (b f , b i , b o , b c ) represent the weight matrices and bias vectors of the forget gate, input gate, output gate, and memory cell respectively; x t and h t represent the input vector and hidden layer vector at time t; σ and tanh are activation functions.

6. The method for extracting threat intelligence entity relationships for unstructured data according to claim 1, characterized in that Step P1 includes the following specific content: P1.1 Input the full text string, entity list, and relationship list; P1.2 For each record in the relationship list, obtain the two entities involved in the relationship from the entity list; P1.3 Set the left pointer to the start index of the entity that comes before the two entities, and the right pointer to the end index of the entity that comes after the two entities; P1.4 Enter a loop until it breaks out using the keyword break. If the left pointer is greater than zero and the left pointer points to a full stop, determine whether the end of the string from the left pointer contains the English sentence ending feature. If it contains, the left pointer index is the start index of the target sentence; similarly, the end index of the target sentence can be obtained; P1.5 Return and output the start index, end index of the target sentence, head entity information, and tail entity information.

7. The threat intelligence entity relationship extraction method for unstructured data according to claim 1, characterized in that Step P2 includes the following specific content: P2.1 Integrate the entity semantic information and entity boundary information output in P1.5, so that the boundary information and type information of the entity are reflected in the labels on both sides of the entity; Given a sentence s containing entities e1 and e2, input it into BERT to output a vector H; For each word vector of the words that make up the entity, calculate the average to obtain the representation vector of the entity. Send the representation vectors of the two entities into the tanh activation function and a fully connected layer, and denote the outputs as H1′ and H2′ respectively. The calculation process is expressed as: Among them, the vector H i to H j represents the word vector of entity e1, and the vector H x to H y represents the word vector of entity e2, W1 = W2 ∈ R d*d , b1 = b2 is the bias vector, and d is the size of the BERT hidden state vector; Send the sentence vector generated by CLS into the tanh activation function and a fully connected layer, which is expressed as: H0′ = W0[tanh(H0)] + b0 where, W0 ∈ R d*d , b0 is the bias vector, and d is the size of the BERT hidden state vector; The CLS represents classification and is a token used by the BERT model for text classification tasks. It obtains the sentence-level information representation through the self-attention mechanism, and its corresponding vector is the sentence vector containing the semantic information of the sentence; After concatenating H0′, H1′, and H2′, send them into a fully connected layer and a SoftMax layer to obtain the output vector P. This process is expressed as: H″ = W3[concat(H0′, H1′, H2′)] + b3 p = softmax(H″) Among them, L is the number of relationship categories, b3 is the bias vector, and the output vector P ∈ R L ; R L is the set of relationships; Take the relationship corresponding to the item with the largest value in P as the output relationship.