Threat Intelligence Information Extraction Method and System Integrating Multiple Models
By building a multi-model fusion information extraction model, the problem of massive network threat intelligence analysis and processing is solved, the structured representation of threat intelligence and knowledge graph construction is realized, and the modeling and analysis of network security is supported.
Patent Information
- Application Number
- CN202211416431.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-12
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-11-12
AI Technical Summary
The difficulty in effectively analyzing and processing massive cyber threat intelligence in the existing technology has caused huge pressure on security analysts, and many alarms have not been processed and become spam data.
The threat intelligence information extraction method with fusion multi-models is adopted. By constructing a multi-model fusion information extraction model, the entity extraction, co-reference removal and relationship extraction models are trained and optimized, and the relationship between entities and entities is obtained, and the knowledge graph is constructed for analysis and reasoning.
Organize scattered and heterogeneous security data, generate structured threat intelligence representations, fill them into the knowledge graph, intuitively display the relationship between entities and entities, and support threat modeling, risk analysis and attack inference in the network security space.
Smart Images

Figure CN116049419B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cyberspace security, and particularly relates to a threat intelligence information extraction method and system integrating multiple models. Background Art
[0002] At present, the development of the Internet has entered a brand-new era, and the Internet of Everything has long become a reality, and the production and living styles of human beings have been affected unprecedentedly. Modern IT infrastructure is suffering from varying degrees of cyber attacks. To cope with this situation, it is necessary to continuously monitor it, collect and process information, and use cyber threat intelligence (CTI) for cyber defense. However, the Internet has complex components, the behaviors of attackers are changeable, and the number of security devices is increasing day by day, resulting in a geometric increase in threat intelligence. At the same time, cyber threat intelligence usually exists in the form of natural language, relevant entities are scattered throughout the article, and there are intricate relationships between entities, which brings challenges to intelligence analysis, utilization, and sharing. The massive amount of alarm data brings huge pressure to security analysts, and many alarms are not processed and become garbage data. Therefore, how to analyze and process threat intelligence has become a key problem to be solved urgently.
[0003] Manual analysis of threat intelligence requires certain professional knowledge of cybersecurity, is time-consuming and laborious, has low evaluation efficiency, and is difficult to cope with the increasing cyber attacks. In view of its importance, many research works are dedicated to extracting structured knowledge from unstructured threat intelligence. This process mainly involves four key technologies: entity extraction, coreference resolution, relation extraction, and knowledge graph construction. The automated analysis of threat intelligence mainly faces the following challenges: (1) Different from the general domain, entities in the threat intelligence domain have strong domain characteristics. For example, threat entities include hacker organizations, attack techniques, malware, etc., and entity extraction models in the general domain are difficult to directly identify; (2) In threat intelligence texts, an entity may appear multiple times in a document, that is, there are multiple mentions. Judging whether a mention points to the same entity requires making full use of context information and extracting semantic knowledge; (3) The document structure of threat intelligence is complex, the sentences are relatively long, and the relationships between entities usually need to be inferred based on multiple sentences. Therefore, there is an urgent need for an information extraction solution to meet the modeling analysis and risk reasoning in the threat intelligence domain. Summary of the Invention
[0004] Therefore, the present invention provides a threat intelligence information extraction method and system integrating multiple models, which can organize scattered, multi-source heterogeneous security data and provide technical support for threat modeling, risk analysis, attack reasoning, etc. in the cyberspace security.
[0005] According to the design solution provided by the present invention, a threat intelligence information extraction method integrating multiple models is provided, which includes the following contents:
[0006] Construct an information extraction model integrated by multiple models and train and optimize each of the multiple models respectively. Among them, the multiple models for integration include an entity extraction model for extracting entity mentions in the input data, a coreference resolution model for fusing entity mentions, and a relation extraction model for extracting relationships between entities;
[0007] Input the threat intelligence document to be processed into the information extraction model. First, use the entity extraction model to perform word segmentation processing and information fusion on the input document to obtain entity mentions in the document; then use the coreference resolution model to determine whether the entity mentions point to the same entity and then enhance the entity mention representation through entity mention fusion; then, use the relation extraction model to obtain entity pair representations and extract relationships between entities through specific relationship probabilities;
[0008] Construct a knowledge graph based on the entities and relationships between entities obtained by the information extraction model, and use this knowledge graph to model, analyze and infer risks in the threat intelligence document.
[0009] As the threat intelligence information extraction method integrating multiple models in the present invention, further, using the entity extraction model to perform word segmentation processing and information fusion on the input document to obtain entity mentions in the document includes: First, obtain the word set and the context representation of words in the document through word segmentation and encoding processing of the input document, and use the Natural Language Toolkit to obtain the part-of-speech sequence of each word in the word set. Generate a part-of-speech enhanced word representation by embedding and linking the context representation of the word and the part-of-speech sequence; then, use the multi-head attention mechanism to obtain the key context embedding of the word by learning the features of different representation subspaces of the word representation; then, input the word representation into the trained BiLSTM model to obtain a feature vector, fuse the key context embedding of the word and the feature vector, and use a linear classifier to obtain a sequence label used as an entity mention.
[0010] As the threat intelligence information extraction method integrating multiple models in the present invention, further, in the word segmentation and encoding processing of the input document, add a position marker at the starting position of the input document, use a tokenizer to obtain the word set of the input document, and obtain the context representation of the word through an encoder.
[0011] As the threat intelligence information extraction method for integrating multiple models in the present invention, further, when inputting the word representation into the trained BiLSTM model to obtain the feature vector, the BiLSTM model includes a forward LSTM layer, a backward LSTM layer, and a connection layer, and in the BiLSTM model, each time step is an LSTM storage unit, and the word feature composed of historical information and future information at the current time is obtained based on the hidden vector at the previous moment, the storage unit vector at the previous moment, and the current input word embedding.
[0012] As the threat intelligence information extraction method for integrating multiple models in the present invention, further, in the entity fusion by using the co-reference resolution model to determine whether entity mentions point to the same entity, a convolutional neural network is used to obtain the entity different-dimensional features represented by each entity mention, the dimensionality reduction and redundancy removal of the entity features are performed through the pooling layer, and the label probability that the entity mentions point to the same entity is calculated by using the tanh activation function, and the context and entity mentions are fused according to the label probability.
[0013] As the threat intelligence information extraction method for integrating multiple models in the present invention, further, the entity pair representation is obtained by using the relation extraction model, and the relation between entities is extracted through the specific relation probability, including: first, mention markers are set at the start and end positions of each entity mention in the input document, and the word representation before the mention marker existing before the entity mention is used as the entity mention representation; then, the width of the entity mention is enhanced by using the trained width embedding matrix, the entity representation is obtained according to the entity mention after width enhancement, the key context of the special entity pair is located through the multi-head attention matrix to obtain the local context embedding of the special entity pair, and the entity representation is enhanced by using the trained entity distance embedding matrix and entity type embedding matrix; then, the enhanced entity representations are semantically grouped and fused to obtain the entity pair representation, and the specific relation probability is obtained by using the non-linear activation function, and the relation between entities is extracted according to the specific relation probability.
[0014] As the threat intelligence information extraction method for integrating multiple models in the present invention, further, in obtaining the entity representation according to the entity mention after width enhancement, the LogSumExp pooling method is used to obtain the entity-level representation, and the specific process is expressed as: Among them, represents the number of entity mentions included in entity e i m j represents the j-th mention of the m-th entity, represents the entity mention m after width enhancement j .
[0015] As the threat intelligence information extraction method for integrating multiple models in the present invention, further, in obtaining the local context embedding of a special entity pair by locating the key context of the special entity pair through the multi-head attention matrix, first, obtain the attention scores between words in the multi-head attention heads, take the attention before the mention marker exists before the entity mention as the attention score of the entity mention, obtain the entity-level attention score by averaging the attention scores of all entity mentions of the same entity, and take this entity-level attention score as the attention of the corresponding entity to all words. Then, use the attention matrix to locate the key context of the special entity pair, and obtain the local context embedding based on the key context.
[0016] Further, the present invention also provides a threat intelligence information extraction system for integrating multiple models, including: a model construction module, an information extraction module, and an information output module, where
[0017] The model construction module is used to construct an information extraction model integrated by multiple models and train and optimize the multiple models respectively. Among them, the multiple models for integration include an entity extraction model for extracting entity mentions in the input data, a coreference resolution model for performing fusion processing on entity mentions, and a relation extraction model for extracting relationships between entities;
[0018] The information extraction module is used to input the threat intelligence document to be processed into the information extraction model. First, use the entity extraction model to perform word segmentation processing and information fusion on the input document to obtain entity mentions in the document; then use the coreference resolution model to determine whether the entity mentions point to the same entity and enhance the entity mention representation through entity mention fusion; then, use the relation extraction model to obtain entity pair representations and extract relationships between entities through specific relation probabilities;
[0019] The information output module is used to construct a knowledge graph based on the entities and relationships between entities obtained by the information extraction model, and use this knowledge graph to model, analyze, and infer risks in the threat intelligence document.
[0020] Advantages of the present invention:
[0021] The present invention can input unstructured threat intelligence text into the model, obtain a structured representation of the text, fill it into the knowledge graph, and can be presented using the Neo4j graph database; it can organize scattered, multi-source heterogeneous security data to construct a knowledge graph, intuitively display entities and the relationships between entities, and provide support for data analysis and knowledge reasoning in threat modeling, risk analysis, attack reasoning, etc. in the cyber security space, and has good application prospects. Brief Description of the Drawings
[0022] Figure 1 It is a schematic diagram of the threat intelligence information extraction process for integrating multiple models in the embodiment;
[0023] Figure 2 Schematic diagram of the information extraction model architecture in the embodiment
[0024] Figure 3 Schematic diagram of the dataset distribution in the embodiment
[0025] Figure 4 Schematic diagram of the threat intelligence knowledge graph in the embodiment Detailed implementation manners
[0026] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and technical solutions
[0027] In the embodiment of this case, refer to Figure 1 as shown, a threat intelligence information extraction method integrating multiple models is provided, including
[0028] S101. Construct an information extraction model fused by multiple models and train and optimize each model respectively. Among them, the multiple models for fusion include an entity extraction model for extracting entity mentions in the input data, a coreference resolution model for fusing entity mentions, and a relation extraction model for extracting relationships between entities
[0029] S102. Input the threat intelligence document to be processed into the information extraction model. First, use the entity extraction model to perform word segmentation processing and information fusion on the input document to obtain entity mentions in the document; then use the coreference resolution model to determine whether the entity mentions refer to the same entity and enhance the entity mention representation through entity mention fusion; then, use the relation extraction model to obtain entity pair representations and extract relationships between entities through specific relationship probabilities
[0030] S103. Construct a knowledge graph based on the entities and relationships between entities obtained by the information extraction model, and use this knowledge graph to model, analyze and infer risks in the threat intelligence document
[0031] Refer to Figure 2 as shown, by integrating entity extraction, coreference resolution, relation extraction, and knowledge graph construction, the input unstructured threat intelligence text is output in a structured manner, and a knowledge graph is generated, which is convenient for storage using the Neo4j graph database, and explicitly shows the entities in the threat intelligence and the relationships between them, so as to provide knowledge support and decision-making support for security analysts to understand attack events and make defense deployments
[0032] As a preferred embodiment, further, an entity extraction model is used to perform word segmentation processing and information fusion on the input document to obtain entity mentions in the document, including: First, the word set in the document and the context representation of the words are obtained by performing word segmentation encoding processing on the input document, and the part-of-speech sequence of each word in the word set is obtained by using the Natural Language Toolkit. The part-of-speech enhanced word representation is generated by embedding and linking the context representation of the words and the part-of-speech sequence; Next, the multi-head attention mechanism is used to obtain the key context embedding of the words by learning the features of different representation subspaces of the word representation; Then, the word representation is input into the trained BiLSTM model to obtain the feature vector, the key context embedding of the words and the feature vector are fused, and a linear classifier is used to obtain the sequence label for use as an entity mention.
[0033] In the entity extraction model, the multi-head self-attention mechanism can be used to obtain the vector representation important for the entity, fuse it with the feature vector generated by the recurrent neural network model, input it into the linear layer to obtain the sequence label, and extract the entities in the text.
[0034] Different from the traditional encoding layer that uses random word embeddings, in the embodiment of this case, on the basis of introducing a pre-trained model to provide rich semantic knowledge, part-of-speech embeddings are incorporated, further enhancing the representation ability of the mention embeddings. The pre-trained model BERT is used as the encoder, and the special tokens "[CLS]" and "[SEP]" are added at the beginning and end positions of the document respectively. For each mention in the document, the special token "*" can be inserted at the beginning and end positions.
[0035] The given document is input into the tokenizer to obtain the tokenized document x t represents the word at position t. Input into the encoder to obtain the context representation H of the document words:
[0036] H = BERT([x1,..., x l ) = [h1,..., h l (1)
[0037] where d1 is the dimension of the hidden layer of the pre-trained model.
[0038] The part-of-speech sequence of the document is obtained by using the Python library Nltk, and the part-of-speech embedding matrix P is constructed:
[0039] P = Pos([x1,..., x l ) = [p1,..., p l (2)
[0040] where d2 is the dimension of the part-of-speech embedding.
[0041] For each word token, the context embeddings generated by the pre-trained model BERT are concatenated with the part-of-speech embeddings to generate part-of-speech enhanced word representations
[0042]
[0043] where denotes the concatenation operation.
[0044] To obtain vector representations important for entities, the entity extraction model incorporates a multi-head self-attention mechanism that can learn the dependency relationships between any two words, assigns different weights to each token representation, and obtains key information. Multiple attention heads can be used to learn the features of different representation subspaces, achieving a significant improvement in model performance. Specifically, the sequence of part-of-speech enhanced word representations is used as the input to the attention layer to obtain important context embeddings for the current word:
[0045]
[0046]
[0047] where Q, K, and V are the query sequence, key vectors, and value vectors respectively, and d k is the dimension of the key vectors, and H is the number of attention heads.
[0048] To obtain the historical and future information of the current word, a BiLSTM model is introduced. In previous work, the BiLSTM encoding layer has demonstrated its effectiveness in capturing the semantic information of words. The BiLSTM consists of a forward LSTM layer, a backward LSTM layer, and a connection layer. Each LSTM contains a set of recurrent connection sub-networks called memory modules. Each time step is a memory module of the LSTM, which is calculated based on the previous hidden vector, the previous cell vector, and the current input word embedding.
[0049] The sequence of part-of-speech enhanced word representations is used as the input to the BiLSTM layer to obtain feature vectors:
[0050]
[0051] The important context embeddings are fused with the feature vectors generated by the BiLSTM and input into a linear classifier to obtain sequence labels.
[0052]
[0053] As a preferred embodiment, further, in the entity fusion by using a coreference resolution model to determine whether entity mentions refer to the same entity, a convolutional neural network is used to obtain entity feature representations of different dimensions for each entity mention. The entity features are reduced in dimension and redundant information is removed through a pooling layer, and a tanh activation function is used to calculate the label probability that the entity mentions refer to the same entity. The context and entity mentions are fused based on the label probability.
[0054] Use a coreference resolution model to fuse context information and mention embeddings to enhance the mention representation. By introducing a convolutional neural network to extract features of different dimensions of mentions, the deficiency of the relatively low recall rate of traditional coreference resolution methods is effectively compensated. In the embodiments of this case, coreference resolution is regarded as a binary classification problem. First, obtain the word representation sequences with enhanced part-of-speech for each mention. To make them of the same length, calculate the average value of the word vectors they contain.
[0055]
[0056]
[0057] The convolutional neural network extracts deep features of the sequence through a sliding window of a certain size to alleviate the problem of long-distance dependence. Usually, a convolutional layer contains a filter, and convolution operations are performed between the convolution kernel and the word vectors. The mention representation is input into the CNN layer to obtain its features of different dimensions. Then, a pooling layer is used to reduce the dimension and compress the features to remove redundant information and prevent overfitting. The model adopts the max-pooling method for pooling, that is, the maximum eigenvalue is selected from the eigenvalues obtained by each filter in the convolutional layer, and the remaining eigenvalues are discarded.
[0058] Mention-Pair i =Conv i (mention1·mention2) (10)
[0059] M=Concat(Mention-Pair1,...,Mention-Pair N ) (11)
[0060] MP=MaxPooling(M) (12)
[0061] Based on the obtained pooled feature vectors of mention pairs, further use the tanh activation function to calculate the label probability, that is, whether the two mentions refer to the same entity.
[0062] y CR =tanh(W2·MP+b′2) (13)
[0063] During prediction, corresponding mentions are extracted based on the sequence tags obtained from the entity extraction model, and the input coreference resolution model is used to predict whether the mentions refer to the same entity.
[0064] As a preferred embodiment, further, a relation extraction model is used to obtain entity pair representations, and the relationships between entities are extracted through specific relation probabilities, including: First, mention markers are set at the start and end positions of each entity mention in the input document, and the word representations before the mention markers in the entity mentions are used as the representations of the entity mentions; Then, the width embedding matrix that has been trained is used to enhance the width of the entity mentions, and based on the entity mentions with enhanced width, entity representations are obtained. The key context of a special entity pair is located through the multi-head attention matrix to obtain the local context embedding of the special entity pair, and the entity distance embedding matrix and entity type embedding matrix that have been trained are used to enhance the entity representations; Then, the enhanced entity representations are semantically grouped and fused to obtain entity pair representations, and a non-linear activation function is used to obtain specific relation probabilities, and the relationships between entities are extracted based on the specific relation probabilities.
[0065] In obtaining the local context embedding of a special entity pair by locating the key context of the special entity pair through the multi-head attention matrix, first, the attention scores between words in the multi-head attention heads are obtained, and the attention before the mention markers in the entity mentions is used as the attention score of the entity mention. The entity-level attention score is obtained by averaging the attention scores of all entity mentions of the same entity, and the entity-level attention score is used as the attention of the corresponding entity to all words. Then, the attention matrix is used to locate the key context of the special entity pair, and the local context embedding is obtained based on the key context.
[0066] In the relation extraction model, multiple features such as part of speech, mention width, entity type, and entity pair distance are incorporated to achieve document-level threat intelligence relation extraction. Document-level relation extraction aims to determine whether there are corresponding relationships between entities, and the present invention regards it as a multi-label classification problem. Additional features are fused in the entity representations to make full use of the document information.
[0067] Specifically, the word representation with enhanced part of speech marked with "*" before the mention is used as the representation of the mention. Experiments have shown that the width of the mention is an important piece of information about the entity. Therefore, a width embedding matrix is trained and fused with the mention representation to generate a mention representation with enhanced width:
[0068]
[0069]
[0070] where d3 is the dimension of the width embedding, and m j represents the j-th mention of the m-th entity.
[0071] For an entity e that contains mentions , it is necessary to integrate mention-level representations to obtain entity-level representations. Traditional methods usually adopt the max-pooling method. This method has good effects when mention pairs can clearly express relationships. However, in actual scenarios, the relationships between mention pairs of different entities are relatively ambiguous. In this paper, a smoothed version of max-pooling, namely LogSumExp pooling, is used to obtain entity-level representations: i
[0072]
[0073] Introduce the multi-head attention matrix A ∈ R of the encoder BERT HD×l×l , where A ijk represents the attention score from word j to word k in the i-th attention head. Take the attention of the mention pre-token "*" as the attention score of this mention, and then average the attentions of all mentions of the same entity to obtain the entity-level attention score representing the attention from the m-th entity to all words. Then use the attention matrix to locate the important context for a specific entity pair (e s , e o ), and calculate the local context embedding:
[0074]
[0075]
[0076] a (s,o) = q (s,o) / 1 T q (s,o)
[0077] c (s,o) = Ha (s,o)
[0078] Experiments prove that the distance between entities and entity types also have certain effects on the performance of relation extraction. Construct a distance embedding matrix and an entity type embedding matrix, and integrate them into the entity representation. In summary, the representation encoding of a specific entity pair is as follows:
[0079]
[0080]
[0081]
[0082] where d4 and d5 are the dimensions of the distance embedding and the type embedding respectively. d soDenote the distance between the first mention of entity s and entity o as e s and e o respectively represent the types of entity s and entity o.
[0083] To reduce the computational overhead, divide the entity representations into k semantic groups of the same size, and then fuse the entity representations to obtain the entity pair representation:
[0084]
[0085]
[0086] Use a non-linear activation function to calculate the probability of a specific relationship:
[0087]
[0088] By integrating the four steps of entity extraction, coreference resolution, relation extraction, and knowledge graph construction, the input unstructured threat intelligence text is output in a structured manner, and a knowledge graph is generated, which can be stored using the Neo4j graph database, explicitly showing the entities in the threat intelligence and the relationships between them, thus providing knowledge support and decision-making support for security analysts to understand attack events and make defense deployments.
[0089] Furthermore, based on the above method, the embodiment of the present invention also provides a threat intelligence information extraction system integrating multiple models, including: a model construction module, an information extraction module, and an information output module, where
[0090] The model construction module is used to construct an information extraction model fused by multiple models and train and optimize the multiple models respectively. Among them, the multiple models for fusion include an entity extraction model for extracting entity mentions in the input data, a coreference resolution model for fusing entity mentions, and a relation extraction model for extracting relationships between entities;
[0091] The information extraction module is used to input the threat intelligence document to be processed into the information extraction model. First, use the entity extraction model to perform word segmentation processing and information fusion on the input document to obtain entity mentions in the document; then use the coreference resolution model to determine whether the entity mentions point to the same entity and enhance the entity mention representation through entity mention fusion; then, use the relation extraction model to obtain the entity pair representation and extract the relationship between entities through the probability of a specific relationship;
[0092] The information output module is used to construct a knowledge graph based on the entities and the relationships between entities obtained by the information extraction model, and use this knowledge graph to model, analyze, and infer the risks in the threat intelligence document.
[0093] To verify the effectiveness of the solution in this case, the following further explanation will be made in combination with experimental data:
[0094] Taking the document to be analyzed as the model input, in the entity extraction model, first, the unstructured text is input into the Bert tokenizer Python library Nltk to obtain word embeddings and part-of-speech embeddings with semantic knowledge respectively. After fusing them, the result is input into the BiLSTM and attention layers to obtain feature vectors and important context embeddings, and a linear layer is used to obtain the document entity labels, that is, entity mentions. In the coreference resolution model, a CNN model is used to obtain features of different dimensions of the mentions, and feature dimensionality reduction is performed through max-pooling operations to remove redundant information, and then it is input into the tanh layer to judge whether the mentions refer to the same entity. In the relation extraction model, for each entity, a Logexpsum operation is used to obtain the entity-level embedding representation. At the same time, additional features such as mention width, entity type, and distance between entity pairs are introduced to enhance the entity representation, and a non-linear activation function is used to calculate the probability of a specific relationship. See Figure 3 the entity type distribution and relation type distribution shown. Using the information extraction model in this case, scattered and multi-source heterogeneous security data can be organized to obtain a structured representation of the text and filled into the knowledge graph. See Figure 4 shown. It can be presented using the Neo4j graph database, which can intuitively display the entities and the relationships between them, providing support for threat modeling, risk analysis, attack reasoning, etc. in the cyber security space in terms of data analysis and knowledge reasoning.
[0095] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0096] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0097] The units and method steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.
[0098] Those of ordinary skill in the art can understand that all or part of the steps in the above method can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as: read-only memory, magnetic disk or optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software function module. The present invention is not limited to any specific form of combination of hardware and software.
[0099] Finally, it should be noted that: the above embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for extracting threat intelligence information by integrating multiple models, characterized in that, It includes the following content: Construct an information extraction model fused by multiple models and train and optimize each model separately. Among them, the multiple models for fusion include an entity extraction model for extracting entity mentions in the input data, a coreference resolution model for fusing entity mentions, and a relation extraction model for extracting relationships between entities; Input the threat intelligence document to be processed into the information extraction model. First, use the entity extraction model to perform word segmentation processing and information fusion on the input document to obtain entity mentions in the document; then use the coreference resolution model to determine whether the entity mentions point to the same entity, and then enhance the entity mention representation through entity mention fusion; then, use the relation extraction model to obtain the entity pair representation, and extract the relationship between entities through the relationship probability; among them, using the relation extraction model to obtain the entity pair representation and extracting the relationship between entities through the relationship probability includes: first, set mention markers at the start and end positions of each entity mention in the input document, and use the word representation before the entity mention with a mention marker as the entity mention representation; then, use the trained width embedding matrix to enhance the entity mention width, obtain the entity representation based on the entity mention after width enhancement, locate the key context of the special entity pair through the multi-head attention matrix to obtain the local context embedding of the special entity pair, and use the trained entity distance embedding matrix and entity type embedding matrix to enhance the entity representation; then, obtain the entity pair representation by performing semantic grouping and fusion on the enhanced entity representation, and use the non-linear activation function to obtain the relationship probability, and extract the relationship between entities based on the relationship probability; and in obtaining the entity representation based on the entity mention after width enhancement, use the LogSumExp pooling method to obtain the entity-level representation, and the specific process is expressed as: Denote entity e i The number of entity mentions contained in, m j Denote the j-th mention of the m-th entity, Denote the entity mention m after width enhancement j ; In obtaining the local context embedding of the special entity pair by locating the key context of the special entity pair through the multi-head attention matrix, first, obtain the attention scores between words in the multi-head attention heads, use the attention before the entity mention with a mention marker as the attention score of the entity mention, obtain the entity-level attention score by averaging the attention scores of all entity mentions of the same entity, and use this entity-level attention score as the attention of the corresponding entity to all words. Then, use the attention matrix to locate the key context of the special entity pair, and obtain the local context embedding based on the key context; Construct a knowledge graph based on the entities and relationships between entities obtained by the information extraction model, and use this knowledge graph to model, analyze, and infer the risks in the threat intelligence document.
2. The method for extracting threat intelligence information by integrating multiple models according to claim 1, characterized in that, Use the entity extraction model to perform word segmentation processing and information fusion on the input document to obtain entity mentions in the document, including: First, obtain the set of words in the document and the context representation of the words through word segmentation and encoding processing of the input document, and use the Natural Language Toolkit to obtain the part-of-speech sequence of each word in the set of words. Generate a part-of-speech enhanced word representation by embedding and linking the context representation of the word and the part-of-speech sequence; Then, use the multi-head attention mechanism to obtain the key context embedding of the word by learning the features of different representation subspaces of the word representation. Then, input the word representation into the trained BiLSTM model to obtain a feature vector, fuse the key context embedding of the word and the feature vector, and use a linear classifier to obtain a sequence label used as an entity mention.
3. The method for extracting threat intelligence information by integrating multiple models according to claim 2, characterized in that, During the word segmentation and encoding process of the input document, add a position marker at the starting position of the input document, use the tokenizer to obtain the set of words in the input document, and obtain the context representation of the words through the encoder.
4. The method for extracting threat intelligence information by integrating multiple models according to claim 2, characterized in that, When inputting the word representation into the trained BiLSTM model to obtain a feature vector, the BiLSTM model includes a forward LSTM layer, a backward LSTM layer, and a connection layer. In the BiLSTM model, each time step is an LSTM storage unit, and the word feature composed of historical information and future information at the current time is obtained based on the hidden vector at the previous moment, the storage unit vector at the previous moment, and the current input word embedding.
5. The method for extracting threat intelligence information by integrating multiple models according to claim 1, characterized in that, When using the coreference resolution model to determine whether entity mentions point to the same entity for entity fusion, use a convolutional neural network to obtain different dimensional features of the entity represented by each entity mention, reduce the dimension and remove redundancy of the entity features through the pooling layer, and use the tanh activation function to calculate the label probability that the entity mentions point to the same entity. Fuse the context and entity mentions based on the label probability.
6. A system for extracting threat intelligence information by integrating multiple models, characterized in that, It includes: a model construction module, an information extraction module, and an information output module. Among them, The model construction module is used to construct an information extraction model fused by multiple models and train and optimize each model separately. Among them, the multiple models for fusion include an entity extraction model for extracting entity mentions in the input data, a coreference resolution model for fusing entity mentions, and a relation extraction model for extracting relationships between entities; An information extraction module is used to input a threat intelligence document to be processed into an information extraction model. First, the entity extraction model is used to perform word segmentation processing and information fusion on the input document to obtain entity mentions in the document. Then, the coreference resolution model is used to determine whether the entity mentions point to the same entity, and then the entity mentions are fused to enhance the entity mention representation. Then, the relationship extraction model is used to obtain the entity pair representation and extract the relationship between entities through the relationship probability. Among them, using the relationship extraction model to obtain the entity pair representation and extract the relationship between entities through the relationship probability includes: First, mention markers are set at the start and end positions of each entity mention in the input document, and the word representation before the entity mention with a mention marker is used as the entity mention representation. Then, the trained width embedding matrix is used to enhance the entity mention width, and the entity representation is obtained based on the entity mention after width enhancement. The key context of the special entity pair is located through the multi-head attention matrix to obtain the local context embedding of the special entity pair, and the trained entity distance embedding matrix and entity type embedding matrix are used to enhance the entity representation. Then, the enhanced entity representations are semantically grouped and fused to obtain the entity pair representation, and the non-linear activation function is used to obtain the relationship probability, and the relationship between entities is extracted based on the relationship probability. And in obtaining the entity representation based on the entity mention after width enhancement, the LogSumExp pooling method is used to obtain the entity-level representation, and the specific process is expressed as: Denote the entity e i The number of entity mentions contained in, m j Denote the j-th mention of the m-th entity, Denote the entity mention m after width enhancement j In obtaining the local context embedding of the special entity pair by locating the key context of the special entity pair through the multi-head attention matrix, first, the attention scores between words in the multi-head attention heads are obtained, and the attention before the entity mention with a mention marker is used as the attention score of the entity mention. The entity-level attention score is obtained by averaging the attention scores of all entity mentions of the same entity, and this entity-level attention score is used as the attention of the corresponding entity to all words. Then, the attention matrix is used to locate the key context of the special entity pair, and the local context embedding is obtained based on the key context; The information output module is used to construct a knowledge graph based on the entities and relationships between entities obtained by the information extraction model, and use this knowledge graph to model, analyze, and infer the risks in the threat intelligence document.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Threat intelligence oriented security knowledge graph construction method and system
CN109857917A
APT organization portrait construction method based on knowledge graph
CN112765366A