Document-level threat intelligence relationship extraction method and system based on feature enhancement
By utilizing BERT model and feature enhancement technology in the field of network security, the scarcity and OOV problems of threat intelligence data sets are solved, and the accuracy of entity relationship extraction is improved, which is suitable for threat intelligence analysis of complex structural documents.
Patent Information
- Application Number
- CN202211416432.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-11-12
AI Technical Summary
The existing technology lacks publicly available threat intelligence data sets in the field of network security, and has OOV problems. The threat intelligence document structure is complex, the entity and relationship frequency is low, the data label distribution is unbalanced, and the existing work mainly focuses on sentence-level text mining, making it difficult to effectively deal with entity relationships between multiple sentences.
The document-level threat intelligence relationship extraction method based on feature enhancement is adopted, and word embedding is generated using the BERT model, and features such as part-of-speech embedding vectors, entity width information and entity type are fused. Through knowledge distillation and multi-head attention mechanism, the accuracy of entity relationship extraction is improved, and the entity information extraction model is constructed and trained and optimized.
It effectively improves the accuracy of threat intelligence relationship extraction, solves OOV problems, improves the processing capabilities of complex structural documents, and provides technical support for threat modeling and risk analysis in network security space.
Smart Images

Figure CN116049343B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of cyberspace security technology, and in particular relates to a document-level threat intelligence relationship extraction method and system based on feature enhancement. Background Art
[0002] With the rapid development of network and information technology, new threat attacks are showing a trend of continuous growth. The increasingly complex attack strategies and ever-changing attack scenarios make it difficult for traditional network defenses such as firewalls and signature registries to resist these new attacks. In order to better understand the threat situation and coordinate the response to unknown threats, security experts propose to use Cyber Threat Intelligence (CTI) for network defense. In 2013, Gartner first proposed that threat intelligence is the knowledge about existing or upcoming threats to assets, including scenarios, mechanisms, indicators, revelations and actionable suggestions, etc. This knowledge can provide the subject with a strategy to deal with the threat.
[0003] The knowledge of threat intelligence comes from security analysis reports, blogs, social networks, vulnerability libraries, threat intelligence libraries, etc., which can provide strong data support for situational awareness and active defense. However, most threat intelligence exists in the form of natural language text, contains a large amount of unstructured data, and it is difficult to visualize the internal connections of attack elements. In order to help researchers quickly understand the semantic associations of attack elements, it is necessary to design corresponding algorithms to mine entities and their relationships from large-scale threat intelligence documents and construct a threat intelligence knowledge graph.
[0004] Relation extraction aims to identify the relationship between entities in a given text. Although relationship extraction in general fields has achieved good results, the following problems still exist in the field of network security: 1) There is a lack of publicly available threat intelligence datasets; 2) Threat intelligence contains a large number of professional terms such as vulnerability names, malware, APT organizations, etc., and there is a serious OOV (Out of Vocabulary) problem; 3) Threat intelligence documents have complex structures, relatively long sentences, low entity and relationship frequencies, and serious imbalance in data label distribution. In addition, existing work mainly focuses on sentence-level text mining. However, in actual scenarios, the same entity may have multiple mentions, and the relationship between entities usually needs to rely on multiple sentences for inference. For this reason, there is an urgent need for a relationship extraction scheme that can meet the analysis and processing of threat intelligence text data with complex document structures. Summary of the invention
[0005] To this end, the present invention provides a document-level threat intelligence relationship extraction method and system based on feature enhancement, which combines additional document information such as part-of-speech sequence and mention width to enhance entity representation features, thereby improving the accuracy of entity relationship extraction and providing technical support for threat modeling, risk analysis, attack reasoning, etc. in the network security space.
[0006] According to the design scheme provided by the present invention, a document-level threat intelligence relationship extraction method based on feature enhancement is provided, which includes the following contents:
[0007] Constructing an entity information extraction model and performing training optimization, wherein the entity information extraction model includes: a BERT model for encoding input text data to obtain a word entity mention context embedding vector of the text data, a fusion processing unit for enhancing the word entity mention context embedding vector by fusing part-of-speech embedding vectors and entity width information and obtaining a global representation of the entity, and an entity relationship extraction model for fusing a given entity pair's local context embedding with the entity's global representation, entity type information, and inter-entity spacing information to obtain a given entity pair's relationship probability through a nonlinear activation function;
[0008] The text data of the threat intelligence document to be processed is input into the trained and optimized entity information extraction model, and the BERT model is used to obtain the word entity mention context embedding vector in the text data. The part-of-speech embedding vector and entity width information are fused with the word entity mention context embedding vector using a fusion processing unit, and the entity relationship of the target entity pair in the text data is obtained through the entity relationship extraction model.
[0009] As a document-level threat intelligence relationship extraction method based on feature enhancement in the present invention, further, in the entity information extraction model training optimization, model training optimization is performed based on knowledge distillation, and the training optimization process includes: collecting sample label data, training a teacher model using the sample label data, and obtaining the teacher model by updating the teacher model parameters; then, performing knowledge distillation on the teacher model using the sample label data to obtain a student model, and using the student model as the entity information extraction model after training optimization.
[0010] As the document-level threat intelligence relationship extraction method based on feature enhancement in the present invention, further, the target loss function in the training optimization is expressed as: L RE =α 1 L AFL +α 2 L KD , where L RE represents the total loss, L AFL represents the adaptive loss for relation extraction as a multi-label separation problem, L KD represents the knowledge distillation loss, α 1, α 2 are the weights of adaptive loss and knowledge distillation loss, respectively.
[0011] As a document-level threat intelligence relationship extraction method based on feature enhancement in the present invention, further, a BERT model is used to obtain a context embedding vector of a word entity mention in text data, comprising: first, a word segmentation process is performed on the input document text data through a word segmenter to obtain a word entity set, and the entity mention is marked with a preset mention symbol; then, a pre-trained BERT model is used as an encoder to encode the word entities in the text data to generate an entity mention context embedding vector.
[0012] As a document-level threat intelligence relationship extraction method based on feature enhancement of the present invention, further, a fusion processing unit is used to fuse the part-of-speech embedding vector and the entity width information with the word entity mention context embedding vector. First, a natural language processing tool is used to obtain the part-of-speech tag, and the part-of-speech embedding vector of the word entity in the text data is generated using the part-of-speech tag, and the part-of-speech embedding enhanced vector representation is generated by fusing it with the context embedding vector; then, the vector representation is enhanced by fusing the entity mention width information; then, for each entity, a pooling operation is performed to obtain the global representation of the entity.
[0013] As a document-level threat intelligence relationship extraction method based on feature enhancement of the present invention, further, for the word entity mention context embedding vector fused with the part-of-speech embedding vector, the attention score of each mention element in the word entity mention context embedding vector is obtained based on the multi-head attention mechanism, the attention score of each mention element is used as the attention of the corresponding entity mention, and the entity-level attention matrix is obtained by averaging the entity mention attention, and the entity-level attention matrix is used to represent the attention score of the corresponding entity to all entity mentions.
[0014] As a document-level threat intelligence relationship extraction method based on feature enhancement of the present invention, further, an entity relationship extraction model is used to obtain the entity relationship of the target entity pair in the text data, including: first, the key context of a given entity pair is located through an entity-level attention matrix, and the local context embedding vector of the given entity pair is obtained based on the context; then, the local context embedding vector is fused with the entity global representation, entity type information and inter-entity distance information to obtain the context embedding representation of the target entity pair; then, a nonlinear activation function is used to obtain the relationship probability of the given target entity pair.
[0015] As a document-level threat intelligence relationship extraction method based on feature enhancement of the present invention, further, in obtaining the relationship probability of a given entity pair using a nonlinear activation function, the context embedding representation of the target entity pair is first grouped and feature fused to obtain an entity pair representation, and then the nonlinear activation function sigmoid is used to obtain the relationship probability of a given target entity pair.
[0016] Furthermore, the present invention also provides a document-level threat intelligence relationship extraction system based on feature enhancement, comprising: a model building module and a relationship extraction module, wherein:
[0017] A model building module, used to build an entity information extraction model and perform training optimization, wherein the entity information extraction model includes: a BERT model for encoding input text data to obtain a word entity mention context embedding vector of the text data, a fusion processing unit for enhancing the word entity mention context embedding vector by fusing the part-of-speech embedding vector and the entity width information and obtaining the entity global representation, and an entity relationship extraction model for fusing the local context embedding of a given entity pair with the entity global representation, entity type information and inter-entity spacing information to obtain the relationship probability of a given entity pair through a non-linear activation function;
[0018] The relationship extraction module is used to input the text data of the threat intelligence document to be processed into the trained and optimized entity information extraction model, obtain the word entity mention context embedding vector in the text data through the BERT model, use the fusion processing unit to fuse the part-of-speech embedding vector and entity width information with the word entity mention context embedding vector, and obtain the entity relationship of the target entity pair in the text data through the entity relationship extraction model.
[0019] Beneficial effects of the present invention:
[0020] In order to solve the OOV problem in threat intelligence, the present invention uses the Bert pre-trained model as an encoder to generate word embeddings and integrates additional features such as entity width, entity distance, entity type, etc., which can make full use of the text information in the document and effectively improve the accuracy of relationship extraction. In order to solve the problems of low frequency of entity relationships in threat intelligence and serious imbalance of data labels, the present invention introduces a teacher-student model, uses soft-label statistical data to calculate the effective information of the data set, retains the inter-class correlation information, removes some invalid redundant information, realizes knowledge distillation, improves the performance of the relationship extraction model, facilitates the processing of threat intelligence documents with complex structures, and provides support for data analysis and knowledge reasoning for threat modeling, risk analysis, attack reasoning, etc. in the network security space, which has good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1The following is a schematic diagram of the document-level threat intelligence relationship extraction process based on feature enhancement in an embodiment;
[0022] Figure 2 This is a schematic diagram of the entity information extraction model architecture in the embodiment;
[0023] Figure 3 This is a diagram of the teacher-student model in the embodiment;
[0024] Figure 4 This is a schematic diagram of the threat intelligence entity in the embodiment;
[0025] Figure 5 The following is a schematic diagram of some threat intelligence knowledge graphs in the embodiments. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of the present invention clearer and more understandable, the present invention is further described in detail below in conjunction with the accompanying drawings and technical solutions.
[0027] For the embodiment of this case, see Figure 1 As shown, a document-level threat intelligence relationship extraction method based on feature enhancement is provided, including:
[0028] S101, constructing an entity information extraction model and performing training optimization, wherein the entity information extraction model includes: a BERT model for encoding input text data to obtain a word entity mention context embedding vector of the text data, a fusion processing unit for enhancing the word entity mention context embedding vector by fusing the part-of-speech embedding vector and the entity width information and obtaining the entity global representation, and an entity relationship extraction model for fusing the local context embedding of a given entity pair with the entity global representation, entity type information and inter-entity spacing information to obtain the relationship probability of a given entity pair through a non-linear activation function;
[0029] S102. Input the text data of the threat intelligence document to be processed into the trained and optimized entity information extraction model, obtain the word entity mention context embedding vector in the text data through the BERT model, fuse the part-of-speech embedding vector and entity width information with the word entity mention context embedding vector using a fusion processing unit, and obtain the entity relationship of the target entity pair in the text data through the entity relationship extraction model.
[0030] For the embodiment of this case, see Figure 2 and 3 As shown in the figure, by constructing a threat intelligence ontology and integrating features such as entity width, entity distance, and entity type, we can make full use of the information in the document and improve the accuracy of relationship extraction. We also use the Bert pre-training model to generate word embeddings to effectively improve the OOV problem.
[0031] As a preferred embodiment, further, in the entity information extraction model training optimization, model training optimization is performed based on knowledge distillation, and the training optimization process includes: collecting sample label data, training the teacher model using the sample label data, and obtaining the teacher model by updating the teacher model parameters; then, performing knowledge distillation on the teacher model using the sample label data to obtain a student model, and using the student model as the entity information extraction model after training optimization.
[0032] The same batch of labeled sample data can be put into two models at the same time. The predicted output of the teacher model is used as the soft label, and the true label is used as the hard label. The two losses of the student model are calculated separately. Finally, the weighted sum of the two losses is used as the final loss to update the network parameters. During actual prediction, only the student model is used.
[0033] Furthermore, the objective loss function in training optimization is expressed as: L RE =α 1 L AFL +α 2 L KD , where L RE represents the total loss, L AFL represents the adaptive loss for relation extraction as a multi-label separation problem, L KD represents the knowledge distillation loss, α 1 , α 2 They are the weights of adaptive loss and knowledge distillation loss respectively. The mean square error loss function can be used to calculate the difference between the logits generated by the student model and the soft labels generated by the teacher model, and then combined with the adaptive Focal Loss loss weight as the overall loss function of the model to further improve the model performance.
[0034] As a preferred embodiment, further, a BERT model is used to obtain a context embedding vector of a word entity mention in text data, comprising: first, performing word segmentation processing on the input document text data through a word segmenter to obtain a word entity set, and marking the entity mention with a preset mention symbol; then, a pre-trained BERT model is used as an encoder to encode the word entities in the text data to generate an entity mention context embedding vector.
[0035] The pre-trained model Bert is used as the document encoder. For a document of length l x t Represents the word at position t in the document. Special symbols can be used to mark entity mentions: you can add "*" before and after the entity mention. Then use Bert to encode the document and generate the context embedding H. The encoding process can be expressed as:
[0036] H=Bert([x 1, ..., x l ])=[h 1 , ..., h l ]
[0037] in d 1 is the hidden layer dimension of the pre-trained model.
[0038] As a preferred embodiment, further, a fusion processing unit is used to fuse the part-of-speech embedding vector and the entity width information with the word entity mention context embedding vector. First, a natural language processing tool is used to obtain the part-of-speech tag, and the part-of-speech tag is used to generate a part-of-speech embedding vector of the word entity in the text data, and the part-of-speech embedding-enhanced vector representation is generated by fusing it with the context embedding vector; then, the vector representation is enhanced by fusing the entity mention width information; then, for each entity, a pooling operation is performed to obtain a global representation of the entity.
[0039] Use the Nltk library in Python to generate part-of-speech tags and generate the part-of-speech embedding matrix P. The specific process can be expressed as:
[0040] P=Pos([x 1 , ..., x l ])=[p 1 , ..., p l ]
[0041] in d 2 is the dimension of part-of-speech embedding.
[0042] It is fused with the context embedding H to generate a token representation enhanced by part-of-speech embedding.
[0043] C=[h 1 |p 1 , ..., h l |p l ]=[c 1 , ..., c l ]
[0044] in [|] indicates a concatenation operation.
[0045] The pre-generated width embedding matrix W and distance embedding matrix D are used to fuse the width information of entity mentions and the distance information between entities.
[0046]
[0047]
[0048] in d3 and d 4 are the dimensions of width embedding and distance embedding respectively.
[0049] The embedding vector of the “*” at the beginning of the entity mention is used as the embedding of the mention, denoted as Fusing it with width embedding, the fusion process can be expressed as:
[0050]
[0051] For including Mentions Entity i , the global representation of the entity is obtained by using the logsumexp pooling operation. The pooling operation process can be expressed as:
[0052]
[0053] As a preferred embodiment, further, for the word entity mention context embedding vector fused with the part-of-speech embedding vector, the attention score of each mention element in the word entity mention context embedding vector is obtained based on the multi-head attention mechanism, the attention score of each mention element is used as the attention of the corresponding entity mention, and the entity-level attention matrix is obtained by averaging the entity mention attentions, and the entity-level attention matrix is used to represent the attention scores of the corresponding entity to all entity mentions.
[0054] Using the pre-trained multi-head attention matrix A∈R HD×l×l , A ijk represents the attention score from token j to token k in the ith attention head, i.e., the mention-level attention. The mention-level attention is averaged to obtain the entity-level attention matrix Represents the attention score of the i-th entity to all tokens.
[0055] As a preferred embodiment, further, the entity relationship of the target entity pair in the text data is obtained through the entity relationship extraction model, including: first, locating the key context of the given entity pair through the entity-level attention matrix, and obtaining the local context embedding vector of the given entity pair based on the context; then, the local context embedding vector is fused with the entity global representation, entity type information and inter-entity distance information to obtain the context embedding representation of the target entity pair; then, a nonlinear activation function is used to obtain the relationship probability of the given target entity pair.
[0056] Generate the entity type embedding matrix T and fuse the entity type information. The fusion process can be expressed as:
[0057]
[0058] in d 5 The dimension to embed for the type.
[0059] For a given entity pair (e s , e o ), the attention matrix can be used to locate its important context and calculate the local context embedding of a specific entity pair. The specific process can be expressed as follows:
[0060]
[0061]
[0062] a (s,o) =q (s,o) / 1 T q (s,o)
[0063] c (s,o) =Ha (s,o)
[0064] The local context embedding is fused with the global entity representation, type embedding, and distance embedding to obtain the embedding representation of a specific entity pair. The fusion process can be expressed as follows:
[0065]
[0066]
[0067] where d so Represents the relative distance between the first mention of entity s and entity o.
[0068] In order to reduce the number of parameters and computational complexity, the entity representations are evenly divided into k groups, and feature fusion is performed to obtain entity pair representations. The fusion process can be expressed as follows:
[0069]
[0070]
[0071]
[0072] The nonlinear activation function sigmoid is used to obtain the relationship probability of a specific entity pair. The specific process can be expressed as:
[0073]
[0074] Relation extraction can be viewed as a multi-label classification problem. Traditional baselines usually use binary cross entropy loss as the loss function and specify a global threshold as the criterion for whether a relationship label exists. However, for different entity pairs, the model may have different confidence levels for the relationship threshold, and it is difficult to meet the requirements using only the global threshold. To address this problem, a learnable adaptive threshold is introduced to effectively reduce the decision errors in the inference process. For each entity pair (e s , e o ), the class with a score greater than the threshold is predicted as the positive class, and the rest is predicted as the negative class. On this basis, for the long-tail class, the adaptive Focal Loss in previous work is introduced, and the loss function can be expressed as follows:
[0075]
[0076] Furthermore, based on the above method, an embodiment of the present invention also provides a document-level threat intelligence relationship extraction system based on feature enhancement, comprising: a model building module and a relationship extraction module, wherein:
[0077] A model building module, used to build an entity information extraction model and perform training optimization, wherein the entity information extraction model includes: a BERT model for encoding input text data to obtain a word entity mention context embedding vector of the text data, a fusion processing unit for enhancing the word entity mention context embedding vector by fusing the part-of-speech embedding vector and the entity width information and obtaining the entity global representation, and an entity relationship extraction model for fusing the local context embedding of a given entity pair with the entity global representation, entity type information and inter-entity spacing information to obtain the relationship probability of a given entity pair through a non-linear activation function;
[0078] The relationship extraction module is used to input the text data of the threat intelligence document to be processed into the trained and optimized entity information extraction model, obtain the word entity mention context embedding vector in the text data through the BERT model, use the fusion processing unit to fuse the part-of-speech embedding vector and entity width information with the word entity mention context embedding vector, and obtain the entity relationship of the target entity pair in the text data through the entity relationship extraction model.
[0079] In order to verify the effectiveness of this solution, the following is an explanation based on specific data:
[0080] In order to address the lack of public data sets in the field of cybersecurity, we collected 227 pieces of threat intelligence and manually annotated them based on a custom ontology. We selected 151 of them as training sets and the remaining 76 as test sets.
[0081] Using the content in this case plan, the pre-trained model Bert is used as the document encoder, and then the Nltk library in Python is used to generate part-of-speech tags, generate part-of-speech embeddings, and fuse them with the context embeddings generated by the encoder. Generate width embeddings and distance embeddings respectively, and fuse the width information of entity mentions and the distance information between entities. Use the logsumexp pooling operation to obtain the global representation of the entity. Generate entity type embeddings and fuse the type information of the entity. Use the attention matrix to locate important contexts and calculate the local context embeddings of specific entity pairs. Then fuse the local context embeddings with the global entity representation, type embeddings, and distance embeddings to obtain the embedding representation of the specific entity pair. Finally, use the classifier to obtain the relationship probability of the specific entity pair. The threat intelligence ontology is as follows Figure 4 Threat intelligence is input into the entity information extraction model, and the relationship between all entity pairs in the text is predicted to fill in the knowledge graph, which is presented using the Neo4j graph database. The results can be shown as follows: Figure 5 As shown. Figure 5 The display can further verify that the solution in this case can be applied to threat intelligence analysis and processing of complex structures, and can provide strong data support for situational awareness and active defense.
[0082] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0083] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0084] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.
[0085] Those skilled in the art will appreciate that all or part of the steps in the above method can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk or an optical disk. Optionally, all or part of the steps in the above embodiment can also be implemented using one or more integrated circuits, and accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or in the form of software function modules. The present invention is not limited to any specific form of combination of hardware and software.
[0086] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A document-level threat intelligence relationship extraction method based on feature enhancement, It is characterized in that Contains the following: Constructing an entity information extraction model and performing model training optimization based on knowledge distillation, wherein the entity information extraction model includes: a BERT model for encoding input text data to obtain a word entity mention context embedding vector of the text data, a fusion processing unit for enhancing the word entity mention context embedding vector by fusing part-of-speech embedding vectors and entity width information and obtaining a global representation of the entity, and an entity relationship extraction model for fusing a given entity pair local context embedding with the entity global representation, entity type information, and entity spacing information to obtain a given entity pair relationship probability through a nonlinear activation function; The text data of the threat intelligence document to be processed is input into the trained and optimized entity information extraction model, and the word entity mention context embedding vector in the text data is obtained through the BERT model. The word entity mention context embedding vector and the entity width information are fused with the word entity mention context embedding vector by the fusion processing unit. For the word entity mention context embedding vector fused with the word entity mention context embedding vector, the attention score of each mention element in the word entity mention context embedding vector is obtained based on the multi-head attention mechanism, and the attention score of each mention element is used as the attention of the corresponding entity mention. The entity-level attention matrix is obtained by averaging the entity mention attention, and the entity-level attention matrix is used to represent the attention score of the corresponding entity to all entity mentions. The key context of a given entity pair is located by the entity-level attention matrix, and the local context embedding vector of the given entity pair is obtained according to the context; then, the local context embedding vector is fused with the entity global representation, entity type information and distance information between entities to obtain the context embedding representation of the target entity pair; then, the context embedding representation of the target entity pair is grouped and feature fused to obtain the entity pair representation, and the nonlinear activation function sigmoid is used to obtain the relationship probability of the given target entity pair.
2. According to the feature-enhanced document-level threat intelligence relationship extraction method of claim 1, It is characterized in that The model training optimization process based on knowledge distillation includes: collecting sample label data, training the teacher model with the sample label data, and obtaining the teacher model by updating the teacher model parameters; then, performing knowledge distillation on the teacher model with the sample label data to obtain the student model, and using the student model as the entity information extraction model after training and optimization.
3. According to the feature-enhanced document-level threat intelligence relationship extraction method of claim 2, It is characterized in that The objective loss function in training optimization is expressed as: L RE =α 1 L AFL +α 2 L KD , where L RE represents the total loss, L AFL represents the adaptive loss for relation extraction as a multi-label separation problem, L KD represents the knowledge distillation loss, α 1 , α 2 are the weights of adaptive loss and knowledge distillation loss, respectively.
4. According to the feature-enhanced document-level threat intelligence relationship extraction method of claim 1, It is characterized in that The BERT model is used to obtain the contextual embedding vector of word entity mentions in text data, including: first, the input document text data is segmented by a tokenizer to obtain a set of word entities, and the entity mentions are marked with preset mention symbols; then, the pre-trained BERT model is used as an encoder to encode the word entities in the text data and generate the contextual embedding vector of the entity mention.
5. According to the feature-enhanced document-level threat intelligence relationship extraction method according to claim 1 or 4, It is characterized in that In the process of fusing the part-of-speech embedding vector and entity width information with the word entity mention context embedding vector using a fusion processing unit, first, a natural language processing tool is used to obtain the part-of-speech tag, and the part-of-speech embedding vector of the word entity in the text data is generated using the part-of-speech tag, and the part-of-speech embedding enhanced vector representation is generated by fusing it with the context embedding vector; then, the vector representation is enhanced by fusing the entity mention width information; and then, for each entity, a pooling operation is performed to obtain the global representation of the entity.
6. A document-level threat intelligence relationship extraction system based on feature enhancement, It is characterized in that Contains: model building module and relationship extraction module, among which, A model building module, used to build an entity information extraction model and perform model training optimization based on knowledge distillation, wherein the entity information extraction model includes: a BERT model for encoding input text data to obtain a word entity mention context embedding vector of the text data, a fusion processing unit for enhancing the word entity mention context embedding vector by fusing the part-of-speech embedding vector and the entity width information and obtaining the entity global representation, and an entity relationship extraction model for fusing the local context embedding of a given entity pair with the entity global representation, entity type information and entity spacing information to obtain the relationship probability of a given entity pair through a non-linear activation function; The relation extraction module is used to input the text data of the threat intelligence document to be processed into the trained and optimized entity information extraction model, obtain the word entity mention context embedding vector in the text data through the BERT model, fuse the part-of-speech embedding vector and the entity width information with the word entity mention context embedding vector by using the fusion processing unit, obtain the attention score of each mention element in the word entity mention context embedding vector based on the multi-head attention mechanism for the word entity mention context embedding vector fused with the part-of-speech embedding vector, use the attention score of each mention element as the attention of the corresponding entity mention, and obtain the entity-level attention matrix by averaging the entity mention attention, use the entity-level attention matrix to represent the attention score of the corresponding entity to all entity mentions, and locate the key context of a given entity pair through the entity-level attention matrix, and obtain the local context embedding vector of the given entity pair according to the context; then, fuse the local context embedding vector with the entity global representation, entity type information and distance information between entities to obtain the context embedding representation of the target entity pair; then, group and feature fuse the context embedding representation of the target entity pair to obtain the entity pair representation, and then use the nonlinear activation function sigmoid to obtain the relationship probability of the given target entity pair.
7. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Threat intelligence information extraction method and system fusing multiple models
CN116049419A