A network threat intelligence information extraction method based on deep learning
By building a joint entity relationship extraction model based on BERT, the difficulty of information extraction in network threat intelligence is solved, and accurate extraction of network threat intelligence information and deep-level context information are achieved, and the prediction performance and computing efficiency of the model are improved.
Patent Information
- Application Number
- CN202510139203.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-08
AI Technical Summary
The prior art has difficulties in accurately mining and analyzing high-value intelligence knowledge from massive cyber threat intelligence reports, especially when dealing with complex sentences or long texts, the reliability of generating results is low and the difficulty of extracting triplets of overlapping entities.
The network threat intelligence information extraction method based on deep learning is adopted to build a joint entity relationship extraction model, and combine BERT's BiGRU encoder and non-autoregressive decoder to train the model through loss function to achieve accurate extraction of network threat intelligence information.
By capturing deep-level context information in network threat intelligence, the prediction performance of the entity relationship joint extraction model is improved, computing efficiency and resource utilization are improved, error propagation problems are solved, and the reliability of extraction results is improved.
Smart Images

Figure CN119578538B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a method for extracting network threat intelligence information based on deep learning. Background Art
[0002] Cyber Threat Intelligence (CTI) has gradually become an important component for resisting cyber attacks and improving security defense capabilities. Cyber Threat Intelligence aims to provide organizations with timely and actionable defense strategies by collecting, analyzing and sharing dynamic information about cyber threats. However, accurately mining and analyzing high-value intelligence knowledge from massive CTI reports remains a major challenge for cybersecurity professionals.
[0003] The core of threat intelligence mining lies in the accurate extraction and analysis of a large amount of information, including the action patterns of both attackers and defenders, attack tools, vulnerability exploits, and the organizational relationships behind the attackers. The effective management of this intelligence knowledge is of great significance for security professionals to make defensive decisions. In the early stages of development, entity recognition and relationship extraction mainly relied on pattern matching and statistical-based machine learning methods based on rules formulated by experts. However, these methods rely on artificial rules and features and are difficult to adapt to new fields. With the advancement of artificial intelligence, natural language processing (NLP) and deep learning technologies have developed rapidly, promoting significant progress in entity and relationship extraction tasks. Researchers developed a network named entity recognizer using pipeline extraction technology and proposed a deep learning model combining recurrent neural network (RNN) and conditional random field (CRF). These models can better capture the dependencies of text contexts, thereby improving the accuracy of entity recognition. After identifying the security entity, the researchers introduced a graph neural network model with an attention mechanism, which directly takes the full dependency tree as input and uses soft pruning technology to automatically learn how to distinguish useful related substructures in relationship extraction tasks. However, pipeline extraction technology has the problem of error propagation. Errors in upstream NER tasks will directly affect the accuracy of subsequent relation extraction. To solve the problem of error propagation, researchers have modeled information extraction as a sequence-to-sequence (Seq2Seq) entity-relation joint extraction problem in recent years, introducing an attention mechanism between end-to-end to capture word dependencies and ultimately generate sequential tags for text sequences. Despite this, this method still has problems with the reliability of the generated results when processing complex sentences or long texts, and faces difficulties in extracting triplets of overlapping entities. Summary of the invention
[0004] In view of the above-mentioned deficiencies in the prior art, the present invention provides a network threat intelligence information extraction method based on deep learning, which solves the problem that network threat intelligence is difficult to accurately extract.
[0005] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: a network threat intelligence information extraction method based on deep learning, comprising:
[0006] S1: Obtain threat text data based on network threat intelligence information;
[0007] S2: Construct entity relationship joint extraction model;
[0008] S3: inputting the threat text data into the entity relationship joint extraction model, and training it using a loss function to obtain a trained entity relationship joint extraction model;
[0009] S4: Analyze the network threat intelligence information using the trained entity relationship joint extraction model to obtain network threat intelligence information extraction results, thereby completing the extraction of network threat intelligence information.
[0010] The beneficial effects of the present invention are: combining the BERT-based BiGRU encoder and the non-autoregressive decoder to construct an entity relationship joint extraction model for predicting network threat intelligence and obtaining network threat intelligence information extraction results. In this way, deep context information in network threat intelligence can be captured to improve the prediction performance of the entity relationship joint extraction model; using a custom data set and manually annotating it, combined with BiGRU, can achieve optimal performance faster, and improve computing efficiency and resource utilization.
[0011] Furthermore, the entity relationship joint extraction model includes:
[0012] A bidirectional encoder layer, used to extract context-aware information of each sequence in the threat text data to obtain corresponding word embedding data;
[0013] A bidirectional recurrent layer, including two layers of GRU units consisting of a forward GRU and a reverse GRU, for obtaining forward and reverse information of the word embedding data to obtain a deep sequence vector;
[0014] A parallel decoding layer, including a non-autoregressive decoder, for fusing sentence information of the deep sequence vector to obtain embedded data;
[0015] The multi-layer perception layer is used to independently decode the entities and relationships embedded in the data to obtain the network threat intelligence information extraction results.
[0016] Furthermore, the expression of the depth sequence vector is:
[0017] ;
[0018] ;
[0019] ;
[0020] ;
[0021] in, represents a depth sequence vector, Represents the first data of the depth sequence vector, The second data representing the depth sequence vector, Indicates the final result of the current sequence after passing through a layer of BiGRU network. Represents the last data of the depth sequence vector, represents the hidden state of the forward GRU unit, represents the hidden state of the backward GRU, represents the computational function of the forward GRU unit, Indicates The feature representation of the input sequence, Indicates The hidden state of the input sequence, represents the computational function of the forward GRU unit, Indicates The hidden state of the input sequence.
[0022] Further, the non-autoregressive decoder comprises:
[0023] A multi-head self-attention sub-layer, used to analyze the relationship between the deep sequence vectors to obtain a query vector;
[0024] A multi-head cross attention sub-layer, for fusing sentence information based on the query vector and the projection of the deep sequence vector to obtain a fused triple;
[0025] The feed-forward neural network sublayer is used to transform the fused triples to obtain embedded data.
[0026] Furthermore, the expression of the fused triple is:
[0027] ;
[0028] ;
[0029] ;
[0030] ;
[0031] in, represents the fusion triple, Represents the calculation function of the multi-head attention mechanism, represents the output of the previous multi-head self-attention sublayer, represents a depth sequence vector, It means concatenating the outputs of all attention heads. represents the first attention head, Indicates A head of attention, Indicates A head of attention, represents the calculation function of the attention mechanism, , , and All represent trainable parameters, , and Both represent initialization to learn embedding, express function, express The dimension of the vector, represents the initial embedding function, Indicates the number of fused triplets.
[0032] Furthermore, the network threat intelligence information extraction result includes the starting position of the head entity, the ending position of the head entity, the starting position of the tail entity, and the ending position of the tail entity, and the expressions thereof are respectively:
[0033] ;
[0034] ;
[0035] ;
[0036] ;
[0037] ;
[0038] in, represents the probability distribution of the starting position of the head entity of the predicted triple, represents the probability distribution of the end position of the head entity of the predicted triple, represents the probability distribution of the starting position of the tail entity of the predicted triple, represents the probability distribution of the end position of the tail entity of the predicted triple, express function, , , and represents the learnable parameters, Represents the deep sequence vector data after triple query conversion, represents a depth sequence vector, , , , , , , and represents the learnable parameters, Represents the relationship label, represents learnable parameters.
[0039] Furthermore, the expression of the bipartite matching loss function is:
[0040] ;
[0041] ;
[0042] in, represents the bipartite matching loss result, represents a real set of triples, Indicates the results of network threat intelligence information extraction. represents the arrangement of elements with the lowest cost, Indicates the probability that the extraction result is a certain relationship category, represents the length of the arrangement space, express relation categories of real triples, Indicates the probability that the extraction result is the starting position of the head entity, express The starting position of the head entity of the real triple, Indicates the probability that the extraction result is the end position of the head entity, express The end position of the head entity of the real triple, Indicates the probability that the extraction result is the starting position of the tail entity, represents the space of all permutations of length N, express The starting position of the tail entity of a real triple, Indicates the probability that the extraction result is the end position of the tail entity, express The end position of the tail entity of a real triple, represents the function that minimizes the objective function. represents the pairwise matching cost, Indicates True triples, express Permuted prediction triplets; where , , Indicates The probability of predicting the relation category of the triple, Indicates The probability distribution of the starting position of the head entity of the predicted triple, Indicates The probability distribution of the end position of the head entity of the predicted triple, Indicates The probability distribution of the starting position of the tail entity of the predicted triple, Indicates The probability distribution of the end position of the tail entity of a predicted triple. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] This specification will be further described in the form of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same number represents the same structure, wherein:
[0044] Figure 1 This is an exemplary flowchart of a method for extracting network threat intelligence information based on deep learning according to some embodiments of this specification. DETAILED DESCRIPTION
[0045] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0046] Example
[0047] Figure 1 This is an exemplary flow chart of a method for extracting network threat intelligence information based on deep learning according to some embodiments of this specification. Figure 1 As shown, the process includes the following steps. In some embodiments, the process can be executed by a processor.
[0048] S1: Obtain threat text data based on network threat intelligence information.
[0049] Cyber threat intelligence information is unstructured cybersecurity threat intelligence data in PDF format.
[0050] Threat text data is text data that reflects the network threat situation. For example, as shown in Table 1, network threat intelligence information can include text with entity-relationship triples. Specifically, entity categories can include operating system (OS), malware (MAL), tools (TOO), threat actors (THR), activities (CAM), vulnerabilities (VUL), mitigation measures (MIT), attack techniques (ATT), groups (GRO), consequences (CON) and organizations (ORG), etc. The corresponding network threat intelligence explanations of entity categories are shown in Table 2.
[0051] Table 1 Textual relation table of entity-relation triples
[0052] text Triple Normal apt20 regularly employs mimikatz toobtain credentials from accounts with elevated privileges. (apt20, THR / TOO / uses, mimikatz) Overlap we attribute the campaign, referred to as spoiledlegacy, to theluckymouse apt group (also known asemissarypanda and apt27). (spoiledlegacy, CAM / THR / attributed_to,luckymouse aptgroup)(emissarypanda, GRO / GRO / Alias_of,luckymouse aptgroup)(apt27, GRO / GRO / Alias_of,luckymouse apt group)
[0053] Table 2 Explanation of network threat intelligence corresponding to entity categories
[0054] Entity Class Specific explanation MAL Unwanted programs or viruses TOO Security testing and attack equipment THR Threat actors launching cyber attacks CAM Cyber attacks or incidents ATT Methods and strategies for launching attacks GRO The team that organized the attack CON Impact or damage of cyber incidents ORG Organizations that may be attacked OS Basic software environment of the device VUL Security flaws in the system MIT Safety solutions to reduce risks
[0055] In some embodiments, the processor can perform data cleaning and preprocessing on the network threat intelligence information to obtain sentence-level text data as threat text data; wherein the threat text data includes multiple sentences , represents sentences in threat text data, Indicates the first word in a sentence. Indicates the second word in the sentence. Indicates the first words, Indicates the last word in a sentence.
[0056] S2: Build an entity relationship joint extraction model.
[0057] The entity-relationship joint extraction model is a neural network model used to analyze entities and corresponding relationships of threat text data. There can be many types of entity-relationship joint extraction models. For example, the type of entity-relationship joint extraction model can include a recurrent neural network model, etc.
[0058] In some embodiments, the input of the entity-relationship joint extraction model may be threat text data, and the output of the entity-relationship joint extraction model may be a network threat intelligence information extraction result.
[0059] In some embodiments, the structure of the entity-relationship joint extraction model is as follows:
[0060] The entity relationship joint extraction model includes a bidirectional encoder representation layer, a bidirectional gated loop layer, a parallel decoding layer, and a multi-layer perception layer. The output of the bidirectional encoder representation layer is used as the input of the bidirectional gated loop layer, the output of the bidirectional gated loop layer is used as the input of the parallel decoding layer, the output of the parallel decoding layer is used as the input of the multi-layer perception layer, and the output of the multi-layer perception layer is used as the final output of the entity relationship joint extraction model.
[0061] The bidirectional encoder layer is used to extract context-aware information of each sequence in the threat text data to obtain corresponding word embedding data. The input of the bidirectional encoder layer may include threat text data, and the output may include word embedding data.
[0062] In some embodiments, the bidirectional encoder layer can be a BERT model built based on a bidirectional encoder of Transformers.
[0063] Word embedding data is word embedding vector data that contains context information for each sequence in a sentence. For example, word embedding data can be represented as , ,in, represents word embedding data, Represents the first data in the word embedding vector, Represents the second data in the word embedding vector, Represents the last data in the word embedding vector, represents the set of real numbers, Indicates the length of the sentence, Represents the number of hidden units in the BERT model.
[0064] In some embodiments, the expression of word embedding data may be:
[0065] ;
[0066] in, represents word embedding data, represents the BERT model, Represents a sentence in the threat text data.
[0067] The bidirectional circulation layer is used to obtain the forward and reverse information of the word embedding data to obtain a deep sequence vector. The input of the bidirectional circulation layer may include the word embedding data, and the output may include the deep sequence vector.
[0068] In some embodiments, the processor may construct a bidirectional recurrent layer based on a BiGRU network comprising two layers of GRU units, wherein each GRU unit is composed of two gating units, namely an update gate and a reset gate.
[0069] In some embodiments, the expression of the GRU unit output data is:
[0070] ;
[0071] ;
[0072] in, Represents the GRU unit Output data for an input sequence, represents the update gate, represents the element-by-element product between vectors of the same dimension, Represents the GRU unit The output data of the input sequence is Indicates candidate states for the input sequence, represents the weight matrix input to the candidate state, Indicates The input data of the input sequence, Indicates The weight matrix from the hidden state of the input sequence to the candidate state, Represents the reset gate, Bias vector representing the candidate state.
[0073] A deep sequence vector is a sequence vector that contains more semantic features of a sentence. For example, a deep sequence vector can be represented as , .
[0074] In some embodiments, the expression of the depth sequence vector may be:
[0075] ;
[0076] ;
[0077] ;
[0078] ;
[0079] in, represents a depth sequence vector, Represents the first data of the depth sequence vector, The second data representing the depth sequence vector, Indicates the final result of the current sequence after passing through a layer of BiGRU network. Represents the last data of the depth sequence vector, represents the hidden state of the forward GRU unit, represents the hidden state of the backward GRU, represents the computational function of the forward GRU unit, Indicates The feature representation of the input sequence, Indicates The hidden state of the input sequence, represents the computational function of the forward GRU unit, Indicates The hidden state of the input sequence.
[0080] The parallel decoding layer is used to fuse the sentence information of the depth sequence vector to obtain embedded data. The input of the parallel decoding layer may include the depth sequence vector, and the output may include embedded data.
[0081] In some embodiments, the processor may construct a parallel decoding layer using a non-autoregressive decoder; wherein the non-autoregressive decoder is used to determine the size (N) of the generated target set, and the input of the decoder is initialized by initializing N learnable embeddings, and each sentence shares the same learnable embedding.
[0082] The embedded data is the deep sequence vector data after the triple query is converted. For example, the embedded data can be represented as .
[0083] In some embodiments, the parallel decoding layer may include a non-autoregressive decoder. The non-autoregressive decoder may include a multi-head self-attention sublayer, a multi-head cross-attention sublayer, and a feed-forward neural network sublayer.
[0084] In some embodiments, the non-autoregressive decoder may be composed of M transformer blocks, each transformer block including a multi-head self-attention sublayer, a multi-head cross-attention sublayer, and a feed-forward neural network sublayer.
[0085] The multi-head self-attention sub-layer is used to analyze the relationship between the deep sequence vectors to obtain a query vector.
[0086] The query vector is the output projection of the multi-head self-attention sub-layer.
[0087] In some embodiments, the processor may utilize a multi-head self-attention sub-layer to analyze the relationships among the triplets in the deep sequence vector to obtain a query vector.
[0088] In some embodiments, the processor may use the encoder-based output projection as the key (K) and value (V) of the multi-head criss-cross attention sub-layer.
[0089] The multi-head cross attention sub-layer is used to fuse sentence information based on the query vector and the projection of the deep sequence vector to obtain a fused triplet.
[0090] The fused triplet is the triplet after the sentence information is fused.
[0091] In some embodiments, the expression of the fused triple may be:
[0092] ;
[0093] ;
[0094] ;
[0095] ;
[0096] in, represents the fusion triple, Represents the calculation function of the multi-head attention mechanism, represents the output of the previous multi-head self-attention sublayer, represents a depth sequence vector, It means concatenating the outputs of all attention heads. represents the first attention head, Indicates A head of attention, Indicates A head of attention, represents the calculation function of the attention mechanism, , , and All represent trainable parameters, , and Both represent initialization to learn embedding, express function, express The dimension of the vector, represents the initial embedding function, Indicates the number of fused triplets.
[0097] The feed-forward neural network sublayer is used to transform the fused triples to obtain embedded data.
[0098] The multi-layer perception layer is used to independently decode the entities and relationships of the embedded data to obtain the network threat intelligence information extraction results. The input of the multi-layer perception layer may include the embedded data, and the output may include the network threat intelligence information extraction results.
[0099] In some embodiments, the processor can use an MLP network to construct a multi-layer perception layer; wherein the MLP network is used to independently decode the entities and relationships of the embedded data, and predict the relationship label, the starting position of the head entity, the ending position of the head entity, the starting position of the tail entity, and the ending position of the tail entity through a softmax classifier.
[0100] The network threat intelligence information extraction result is a triple extraction result reflecting the network threat situation in the threat text data. For example, the network threat intelligence information extraction result may include an entity-relationship triple with a starting position of a head entity, an ending position of a head entity, a starting position of a tail entity, and an ending position of a tail entity.
[0101] In some embodiments, the expression of the network threat intelligence information extraction result can be:
[0102] ;
[0103] ;
[0104] ;
[0105] ;
[0106] ;
[0107] in, represents the probability distribution of the starting position of the head entity of the predicted triple, represents the probability distribution of the end position of the head entity of the predicted triple, represents the probability distribution of the starting position of the tail entity of the predicted triple, represents the probability distribution of the end position of the tail entity of the predicted triple, express function, , , and represents the learnable parameters, Represents the deep sequence vector data after triple query conversion, represents a depth sequence vector, , , , , , , and represents the learnable parameters, Represents the relationship label, represents learnable parameters.
[0108] S3: Input the threat text data into the entity relationship joint extraction model, and train it using a loss function to obtain a trained entity relationship joint extraction model.
[0109] In some embodiments, the entity relationship joint extraction model can be obtained by training multiple labeled training samples. For example, multiple labeled training samples can be input into the initial entity relationship joint extraction model, and the loss function can be constructed by using the labels and the results of the initial entity relationship joint extraction model. The parameters of the initial entity relationship joint extraction model can be iteratively updated by gradient descent or other methods based on the binary matching loss function. When the preset conditions are met, the model training is completed and a trained entity relationship joint extraction model is obtained. The preset conditions can be that the binary matching loss function converges, the number of iterations reaches a threshold, etc.
[0110] In some embodiments, the training samples may include historical threat text data. The labels may be corresponding real triples. The labels may be manually annotated.
[0111] In some embodiments, the specific content of the label is shown in Table 3 and Table 4.
[0112] Table 3 Real triple table
[0113] Head Entity relation Tail Entity CAM CAM / ORG / target ORG THR THR / TOO / uses TOO GRO GRO / ORG / target ORG MAL MAL / TOO / includes TOO GRO GRO / MAL / associate MAL THR THR / ATT / uses ATT MAL MAL / ORG / target ORG ATT ATT / TOO / uses TOO ATT ATT / MAL / uses MAL THR THR / ORG / target ORG GRO GRO / TOO / uses TOO CAM CAM / THR / attributed_to THR CAM CAM / TOO / uses TOO GRO GRO / ATT / uses ATT MAL MAL / MAL / Alias_of MAL ATT ATT / VUL / exploits VUL
[0114] Table 4 Real triple table
[0115] Head Entity relation Tail Entity CAM CAM / MAL / uses MAL CAM CAM / ATT / uses ATT GRO GRO / GRO / Alias_of GRO GRO GRO / GRO / includes GRO GRO GRO / THR / includes THR HTR THR / VUL / exploits VUL MOTH MAL / VUL / exploits VUL TO ATT / ATT / causes TO TO ATT / CON / causes CON MIT MIT / TOO / mitigates TOO CAM CAM / VUL / exploits VUL MOTH MAL / CON / causes CON VUL VUL / CON / causes CON MIT MIT / MAL / mitigates MOTH MIT MIT / ATT / mitigates TO Olympics OS / VUL / exploits VUL
[0116] In some embodiments, the expression of the bipartite matching loss function may be:
[0117] ;
[0118] ;
[0119] in, represents the bipartite matching loss result, represents a real set of triples, Indicates the results of network threat intelligence information extraction. represents the arrangement of elements with the lowest cost, Indicates the probability that the extraction result is a certain relationship category, represents the length of the arrangement space, express relation categories of real triples, Indicates the probability that the extraction result is the starting position of the head entity, express The starting position of the head entity of the real triple, Indicates the probability that the extraction result is the end position of the head entity, express The end position of the head entity of the real triple, Indicates the probability that the extraction result is the starting position of the tail entity, represents the space of all permutations of length N, express The starting position of the tail entity of a real triple, Indicates the probability that the extraction result is the end position of the tail entity, express The end position of the tail entity of a real triple, represents the function that minimizes the objective function. represents the pairwise matching cost, Indicates True triples, express Permuted prediction triplets; where , , Indicates The probability of predicting the relation category of the triple, Indicates The probability distribution of the starting position of the head entity of the predicted triple, Indicates The probability distribution of the end position of the head entity of the predicted triple, Indicates The probability distribution of the starting position of the tail entity of the predicted triple, Indicates The probability distribution of the end position of the tail entity of a predicted triple.
[0120] In some embodiments, the processor may obtain the pairwise matching cost based on the real triple set and the network threat intelligence information extraction result, so as to analyze the computational efficiency of the entity relationship joint extraction model:
[0121] ;
[0122] in, represents the pairwise matching cost.
[0123] S4: Analyze the network threat intelligence information using the trained entity relationship joint extraction model to obtain network threat intelligence information extraction results, thereby completing the extraction of network threat intelligence information.
[0124] In some embodiments, the processor can input the threat text data extracted from the network threat intelligence information into a trained entity-relationship joint extraction model, use a bidirectional encoder layer to extract the context-aware information of each sequence in the threat text data, and obtain the corresponding word embedding data, use a bidirectional recurrent layer to obtain the forward and reverse information of the word embedding data, and obtain a deep sequence vector, use a parallel decoding layer to fuse the sentence information of the deep sequence vector to obtain embedded data, use a multi-layer perception layer to independently decode the entities and relationships of the embedded data, obtain the network threat intelligence information extraction results, and complete the extraction of network threat intelligence information.
[0125] In some embodiments of this specification, a BERT-based BiGRU encoder and a non-autoregressive decoder are combined to construct an entity relationship joint extraction model for predicting network threat intelligence and obtaining network threat intelligence information extraction results. In this way, deep contextual information in network threat intelligence can be captured to improve the prediction performance of the entity relationship joint extraction model; using custom data sets and manual annotations combined with BiGRU can achieve optimal performance faster and improve computing efficiency and resource utilization.
Claims
1. A network threat intelligence information extraction method based on deep learning, characterized in that: include: S1: Obtain threat text data based on network threat intelligence information; S2: Construct entity relationship joint extraction model; S3: Input the threat text data into the entity relationship joint extraction model, and train it using the loss function to obtain a trained entity relationship joint extraction model; wherein the entity relationship joint extraction model includes: A bidirectional encoder layer, used to extract context-aware information of each sequence in the threat text data to obtain corresponding word embedding data; A bidirectional recurrent layer, including two layers of GRU units consisting of a forward GRU and a reverse GRU, for obtaining forward and reverse information of the word embedding data to obtain a deep sequence vector; A parallel decoding layer, including a non-autoregressive decoder, for fusing sentence information of the deep sequence vector to obtain embedded data; Multi-layer perception layer, used to independently decode entities and relationships embedded in data to obtain network threat intelligence information extraction results; S4: Analyze the network threat intelligence information using the trained entity relationship joint extraction model to obtain network threat intelligence information extraction results, thereby completing the extraction of network threat intelligence information.
2. The method for extracting network threat intelligence information based on deep learning according to claim 1 is characterized in that: The expression of the depth sequence vector is: ; ; ; ; in, represents a depth sequence vector, Represents the first data of the depth sequence vector, The second data representing the depth sequence vector, Indicates the final result of the current sequence after passing through a layer of BiGRU network. Represents the last data of the depth sequence vector, represents the hidden state of the forward GRU unit, represents the hidden state of the backward GRU, represents the computation function of the forward GRU unit, Indicates The feature representation of the input sequence, Indicates The hidden state of the input sequence, represents the computation function of the forward GRU unit, Indicates The hidden state of the input sequence.
3. The method for extracting network threat intelligence information based on deep learning according to claim 1 is characterized in that: The non-autoregressive decoder comprises: A multi-head self-attention sub-layer, used to analyze the relationship between the deep sequence vectors to obtain a query vector; A multi-head cross attention sub-layer, for fusing sentence information based on the query vector and the projection of the deep sequence vector to obtain a fused triple; The feed-forward neural network sublayer is used to transform the fused triples to obtain embedded data.
4. The method for extracting network threat intelligence information based on deep learning according to claim 3 is characterized in that: The expression of the fusion triple is: ; ; ; ; in, represents the fusion triple, Represents the calculation function of the multi-head attention mechanism, represents the output of the previous multi-head self-attention sublayer, represents a depth sequence vector, It means concatenating the outputs of all attention heads. represents the first attention head, Indicates Attention head, Indicates Attention head, represents the calculation function of the attention mechanism, , , and All represent trainable parameters, , and Both represent initialization to learn embedding, express function, express The dimension of the vector, represents the initial embedding function, Indicates the number of fused triplets.
5. The method for extracting network threat intelligence information based on deep learning according to claim 1 is characterized in that: The network threat intelligence information extraction result includes the starting position of the head entity, the ending position of the head entity, the starting position of the tail entity, and the ending position of the tail entity, and the expressions thereof are respectively: ; ; ; ; ; in, represents the probability distribution of the starting position of the head entity of the predicted triple, represents the probability distribution of the end position of the head entity of the predicted triple, represents the probability distribution of the starting position of the tail entity of the predicted triple, represents the probability distribution of the end position of the tail entity of the predicted triple, express function, , , and represents the learnable parameters, Represents the deep sequence vector data after triple query conversion, represents a depth sequence vector, , , , , , , and represents the learnable parameters, Represents the relationship label, represents learnable parameters.
6. The method for extracting network threat intelligence information based on deep learning according to claim 1 is characterized in that: The expression of the loss function is: ; ; in, Indicates the loss result, represents a real set of triples, Indicates the results of network threat intelligence information extraction. represents the arrangement of elements with the lowest cost, Indicates the probability that the extraction result is a certain relationship category, represents the length of the arrangement space, express relation categories of real triples, Indicates the probability that the extraction result is the starting position of the head entity, express The starting position of the head entity of the real triple, Indicates the probability that the extraction result is the end position of the head entity, express The end position of the head entity of the real triple, Indicates the probability that the extraction result is the starting position of the tail entity, represents the space of all permutations of length N, express The starting position of the tail entity of a real triple, Indicates the probability that the extraction result is the end position of the tail entity, express The end position of the tail entity of a real triple, represents the function that minimizes the objective function. represents the pairwise matching cost, Indicates True triples, express Permuted prediction triplets; where , , Indicates The probability of predicting the relation category of the triple, Indicates The probability distribution of the starting position of the head entity of the predicted triple, Indicates The probability distribution of the end position of the head entity of the predicted triple, Indicates The probability distribution of the starting position of the tail entity of the predicted triple, Indicates The probability distribution of the end position of the tail entity of a predicted triple.
Citation Information
Patent Citations
Network threat intelligence relation triple combined extraction method based on deep learning
CN117787403A