Network security named entity recognition method and device based on graph propagation and dynamic attention
By adopting a method based on graph propagation and dynamic attention in the field of network security, building a heterogeneous graph model and combining a variety of deep learning technologies, the shortcomings of traditional named entity recognition technology in the processing of complex entity associations and contextual relationships are solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411100914.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-12
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-08-12
AI Technical Summary
Traditional named entity recognition technology is difficult to effectively deal with complex entity associations and contextual relationships in the field of network security, resulting in insufficient recognition accuracy and robustness.
Using a method based on graph propagation and dynamic attention, a heterogeneous graph model is constructed, and graph attention network propagation and aggregate feature representation is used, and character and word features are dynamically balanced through a gated fusion mechanism, combined with relative attention mechanism and dynamic convolutional layer, and finally, beam search is introduced in the double-layer CRF decoding to solve the global optimality.
It significantly improves the accuracy and robustness of named entity recognition, can effectively handle complex entity associations and contextual relationships, and improves the sensitivity and accuracy of entity recognition in network security text.
Smart Images

Figure CN119106683B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a network security named entity recognition method and device based on graph propagation and dynamic attention. Background Art
[0002] With the current trend of network development and the increasing number of network security threats, it is urgent to quickly and accurately identify key information from a large number of network security related texts. These texts usually include threat intelligence reports, security event logs, vulnerability descriptions, etc., which contain a large amount of entity information, such as malware names, attack techniques, system vulnerabilities, etc. The application of traditional named entity recognition (NER) technology in the field of network security faces some unique challenges.
[0003] The entity types in the cybersecurity field are far more complex than general texts, involving multi-level and nested entity structures. For example, an attack description may contain attackers, exploited tools, and affected system components, and there are complex associations between these entities. Entities in cybersecurity-related texts often do not follow fixed naming rules, resulting in unclear boundaries of entities. For example, the name of a software may only be recognized as an entity in a specific context. Effective entity recognition needs to consider long-range semantic dependencies and contextual relationships between entities, which are often difficult to handle in traditional models.
[0004] Based on the above situation, a network security named entity recognition method is proposed that can effectively utilize the association relationship between entities and dynamically adjust context information. By constructing a heterogeneous graph model, the dictionary boundary information and context semantic information can be fully utilized. The graph attention network is applied on the heterogeneous graph to propagate and aggregate the feature representation of characters and words. At the same time, the gate mechanism is used to dynamically balance the character features and word features. The relative attention module is input into the fused feature sequence to capture long-distance dependencies. Finally, beam search is introduced in the double-layer CRF decoding. The global optimal path is found as the prediction result through dynamic planning. The global sequence score is comprehensively considered to improve the accuracy and robustness of named entity recognition. Summary of the invention
[0005] Purpose of the invention: In view of the above problems, the present invention provides a network security named entity recognition method and device based on graph propagation and dynamic attention, constructs a heterogeneous graph model, introduces a dynamic attention mechanism, and introduces beam search in a two-layer decoding architecture, which effectively utilizes the relationship information between entities in network security texts; dynamically adjusts the model's attention to different parts according to the context, uses beam search to approximate the global optimum, avoids the error accumulation of greedy solutions, and can improve the accuracy and robustness of named entity recognition.
[0006] Technical solution: The present invention proposes a network security named entity recognition method based on graph propagation and dynamic attention, which includes the following steps:
[0007] Step 1: Preprocess and initialize the cybersecurity text data, and use the pre-trained language model to obtain the comprehensive feature output sequence H of each character c ;
[0008] Step 2: According to the network security dictionary, match the input text sequence with words to obtain a set of matching words, build a heterogeneous graph based on the character nodes of the character sequence and the matched word nodes, and mark the boundary relationship;
[0009] Step 3: Pre-train the words matched in step 2 to obtain pre-trained vectors; propagate and aggregate the feature representations of characters and words through the graph attention network on the constructed heterogeneous graph, integrate the word features of semantic and boundary information, and obtain the final representation of the word nodes in the heterogeneous graph;
[0010] Step 4: Use the gated fusion mechanism to dynamically fuse character features and word features, capture the contextual information of character and word sequences, and obtain the fused feature sequence M f ;
[0011] Step 5: Perform relative position encoding on the fused feature sequence, combine the dynamic convolution layer and relative attention mechanism to obtain the hidden state sequence after global context encoding;
[0012] Step 6: Introduce beam search in conditional random field decoding to calculate the hidden state sequence of the global context obtained in step 5 and perform sequence labeling to obtain the optimal labeled sequence.
[0013] Furthermore, the specific method of step 1 is:
[0014] Step 1.1: Preprocess the text data, including removing useless format information, error correction, and text standardization;
[0015] Step 1.2: Format the processed text data into an input format acceptable to the model and define the input text sequence X = {x1, x2, ..., x n}, where x i represents the i-th character, n is the number of characters, and X is a preprocessed single network security report, threat intelligence text, or a collection of texts from multiple data sources;
[0016] Step 1.3: For each character x in the text sequence i By querying the character map C, we can get the character embedding vector F xi =C(x i), where C is a character-to-vector mapping table, which is randomly initialized;
[0017] Step 1.4: Embed the characters into vector F xi Input into the pre-trained language model BERT to obtain character x i The corresponding character vector represents h i =BERT(F xi );
[0018] Step 1.5: Iterate through multiple encoder layers of the pre-trained model, each layer may adjust and optimize the feature representation, and finally obtain the feature output sequence H for each character c ={h1,h2,...,h n}, where h i Corresponding to x i A character vector representation of .
[0019] Furthermore, the specific method of step 2 is:
[0020] Step 2.1: Define the network security dictionary as Dict, initialize a prefix tree T, and set the root node to be empty;
[0021] Step 2.2: For each word in the cybersecurity dictionary Dict, starting from the root node, add each character of each word as a node to T, and mark the node corresponding to the last character of each word as the terminal node;
[0022] Step 2.3: Initialize an empty matching word set W = {};
[0023] Step 2.4: For the input text sequence X = {x1, x2, ..., x n}, where x i Represents the i-th character in the text. Starting from the first character of the input text sequence X, try to positively match a longer word on the T tree;
[0024] Step 2.4.1: If the match fails, add the current character as a new word w to W. Otherwise, add the longest matching word w to W.
[0025] Step 2.4.2: Move the cursor to the next unmatched character and continue the matching process in step 2.4 until the end of the input text X is reached;
[0026] Step 2.4.3: Return the final set of matching words W = {w1,w2,..,w m}, m is the total number of matched words, w m is the matched word;
[0027] Step 2.5: Define the heterogeneous graph G G =(V G ,E G ), node set V G Contains two node types character node set v c and word node set v w , E G For edge sets, set edge labels to B, M, F, S;
[0028] Step 2.6: For an input text sequence X, if character x i It is a word w j ∈W, add word node w j to v w , x i Add to v c , and connect the edges, setting the edge label to B;
[0029] Step 2.7: If character x i It is a word w j ∈W, add word node w j to v w , x i Add to v c , and connect the edges, setting the edge label to M;
[0030] Step 2.8: If character x i It is a word w j ∈W at the end: add word node w j to v w , x i Add to v c , and connect the edges, setting the edge label to F;
[0031] Step 2.9: If character x i It is a single-character word in W: i Add to v respectively w and v c , and connect the edges, setting the edge label to S;
[0032] Step 2.10: Traverse the input sequence X to obtain the heterogeneous graph G G =(V G ,E G ).
[0033] Furthermore, the specific method of step 3 is:
[0034] Step 3.1: For the words W = {w1,w2,..,w m}Use the pre-trained model to train and get the pre-trained vector HW = {h wi |w i ∈W};
[0035] Step 3.2: Initialize node representation, character node representation: Word nodes represent: where h c ∈H c is the pre-trained vector of the character, H c is the comprehensive feature output sequence in step 1, h w ∈HW is the pre-trained vector of the word;
[0036] Step 3.3: For heterogeneous graph G G =(V G ,E G ) in the node v, the self-attention mechanism is used to calculate the attention weight between nodes: the attention score is The attention weight is Where N(v) is the set of neighbor nodes of node v, is the representation vector of node v’s neighbor node u in the previous layer, l represents the number of layers, It is a learnable weight matrix used to linearly transform node representation. The representation vectors of nodes v and u are linearly transformed through the weight matrix, and the LeakyReLU activation function is used to normalize the result by softmax to obtain the attention weight of node v to neighbor node u.
[0037] Step 3.4: Using the attribute information of the edges in the heterogeneous graph structure, each edge attribute label represents an embedded vector where d r is the edge attribute embedding dimension, converting the original attention score Embedding vector with edge attributes Through the learnable transformation matrix Get the new attention score of the fused edge attributes Use the softmax function to convert the new attention score into the attention weight of the fused edge attribute
[0038] Step 3.5: In the lth layer of the graph attention network GAT, for each node v, calculate the weighted sum representation of its attention from neighboring nodes u and edge attributes: is the attention weight of node v to neighbor node u and fusion edge attributes;
[0039] Step 3.6: Perform multi-head attention fusion in is the value transformation matrix of the g-th attention head, is the attention weight of node v of the g-th attention head to the neighbor node u and the fused edge attribute. After fusing the multi-head attention, the final representation of node v in the l-th layer is obtained in represents the output linear transformation matrix, which concatenates the outputs of G attention heads and transforms them, and σ is a nonlinear activation function;
[0040] Step 3.7: After l layers of propagation, each word node will have a representation that comprehensively considers multi-hop neighbor information, and outputs a heterogeneous graph G G The final representation of the word node in
[0041] Furthermore, the specific method of step 4 is:
[0042] Step 4.1: Transform character features and word features into a new vector space t through a linear mapping and tanh nonlinear activation function respectively. c and t w , the calculation formula is as follows: c =tanh(H c W c +b c ), t w =tanh(H w W w +b w ), where W c 、b c , W w 、b w are all trainable parameters, H c is the character comprehensive feature output sequence in step 1, H w is the word feature sequence after graph propagation;
[0043] Step 4.2: Set t c and t w For splicing, after a linear mapping W p And the sigmoid activation function, get a gate value p between 0 and 1, the calculation formula is as follows: p = σ (([t c ;t w ])W p ), where W p Trainable parameters, σ is the sigmoid activation function;
[0044] Step 4.3: Use the gating weight p to t c and t w Perform weighted summation to obtain the final fusion feature sequence M f , the calculation formula is as follows: f =pt c+(1-p)t w , when p is close to 1, M f Mainly composed of character features c decision; when p is close to 0, M f Mainly composed of word features w Decide;
[0045] Furthermore, the specific method of step 5 is:
[0046] Step 5.1: In order to capture the relative position information between words / characters in the sequence, the input fusion feature M f Add relative position encoding PE, using relative position encoding based on sine / cosine function: Where pos is the position of the element in the sequence, I is the dimension index in the positional encoding vector, and d model is the dimension of the feature vector;
[0047] Step 5.2: Combine the position encoding PE with the fusion feature M f Add together to generate a feature representation M′ that enhances the position information f ;
[0048] Step 5.3: By converting M′ f They are mapped to Q (query), K (key), and V (value) respectively, and the calculation formula is: Q = M' f W Q ,K=M′ f W K ,V=M′ f W V , where Q, K, V are query, key and value vectors respectively, and W Q , W K , W V are the corresponding weight matrices respectively;
[0049] Step 5.4: Relative attention calculation, the calculation formula is as follows: (QK T ) ij represents the dot product operation of the query vector and the key vector of the i-th element and the j-th element in the sequence, pos i -pos j represents the relative position difference between the i-th element and the j-th element in the sequence, A is the trainable relative position attention bias matrix, and d k is the dimension of the key vector, Q, K, V are the query, key, and value vectors respectively;
[0050] Step 5.5: Use relative attention layers and introduce dynamic convolution layers for alternate stacking, add feedforward network FFN to enhance nonlinear modeling capabilities, use the output of each layer as the input of the next layer, and use residual connections and normalization to prevent gradient disappearance / explosion;
[0051] Step 5.6: After several layers of stacking, the context-encoded hidden state sequence U is output.
[0052] Furthermore, the specific method of step 6 is:
[0053] Step 6.1: For The hidden state sequence output by the context encoding module corresponds to the n character positions of the original text; the hidden state sequence is calculated using a linear layer as the normalized emission score: P (1) =UW (1) +b (1) , where W (1) and b (1) are the weights and biases of the linear layer, P (1) is a score matrix;
[0054] Step 6.2: Initialize the bundle size to R1. For each position i in the sequence, i = 1, 2, ..., n, according to Keep the R1 markers with the highest scores is the score matrix corresponding to each character position i of the original text, for each path y1:i (1) Calculate the cumulative score, including the emission score and the transfer score: in Indicates hidden state and Mark The probability of emission, Indicates from the mark Transfer to Mark probability;
[0055] Step 6.3: Prune. After each position i, keep the R1 paths with the highest scores.
[0056] Step 6.4: Repeat steps 6.2 and 6.3 until the entire sequence is processed, and output the labeled path with the highest score as the first layer prediction
[0057] Step 6.5: Predict using the first beam search Generate enhanced hidden sequence state U (2) : in represents the augmented vector at position i in the second layer CRF decoder, is the predicted token for position i in the first decoder Embedded representation of ;
[0058] Step 6.6: Put U (2) Input a linear layer to get the unnormalized emission score P (2) ;
[0059] Step 6.7: Initialize the second layer bundle size to R2, based on Keep the R2 highest-scoring markers for each position i For each path y1:i (2) , calculate the cumulative score, including the emission score and the transfer score, and perform pruning after each position i, retaining only the R2 paths with the highest cumulative score;
[0060] Step 6.8: Repeat steps 6.6 and 6.7 until the entire sequence is processed, and output the highest-scoring annotation path as the second-layer prediction
[0061] The present invention also discloses a network security named entity recognition device based on graph propagation and dynamic attention, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, the network security named entity recognition method based on graph propagation and dynamic attention is implemented.
[0062] Beneficial effects:
[0063] The present invention effectively utilizes the relationship information between entities in network security text by constructing a heterogeneous graph model. The heterogeneous graph can significantly enhance the semantic connection between entities, so that the model can not only identify isolated entities, but also accurately handle complex dependencies and associations between entities. The introduced dynamic attention mechanism can dynamically adjust the model's attention to different parts according to the context. The dynamic adjustment can improve the model's sensitivity to entity boundaries and entity type recognition, especially in the field of network security where context information is constantly changing. The beam search is introduced in the two-layer decoding architecture to first identify simple entities and then model the internal structure of complex entities. The beam search is used to approximately solve the global optimum to avoid the error accumulation of greedy solutions, which can improve the accuracy and robustness of named entity recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is the overall flow chart of the present invention;
[0065] Figure 2 is a flowchart of the character-word heterogeneous graph;
[0066] Figure 3 Flowchart for graph propagation on heterogeneous graphs;
[0067] Figure 4 Flowchart of named entity recognition for network security. DETAILED DESCRIPTION
[0068] The present invention is further explained below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0069] Step 1: Preprocess and initialize the cybersecurity text data, and use the pre-trained language model to obtain the comprehensive feature output sequence H of each character c , the specific method is:
[0070] Step 1.1: Preprocess the text data, including removing useless format information, error correction, and text standardization;
[0071] Step 1.2: Format the processed text data into an input format acceptable to the model and define the input text sequence X = {x1, x2, ..., x n}, where x i represents the i-th character, n is the number of characters, and X is a preprocessed single network security report, threat intelligence text, or a collection of texts from multiple data sources.
[0072] Step 1.3: For each character x in the text sequence i By querying the character map C, we can get the character embedding vector F xi =C(x i ), where C is a character-to-vector mapping table, which is randomly initialized.
[0073] Step 1.4: Embed the characters into vector F xi Input into the pre-trained language model BERT to obtain character x i The corresponding character vector represents h i =BERT(F xi ).
[0074] Step 1.5: Iterate through multiple encoder layers of the pre-trained model, each layer may adjust and optimize the feature representation, and finally obtain the feature output sequence H for each character c ={h1,h2,...,h n}, where h i Corresponding to x i A character vector representation of .
[0075] Step 2: According to the network security dictionary, the character nodes of the character sequence and the matched word nodes are used to construct a heterogeneous graph, and the boundary relationships are marked respectively. The specific method is as follows:
[0076] Step 2.1: Define the network security dictionary as Dict, initialize a prefix tree T, and set the root node to be empty.
[0077] Step 2.2: For each word in the cybersecurity dictionary Dict, starting from the root node, each character of each word is added as a node to T, and the node corresponding to the last character of each word is marked as the terminal node.
[0078] Step 2.3: Initialize an empty matching word set W = {}.
[0079] Step 2.4: For the input text sequence X = {x1, x2, ..., x n}, where x i Represents the i-th character in the text. Starting from the first character of the input text sequence X, try to positively match a longer word on the T tree.
[0080] Step 2.4.1: If the match fails, the current character is added to W as a new word w. Otherwise, the longest matching word w is added to W.
[0081] Step 2.4.2: Move the cursor to the next unmatched character and continue the above matching process until the end of the input text X is reached.
[0082] Step 2.4.3: Return the final set of matching words W = {w1,w2,..,w m}, m is the total number of matched words, w m The words that matched.
[0083] Step 2.5: Define the heterogeneous graph G G =(V G ,E G ), node set V G Contains two node types character node set v c and word node set v w , E G For the edge set, set the edge labels to B, M, F, S.
[0084] Step 2.6: For an input text sequence X, if character x i It is a word w j ∈W: add word node w j to v w , x i Add to v c , and connect the edges, setting the edge label to B.
[0085] Step 2.7: If character x i It is a word w j∈W middle part: add word node w j to v w , x i Add to v c , and connect the edges, setting the edge label to M.
[0086] Step 2.8: If character x i It is a word w j ∈W at the end: add word node w j to v w , x i Add to v c , and connect the edges, setting the edge label to F.
[0087] Step 2.9: If character x i It is a single-character word in W: i Add to v respectively w and v c , and connect the edges, setting the edge label to S.
[0088] Step 2.10: Traverse the input sequence X to obtain the heterogeneous graph G G =(V G ,E G ).
[0089] Step 3: The constructed heterogeneous graph is propagated and aggregated through the graph attention network to represent the features of characters and words, and the word features of semantic and boundary information are integrated. The specific method is as follows:
[0090] Step 3.1: For the words W = {w1,w2,..,w m}Use the pre-trained model to train and get the pre-trained vector HW = {h wi |w i ∈W}, where m represents the number of matched words.
[0091] Step 3.2: Initialize node representation, character node representation: Word nodes represent: where h c ∈H c is the pre-trained vector of the character, H c is the comprehensive feature output sequence in step 1, h w ∈HW is the pre-trained vector of the word.
[0092] Step 3.3: For heterogeneous graph G G =(V G ,E G ) in the node v, the self-attention mechanism is used to calculate the attention weight between nodes: the attention score is The attention weight is Where N(v) is the set of neighbor nodes of node v, is the representation vector of node v’s neighbor node u in the previous layer, l represents the number of layers, It is a learnable weight matrix used to linearly transform node representation. The representation vectors of nodes v and u are linearly transformed through the weight matrix, and the LeakyReLU activation function is used to normalize the result by softmax to obtain the attention weight of node v to neighbor node u.
[0093] Step 3.4: Using the attribute information of the edges in the heterogeneous graph structure, each edge attribute label represents an embedded vector where d r is the edge attribute embedding dimension, converting the original attention score Embedding vector with edge attributes Through the learnable transformation matrix Get the new attention score of the fused edge attributes Use the softmax function to convert the new attention score into the attention weight of the fused edge attribute
[0094] Step 3.5: In the lth layer of the graph attention network GAT, for each node v, calculate the weighted sum representation of its attention from neighboring nodes u and edge attributes: is the attention weight of node v to its neighbor node u and the fused edge attributes.
[0095] Step 3.6: Perform multi-head attention fusion in is the value transformation matrix of the g-th attention head, is the attention weight of node v of the g-th attention head to the neighbor node u and the fused edge attribute. After fusing the multi-head attention, the final representation of node v in the l-th layer is obtained in represents the output linear transformation matrix, which concatenates the outputs of G attention heads and transforms them, and σ is a nonlinear activation function.
[0096] Step 3.7: After l layers of propagation, each word node will have a representation that comprehensively considers multi-hop neighbor information, and outputs a heterogeneous graph G G The final representation of the word node in
[0097] Step 4: Use the gated fusion mechanism to dynamically fuse character features and word features, capture the contextual information of character and word sequences, and obtain the fused feature sequence. The specific method is as follows:
[0098] Step 4.1: Transform character features and word features into a new vector space t through a linear mapping and tanh nonlinear activation function respectively. c and t w , the calculation formula is as follows: c =tanh(H c W c +b c ), t w =tanh(H w W w +b w ), where Wc, bc, Ww, bw are all trainable parameters, H c is the character comprehensive feature output sequence in step 1, H w is the word feature sequence after graph propagation.
[0099] Step 4.2: Set t c and t w After concatenation, a linear mapping Wp and sigmoid activation function are used to obtain a gate value p between 0 and 1. The calculation formula is as follows: p = σ(([t c ;t w ])W p ), where σ is the sigmoid activation function.
[0100] Step 4.3: Use the gating weight p to t c and t w Perform weighted summation to obtain the final fusion feature sequence M, the calculation formula is as follows: M = pt c +(1-p)t w , when p is close to 1, M is mainly composed of character features t c When p is close to 0, M is mainly determined by word feature t w Decide.
[0101] Step 5: Combine the dynamic convolution layer and the relative attention mechanism to encode the fused feature sequence to obtain the hidden state representation of the global context. The specific method is:
[0102] Step 5.1: In order to capture the relative position information between words / characters in the sequence, the input fusion feature M f Add relative position encoding PE, using relative position encoding based on sine / cosine function: Where pos is the position of the element in the sequence, I is the dimension index in the positional encoding vector, and d model is the dimension of the feature vector.
[0103] Step 5.2: Combine the position encoding PE with the fusion feature M fAdd together to generate a feature representation M′ that enhances the position information f .
[0104] Step 5.3: By converting M′ f They are mapped to Q (query), K (key), and V (value) respectively, and the calculation formula is: Q = M' f W Q ,K=M′ f W K ,V=M′ f W V , where Q, K, V are query, key and value vectors respectively, and W Q , W K , W V are the corresponding weight matrices respectively.
[0105] Step 5.4: Relative attention calculation, the calculation formula is as follows: (QK T ) ij represents the dot product operation of the query vector and the key vector of the i-th element and the j-th element in the sequence, pos i -pos j represents the relative position difference between the i-th element and the j-th element in the sequence, A is the trainable relative position attention bias matrix, and d k is the dimension of the key vector, Q, K, and V are the query, key, and value vectors respectively.
[0106] Step 5.5: Use relative attention layers and introduce dynamic convolution layers for alternate stacking, add feedforward network FFN to enhance nonlinear modeling capabilities, use the output of each layer as the input of the next layer, and use residual connections and normalization to prevent gradient disappearance / explosion.
[0107] Step 5.6: After several layers of stacking, the context-encoded hidden state sequence U is output.
[0108] Step 6: Introduce beam search in conditional random field decoding to calculate the comprehensive feature information obtained in step 5 and perform sequence labeling to obtain the optimal labeling sequence. The specific method is:
[0109] Step 6.1: For The hidden state sequence output by the context encoding module corresponds to the n character positions of the original text; the hidden state sequence is calculated using a linear layer as the normalized emission score: P (1) =UW (1) +b (1) , where W (1) and b (1) are the weights and biases of the linear layer, P (1) is a score matrix;
[0110] Step 6.2: Initialize the bundle size to R1, and for each position i in the sequence, Keep the R1 markers with the highest scores For each path y1:i (1) Calculate the cumulative score, including the emission score and the transfer score: in Indicates hidden state and Mark The probability of emission, Indicates from the mark Transfer to Mark probability.
[0111] Step 6.3: Prune. After each position i, keep the R1 paths with the highest scores.
[0112] Step 6.4: Repeat steps 6.2 and 6.3 until the entire sequence is processed, and output the labeled path with the highest score as the first layer prediction
[0113] Step 6.5: Predict using the first beam search Generate enhanced hidden sequence state U (2) : in represents the augmented vector at position i in the second layer CRF decoder, is the predicted token for position i in the first decoder The embedded representation of .
[0114] Step 6.6: Put U (2) Input a linear layer to get the unnormalized emission score P (2) .
[0115] Step 6.7: Initialize the second layer bundle size to R2, based on Keep the R2 highest-scoring markers for each position i For each path y1:i (2) , calculate the cumulative score, including the emission score and the transfer score, and perform pruning after each position i, retaining only the R2 paths with the highest cumulative scores.
[0116] Step 6.8: Repeat steps 6.6 and 6.7 until the entire sequence is processed, and output the highest-scoring annotation path as the second-layer prediction
[0117] The above embodiments are only used to illustrate the technical ideas and features of the present invention, and the purpose is to enable professionals familiar with the technology to understand the embodiments of the present invention. The present invention is not limited to the above embodiments, and any equivalent transformation or improvement based on the core idea of the present invention should be included in the protection scope of the present invention.
Claims
1. A network security named entity recognition method based on graph propagation and dynamic attention, characterized in that: The steps include: Step 1: Preprocess and initialize the cybersecurity text data, and use the pre-trained language model to obtain the comprehensive feature output sequence H of each character c ; Step 2: According to the network security dictionary, match the input text sequence with words to obtain a set of matching words, build a heterogeneous graph based on the character nodes of the character sequence and the matched word nodes, and mark the boundary relationship; Step 2.1: Define the network security dictionary as Dict, initialize a prefix tree T, and set the root node to be empty; Step 2.2: For each word in the cybersecurity dictionary Dict, starting from the root node, add each character of each word as a node to T, and mark the node corresponding to the last character of each word as the terminal node; Step 2.3: Initialize an empty matching word set W = {}; Step 2.4: For the input text sequence X = {x1, x2, ..., x n }, i ranges from 1 to n, where x i Represents the i-th character in the text. Starting from the first character of the input text sequence X, try to positively match a longer word on the T tree; Step 2.4.1: If the match fails, add the current character as a new word w to W. Otherwise, add the longest matching word w to W. Step 2.4.2: Move the cursor to the next unmatched character and continue the matching process in step 2.4 until the end of the input text X is reached; Step 2.4.3: Return the final set of matching words W = {w1,w2,..,w m }, m is the total number of matched words, w m is the matched word; Step 2.5: Define the heterogeneous graph G G =(V G ,E G ), node set V G Contains two node types character node set v c and word node set v w , E G For edge sets, set edge labels to B, M, F, S; Step 2.6: For an input text sequence X, if character x i It is a word w j ∈W, add word node w j to v w , x i Add to v c , and connect the edges, setting the edge label to B; Step 2.7: If character x i It is a word w j ∈W, add word node w j to v w , x i Add to v c , and connect the edges, setting the edge label to M; Step 2.8: If character x i It is a word w j ∈W at the end: add word node w j to v w , x i Add to v c , and connect the edges, setting the edge label to F; Step 2.9: If character x i It is a single-character word in W: i Add to v respectively w and v c , and connect the edges, setting the edge label to S; Step 2.10: Traverse the input sequence X to obtain the heterogeneous graph G G =(V G ,E G ); Step 3: Pre-train the words matched in step 2 to obtain pre-trained vectors; propagate and aggregate the feature representations of characters and words through the graph attention network on the constructed heterogeneous graph, integrate the word features of semantic and boundary information, and obtain the final representation of the word nodes in the heterogeneous graph; Step 4: Use the gated fusion mechanism to dynamically fuse character features and word features, capture the contextual information of character and word sequences, and obtain the fused feature sequence M f ; Step 5: Perform relative position encoding on the fused feature sequence, combine the dynamic convolution layer and relative attention mechanism to obtain the hidden state sequence after global context encoding; Step 6: Introduce beam search in conditional random field decoding to calculate the hidden state sequence of the global context obtained in step 5 and perform sequence labeling to obtain the optimal labeled sequence.
2. The network security named entity recognition method based on graph propagation and dynamic attention according to claim 1 is characterized in that: The specific method of step 1 is: Step 1.1: Preprocess the text data, including removing useless format information, error correction, and text standardization; Step 1.2: Format the processed text data into an input format acceptable to the model and define the input text sequence X = {x1, x2, ..., x n }, i ranges from 1 to n, x i represents the i-th character, n is the number of characters, and X is a preprocessed single network security report, threat intelligence text, or a collection of texts from multiple data sources; Step 1.3: For each character x in the text sequence i By querying the character map C, we get the character embedding vector F xi =C(x i ), where C is a character-to-vector mapping table, which is randomly initialized; Step 1.4: Embed the characters into vector F xi Input into the pre-trained language model BERT to obtain character x i The corresponding character vector represents h i =BERT(F xi ); Step 1.5: Iterate through multiple encoder layers of the pre-trained model, each layer may adjust and optimize the feature representation, and finally obtain the feature output sequence H for each character c ={h1,h2,...,h n }, i ranges from 1 to n, where h i Corresponding to x i A character vector representation of .
3. The network security named entity recognition method based on graph propagation and dynamic attention according to claim 1 is characterized in that: The specific method of step 3 is: Step 3.1: For the words W = {w1,w2,..,w m }Use the pre-trained model to train and get the pre-trained vector of the word Step 3.2: Initialize node representation, character node representation: Word nodes represent: where h c ∈H c is the pre-trained vector of the character, H c is the comprehensive feature output sequence in step 1, h w ∈HW is the pre-trained vector of the word; Step 3.3: For heterogeneous graph G G =(V G ,E G ) in the node v, the self-attention mechanism is used to calculate the attention weight between nodes: the attention score is The attention weight is Where N(v) is the set of neighbor nodes of node v, is the representation vector of node v’s neighbor node u in the previous layer, l represents the number of layers, It is a learnable weight matrix used to linearly transform node representation. The representation vectors of nodes v and u are linearly transformed through the weight matrix, and the LeakyReLU activation function is used to normalize the result by softmax to obtain the attention weight of node v to neighbor node u. Step 3.4: Using the attribute information of the edges in the heterogeneous graph structure, each edge attribute label represents an embedded vector where d r is the edge attribute embedding dimension, converting the original attention score Embedding vector with edge attributes Through the learnable transformation matrix Get the new attention score of the fused edge attributes Use the softmax function to convert the new attention score into the attention weight of the fused edge attribute Step 3.5: In the lth layer of the graph attention network GAT, for each node v, calculate the weighted sum representation of its attention from neighboring nodes u and edge attributes: is the attention weight of node v to neighbor node u and fusion edge attributes; Step 3.6: Perform multi-head attention fusion in is the value transformation matrix of the g-th attention head, is the attention weight of node v of the g-th attention head to the neighbor node u and the fused edge attribute. After fusing the multi-head attention, the final representation of node v in the l-th layer is obtained in represents the output linear transformation matrix, which concatenates the outputs of G attention heads and transforms them, and σ is a nonlinear activation function; Step 3.7: After l layers of propagation, each word node will have a representation that comprehensively considers multi-hop neighbor information, and outputs a heterogeneous graph G G The final representation of the word node in 4. The network security named entity recognition method based on graph propagation and dynamic attention according to claim 3 is characterized in that: The specific method of step 4 is: Step 4.1: Transform the character features and word features into the new vector space t through a linear mapping and tanh nonlinear activation function respectively. c and t w , the calculation formula is as follows: c =tanh(H c W c +b c ), t w =tanh(H w W w +b w ), where W c , b c , W w , b w are all trainable parameters, H c is the character comprehensive feature output sequence in step 1, H w is the word feature sequence after graph propagation; Step 4.2: Set t c and t w For splicing, after a linear mapping W p And the sigmoid activation function, get a gate value p between 0 and 1, the calculation formula is as follows: p = σ (([t c ;t w ])W p ), where W p Trainable parameters, σ is the sigmoid activation function; Step 4.3: Use the gating weight p to t c and t w Perform weighted summation to obtain the final fusion feature sequence M f , the calculation formula is as follows: f =pt c +(1-p)t w , when p is close to 1, M f Mainly composed of character features c decision; when p is close to 0, M f Mainly composed of word features w Decide.
5. The network security named entity recognition method based on graph propagation and dynamic attention according to claim 1 is characterized in that: The specific method of step 5 is: Step 5.1: In order to capture the relative position information between words / characters in the sequence, the input fusion feature M f Add relative position encoding PE, using relative position encoding based on sine / cosine function: Where pos is the position of the element in the sequence, I is the dimension index in the positional encoding vector, and d model is the dimension of the feature vector; Step 5.2: Combine the position encoding PE with the fusion feature M f Add together to generate a feature representation M′ that enhances the position information f ; Step 5.3: By converting M′ f They are mapped to Q, K, and V respectively, and the calculation formula is: Q = M' f W Q ,K=M′ f W K ,V=M′ f W V , where Q, K, V are query, key and value vectors respectively, and W Q , W K , W V are the corresponding weight matrices respectively; Step 5.4: Relative attention calculation, the calculation formula is as follows: (QK T ) ij represents the dot product operation of the query vector and the key vector of the i-th element and the j-th element in the sequence, pos i -pos j represents the relative position difference between the i-th element and the j-th element in the sequence, A is the trainable relative position attention bias matrix, and d k is the dimension of the key vector, Q, K, V are the query, key, and value vectors respectively; Step 5.5: Use relative attention layers and introduce dynamic convolution layers for alternate stacking, add feedforward network FFN to enhance nonlinear modeling capabilities, use the output of each layer as the input of the next layer, and use residual connections and normalization to prevent gradient disappearance / explosion; Step 5.6: After several layers of stacking, the context-encoded hidden state sequence U is output.
6. The network security named entity recognition method based on graph propagation and dynamic attention according to claim 5 is characterized in that: The specific method of step 6 is: Step 6.1: For The hidden state sequence output by the context encoding module corresponds to the n character positions of the original text; A linear layer is applied to the hidden state sequence to calculate the normalized emission score: P (1) =UW (1) +b (1) , where W (1) and b (1) are the weights and biases of the linear layer, P (1) is a score matrix; Step 6.2: Initialize the bundle size to R1. For each position i in the sequence, i = 1, 2, ..., n, according to P i (1) Keep the R1 markers with the highest scores P i (1) is the score matrix corresponding to each character position i of the original text, for each path y1:i (1) Calculate the cumulative score, including the emission score and the transfer score: in Indicates hidden state and Mark The probability of emission, Indicates from the mark Transfer to Mark probability; Step 6.3: Prune. After each position i, keep the R1 paths with the highest scores. Step 6.4: Repeat steps 6.2 and 6.3 until the entire sequence is processed, and output the labeled path with the highest score as the first layer prediction Step 6.5: Predict using the first beam search Generate enhanced hidden sequence state U (2) : in represents the augmented vector at position i in the second layer CRF decoder, is the predicted token for position i in the first decoder Embedding representation of ; Step 6.6: Put U (2) Input a linear layer to get the unnormalized emission score P (2) ; Step 6.7: Initialize the second layer bundle size to R2, based on P i (2) Keep the R2 highest-scoring markers for each position i For each path y1:i (2) , calculate the cumulative score, including the emission score and the transfer score, and perform pruning after each position i, retaining only the R2 paths with the highest cumulative score; Step 6.8: Repeat steps 6.6 and 6.7 until the entire sequence is processed, and output the highest-scoring annotation path as the second-layer prediction 7. A network security named entity recognition device based on graph propagation and dynamic attention, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, the method for network security named entity recognition based on graph propagation and dynamic attention is implemented according to any one of claims 1 to 6.
Citation Information
Patent Citations
Chinese medical named entity recognition method based on improved graph attention network
CN115879473A