A method for named entity recognition in network data based on attention mechanism optimization

By introducing the Transformer-XL model of the BERT model into the network data naming entity recognition method, the self-attention of the position-content and content-position is increased, and the nonlinear relative distance is calculated, the problem of low accuracy of network security data naming entity recognition in the prior art is solved, and high-precision naming entity recognition is achieved.

CN119272770BActive Publication Date: 2025-05-16HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411190943.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2025-05-16
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

In the prior art, the network data naming entity recognition method does not take into account the characteristics of network security data, resulting in low recognition results accuracy.

Method used

Using the network data naming entity recognition method based on attention mechanism optimization, the self-attention of position-content and content-position is increased by introducing the BERT model, non-linear relative distance is calculated, and the dependence between position and content is enhanced.

Benefits of technology

It improves the accuracy of named entity recognition, can better capture the relationships between different entities, and supports high-precision named entity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119272770B_ABST
    Figure CN119272770B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for named entity recognition of network data based on attention mechanism optimization, and belongs to the technical field of pre-training model optimization of named entity recognition. The method solves the problem of low recognition result accuracy caused by failure to consider the characteristics of network security data in the traditional network data named entity recognition method in the prior art; the present invention gives an input sequence, inputs it into the BERT model, generates three embeddings and adds them, obtains the final input of the word, inputs it into the Transformer-XL model that introduces the BERT model, sets the basic matrix, introduces the content embedding matrix and the position embedding matrix, obtains the content embedding basic matrix and the position embedding basic matrix; obtains the attention mechanism score between any two words in the sentence, normalizes the sum of all attention mechanism scores, and obtains the normalized attention mechanism score. The present invention effectively improves the accuracy of named entity recognition and can be applied to entity recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a network data named entity recognition method, and in particular to a network data named entity recognition method based on attention mechanism optimization, belonging to the technical field of pre-training model optimization for named entity recognition. Background Art

[0002] Named entity recognition is the basis for establishing each piece of knowledge in the knowledge graph. In the past general field named entity recognition research, a method combining dictionaries and deep learning models was generally adopted. For example, unique dictionaries and rules are built for software entities. The dictionary content can come from the matrix structure ATT&CK and tree structure CAPEC knowledge base.

[0003] In the prior art, based on dictionary matching, the BERT-BiLSTM-CRF network model is used to train the labeled data to finally obtain the result of named entity recognition. However, under the above research framework, a pre-trained BERT model is often used as the embedding layer, and the above pre-trained model is not adjusted according to the characteristics of network security data, which will lead to a certain degree of loss of accuracy.

[0004] In summary, a method for named entity recognition of network data is needed that optimizes and improves the pre-trained model so as to improve the accuracy of the named entity recognition results. Summary of the invention

[0005] A brief overview of the present invention is provided below in order to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is merely to present certain concepts in a simplified form as a prelude to a more detailed description discussed later.

[0006] In view of this, in order to solve the problem of low recognition accuracy of traditional network data named entity recognition methods in the prior art due to failure to consider the characteristics of network security data, the present invention provides a network data named entity recognition method based on attention mechanism optimization.

[0007] The technical solution is as follows: A method for named entity recognition based on network data optimized by attention mechanism, comprising the following steps:

[0008] S1. Given an input sequence, feed it into the BERT model, generate three embeddings and add them together to get the final input of the word;

[0009] S2. Build a Transformer-XL model that introduces the BERT model, input the final input of the word, set the basic matrix, introduce the content embedding matrix and the position embedding matrix, and obtain the content embedding basic matrix and the position embedding basic matrix;

[0010] S3. Obtain the attention mechanism score between any two words in the sentence according to the content embedding basis matrix and the position embedding basis matrix, normalize the sum of all attention mechanism scores, and obtain the normalized attention mechanism score;

[0011] S4. According to the normalized attention mechanism score, random masking and attention relaxation are performed through the MLM model where the BERT model is located to realize named entity recognition of network data.

[0012] Furthermore, in S1, the input sequence is w, w=(x 1 ,x 2 ,……x n ), x i is the i-th input vector, i=1,2,…,n, and the three embeddings generated are word embedding h i , position embedding p i and segment embedding d i , add the three embeddings together to get the final input X of the word;

[0013] The element x in the final input X of the word i ' is expressed as:

[0014] x i '=h i +p i +d i .

[0015] Further, in said S2, the basic matrix includes a query matrix Q, a key matrix K and a value matrix V;

[0016] The query matrix Q is expressed as:

[0017] Q=XW Q

[0018] The bond matrix K is expressed as:

[0019] K=XW K

[0020] The value matrix V is expressed as:

[0021] V=XW V

[0022] Among them, W Q is the parameter matrix that needs to be learned for the query matrix Q, W Kis the parameter matrix that needs to be learned for the key matrix K, W V is the parameter matrix that needs to be learned for the value matrix V;

[0023] The content embedding matrix H is introduced into the query matrix Q, key matrix K and value matrix V. The content embedding matrix H is composed of word embedding h i The obtained content embedding basic matrix includes the content embedding query matrix Q c , content embedding key matrix K c and content embedding value matrix V c ;

[0024] Q c =HW q,c

[0025] Among them, W q,c Represents the content embedding query matrix Q c The parameter matrix to be learned;

[0026] K c =HW k,c

[0027] Among them, W k,c Represents the content embedding key matrix K c The parameter matrix to be learned;

[0028] Content embedding value matrix V c It is expressed as:

[0029] V c =HW v,c

[0030] Among them, W v,c Represents the content embedding value matrix V c The parameter matrix to be learned;

[0031] The position embedding matrix P is introduced into the query matrix Q and the key matrix K. The position embedding matrix P is composed of the position embedding p i The obtained position embedding basic matrix includes the position embedding query matrix Q r and the position embedding key matrix K r ;

[0032] Position embedding query matrix Q r It is expressed as:

[0033] Q r =PW q,r

[0034] Position embedding key matrix K r It is expressed as:

[0035] K r =PW k,r

[0036] Among them, W q,r Represents the position embedding query matrix Q r The parameter matrix to be learned, W k,r Represents the position embedding key matrix K r The parameter matrix to be learned.

[0037] Furthermore, in S3, the normalized attention mechanism score H 0 The calculation process is expressed as:

[0038]

[0039] in, is the attention mechanism score matrix of two words at the first position i and the second position j in the sentence The attention mechanism score for each element in , is the content embedding query matrix at the first position i in the sentence, is the content embedding key matrix at the second position j in the sentence, The key matrix is ​​the position embedding matrix for the first position i and the second position j at the word distance σ(i,j) in the sentence, is the position embedding query matrix of the first position i and the second position j at the word distance σ(i,j) in the sentence, T is the matrix transpose, · is the dot product, and softmax is the activation function;

[0040] The word distance σ(i,j) between the first position i and the second position j is expressed as:

[0041] σ(i,j)=[K'*sigmoid(ij)]

[0042] Among them, σ is the distance function, K' is a custom parameter, and [] represents the rounding operation.

[0043] The beneficial effects of the present invention are as follows: the present invention intends to study the pre-training model construction technology for network security named entity recognition, constructs a pre-training model suitable for the network security field according to the characteristics of network security data, and combines the subsequent fine-tuning process to perform more accurate entity extraction; the present invention enhances the dependency between position and content by increasing the self-attention of position-content and content-position. In the pre-training process of the Transformer-XL model that introduces the BERT model, the relative distance between entities is a relatively critical information. The introduction of relative distance can enable the model to better capture the relationship between different entities. Traditional methods generally use linear formulas to calculate the relative distance between entities. However, in actual application environments, the relationship between entities and relative distances are often not linearly reduced. The present invention introduces nonlinear functions and increases parameters in the relative distance calculation process to obtain a relative distance calculation method with better results, providing data support for high-precision named entity recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0045] Figure 1 A flowchart of a method for named entity recognition of network data based on optimization of attention mechanism;

[0046] Figure 2 This is a schematic diagram of the structure of the Transformer-XL model that introduces the BERT model. DETAILED DESCRIPTION

[0047] In order to make the technical solutions and advantages of the embodiments of the present invention more clearly understood, the exemplary embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than an exhaustive list of all the embodiments. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.

[0048] refer to Figure 1 and Figure 2 The present embodiment is described in detail, a method for identifying named entities in network data based on attention mechanism optimization, specifically comprising the following steps:

[0049] S1. Given an input sequence, feed it into the BERT model, generate three embeddings and add them together to get the final input of the word;

[0050] S2. Build a Transformer-XL model that introduces the BERT model, input the final input of the word, set the basic matrix, introduce the content embedding matrix and the position embedding matrix, and obtain the content embedding basic matrix and the position embedding basic matrix;

[0051] S3. Obtain the attention mechanism scores between any two words in the sentence according to the content embedding basis matrix and the position embedding basis matrix, normalize all the attention mechanism scores, and obtain the normalized attention mechanism scores;

[0052] S4. According to the normalized attention mechanism score, random masking and attention relaxation are performed through the MLM model where the BERT model is located to realize named entity recognition of network data.

[0053] Furthermore, in S1, the input sequence is w, w=(x 1 ,x 2 ,……x n ), x i is the i-th input vector, i=1,2,…,n, and the three embeddings generated are word embedding h i , position embedding p i and segment embedding d i , add the three embeddings together to get the final input X of the word;

[0054] The element x in the final input X of the word i ' is expressed as:

[0055] x i '=h i +p i +d i .

[0056] Specifically, the basic idea of ​​the attention mechanism is to i A weight is calculated and then the input vectors are weighted summed using these weights to produce a new representation.

[0057] Furthermore, in S2, the basic matrix includes a query matrix Q (Query), a key matrix K (Key) and a value matrix V (Value);

[0058] The query matrix Q is expressed as:

[0059] Q=XW Q

[0060] The bond matrix K is expressed as:

[0061] K=XW K

[0062] The value matrix V is expressed as:

[0063] V=XW V

[0064] Among them, W Q is the parameter matrix that needs to be learned for the query matrix Q, W K is the parameter matrix that needs to be learned for the key matrix K, W V is the parameter matrix that needs to be learned for the value matrix V;

[0065] The content embedding matrix H is introduced into the query matrix Q, key matrix K and value matrix V. The content embedding matrix H is composed of word embedding h i The obtained content embedding basic matrix includes the content embedding query matrix Q c , content embedding key matrix K c and content embedding value matrix V c ;

[0066] Q c =HW q,c

[0067] Among them, W q,c Represents the content embedding query matrix Q c The parameter matrix to be learned;

[0068] K c =HW k,c

[0069] Among them, W k,c Represents the content embedding key matrix K c The parameter matrix to be learned;

[0070] Content embedding value matrix V c It is expressed as:

[0071] V c =HW v,c

[0072] Among them, W v,c Represents the content embedding value matrix V c The parameter matrix to be learned;

[0073] The position (distance) embedding matrix P is introduced into the query matrix Q and the key matrix K. The position embedding matrix P is composed of the position embedding p i The obtained position embedding basic matrix includes the position embedding query matrix Q r and the position embedding key matrix K r ;

[0074] Position embedding query matrix Q r It is expressed as:

[0075] Q r =PWq,r

[0076] Position embedding key matrix K r It is expressed as:

[0077] K r =PW k,r

[0078] Among them, W q,r Represents the position embedding query matrix Q r The parameter matrix to be learned, W k,r Represents the position embedding key matrix K r The parameter matrix to be learned;

[0079] Specifically, the introduced content embedding matrix H represents the meaning of the word itself, which can affect the calculation of the attention mechanism. In the same sentence, the position of the word is also important, so the position embedding matrix P is introduced. However, the position itself has no meaning and is just a simple number, so it is not substituted into the value matrix V.

[0080] Furthermore, in S3, the normalized attention mechanism score H 0 The calculation process is expressed as:

[0081]

[0082] in, is the attention mechanism score matrix of two words at the first position i and the second position j in the sentence The attention mechanism score of each element in is actually the sum of the three attention mechanism scores calculated by the dot product method between the content-content, content-position and position-content of the two words. is the content embedding query matrix at the first position i in the sentence, is the content embedding key matrix at the second position j in the sentence, The key matrix is ​​the position embedding matrix for the first position i and the second position j at the word distance σ(i,j) in the sentence, is the position embedding query matrix of the first position i and the second position j at the word distance σ(i,j) in the sentence, T is the matrix transpose, · is the dot product, and softmax is the activation function;

[0083] The word distance σ(i,j) between the first position i and the second position j is expressed as:

[0084] σ(i,j)=[K'*sigmoid(ij)]

[0085] Where σ is the distance function, K' is a custom parameter which can be 1, and [] indicates a rounding operation.

[0086] Specifically, the traditional relative distance linear function σ(i,j)' with a threshold of k is expressed as:

[0087]

[0088] The present invention modifies the above formula into a nonlinear function, that is, the word distance σ(i, j) between the first position i and the second position j, and introduces a custom parameter K' and determines the optimal value. The normalized attention mechanism score H 0 It is a scoring matrix After an activation function softmax, it is multiplied by the content embedding value matrix V c The result obtained is that we can now calculate the attention score of any word in a sentence to another word. For example, in the sentence: this paper is about, the score of the attention mechanism for the word this is i=1, j=5, otherwise it is i=5, j=1, and finally a 5*5 attention mechanism matrix is ​​formed;

[0089] refer to Figure 2 , the blue part is the calculation layer of the attention mechanism structure, Add-Normalize is the residual addition and normalization processing layer, the self-attention layer of the attention mechanism structure is disassembled, matmul(A,V) is the matrix product of A and the value matrix V, A is the attention matrix, encoder is the encoder, mask is the random mask, scale is a one-step scaling operation to prevent the product of the query matrix Q and the key matrix K from being too large, the introduction point of the absolute position information K-position, so that the feedforward neural network in the decoder can better capture the absolute position information, which can be adjusted as needed, the absolute position information K-position, that is, the position embedding matrix P is injected into the random mask process, the absolute position information K-position is used to replace all elements in the matrix with a probability of 15%, or the absolute position information K-position is injected into the process after calculating all attention mechanism scores, and it can be simply scaled directly by K'. The residual addition and normalization processing layer Add-Normalize after the self-attention layer self-attention will be normalized.

[0090] Although the present invention has been described according to a limited number of embodiments, it will be apparent to those skilled in the art, with the benefit of the above description, that other embodiments may be envisioned within the scope of the invention thus described. In addition, it should be noted that the language used in this specification is selected primarily for readability and teaching purposes, rather than for explaining or defining the subject matter of the present invention. Therefore, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended claims. The disclosure of the present invention is illustrative, not restrictive, with respect to the scope of the present invention, which is defined by the appended claims.

Claims

1. A method for named entity recognition based on network data optimized by attention mechanism, characterized in that: The following steps are involved: S1. Given an input sequence, feed it into the BERT model, generate three embeddings and add them together to get the final input of the word; S2. Build a Transformer-XL model that introduces the BERT model, input the final input of the word, set the basic matrix, introduce the content embedding matrix and the position embedding matrix, and obtain the content embedding basic matrix and the position embedding basic matrix; S3. Obtain the attention mechanism score between any two words in the sentence according to the content embedding basis matrix and the position embedding basis matrix, normalize the sum of all attention mechanism scores, and obtain the normalized attention mechanism score; S4. According to the normalized attention mechanism score, random masking and attention unbundling are performed through the MLM model where the BERT model is located to realize network data named entity recognition; In S3, the calculation process of the normalized attention mechanism score H0 is expressed as: in, is the attention mechanism score matrix of two words at the first position i and the second position j in the sentence The attention mechanism score for each element in , is the content embedding query matrix at the first position i in the sentence, is the content embedding key matrix at the second position j in the sentence, The key matrix is ​​the position embedding matrix for the first position i and the second position j at the word distance σ(i,j) in the sentence, is the position embedding query matrix of the first position i and the second position j at the word distance σ(i,j) in the sentence, T is the matrix transpose, · is the dot product, and softmax is the activation function; The word distance σ(i,j) between the first position i and the second position j is expressed as: σ(i,j)=[K'*sigmoid(ij)] Among them, σ is the distance function, K' is a custom parameter, and [] represents the rounding operation.

2. According to the method of network data named entity recognition based on attention mechanism optimization according to claim 1, it is characterized in that: In S1, the input sequence is w, w=(x1,x 2, ……x n ), x i is the i-th input vector, i=1,2,…,n, and the three embeddings generated are word embedding h i , position embedding p i and segment embedding d i , add the three embeddings together to get the final input X of the word; The element x in the final input X of the word i ' is expressed as: x i '=h i +p i +d i 。 3. According to claim 2, a method for network data named entity recognition based on attention mechanism optimization is characterized in that: In S2, the basic matrix includes a query matrix Q, a key matrix K and a value matrix V; The query matrix Q is expressed as: Q=XW Q The bond matrix K is expressed as: K=XW K The value matrix V is expressed as: V=XW V Among them, W Q is the parameter matrix that needs to be learned for the query matrix Q, W K is the parameter matrix that needs to be learned for the key matrix K, W V is the parameter matrix that needs to be learned for the value matrix V; The content embedding matrix H is introduced into the query matrix Q, key matrix K and value matrix V. The content embedding matrix H is composed of word embedding h i The obtained content embedding basic matrix includes the content embedding query matrix Q c , content embedding key matrix K c and content embedding value matrix V c ; Q c =HW q,c Among them, W q,c Represents the content embedding query matrix Q c The parameter matrix to be learned; K c =HW k,c Among them, W k,c Represents the content embedding key matrix K c The parameter matrix to be learned; Content embedding value matrix V c It is expressed as: V c =HW v,c Among them, W v,c Represents the content embedding value matrix V c The parameter matrix to be learned; The position embedding matrix P is introduced into the query matrix Q and the key matrix K. The position embedding matrix P is composed of the position embedding p i The obtained position embedding basic matrix includes the position embedding query matrix Q r and the position embedding key matrix K r ; Position embedding query matrix Q r It is expressed as: Q r =PW q,r Position embedding key matrix K r It is expressed as: K r =PW k,r Among them, W q,r Represents the position embedding query matrix Q r The parameter matrix to be learned, W k,r Represents the position embedding key matrix K r The parameter matrix to be learned.

Citation Information

Patent Citations

  • Small sample named entity recognition method based on multiple tasks and prompt learning

    CN116151256A

  • Relationship classification method for texts in food safety field

    CN116578709A