A data augmentation method for network security named entity recognition based on pre-training model
By using a data augmentation method based on pretrained model in network security named entity recognition tasks, the problem of text semantic errors in the prior art is solved, and more accurate named entity recognition and better model generalization capabilities are achieved.
Patent Information
- Application Number
- CN202411190945.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-08-28
AI Technical Summary
Existing methods for augmenting data of network security named entity recognition can easily lead to text semantic errors and fail to fully consider context semantics.
The network security named entity identification data augmentation method based on pre-trained models is adopted. By performing sentence-by-section processing of the input sequence, truncating fragments beyond the preset length, masking the fragment set, and using the BERT model to predict, the augmented data set is generated.
By considering context semantics, reducing the introduction of noise, improving text diversity, and reducing overfitting, the accuracy of named entity recognition and the generalization ability of pre-trained models are improved.
Smart Images

Figure CN119204011B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a network security named entity recognition data augmentation method, in particular to a network security named entity recognition data augmentation method based on a pre-training model, and belongs to the technical field of network data security. Background Art
[0002] In the field of natural language processing (NLP), pre-trained models have become an extremely important technology. Especially in the past few years, such models have shown excellent performance in various language understanding and generation tasks. A common method is to directly fine-tune the pre-trained model to adapt to specific downstream tasks, which has achieved good results in previous studies. However, the above method does not take into account the adaptability to the domain, but assumes that the characteristics of the domain are learned during the fine-tuning process.
[0003] For the named entity recognition (NER) task, the synonym replacement method is a simple and effective augmentation method, which mainly replaces the entity part. That is, for each entity that appears in the text segment, a binomial distribution with a replacement probability of p (generally 0.1-0.7) is used to randomly decide whether it should be replaced. If it is decided to replace it, another entity of the same entity type as the named entity is randomly selected for replacement. Figure 3 , the corresponding BIO label sequence should also be replaced together. The specific implementation of replacing the entity part is as follows: first traverse all the original data, record the text sequence and the corresponding label sequence, and record the entities by category. Then traverse the text sequence. Whenever an entity is encountered, randomly generate a number between 0 and 1. When the randomly generated number is less than the replacement probability p, randomly select other entities in its corresponding entity set for replacement. During the replacement process, since the length of the entity (number of words) may be inconsistent, it is necessary not only to replace the corresponding content of the entity in the text sequence, but also to replace the corresponding label sequence. The new text sequence and label sequence are formatted to obtain new training data. Refer to Figure 3 , the execution process is as shown in the replacement augmentation process algorithm mentioned above, where, Refers to a randomly selected entity or non-entity part of length n', D synthesis This is the final augmented data set after synonym replacement. foreach is to operate on each entity, and choose is to select an entity. is a randomly selected entity or non-entity part of length m', and rand is randomly generated; however, the synonym replacement method does not consider the contextual semantics, and the replacement effect is poor.
[0004] In summary, there is a need for a data augmentation method for network security named entity recognition that performs domain-adaptive pre-training and improves text diversity. Summary of the invention
[0005] A brief overview of the present invention is provided below in order to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is merely to present certain concepts in a simplified form as a prelude to a more detailed description discussed later.
[0006] In view of this, in order to solve the problem that the traditional network security named entity recognition data augmentation method in the prior art is prone to cause semantic errors in the recognition of text, the present invention provides a network security named entity recognition data augmentation method based on a pre-training model.
[0007] The technical solution is as follows: A method for augmenting network security named entity recognition data based on a pre-training model, comprising the following steps:
[0008] S1. Given a label set and an input sequence, generate a label sequence according to the labeling rules, perform sentence processing on the input sequence, and obtain subsequences and labeled subsequences;
[0009] S2. Replace the length of the text fragment of the input sequence, that is, truncate the subsequent content of the fragment that exceeds the preset maximum length of the input text fragment according to the subsequence and the marked subsequence to obtain a fragment set;
[0010] S3. Perform mask operation on the segment set according to the annotation set to obtain a new segment set, and use the BERT model to predict the new segment set to obtain an augmented data set;
[0011] S4. Based on the augmented data set, generate adjacent sentence pairs to obtain the adjacent sentence probability matrix, use the BERT model to calculate the continuous probability and search the adjacent sentence probability matrix to disrupt the sentence order to obtain the final augmented data set.
[0012] Furthermore, in S1, the annotation set is L, L=(Malware, Organization, System, O…), where Malware is the first entity type of the annotated data, Organization is the second entity type of the annotated data, System is the third entity type of the annotated data, and O indicates that the annotated data does not belong to any entity;
[0013] The input sequence is X, X=(x 1 ,x 2 ,……x n), where x i is the i-th word, i=1,2,…,n;
[0014] The label sequence is Y, Y=(y 1 ,y 2 ,……y n ), where y i is the i-th mark, y i ∈T, T is a temporary segment;
[0015] The input sequence X is processed according to natural sentences to obtain sentence X i , integrate to get the subsequence {X 1 ,X 2 ,……,X n}, that is, X=X 1 +X 2 +X 3 ..., the subsequence of the input sequence X corresponds to the labeled subsequence {Y 1 ,Y 2 ,……,Y n}.
[0016] Furthermore, in S2, the preset maximum length of the input text segment max_length is 512, and len() represents the number of feature tokens in the segment set;
[0017] Initialize the segment set S and temporary segment T, S = {}, T = {}, for each sentence X i , if len(T+X i )≤max_length, then T=T+X i Otherwise, add the temporary fragment T to the fragment set S and initialize the temporary fragment T, T = {X i}, finally, add the remaining temporary fragment T to the fragment set S to obtain the fragment set containing element s i All sentences are X i The combined fragment set S, and len(s i )≤max_length.
[0018] Furthermore, in S3, for element s i For each word x in i , if its corresponding mark y i If it is an entity, no mask operation is performed. If the corresponding tag y i does not belong to any entity, then generate a random number r, r∈[0,1], if the random number r is less than the replacement probability P, then replace the word x iReplace with mask [MASK], and get a new segment set S containing mask [MASK] mask ;
[0019] The replacement probability P is expressed as:
[0020]
[0021] Augmented dataset It is expressed as:
[0022]
[0023] Furthermore, in S4, for the augmented data set Each sentence in Generate all adjacent sentence pairs Among them, i≠j and j≠i+1, the BERT model is used to calculate the continuous probability P of all combinations of two sentences into adjacent sentence pairs ij , integrating all continuous probabilities P ij Get the adjacent sentence probability matrix P';
[0024] Continuous probability P ij The calculation process is expressed as:
[0025]
[0026] P ij =softmax(BERT(Input))
[0027] Among them, [CLS] and [SEP] are special feature tokens of the BERT model. [CLS] indicates the beginning of the text segment, i.e., the first sentence of a paragraph. [SEP] indicates the clauses in the text segment except the first sentence. Softmax is the activation function. Input indicates the current input of the BERT model. For the first sentence, For The adjacent second sentence;
[0028] Use the BERT model to search the first row of the adjacent sentence probability matrix P' and select the maximum continuous probability P in the first row ijmax , with the second sentence As the sentence with the maximum continuous probability in the first row The next sentence, and the maximum continuous probability P ijmax Set to 0, repeat the above search operation, traverse the j-th row of the adjacent sentence probability matrix P', and select the maximum continuous probability P in the j-th row jk , the adjacent sentences As the maximum continuous probability P in the jth row jkThe corresponding sentence is the second sentence The next sentence, and the maximum continuous probability P in the jth row jk Set to 0 and repeat the above search operation until the augmented data set Each element of is redistributed in order, and the final augmented data set is obtained by combining the input sequence X.
[0029] The beneficial effects of the present invention are as follows: the present invention is inspired by the synonym replacement method and is mainly aimed at network security NER data. The Masked Language Model (MLM) task is similar to the synonym replacement and cannot act on the entity part, but its replacement effect is significantly better than the synonym replacement method. The main reason is that the synonym replacement does not consider the semantics of the context, but simply uses words with similar meanings for replacement, which will cause semantic errors in the text. The use of a pre-trained model, namely the BERT model, for prediction, since the model has been trained on a large amount of corpus data, can fully consider the influence of the context to generate semantically consistent replacement words, reducing the introduction of noise. In addition, the prediction results of the pre-trained model may not be synonyms, but other semantically consistent words, which can further improve the diversity of the text and reduce overfitting. The present invention chooses to abandon the NSP pre-training task and further improves the quality of the augmented data generated by the masked prediction augmentation method by using the model trained by domain adaptability, thereby further improving the generalization ability and performance of the pre-trained model. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0031] Figure 1 A flowchart of a method for data augmentation of network security named entity recognition based on a pre-trained model;
[0032] Figure 2 It is a schematic diagram of the replacement effect of the synonym replacement method;
[0033] Figure 3 A pseudo code diagram for the replacement augmentation process is shown below;
[0034] Figure 4 A schematic diagram of the results of an embodiment of masking operation for a BERT model;
[0035] Figure 5 This is a pseudo code diagram of the mask prediction augmentation algorithm;
[0036] Figure 6 This is a pseudocode diagram of the data augmentation process algorithm. DETAILED DESCRIPTION
[0037] In order to make the technical solutions and advantages of the embodiments of the present invention more clearly understood, the exemplary embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than an exhaustive list of all the embodiments. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.
[0038] refer to Figure 1-Figure 6 The present embodiment is described in detail, a method for augmenting network security named entity recognition data based on a pre-training model, specifically comprising the following steps:
[0039] S1. Given a label set and an input sequence, generate a label sequence according to the labeling rules, perform sentence processing on the input sequence, and obtain subsequences and label subsequences;
[0040] S2. Replace the length of the text fragment of the input sequence, that is, truncate the subsequent content of the fragment that exceeds the preset maximum length of the input text fragment according to the subsequence and the marked subsequence to obtain a fragment set;
[0041] S3. Perform mask operation on the segment set according to the annotation set to obtain a new segment set, and use the BERT model to predict the new segment set to obtain an augmented data set;
[0042] S4. Based on the augmented data set, generate adjacent sentence pairs to obtain an adjacent sentence probability matrix, use the BERT model to calculate the continuous probability and search the adjacent sentence probability matrix to disrupt the sentence order, and obtain the final augmented data set;
[0043] Specifically, the BERT model is a pre-trained model derived from the Transformer architecture component. The BERT model uses a deep bidirectional Transformer architecture component to build the model, breaking the limitation of unidirectional fusion context and generating a deep bidirectional language representation that integrates context information. The BERT model captures deep language features through pre-training and can be fine-tuned for various natural language processing (NLP) tasks. It has achieved remarkable results in many tasks such as question-answering systems, text summarization, and sentiment analysis. The BERT model specifically introduces the masked language model MLM task, which can help the model better understand the pre-training method of language context. The content of the MLM task is: in a sentence, a certain percentage of feature tokens are randomly selected, the above feature tokens are replaced with masks [MASK], and then the classification model is used to predict what word category the mask [MASK] actually belongs to. Refer to Figure 4, randomly mask the sentence and give the prediction result. Similar to synonym replacement, the MLM task cannot be applied to the entity part, mainly because it is essentially a classification task and can only select the most suitable word in the vocabulary, while entities usually do not exist in the predefined vocabulary. "Proofpoint" belongs to the Organization entity class. After masking, the prediction result of the model becomes "he", which obviously does not belong to Organization, which will lead to the error of entity category. However, the BERT model considers the semantics of the context and improves the accuracy of named entity recognition.
[0044] In this embodiment, reference Figure 5 ,in, Represents the segment after division, Output is the output, replace is the replacement, and the MLM task, i.e., steps S1-S3, is specifically implemented as follows: First, the length of the replacement text segment needs to be considered. A longer segment can provide more contextual information and improve the accuracy of the prediction result. However, if it exceeds the maximum input length of the model, its subsequent content will be truncated, resulting in incomplete sentences and affecting the semantics. The segment length should be guaranteed to be long and not exceed the maximum input length of the pre-trained model, i.e., the BERT model. Therefore, the original text is segmented. For each segment, every time a sentence is added, the length of the input text segment is judged. If the length is greater than the maximum length of the input text segment preset by the model, max_length, the corresponding sentence text is abandoned and the current text segment is saved. After the text segment is divided, the segment that needs to be augmented is traversed. If it is a non-entity part and the random number 0-1 is less than the replacement probability P, it is modified to a mask [MASK]. The BERT model is used to predict the replaced masked text segment, and the prediction result is used to replace the corresponding content in the original text.
[0045] In the implementation of the original BERT model, the pre-training tasks mainly include the Masked Language Model (MLM) task and the Next Sentence Prediction (NSP) task. The NSP task means that given two sentences A and B, the model must predict whether sentence B is the logical continuation of sentence A. The purpose is to help the model understand the relationship between sentences. It plays an important role in question answering and natural language reasoning. However, in the subsequent research process, some researchers found that the performance of the model trained without the NSP task on multiple tasks did not decrease, but improved. The MLM task performed by the present invention is the key to data augmentation. Introducing the NSP task for comprehensive evaluation will affect the performance of the pre-trained model on the MLM task. For this reason, when continuing pre-training adapted to the field of network security, it is chosen to abandon the NSP pre-training task and only retain the MLM task. Figure 6 , based on continued pre-training, the data augmentation process algorithm is obtained.
[0046] Furthermore, in S1, the annotation set is L, L=(Malware, Organization, System, O…), where Malware is the first entity type of the annotated data, Organization is the second entity type of the annotated data, System is the third entity type of the annotated data, and O indicates that the annotated data does not belong to any entity;
[0047] The input sequence is X, X=(x 1 ,x 2 ,……x n ), where x i is the i-th word, i=1,2,…,n;
[0048] The label sequence is Y, Y=(y 1 ,y 2 ,……y n ), where y i is the i-th mark, y i ∈T, T is a temporary segment;
[0049] The input sequence X is processed according to natural sentences, that is, it is split according to periods to obtain sentence X i , integrate to get the subsequence {X 1 ,X 2 ,……,X n}, that is, X=X 1 +X 2 +X 3 ..., the subsequence of the input sequence X corresponds to the labeled subsequence {Y1 ,Y 2 ,……,Y n}.
[0050] Furthermore, in S2, the preset maximum length of the input text segment max_length is 512, and len() represents the number of feature tokens in the segment set;
[0051] Initialize the segment set S and temporary segment T, S = {}, T = {}, at this time the segment set S is empty, for each sentence X i , if len(T+X i )≤max_length, then T=T+X i Otherwise, add the temporary fragment T to the fragment set S and initialize the temporary fragment T, T = {X i}, finally, add the remaining temporary fragment T to the fragment set S to obtain the fragment set containing element s i All sentences are X i The combined fragment set S, and len(s i )≤max_length.
[0052] Furthermore, in S3, the fragment set S is masked, and the element s i All sentences are X i The combination contains multiple words x i , for element s i For each word x in i , if its corresponding mark y i If it is an entity, no mask operation is performed. If the corresponding tag y i does not belong to any entity, then generate a random number r, r∈[0,1], if the random number r is less than the replacement probability P, then replace the word x i Replace with mask [MASK], and get a new segment set S containing mask [MASK] mask ;
[0053] The replacement probability P is expressed as:
[0054]
[0055] Augmented dataset It is expressed as:
[0056]
[0057] Specifically, the new segment set S containing the mask [MASK] mask As the input of the BERT model, the BERT model can automatically predict the mask [MASK] based on the context information.
[0058] Furthermore, in S4, for the augmented data set Each sentence in Generate all possible adjacent sentence pairs Among them, i≠j and j≠i+1, the BERT model is used to calculate the continuous probability P of all combinations of two sentences into adjacent sentence pairs ij , integrating all continuous probabilities P ij Get the adjacent sentence probability matrix P';
[0059] Continuous probability P ij The calculation process is expressed as:
[0060]
[0061] P ij =softmax(BERT(Input))
[0062] Among them, [CLS] and [SEP] are special feature tokens of the BERT model. [CLS] indicates the beginning of the text segment, i.e., the first sentence of a paragraph. [SEP] indicates the clauses in the text segment except the first sentence. Softmax is the activation function. Input indicates the current input of the BERT model. For the first sentence, For The adjacent second sentence;
[0063] Use the BERT model to search the first row of the adjacent sentence probability matrix P' and select the maximum continuous probability P in the first row ijmax , with the second sentence As the sentence with the maximum continuous probability in the first row The next sentence, and the maximum continuous probability P ijmax Set to 0, repeat the above search operation, traverse the j-th row of the adjacent sentence probability matrix P', and select the maximum continuous probability P in the j-th row jk , the adjacent sentences As the maximum continuous probability P in the jth row jk The corresponding sentence is the second sentence The next sentence, and the maximum continuous probability P in the jth row jk Set to 0 and repeat the above search operation until the augmented data set After the order of each element of is redistributed, the final augmented data set is obtained by combining the input sequence X. At this time, the words in the final augmented data set are replaced and the order of sentences is also disrupted, thus completing the entire augmentation process;
[0064] Specific, augmented dataset Only the words are replaced, but the order remains unchanged. The augmented dataset The sentence is The BERT model’s sentence continuity (NSP) function is used to disrupt the order of sentences. The probability P is calculated for each sentence and its non-adjacent sentences. ij , if the sentence In the sentence i=j or j=i+1, Recorded as 0.
[0065] Although the present invention has been described according to a limited number of embodiments, it will be apparent to those skilled in the art, with the benefit of the above description, that other embodiments may be envisioned within the scope of the invention thus described. In addition, it should be noted that the language used in this specification is selected primarily for readability and didactic purposes, rather than for explaining or defining the subject matter of the present invention. Therefore, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended claims. The disclosure of the present invention is illustrative, not restrictive, with respect to the scope of the present invention, which is defined by the appended claims.
Claims
1. A method for augmenting network security named entity recognition data based on a pre-training model, characterized in that: The following steps are involved: S1. Given a label set and an input sequence, generate a label sequence according to the labeling rules, perform sentence processing on the input sequence, and obtain subsequences and label subsequences; S2. Replace the length of the text fragment of the input sequence, that is, truncate the subsequent content of the fragment that exceeds the preset maximum length of the input text fragment according to the subsequence and the marked subsequence to obtain a fragment set; S3. Perform mask operation on the segment set according to the annotation set to obtain a new segment set, and use the BERT model to predict the new segment set to obtain an augmented data set; S4. Based on the augmented data set, generate adjacent sentence pairs to obtain an adjacent sentence probability matrix, use the BERT model to calculate the continuous probability and search the adjacent sentence probability matrix to disrupt the sentence order, and obtain the final augmented data set; In S3, for element s i For each word x in i , if its corresponding mark y i If it is an entity, no mask operation is performed. If the corresponding tag y i does not belong to any entity, then generate a random number r, r∈[0,1], if the random number r is less than the replacement probability P, then replace the word x i Replace with mask [MASK], and get a new segment set S containing mask [MASK] mask ; The replacement probability P is expressed as: Augmented dataset It is expressed as: In S4, for the augmented dataset Each sentence in Generate all adjacent sentence pairs Among them, i≠j and j≠i+1, the BERT model is used to calculate the continuous probability P of all combinations of two sentences into adjacent sentence pairs ij , integrating all continuous probabilities P ij Get the adjacent sentence probability matrix P'; Continuous probability P ij The calculation process is expressed as: P ij =softmax(BERT(Input)) Among them, [CLS] and [SEP] are special feature tokens of the BERT model. [CLS] indicates the beginning of the text segment, i.e., the first sentence of a paragraph. [SEP] indicates the clauses in the text segment except the first sentence. Softmax is the activation function. Input indicates the current input of the BERT model. For the first sentence, For The adjacent second sentence; Use the BERT model to search the first row of the adjacent sentence probability matrix P' and select the maximum continuous probability P in the first row ijmax , with the second sentence As the sentence with the maximum continuous probability in the first row The next sentence, and the maximum continuous probability P ijmax Set to 0, repeat the above search operation, traverse the j-th row of the adjacent sentence probability matrix P', and select the maximum continuous probability P in the j-th row jk , the adjacent sentences As the maximum continuous probability P in the jth row jk The corresponding sentence is the second sentence The next sentence, and the maximum continuous probability P in the jth row jk Set to 0 and repeat the above search operation until the augmented data set Each element of is redistributed in order, and the final augmented data set is obtained by combining the input sequence X.
2. According to the method for augmenting network security named entity recognition data based on a pre-training model according to claim 1, it is characterized in that: In S1, the annotation set is L, L = (Malware, Organization, System, O...), where Malware is the first entity type of the annotated data, Organization is the second entity type of the annotated data, System is the third entity type of the annotated data, and O indicates that the annotated data does not belong to any entity; The input sequence is X, X=(x1,x 2, ……x n ), where x i is the i-th word, i=1,2,…,n; The label sequence is Y, Y = (y1, y 2, ……y n ), where y i is the i-th mark, y i ∈T, T is a temporary segment; The input sequence X is processed according to natural sentences to obtain sentence X i , integrate to get the subsequences {X1,X2,……,X n }, that is, X=X1+X2+X3……, the labeled subsequence corresponding to the subsequence of the input sequence X is {Y1,Y2,……,Y n }.
3. According to the method for augmenting network security named entity recognition data based on a pre-training model according to claim 2, it is characterized in that: In S2, the preset maximum length of the input text segment max_length is 512, and len() represents the number of feature tokens in the segment set; Initialize the segment set S and temporary segment T, S = {}, T = {}, for each sentence X i , if len(T+X i )≤max_length, then T=T+X i Otherwise, add the temporary fragment T to the fragment set S and initialize the temporary fragment T, T = {X i }, finally, add the remaining temporary fragment T to the fragment set S to obtain the fragment set containing element s i All sentences are X i The combined fragment set S, and len(s i )≤max_length.
Citation Information
Patent Citations
Judicial text named entity recognition-oriented method and system
CN113869053A
Noise label correction method for fine-grained entity classification
CN114912436A
Cited By
Large model data enhanced named entity recognition and RAG system integration method
CN121168654A