A threat intelligence named entity recognition method based on machine reading comprehension
Through the threat intelligence named entity recognition method based on machine reading comprehension, using Q&A pair annotation and ALBERT model feature extraction, the problem of naming entity recognition in threat intelligence is solved, achieving higher recognition accuracy and lower data requirements.
Patent Information
- Application Number
- CN202210375786.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-11
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-04-11
AI Technical Summary
It is difficult for the prior art to effectively identify and analyze named entities in threat intelligence, especially when facing analytics and multi-layer nested entities, and there is little research on the field of cybersecurity.
The threat intelligence named entity recognition method based on machine reading comprehension is adopted, and the transformed entity recognition is a classification matching problem through technical means such as sentence processing, professional network security thesaurus filtering, question-and-answer marking, ALBERT model feature extraction, multiple binary classifiers and binary classifier matching degree calculation.
It effectively solves the problems of fuzzy classification and nested entities of threat intelligence entities, improves the accuracy of naming entity recognition, and reduces the requirements for the number of sentences.
Smart Images

Figure CN114757193B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, and in particular relates to a threat intelligence named entity recognition method based on machine reading comprehension. Background Art
[0002] With the rapid development of network information technology, the security situation in cyberspace is becoming increasingly severe. To protect network security, we must not only be stronger than attackers in technology, but also be ahead of attackers in awareness, pay enough attention to the "unknown unknowns", and plan ahead. Only in this way can we be impregnable and not give hackers any chance to take advantage of. Therefore, effective analysis of threat intelligence has important practical significance and scientific research value for defending against organized, premeditated and various security threats.
[0003] Threat intelligence analysis must be based on knowledge extraction to identify the corresponding entities and relationships. Currently, there are mainly the following named entity recognition (NER): (1) LSTM-CRF model based on unidirectional long short-term memory (LSTM) neural network. Based on the excellent sequence modeling function of LSTM, LSTM-CRF has become one of the basic frameworks for named entity recognition. Many methods use LSTM-CRF as the main framework and integrate various related functions. For example, adding manual spelling features, using text CNN to extract text features, or using character-level LSTM. (2) CNN-based entity recognition schemes, such as CNN-CRF structure, or based on CNN-CRF, using the enhanced model proposed by character CNN. (3) Entity recognition scheme based on atrous convolutional network (IDCNN-CRF), which can speed up training while extracting sequence information. (4) Based on the BiLSTM-CRF model, the attention mechanism is used to obtain the word context within the entire text, or the GRU calculation unit is used to propose a bidirectional GRU-based entity recognition method.
[0004] On the one hand, the common problems of these traditional methods are that they cannot represent the polysemy of words and it is difficult to identify multi-layer nested entities. However, there are a lot of such cases in threat intelligence and the classification of threat intelligence entities is relatively vague, which is also a major difficulty in threat intelligence analysis. On the other hand, the current field-oriented NER research is relatively concentrated in scientific fields such as medicine and biology, and there is still little research in the field of threat intelligence for network security. With the exponential increase in the amount of data in various fields, the automated processing and analysis of professional data information in the field has become a mainstream trend. Therefore, text recognition in the field of threat intelligence has great research value and space, and it also plays a very critical role in threat intelligence analysis. Summary of the invention
[0005] The purpose of the present invention is to provide a threat intelligence named entity recognition method based on machine reading comprehension to improve the accuracy of named entity recognition.
[0006] To achieve the above object, the technical solution adopted by the present invention is:
[0007] A method for identifying named entities of threat intelligence based on machine reading comprehension, the method for identifying named entities of threat intelligence based on machine reading comprehension comprising:
[0008] Step 1: Sentence processing is performed on the threat intelligence, and sentences that do not contain professional cybersecurity vocabulary are filtered out based on a professional cybersecurity vocabulary to obtain a sentence set after filtering;
[0009] Step 2: Take sentences from the sentence set one by one and label each entity in the sentence with a question-answer pair;
[0010] Step 3: Use the annotated sentence set to train the recognition model, including:
[0011] Step 3.1, take a class of entities in the sentence for training, concatenate the question and the corresponding sentence in the question-answer pair annotated with the entity, perform word segmentation on the concatenated sentence to obtain a word sequence, and perform feature extraction based on the word sequence to obtain a text feature matrix;
[0012] Step 3.2: Use two binary classifiers to classify each word in the text feature matrix. The first binary classifier is used to determine whether the word is the starting word of the answer corresponding to the question, and the second binary classifier is used to determine whether the word is the ending word of the answer corresponding to the question. The probability of each word in the sentence being the starting word and the probability of each word being the ending word are obtained.
[0013] Step 3.3: From the probability of each word in the sentence belonging to the start word, select the index corresponding to the word with the maximum probability as the start index, and from the probability of each word in the sentence belonging to the end word, select the index corresponding to the word with the maximum probability as the end index;
[0014] Step 3.4: Based on the start index and the end index, use a binary classifier to calculate the matching degree between the start word and the end word, and output the start index and the end index with a matching degree higher than the threshold as the predicted answer;
[0015] Step 3.5: Calculate the loss function based on the output predicted answer and the actual answer to update the parameters of each binary classifier, and return to step 3.1 to continue the recognition model training until convergence;
[0016] Step 4: Use the trained recognition model to perform named entity recognition on threat intelligence.
[0017] Several optional methods are also provided below, but they are not intended to be additional limitations on the above-mentioned overall solution, but are merely further supplements or preferences. Under the premise that there are no technical or logical contradictions, each optional method can be combined with the above-mentioned overall solution separately, and multiple optional methods can also be combined.
[0018] Preferably, labeling a question-answer pair for each entity in the sentence includes:
[0019] Assign an entity type to the entity. Each entity type is preset to correspond to a fixed question. The question corresponding to the entity type is used as the question when the entity is labeled.
[0020] Taking words as units, the start and end positions of the entity in the sentence are used as the answer for annotation.
[0021] Preferably, the question corresponding to the entity type is a glossary of the entity type.
[0022] Preferably, the feature extraction based on the word sequence to obtain the text feature matrix includes:
[0023] Add a special marker to the head of the word sequence to indicate the start of the sequence, and add a marker separator between the question and the sentence to get the complete word sequence;
[0024] The ALBERT model is used to extract features from the complete word sequence to obtain a text feature matrix, which is represented by E∈R n*d , where n is the sentence length, d is the vector dimension of the features extracted by the last layer of the ALBERT model, that is, each row of the text feature matrix E represents the feature vector corresponding to a word.
[0025] Preferably, the loss function is calculated according to the output predicted answer and the actual answer, comprising:
[0026] There are three loss functions in the training phase: starting position loss function L start , end position loss function L end , entity matching loss function L span , the loss function L of the final recognition model is the sum of the losses in each stage of training, and the calculation formula is as follows:
[0027] L start =CE(P start , Y start )
[0028] L end =CE(P end , Y end )
[0029] Lspan =CE(P start_end , Y start_end )
[0030] L=αL start +βL end +γL span
[0031] Where P start is the predicted probability that each word in the sentence belongs to the start word, Y start is the actual probability that each word in the sentence belongs to the start word of the entity, P end is the predicted probability that each word in the sentence belongs to the end word, Y end is the actual probability that each word in the sentence belongs to the end word of the entity, P start_end Indicates the predicted matching degree between the start word and the end word, Y start_end It represents the actual matching degree between the start word and the end word, CE() is the cross entropy loss function, α, β, γ are weight coefficients, α, β, γ∈[0,1].
[0032] The threat intelligence named entity recognition method based on machine reading comprehension provided by the present invention can effectively solve the problem of fuzzy classification and nested entities of threat intelligence entities; the hidden information of entities in the constructed questions can effectively improve the recognition accuracy; the entity recognition is converted from a sequence labeling problem to a classification matching problem, so a sentence with multiple entities can generate multiple training samples, thereby reducing the requirement on the number of sentences. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a flowchart of the threat intelligence named entity recognition method based on machine reading comprehension of the present invention;
[0034] Figure 2 Schematic diagram of input and output of the ALBERT model of the present invention;
[0035] Figure 3 The present invention is a flowchart of using the recognition model to perform named entity recognition. DETAILED DESCRIPTION
[0036] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0038] In order to overcome the difficulties in threat intelligence analysis in the prior art, this embodiment provides a threat intelligence named entity recognition method based on machine reading comprehension, which achieves high-accuracy named entity recognition when training samples are scarce.
[0039] Specifically, Figure 1 As shown, this embodiment provides a threat intelligence named entity recognition method based on machine reading comprehension, including the following steps:
[0040] Step 1: Sentence the threat intelligence and filter out sentences that do not contain cybersecurity patent vocabulary based on the cybersecurity professional vocabulary to obtain a sentence set.
[0041] Since there is currently no relevant international or domestic network security professional vocabulary library, this embodiment can refer to the existing network security professional vocabulary summary library for filtering, such as "Network security words", "Network security professional vocabulary collection", etc., and of course it can also be a collection of existing network security professional vocabulary summary libraries to achieve the purpose of comprehensive filtering and avoid sentences containing non-professional vocabulary from affecting the subsequent model training accuracy and recognition efficiency.
[0042] Step 2: Take sentences from the sentence set one by one and label each entity in the sentence with a question-answer pair.
[0043] Question and answer j , A j ) in question Q j is used to query entity E j A description of j It is entity E j In sentence S i where i is the index of the sentence in the sentence set, S i is the i-th sentence in the sentence set, j is the index of the entity in the sentence, E j is the jth entity in the sentence. Given a case, sentence S i="We further tested TweetyChat and saw red flags indicating their targets of interest: verificationemails with a physical address whose postal code is assigned to a provincial capital that also appears (upon logging in) as a chat channel in Tweety Chat.", the entity to be marked E j = "Tweety Chat", the specific marking process is as follows:
[0044] 1) Question annotation: manually annotate entity E based on expert knowledge j Assign an entity type. Each entity type is preset to correspond to a fixed question. The definition of the question can directly refer to the definition of the entity type.
[0045] For example, Table 1 gives the definitions of some entity types in the field of cyber attacks in STIX, which can be used directly as entity annotation issues. j ("Tweety Chat") is assigned an entity type of malware.backdoor, then question Q j ="A malicious program that allows an attacker to perform actions on a remote system, such as transferring files, acquiring passwords, or executing arbitrary commands."
[0046] Table 1 Glossary of some entity types in the field of cyber attacks in STIX
[0047]
[0048] 2) Answer: In sentence S i Marked entity E j The position of the appearance, in the form of [start j , end j ], in units of words. Among them, start j is the position of the starting word (referred to as the starting position), end jis the position of the end word (referred to as the end position), entity E j There may be multiple locations where entity E j The occurrence positions of ("Tweety Chat") are [3, 4] and [43, 44]. In this embodiment, the position of a word is understood as the index of the word in the sentence.
[0049] Step 3: Use the annotated sentence set to train the recognition model. The recognition model in this embodiment includes three binary classifiers, which transform entity recognition from a sequence labeling problem to a classification matching problem. Specifically, the model training process is as follows:
[0050] Step 3.1, input questions and sentences: Figure 2 As shown, we take a class of entities in the sentence for training and put the question Q in the indirect answer of the entity annotation j and the corresponding sentence S i The concatenated sentences are segmented to obtain word sequences, and feature extraction is performed based on the word sequences to obtain a text feature matrix.
[0051] After the word sequence is obtained, a special marker [CLS] is added to the beginning of the word sequence to mark the beginning of the sequence. The question and the original sentence are separated by the marker [SEP]. The final complete word sequence is represented as {[CLS], q 1 ,q 2 , ..., q m ,[SEP],X 1 , X 2 , ..., X n}. Among them, q v For question Q j The vth word in X k The original sentence S i The kth word in .
[0052] It should be noted that the above complete word sequence is a representation provided by this embodiment. Under the premise of having special tags and separators, tags can be replaced or further added without affecting the feature extraction of the model, for example Figure 2 As shown in , a special marker is added at the end of the word sequence to mark the end of the sequence. The final complete word sequence can be expressed as {[CLS], q 1 ,q 2 ,......,q m ,[SEP1],X 1 , X 2 ,......,X n , [SEP2]}.
[0053] After obtaining the complete word sequence, the complete word sequence is input into the ALBERT model for feature extraction to obtain the text feature matrix E∈R n*d . Where n is the sentence length, d is the vector dimension of the feature matrix extracted by the last layer of the ALBERT model, that is, each row of the text feature matrix E represents the feature vector corresponding to a word.
[0054] When there is some annotated data in this field, the question-answering nature of machine reading comprehension combined with the ALBERT model can be used to expand knowledge learning. Therefore, other scenarios in this field can be used normally even if there is a lack of annotated data, thus overcoming the problem of scarce training samples.
[0055] Step 3.2, multiple binary classification: Use two binary classifiers to classify each word X in the text feature matrix k For classification, the first two classifier is used to judge word X k Is it the starting word of the answer corresponding to the question? The second binary classifier is used to judge whether word X k Is it the ending word of the answer corresponding to the question, and get the probability of each word in the sentence belonging to the starting word and the probability of each word belonging to the ending word.
[0056] Specifically, the probability P that a word in a sentence belongs to the start word is start and the probability P of belonging to the end word end The calculations of are shown in formula (1) and formula (2), where E is the text feature matrix output by the ALBERT model, T start ∈R d*2 and T end ∈R d*2 is the parameter that the recognition model needs to learn, P start and P end It is a matrix, each row of the matrix represents the probability distribution of the index corresponding to each word (for example, k) as the start / end position of the entity for a given query, and the two columns of the matrix represent the probabilities of true and false output by the binary classifier, respectively.
[0057] P start =softmax each row (E.T start )∈R n*2 (1)
[0058] P end =softmax each row (E.T end )∈R n*2 (2)
[0059] Step 3.3, predict index: From the probability of each word in the sentence belonging to the start word, select the index corresponding to the word with the maximum probability as the start index, and from the probability of each word in the sentence belonging to the end word, select the index corresponding to the word with the maximum probability as the end index.
[0060] Specifically, for the probability matrix P start and P end Use the argmax function on each line to get the start index that may correspond to the start word (as shown in formula (3)) and the end index corresponding to the end word (As shown in formula (4)). The superscripts (i) and (j) represent the i-th and j-th rows respectively, and argmax() is the index of the maximum output score.
[0061]
[0062]
[0063] The probability matrix P representing the start word start The i-th row of The probability matrix P representing the end word end Since there may be multiple entities in a matrix, the probability matrix may output multiple start or end indices after argmax() processing, that is, and Equivalent to a collection.
[0064] Step 3.4, perform binary linking: Based on the start index and the end index, use a binary classifier to calculate the matching degree between the start word and the end word, and output the start index and the end index with a matching degree higher than the threshold as the predicted answer. The specific steps are as follows:
[0065] According to the start index and end index obtained in step 3.3 and Use a binary classifier to calculate the probability that the start word corresponding to the start index matches the end word corresponding to the end index (i.e., the degree of matching, as shown in formula (5)), where E is the text feature matrix output by the ALBERT model, Concat() is vector concatenation, and index index m∈R 1*2d is the parameter that the recognition model needs to learn. When the model determines X(i start :j end ) is an entity, otherwise it is not an entity.
[0066]
[0067] In the formula, For sentence S i Index i start The starting word at position, For sentence S i Index j end The end word at position, sigmoid() is the sigmoid function, X(i start :j end ) indicates that the index i in the sentence start to end The start word and end word of the same entity must have a higher matching probability, so this embodiment traverses and The index in uses the matching probability to filter the predicted entities.
[0068] Step 3.5: Calculate the loss function based on the output predicted answer and the actual answer to update the parameters of each binary classifier, and return to step 3.1 to continue the recognition model training until convergence.
[0069] There are three loss functions in the training phase: starting position loss function L start , end position loss function L end , entity matching loss function L span , the loss function L of the final recognition model is the sum of the losses in each stage of training, and the calculation formula is as follows:
[0070] L start =CE(P start , Y start ) (6)
[0071] L end =CE(P end , Y end ) (7)
[0072] L span =CE(P start_end , Y start_end ) (8)
[0073] L=αL start +βL end +γL span (9)
[0074] Where P start is the predicted probability that each word in the sentence belongs to the start word, Y start is the actual probability that each word in the sentence belongs to the start word of the entity, P end is the predicted probability that each word in the sentence belongs to the end word, Yend is the actual probability that each word in the sentence belongs to the end word of the entity, P start_end Indicates the predicted matching degree between the start word and the end word, Y start_end represents the actual matching degree between the start word and the end word, CE() is the cross entropy loss function, α, β, γ are weight coefficients, α, β, γ∈[0,1]. This embodiment uses two corresponding sequences (for example, P start , Y start ) to calculate the loss function.
[0075] Since the matching degree is based on the starting index and end index Calculation is performed, so P in this embodiment start_end Understood as The matching probability of each start word and end word is Y start_end Empathy.
[0076] Step 4: Use the trained recognition model to perform named entity recognition on threat intelligence.
[0077] Given a sentence in threat intelligence, refer to Figure 3 , the specific identification steps are as follows:
[0078] Step 4.1: Input the text to be predicted and the question: Concatenate S according to the method in step 3.1 i and the type of entity to be identified, add special tags and separators and input them into the ALBERT model for feature extraction to obtain the text feature matrix E∈R n*d .
[0079] Step 4.2, calculate the position probability: use the binary classifier trained in steps 3.2 and 3.3 to calculate the corresponding start index and end index.
[0080] Step 4.3, perform position matching: Use the start-end matcher (i.e., binary classifier) trained in step 3.4 to obtain the paired start-end index, which is the specific string for each type of entity that needs to be obtained, and the string is the corresponding answer.
[0081] In this embodiment, the threat intelligence named entity recognition based on machine reading comprehension can effectively solve the problem of fuzzy threat intelligence entity classification and nested entities; the constructed questions with hidden entity information can effectively improve the recognition accuracy; the entity recognition is transformed from a sequence labeling problem to a classification matching problem, so a sentence with multiple entities can generate multiple training samples, thereby reducing the requirement on the number of sentences.
[0082] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0083] The above-mentioned embodiments only express several implementation modes of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.
Claims
1. A threat intelligence named entity recognition method based on machine reading comprehension, It is characterized in that The threat intelligence named entity recognition method based on machine reading comprehension includes: Step 1: Sentence processing is performed on the threat intelligence, and sentences that do not contain professional cybersecurity vocabulary are filtered out based on a professional cybersecurity vocabulary to obtain a sentence set after filtering; Step 2: Take sentences from the sentence set one by one and label each entity in the sentence with a question-answer pair; Step 3: Use the annotated sentence set to train the recognition model, including: Step 3.1, take a class of entities in the sentence for training, concatenate the question and the corresponding sentence in the question-answer pair annotated with the entity, perform word segmentation on the concatenated sentence to obtain a word sequence, and perform feature extraction based on the word sequence to obtain a text feature matrix; wherein the feature extraction based on the word sequence to obtain the text feature matrix includes: Add a special marker to the head of the word sequence to indicate the start of the sequence, and add a marker separator between the question and the sentence to get the complete word sequence; The ALBERT model is used to extract features from the complete word sequence to obtain a text feature matrix, which is represented by E∈R n*d , where n is the sentence length, d is the vector dimension of the features extracted by the last layer of the ALBERT model, that is, each row of the text feature matrix E represents the feature vector corresponding to a word; Step 3.2: Use two binary classifiers to classify each word in the text feature matrix. The first binary classifier is used to determine whether the word is the starting word of the answer corresponding to the question, and the second binary classifier is used to determine whether the word is the ending word of the answer corresponding to the question. The probability of each word in the sentence being the starting word and the probability of each word being the ending word are obtained. Step 3.3: From the probability of each word in the sentence belonging to the start word, select the index corresponding to the word with the maximum probability as the start index, and from the probability of each word in the sentence belonging to the end word, select the index corresponding to the word with the maximum probability as the end index; Step 3.4: Based on the start index and the end index, use a binary classifier to calculate the matching degree between the start word and the end word, and output the start index and the end index with a matching degree higher than the threshold as the predicted answer; Step 3.5: Calculate the loss function based on the output predicted answer and the actual answer to update the parameters of each binary classifier, and return to step 3.1 to continue the recognition model training until convergence; wherein, the calculation of the loss function based on the output predicted answer and the actual answer includes: There are three loss functions in the training phase: starting position loss function L start , end position loss function L end , entity matching loss function L span , the loss function L of the final recognition model is the sum of the losses in each stage of training, and the calculation formula is as follows: IT start =CE(P start ,Y start ) IT end =CE(P end ,Y end ) IT span =CE(P start_end ,Y start_end ) L=αL start +βL end +γL span Where P start is the predicted probability that each word in the sentence belongs to the start word, Y start is the actual probability that each word in the sentence belongs to the start word of the entity, P end is the predicted probability that each word in the sentence belongs to the end word, Y end is the actual probability that each word in the sentence belongs to the end word of the entity, P start_end Indicates the predicted matching degree between the start word and the end word, Y start_end represents the actual matching degree between the start word and the end word, CE() is the cross entropy loss function, α, β, γ are weight coefficients, α, β, γ∈[0, 1|; Step 4: Use the trained recognition model to perform named entity recognition on threat intelligence.
2. The threat intelligence named entity recognition method based on machine reading comprehension as claimed in claim 1, It is characterized in that Each entity in the sentence is labeled with a question-answer pair, including: Assign an entity type to the entity. Each entity type is preset to correspond to a fixed question. The question corresponding to the entity type is used as the question when the entity is labeled. Taking words as units, the start and end positions of the entity in the sentence are used as the answer for annotation.
3. The threat intelligence named entity recognition method based on machine reading comprehension as claimed in claim 2, It is characterized in that The question corresponding to the entity type is the definition of the entity type.
Citation Information
Patent Citations
Threat intelligence oriented entity identification method and system
CN109858018A
Custom problem data automatic generation method of machine reading understanding system
CN112307773A