An event extraction, classification and fusion method for network security
By proposing event extraction, classification and fusion methods in the field of network security, using the dual attention mechanism and event chain method, the shortcomings of event correlation analysis in the existing technology are solved, and effective correlation analysis and law mining of network security events are achieved.
Patent Information
- Application Number
- CN202210432552.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-04-24
AI Technical Summary
After the prior art event extraction in the field of network security, it lacks effective correlation analysis methods, making it difficult to explore event timing relationships and development laws.
A method of event extraction, classification and fusion for the field of network security is proposed. By defining event categories and argument templates in the field of network security, the event classification is used by the dual attention mechanism, and the event chain method is used for event fusion.
The correlation analysis of network security events is realized, the laws of event development and change can be mined, the pressure of data labeling is reduced, and the accuracy of event classification is improved.
Smart Images

Figure CN114860903B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to natural language processing technology, and in particular to an event extraction, classification and fusion method for the network security field. Background Art
[0002] An event is a description of something that has happened, including the time, location, content, and roles of the event. It is usually expressed in unstructured text described in natural language. With the rapid development of the Internet, the data content generated on the network has grown explosively, and it is very difficult to manually process, analyze, and associate data. Therefore, it is very important to automatically extract event information and analyze the correlation characteristics between events. Most of the existing work focuses on the extraction of events, and there is little research on further correlation analysis after event extraction. However, the correlation analysis of events is very valuable and is crucial for studying the temporal relationship of events and exploring the laws of event development. Summary of the invention
[0003] The purpose of the present invention is to propose an event extraction, classification and fusion method for the network security field.
[0004] The technical solution to achieve the purpose of the present invention is: a method for event extraction, classification and fusion in the field of network security, characterized by comprising the following steps:
[0005] Step 1: According to the completeness of event element information, select several representative events from each event chain in the historical database;
[0006] Step 2: Define the event categories and argument templates in the network security field, and perform structured extraction of meta-events according to the templates for the input unstructured network security text;
[0007] Step 3: Build an event classification model, combine all the extracted meta-events with the representative events in the event chain to form event pairs, and use the dual attention mechanism to determine whether the events belong to the same category from the perspective of text semantic similarity and event argument and role similarity;
[0008] Step 4: Train the event classification model. Based on the event classification results, use the event chain approach to integrate the meta-events into the event chain by calculating the votes and similarity scores of the representative events on the event chain.
[0009] A network security-oriented event extraction, classification and fusion system realizes network security-oriented event extraction, classification and fusion based on the network security-oriented event extraction, classification and fusion method.
[0010] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, based on the event extraction, classification and fusion method for the network security field, event extraction, classification and fusion for the network security field are realized.
[0011] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements event extraction, classification and fusion in the network security field based on the event extraction, classification and fusion method in the network security field.
[0012] Compared with the prior art, the present invention has the following significant advantages: 1) It proposes a dual attention model based on text and arguments, which can comprehensively judge whether an event pair belongs to the same type of event from the perspective of text semantic similarity and event argument role similarity. 2) It proposes a novel data sampling method, which can automatically generate event classification and annotation data based on the event extraction data set, greatly reducing the pressure of data annotation. 3) It adopts the event chain method, through event classification and fusion strategy, associates and analyzes existing events with historical events, and can explore the laws of event development and change. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a flow chart of the incident service framework for the network security field;
[0014] Figure 2 It is the structure diagram of the meta-event extraction model;
[0015] Figure 3 It is the structure diagram of the event classification model. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0017] The present invention proposes an event extraction, classification and fusion method. The method extracts event elements from a large number of unstructured network security field texts, classifies different events and fuses events belonging to the same category, thereby realizing the correlation analysis function of network security events. It includes an event extraction model for extracting meta-events from unstructured network security field texts, a meta-event classification model and its training and prediction methods, and an event fusion strategy based on event chains. The specific steps are as follows:
[0018] Step 1: Select representative events from each event chain in the event database. Since the amount of event data accumulates over time, it is unreasonable to comprehensively consider all events on the event chain in the database. Therefore, in order to reduce the computational cost and speed up the response speed of model prediction, it is necessary to select representative events from each event chain. The principle of selecting representative events is that the more complete the event element information is, the more obvious the theme characteristics of the event are, and the more representative the event chain is. In the present invention, the event elements include the category, argument and role information of the event. The specific method is to traverse the event database, sort each event chain according to the event category, argument and role information richness, and select the top K data in the cumulative value ranking as representatives. In order to speed up the calculation speed, the representative information of each event chain is cached. When the event chain is updated, the representative information needs to be recalculated.
[0019] Step 2: Extract meta-events from the input unstructured network security text. First, the event definitions in the network security field in the dataset used in the present invention are introduced. The event definitions in the network security field are mainly divided into event type definitions and event role label definitions. For specific definition contents, refer to Table 1.
[0020] Table 1 Definition of network security event types and event roles
[0021]
[0022] The present invention adopts a meta-event extraction model based on sequence annotation for network security text. The meta-event extraction model receives unstructured network security text as input and outputs event type, event role and event argument results. The structure of the model is as follows: Figure 2 The following is an introduction to the working principles of each part of the model.
[0023] Step 2.1: Use BERT to encode the input text. The input text is a set of characters, and BERT is used to map each character in the text into a character vector. The specific calculation formula is as follows:
[0024] s={c1,c2,c3...c n} (1)
[0025]
[0026] Among them, s represents the input sentence, c i Represents the characters in the sentence, Represents a character vector after BERT encoding. c Indicates a character sequence for differentiation. i Indicates the position of the current character in the character set.
[0027] Step 2.2: Use the fully connected layer and CRF layer to calculate the event role label probability. The input is a set of character vectors, and the output is the role label probability. The specific calculation formula is as follows:
[0028] h=Wx+b (3)
[0029] P=CRF(h) (4)
[0030] Among them, h represents the calculation result of the fully connected layer on the character vector, x represents the character vector set, W and b represent trainable parameters, P represents the character label probability, and CRF represents the conditional random field method.
[0031] According to the probability of event role labels, arguments and role labels are extracted. Here, the role label of the event and the type label of the event are bound, and the type of the event can be determined by determining the role label.
[0032] Step 3: Classify all the meta-events extracted and the representative events in the event chain into event pairs. For the current input text, the meta-events extracted using the method in step 2 will be combined with the N representative events on each event chain, and a binary event classifier will be used to determine whether the two belong to the same event chain. Regarding the event classifier, the present invention proposes a dual attention model based on text and arguments, which comprehensively determines whether event pairs belong to the same type of events from the perspectives of text semantic similarity and event argument role similarity. The model structure diagram is shown in the figure below. Figure 3 The following is a detailed introduction to each module of the model.
[0033] Step 3.1: Encode the meta-events, input text representing the events, and event arguments. For the input text, use BERT to map each character in the text into a character vector; for the event arguments, use the word embedding matrix for encoding. The specific calculation formula is as follows:
[0034] s1={c 1 1,c 1 2,c 1 3...c 1 n} (5)
[0035] s2={c 2 1,c 2 2,c 2 3...c 2 n} (6)
[0036] a1={w 1 1,w 1 2,w 1 3...w 1 n} (7)
[0037] a2={w 2 1,w 2 2,w 2 3...w 2 n} (8)
[0038]
[0039]
[0040]
[0041]
[0042] Among them, s1 and s2 are the texts of two events, a1 and a2 are the arguments of two events, and x 1 and x 2 is a BERT-encoded character vector, h 1 and h 2 is the vector after argument encoding. The superscripts 1 and 2 are used to distinguish two events, and the subscript i Refers to the position of the current character or character vector in the set. Since the input is an event pair, the text and arguments of the two events need to be encoded separately.
[0043] Step 3.2: Use BiLSTM to calculate the temporal information of the meta-event and the input text representing the event. The specific calculation formula is as follows:
[0044]
[0045] in, and It is the result of BiLSTM calculation. The superscripts 1 and 2 are used to distinguish two events, and the superscript ' is only used for distinction and has no practical meaning.
[0046] Step 3.3: Based on the BiLSTM calculation results, calculate the attention scores of the input text and arguments, and update the vector weights. By using the attention mechanism, we focus on the focus information in the input text and have a clearer representation of the semantic information contained in the text. The specific calculation formula is as follows:
[0047] First calculate the text vector attention score matrix:
[0048]
[0049]
[0050]
[0051] Where x_score is the text vector attention score matrix. The superscript is only used for distinction.
[0052] The matrix elements are accumulated and averaged by row and column respectively to calculate the attention weight of the text vector:
[0053]
[0054]
[0055] in, and Respectively and The subscript indicates the position of the current vector in the set, and the superscript has no actual meaning and is only used for distinction.
[0056] Update the text vectors for two events:
[0057]
[0058]
[0059] Similarly, the attention score of the event argument vector is calculated and the argument vector is updated. The calculation steps are as follows:
[0060]
[0061]
[0062]
[0063]
[0064]
[0065]
[0066]
[0067] Among them, a_score is the argument attention score matrix, and Yes 1 and h 2 The argument vector attention weight of . The superscripts 1 and 2 are used to distinguish event 1 from event 2, and * is only used for distinction.
[0068] Step 3.4: Calculate the text vector distance and argument vector distance features between the meta-event and the representative event to determine whether the two events belong to the same event type. The specific calculation steps are as follows:
[0069]
[0070]
[0071] P = soft max(W3[f a ;f s ]+b3) (31)
[0072] Among them, f a and f s Represent the argument distance feature and text distance feature respectively, W1, b1, W2, b2, W3, b3 are trainable parameters, and P is the probability that two events belong to the same type of event. Subscripts 1, 2, and 3 are only used to distinguish. a and s Indicates the argument level and text level respectively. Superscript 1 and 2 are used to distinguish event 1 from event 2, and * is only used for distinction.
[0073] Regarding the training method of the event classifier, due to the lack of labeled data and the high cost of manual labeling, the present invention proposes a sampling method that can train the event classification model using only event extraction and labeling data. First, each sample in the event extraction and labeling data is segmented according to sentences, and the labeled event type, event argument, and role information are divided into their respective sentences. After this step, the original event labeling sample is divided into several sub-event labeling samples according to sentences. Secondly, all sub-events are traversed, and for each sub-event, other sub-events that originally belonged to the same event are selected as positive samples; any other event that is different from the current event is randomly selected, and a sub-event is randomly selected from them as a negative sample. According to the above sampling method, the training data of the event classification model can be obtained, that is, the event classification model can be trained and whether the event pair belongs to the same category can be predicted.
[0074] Step 4: Based on the event classification results, the event fusion strategy is used to integrate the meta-event into the event chain. First, the meta-event is classified with the representative events selected on each event chain, and the score of the meta-event belonging to a certain event chain is calculated by voting. The calculation steps are as follows:
[0075]
[0076] This formula calculates the voting results of the meta-event and the representative events on the event chain. Among them, K represents the number of representative events on the event chain, and f classify represents the event classifier, e * Represents a meta-event, e i Indicates a representative event in an event chain. * For distinguishing purposes, the following table i Indicates the sequence number of the event in the event chain, the same below.
[0077]
[0078] This formula calculates the text similarity score between the current meta-event and the representative event. sim Indicates the cosine similarity calculation method.
[0079] Calculate the final score of the meta event and event chain:
[0080] score = αsim + (1-α)vote (34)
[0081] Among them, α is the proportional coefficient for adjusting the weight of text similarity score and voting score.
[0082] According to the scores of the meta-event and each event chain, the event chain with the highest score is selected. If the score exceeds the given threshold, the meta-event is integrated into the target event chain and the event chain representative event is updated; if the score is lower than the threshold, the meta-event is created as a new event chain.
[0083] Example
[0084] In order to verify the effectiveness of the scheme of the present invention, the following examples are carried out.
[0085] Input: A text that reads "A phishing attack is spreading rapidly and has reached 1 million Gmail users. The phishing attack is disguised as a virtual application that looks like Google Docs. The recipient is invited to click a blue box that says 'Open in Docs'. Clicking the blue box will take them to a Google account page, where the phishing software will gain access to the recipient's Gmail."
[0086] Step 1: Select the representative event of the event chain. Here we take a representative event on an event chain in the database as an example.
[0087] The event text reads: "Colorado's state-owned computers are being held for ransom. According to the governor's office, some computers at the Colorado Department of Transportation were maliciously installed with ransomware for the first time on Wednesday."
[0088] The extraction results of representative events are:
[0089] {
[0090] "Event Type": "Cyber Ransomware",
[0091] "Attack mode": "Malicious installation of ransomware",
[0092] "Victim Device": "Computer",
[0093] "Location": "Colorado",
[0094] "Victim Organization": "Colorado Department of Transportation",
[0095] "time": "Wednesday"
[0096] }
[0097] Step 2: Meta-event extraction. After preprocessing the input text, event extraction is performed, and the extraction results are:
[0098] {
[0099] "Event Type": "Phishing",
[0100] "Attack Mode": "Click on a blue box that says 'Open in Documents'"
[0101] "Number of victims": "1 million",
[0102] "Motivation": "Get access to the recipient's Gmail account",
[0103] "Virtual Application"
[0104] "Trusted Entity": "Google Docs",
[0105] “Victim”: “Google Mail User”
[0106] }
[0107] Step 3: Meta-event and event chain classification and fusion. The average text similarity score between the meta-event and the representative event on the event chain is 0.11, the average voting score is 0.1, α is set to 0.8, and the final meta-event and event chain score is 0.108. The event fusion threshold is set to 0.5. Since the meta-event and event chain score is less than the threshold, the meta-event does not belong to the event chain.
[0108] Output: The meta-event extraction results are stored in the database as a new event chain.
[0109] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0110] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for event extraction, classification and fusion in the field of network security, characterized in that: The steps include: Step 1: According to the completeness of event element information, select several representative events from each event chain in the historical database; Step 2: Define the event categories and argument templates in the network security field, and perform structured extraction of meta-events according to the templates for the input unstructured network security text; Step 3: Build an event classification model, combine all the extracted meta-events with the representative events in the event chain to form event pairs, and use the dual attention mechanism to determine whether the events belong to the same category from the perspective of text semantic similarity and event argument and role similarity; Step 4: train the event classification model. Based on the event classification results, the event chain is adopted to integrate the meta-event into the event chain by calculating the votes and similarity scores of the representative events on the event chain. Step 3: Build an event classification model, combine all the extracted meta-events with the representative events in the event chain to form event pairs, and use the dual attention mechanism to determine whether the events belong to the same category from the perspective of text semantic similarity and event argument and role similarity. The specific method is as follows: Step 3.1: Encode the input text and event arguments of the meta-event and the representative event. For the input text, use BERT to map each character in the text into a character vector; for the event arguments, use the word embedding matrix for encoding. The specific calculation formula is as follows: s1={c 1 1,c 1 2,c 1 3...c 1 n } (5) s2={c 2 1,c 2 2,c 2 3...c 2 n } (6) a1={in 1 1,in 1 2,in 1 3...in 1 n } (7) a2={w 2 1,w 2 2,w 2 3...w 2 n } (8) Among them, s1 and s2 are the texts of two events, a1 and a2 are the arguments of two events, and x 1 and x 2 is a BERT-encoded character vector, h 1 and h 2 is the argument-encoded vector. The superscripts 1 and 2 are used to distinguish two events. i Refers to the position of the current character or character vector in the collection; Step 3.2: Use BiLSTM to calculate the temporal information of the meta-event and the input text representing the event. The specific calculation formula is as follows: x' 1 =BiLSTM(x 1 )(13)x' 2 =BiLSTM(x 2 )(14) Where x' 1 and x' 2 It is the result of BiLSTM calculation. The superscripts 1 and 2 are used to distinguish two events. The superscript ' is only used for distinction and has no practical meaning. Step 3.3: Based on the BiLSTM calculation results, use the attention mechanism to calculate the attention scores of the input text and arguments and update the vector weights. The specific calculation formula is as follows: First calculate the text vector attention score matrix: Among them, x_score is the text vector attention score matrix, and the superscript is only used for distinction; The matrix elements are accumulated and averaged by row and column respectively to calculate the attention weight of the text vector: in, and Respectively represent x' 1 and x' 2 The vector attention weight of , the subscript indicates the position of the current vector in the set, and the superscript has no actual meaning, only for distinction; Update the text vectors for two events: Similarly, the attention score of the event argument vector is calculated and the argument vector is updated. The calculation steps are as follows: Among them, a_score is the argument attention score matrix, and Yes 1 and h 2 The argument vector attention weights, superscripts 1 and 2 are used to distinguish event 1 from event 2, and * is only used for distinction; Step 3.4: Calculate the text vector distance and argument vector distance features between the meta-event and the representative event to determine whether the two events belong to the same event type. The specific calculation steps are as follows: f a =W1[x *1 ;x *2 ;x *1 -x *2 ]+b1 (29) f s =W2[h *1 ;h *2 ;h *1 -h *2 ]+b2 (30) P=soft max(W3[f a ;f s ]+b3) (31) Among them, f a and f s They represent argument vector distance feature and text vector distance feature respectively. W1, b1, W2, b2, W3, and b3 are trainable parameters. P is the probability that two events belong to the same type of event. Subscripts 1, 2, and 3 are only used for distinction. a and s represent the argument level and text level respectively. Superscripts 1 and 2 are used to distinguish event 1 from event 2. * is only used for distinction.
2. The event extraction, classification and fusion method for network security field according to claim 1 is characterized in that: Step 1: According to the completeness of event element information, select several representative events from each event chain in the historical database, where the event element includes the event category, argument and role information. When selecting representative events, sort each event chain according to the cumulative value of event category, argument and role information, select the top K data as representative events, and cache the representative information. When the event chain is updated, the representative information needs to be recalculated.
3. The event extraction, classification and fusion method for network security field according to claim 1 is characterized in that: Step 2: define the event type, event role label and argument template of network security events, and perform meta-event structured extraction according to the argument template for the input unstructured network security text. The specific definition of event type and event role label is shown in Table 1, and the arguments and event role labels correspond one to one. Table 1 Definition of network security event types and event roles 4. The event extraction, classification and fusion method for network security field according to claim 1 is characterized in that: Step 2: Define the event type, event role label and argument role template of network security events, and perform meta-event structured extraction according to the argument role template for the input unstructured network security text. The specific method of meta-event structured extraction is as follows: Step 2.1: Use BERT to encode the input text and map each character in the text into a character vector. The specific calculation formula is as follows: s={c1,c2,c3...c n } (1) Among them, s represents the input sentence, c i Represents the characters in the sentence, Represents a character vector after BERT encoding, with superscript c Indicates a character sequence, used to distinguish, subscript i Indicates the position of the current character in the character set; Step 2.2: Use the fully connected layer and CRF layer to calculate the probability of the event role label corresponding to the character vector set. The specific calculation formula is as follows: h=Wx+b(3)P=CRF(h)(4)where h represents the calculation result of the fully connected layer on the character vector, x represents the character vector set, W and b represent trainable parameters, P represents the character label probability, and CRF represents the conditional random field model; Step 2.3: Extract arguments and event role labels based on the role label probability, determine the event type based on the event role label, and complete the meta-event structured extraction based on this.
5. The event extraction, classification and fusion method for network security field according to claim 1 is characterized in that: Step 4: Training the event classification model based on the event extraction and annotation data. First, each sample in the event extraction and annotation data is segmented according to sentences, and the annotated event type, event argument, and event role label are divided into their respective sentences. After this step, the original event annotation sample is segmented into several sub-event annotation samples according to sentences. Secondly, traverse all sub-events, and for each sub-event, select other sub-events that originally belonged to the same event as positive samples; randomly select any other event different from the current event, and randomly select a sub-event from them as negative samples to obtain the training data of the event classification model, which is used to train the event classification model to predict whether the event pair belongs to the same category.
6. The event extraction, classification and fusion method for network security field according to claim 1 is characterized in that: Step 4: Based on the event classification results, the event chain is used to integrate the meta-event into the event chain by calculating the votes and similarity scores of the representative events on the event chain. The specific method is as follows: First, classify the meta-event with the representative events selected on each event chain, and vote to calculate the score of the meta-event belonging to a certain event chain: Among them, K represents the number of representative events on the event chain, f classify represents the event classifier, e * Represents a meta-event, e i Indicates the representative event in the event chain, superscript * For distinguishing purposes, in the following table, i represents the sequence number of an event in the event chain, the same below; Then, calculate the text similarity score between the current meta event and the representative event: Among them, f sim Indicates the cosine similarity calculation method; Next, calculate the final score of the meta event and event chain: score = αsim + (1-α)vote (34) Among them, α is the proportional coefficient for adjusting the weight of text similarity score and voting score; Finally, based on the scores of the meta-event and each event chain, the event chain with the highest score is selected. If the score exceeds a given threshold, the meta-event is integrated into the target event chain and the event chain representative event is updated; if the score is lower than the threshold, the meta-event is created as a new event chain.
7. An event extraction, classification and fusion system for network security, characterized in that: Based on the event extraction, classification and fusion method for the network security field as described in any one of claims 1 to 6, event extraction, classification and fusion for the network security field are implemented.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, based on the event extraction, classification, and fusion method for the network security field as described in any one of claims 1 to 6, event extraction, classification, and fusion for the network security field are implemented.
9. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, based on the event extraction, classification and fusion method for the network security field as described in any one of claims 1 to 6, event extraction, classification and fusion for the network security field is implemented.
Citation Information
Patent Citations
Event argument role extraction method based on multi-head attention mechanism
CN110134757A
Social event classification method and device based on local aggregation graph attention network
CN113449204A