Knowledge extraction method for network security intention understanding, electronic equipment and storage medium
Through the method of preprocessing and fine-tuning of the model of network security threat analysis reports, the problem of difficulty in extracting network security knowledge in the existing technology is solved, and efficient and accurate knowledge extraction and model applicability are achieved.
Patent Information
- Application Number
- CN202510242140.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The prior art is difficult to effectively extract knowledge from network security threat analysis reports, making it difficult for computers to understand and process these data intelligently.
A knowledge extraction method for understanding network security intentions is proposed. By pre-processing the data of the network security threat analysis report, including IOC word substitution and structural word insertion, and using pre-trained models for mask language model task training, adding entity extraction sub-models and relation extraction sub-models, and performing targeted fine-tuning to meet the needs of the field of network security.
It realizes efficient and accurate knowledge extraction of network security threat analysis reports, and improves the applicability and language understanding capabilities of the model in the field of network security.
Smart Images

Figure CN120163149A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence network security, and specifically relates to a knowledge extraction method, an electronic device, and a storage medium for network security intention understanding. Background Art
[0002] The network space security situation in today's world is complex and ever-changing, and all aspects such as politics, economy, military, and culture are constantly affected by network space security. In recent years, network security incidents have occurred frequently, and increasingly large-scale and dynamic attack actions have brought huge challenges to the national network space security protection work.
[0003] At the same time, with the booming development of artificial intelligence technology, NLP technologies such as knowledge extraction have also provided new ideas for network security-related practitioners. There is a large amount of network security-related data and text on the Internet, such as security threat analysis reports released by Internet security forums and manufacturers. However, the knowledge existing in the text depends on manual processing and cannot be directly understood and utilized by computers. Through knowledge extraction technology, relevant information can be automatically discovered and extracted from multi-source heterogeneous texts, enabling computers to understand and process them more intelligently. Furthermore, scattered unstructured network security data can be integrated into structured network security intelligence knowledge for effective management, understanding, and organization of massive intelligence information.
[0004] Currently, most of the existing knowledge extraction technologies on the market focus on texts in the general field, and the pre-training-fine-tuning paradigm has become the widely adopted mainstream methodology. The pre-trained model is first pre-trained on a large corpus dataset in the general field. Through a large amount of pre-training, it has the basic ability to extract data features. In the fine-tuning stage, it can achieve good results with not too much labeled corpus, and at the same time, a large amount of training time can be saved. However, there are certain differences between the general field and the network security field corpora. For example, there are professional terms and irregular characters such as hash values. In addition, the corpora in the general field only contain one or two sentences, and the pre-training stage ignores the possible discourse structures in network security reports. This may lead to poor prediction results when directly applying this model to network security threat analysis reports. Therefore, it is particularly important to develop a knowledge extraction method for network security intention understanding. Summary of the Invention
[0005] The problem to be solved by the present invention is to accurately extract the knowledge of network security threat analysis reports, and a knowledge extraction method, an electronic device, and a storage medium for network security intention understanding are proposed.
[0006] To achieve the above object, the present invention is realized through the following technical solutions:
[0007] A knowledge extraction method for network security intent understanding, comprising the following steps:
[0008] S1. Collect the original text of network security threat analysis reports, and divide it into training data and prediction data;
[0009] S2. Perform data preprocessing on the training data of the original text of the collected network security analysis reports, including IOC token replacement and structural token insertion, and then use a tokenizer to divide the processed original text to obtain a processed token sequence as the training set;
[0010] S3. Select a pre-trained model, input the processed token sequence obtained in step S2 into the selected pre-trained model for masked language model (MLM) task training to obtain the model weights after continued pre-training;
[0011] S4. Based on the pre-trained model selected in step S3, perform model adjustment, add an entity extraction sub-model and a relationship extraction sub-model downstream of the pre-trained model, and use the model weights obtained after continued pre-training in step S3 for model training to obtain a trained knowledge extraction model for network security threat analysis reports;
[0012] S5. Use the prediction data of the original text of the network security threat analysis report obtained in step S1 to perform model prediction on the trained knowledge extraction model for network security threat analysis reports, and conduct result verification and optimization.
[0013] Further, the specific implementation method of step S2 includes the following steps:
[0014] S2.1. Perform IOC token replacement on the training data of the original text of the collected network security analysis reports, and replace the hash values, URLs, domain names, and IP addresses mentioned in the text with the corresponding tokens [HASH], [URL], [DOMAIN], and [IP] through regular replacement;
[0015] S2.2. Then perform structural token insertion, insert a token representing the structure of the line in the text at the beginning of each line of the text, and set [h1], [h2], [p], and [tr] to represent the first-level heading, second-level heading, body paragraph, and table row respectively;
[0016] S2.3. Use a tokenizer to divide the processed text into a token sequence T = {w1, w2, w3,..., w n}, where w n is the nth token, and the sequence length does not exceed the maximum length of the pre-trained model.
[0017] Further, in step S2.1, the original text of the network security analysis report that exceeds the maximum length of the pre-trained model is first divided into multiple segments that meet the length requirements.
[0018] Further, the specific implementation method of step S3 includes the following steps:
[0019] S3.1. Select BERT as the pre-trained model and change the tokenizer of the pre-trained model for dividing the token sequence;
[0020] S3.2. Set the input text as the processed token sequence T obtained in step S2, perform a masking operation, define the masking function as M(T, p), which accepts the processed token sequence T and the masking probability p as inputs. The masking function replaces the tokens in the text with the verification token [MASK] with the masking probability p. The masked sequence T’ is expressed as:
[0021] T’ = {w’1, w’2, w’3, …, w’ n}
[0022] where w’ n is the nth masked token;
[0023] S3.3. Use the pre-trained model selected in step S3.1, accept the masked sequence as the input, perform the masked language model MLM task training, and output the prediction distribution of the masked word. The expression is:
[0024] P(w i ∣T’) = f(T’)[i]
[0025] where f(T’)[i] represents the prediction distribution output by the pre-trained model at position i, and P(w i ∣T’) is the prediction distribution of the ith masked token when the input is T’;
[0026] For each masked token, calculate the cross-entropy loss between the prediction distribution of the pre-trained model and the true word distribution. The expression is:
[0027] L i = -log P(w i ∣T’)
[0028] where L i is the cross-entropy loss at position i;
[0029] Then sum the losses at all masked positions to obtain the loss function L. The expression is:
[0030] L = ∑L i ;
[0031] Then, according to the loss function L, update the parameters of the pre-trained model through the backpropagation algorithm to obtain the model weights after continued pre-training.
[0032] Further, the specific implementation method of step S4 includes the following steps:
[0033] S4.1. Add an entity extraction sub-model downstream of the pre-trained model, and use the Conditional Random Field (CRF) as the entity extraction sub-model. The expression is:
[0034]
[0035] where x is the input sequence, and P(y|x) is the conditional probability distribution of the output sequence y given the input sequence x; f k (y i-1 , y i , x, i) is the feature function, and λ k is the weight of the feature function; Z(x) is the normalization factor;
[0036] Set the input sequence x = BERT(T) as the output feature vector sequence of the pre-trained model, y as the BIO label corresponding to each token, and the normalization factor is obtained by summing over all possible output sequences. The expression is:
[0037]
[0038] where y′ represents all possible values of the output label sequence, and y′ i is the label of the i-th token;
[0039] S4.2. Use a Feed-Forward Neural Network (FFNN) as the relation extraction sub-model. Obtain the positions of each entity according to the output of the entity extraction sub-model, calculate the average value x ei of the feature vectors of all tokens of each entity, and concatenate them pairwise with the feature vector corresponding to the first token x [CLS] to obtain the input x of the relation extraction sub-model. The expression is:
[0040] x = x [CLS] || x ei || x ej ;
[0041] Then, output the distribution of relation types between entity pairs through the feed-forward neural network. The expression is:
[0042] F(x) = max(0, xW1 + b1)W2 + b2
[0043] Among them, W1 and W2 are the weight matrices of the first hidden layer and the second hidden layer of the feedforward neural network model respectively, and b1 and b2 are the bias terms of the first hidden layer and the second hidden layer of the feedforward neural network model respectively;
[0044] S4.3. Train the entity extraction submodel using the labeled dataset, and set the CRF loss function L CRF , and the expression is:
[0045] L CRF = -logP(y∣x);
[0046] Set the FFNN loss function L FFNN , and the expression is:
[0047] L FFNN = -∑ i y i ×log(y’ i );
[0048] Among them, y i is the true relationship type, and y’ i is the relationship type predicted by FFNN;
[0049] Finally, calculate the weighted average L' based on the CRF loss function L CRF and the FFNN loss function L FFNN , and the expression is:
[0050] L' = λ'L CRF +(1 - λ')L FFNN
[0051] Among them, λ' is the weight corresponding to the weighted average.
[0052] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the described knowledge extraction method for network security intention understanding are implemented.
[0053] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the described knowledge extraction method for network security intention understanding is implemented.
[0054] Advantages of the present invention:
[0055] A knowledge extraction method for network security intent understanding according to the present invention is based on an advanced pre-trained model. By continuing pre-training and fine-tuning on the corpus in the network security field to meet the specific requirements of network security threat analysis reports, efficient and accurate knowledge extraction from network security threat analysis reports is achieved. This method not only utilizes the powerful language understanding and generation capabilities of the pre-trained model but also enhances the applicability of the model in the network security field through targeted continued pre-training and fine-tuning.
[0056] A knowledge extraction method for network security intent understanding according to the present invention eliminates the gap between the general corpus and the network security corpus by preprocessing the data and adding a continued pre-training step. That is, the model first continues the pre-training task on the unannotated corpus in the preprocessed network security vertical field, and then fine-tunes for the downstream task to obtain the final prediction model. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 is a flowchart of a knowledge extraction method for network security intent understanding according to the present invention;
[0058] Figure 2 is a structural block diagram of a knowledge extraction method for network security intent understanding according to the present invention;
[0059] Figure 3 is a schematic diagram of the structure of a knowledge extraction model for network security threat analysis reports according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0060] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only a part of the embodiments of the present invention, rather than all of the specific embodiments. The components of the specific embodiments of the present invention usually described and shown in the drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.
[0061] Therefore, the detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents the selected specific embodiments of the present invention. All other specific embodiments obtained by those skilled in the art based on the specific embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0062] To further understand the content, features and effects of the present invention, the following specific embodiments are exemplified and coordinated with the attached Figure 1 - AttachedFigure 3 The detailed description is as follows:
[0063] Example 1:
[0064] A knowledge extraction method for network security intention understanding, comprising the following steps:
[0065] S1. Collect the original text of network security threat analysis reports, and divide it into training data and prediction data;
[0066] S2. Perform data preprocessing on the training data of the original text of the collected network security analysis reports, including IOC token replacement and structural token insertion, and then use a tokenizer to divide the processed original text to obtain a processed token sequence as the training set;
[0067] Further, the specific implementation method of step S2 includes the following steps:
[0068] S2.1. Perform IOC token replacement on the training data of the original text of the collected network security analysis reports. For the hash values, URLs, domain names, and IP addresses mentioned in the text, replace them with the corresponding tokens [HASH], [URL], [DOMAIN], and [IP] through regular replacement;
[0069] Further, in step S2.1, first divide the original text of the network security analysis report that exceeds the maximum length of the pre-trained model into multiple segments that meet the length requirements;
[0070] S2.2. Then perform structural token insertion, and insert a token representing the structure of the line in the text at the beginning of each line of the text. Set [h1], [h2], [p], and [tr] to represent the first-level heading, second-level heading, body paragraph, and table row respectively;
[0071] S2.3. Use a tokenizer to divide the processed text into a token sequence T = {w1, w2, w3,..., w n}, where w n is the nth token, and the sequence length does not exceed the maximum length of the pre-trained model;
[0072] S3. Select a pre-trained model, input the processed token sequence obtained in step S2 into the selected pre-trained model for masked language model (MLM) task training, and obtain the model weights after continued pre-training;
[0073] Further, the specific implementation method of step S3 includes the following steps:
[0074] S3.1. Select BERT as the pre-trained model and change the tokenizer of the pre-trained model for dividing the token sequence;
[0075] Furthermore, the selected pre-trained model is BERT, a bidirectional pre-trained language representation model based on Transformer. For each input token sequence T, its output BERT(T) has a one-to-one correspondence with the input sequence, that is, each output vector represents the representation information of its corresponding token;
[0076] S3.2. Set the input text as the processed token sequence T obtained in step S2, and perform a masking operation. Define the masking function as M(T, p), which accepts the processed token sequence T and the masking probability p as inputs. The masking function replaces the tokens in the text with the verification token [MASK] with the masking probability p. The masked sequence T’ is expressed as:
[0077] T’ = {w’1, w’2, w’3, …, w’ n}
[0078] where w’ n is the nth masked token;
[0079] S3.3. Use the pre-trained model selected in step S3.1, accept the masked sequence as the input, and perform masked language model (MLM) task training to output the prediction distribution of the masked word. The expression is:
[0080] P(w i ∣T’) = f(T’)[i]
[0081] where f(T’)[i] represents the prediction distribution output by the pre-trained model at position i, and P(w i ∣T’) is the prediction distribution of the ith masked token when the input is T’;
[0082] For each masked token, calculate the cross-entropy loss between the prediction distribution of the pre-trained model and the true word distribution. The expression is:
[0083] L i = -log P(w i ∣T’)
[0084] where L i is the cross-entropy loss at position i;
[0085] Then sum the losses at all masked positions to obtain the loss function L. The expression is:
[0086] L = ∑L i ;
[0087] Then, according to the loss function L, update the parameters of the pre-trained model through the backpropagation algorithm to obtain the model weights after continued pre-training.
[0088] S4. Based on the pre-trained model selected in step S3, perform model adjustment. Add an entity extraction sub-model and a relation extraction sub-model downstream of the pre-trained model. After using the model weights obtained from continued pre-training in step S3 for model training, obtain a trained knowledge extraction model for network security threat analysis reports;
[0089] Further, task definition: Define the specific entity and relation types to be identified in the knowledge extraction task, such as entities like "software", "intruder", "attack means", "defense means", "identity", "invasion indicator", etc., and relations like "use", "target", "impersonate", "possibly related", etc.
[0090] Further, data annotation: Mark the positions of the entities defined in the task and the relation types between entities on the training corpus, and perform preprocessing for model training.
[0091] Further, downstream model construction: Add a new model layer downstream of the pre-trained model to decode the hidden state of the pre-trained model and output the identified entity relation triples.
[0092] Further, the specific implementation method of step S4 includes the following steps:
[0093] S4.1. Add an entity extraction sub-model downstream of the pre-trained model, and use the conditional random field CRF as the entity extraction sub-model. The expression is:
[0094]
[0095] where x is the input sequence, and P(y|x) is the conditional probability distribution of the output sequence y given the input sequence x; f k (y i-1 ,y i ,x,i) is the feature function, and λ k is the weight of the feature function; Z(x) is the normalization factor;
[0096] Set the input sequence x = BERT(T) as the output feature vector sequence of the pre-trained model, y as the BIO label corresponding to each token, and the normalization factor is obtained by summing over all possible output sequences. The expression is:
[0097]
[0098] where y′ represents all possible values of the output label sequence, and y′ i is the label of the i-th token;
[0099] S4.2. Use the feed-forward neural network FFNN as the relation extraction sub-model. Based on the output of the entity extraction sub-model, obtain the positions of each entity, and calculate the average value x of the feature vectors of all the tokens of each entity. ei , and combine them pairwise and concatenate them with the feature vector corresponding to the first token x [CLS] to obtain the input x of the relation extraction sub-model. The expression is:
[0100] x = x [CLS] || x ei || x ej ;
[0101] Then, output the distribution of relation types between entity pairs through the feed-forward neural network. The expression is:
[0102] F(x) = max(0, xW1 + b1)W2 + b2
[0103] where W1 and W2 are the weight matrices of the first hidden layer and the second hidden layer of the feed-forward neural network model respectively, and b1 and b2 are the bias terms of the first hidden layer and the second hidden layer of the feed-forward neural network model respectively;
[0104] S4.3. Use the labeled dataset to train the entity extraction sub-model, and set the CRF loss function L CRF . The expression is:
[0105] L CRF = -logP(y|x);
[0106] Set the FFNN loss function L FFNN . The expression is:
[0107] L FFNN = -∑ i y i × log(y' i );
[0108] where y i is the true relation type, and y' i is the relation type predicted by FFNN;
[0109] Finally, calculate the weighted average L' based on the CRF loss function L CRF and the FFNN loss function L FFNN . The expression is:
[0110] L' = λ'L CRF + (1 - λ')L FFNN
[0111] where λ' is the weight value corresponding to the weighted average.
[0112] Further, backpropagation: Update the parameters of the model through the backpropagation algorithm according to the weighted average L'.
[0113] S5. Use the prediction data of the original text of the network security threat analysis report obtained in step S1 to perform model prediction on the trained knowledge extraction model for the network security threat analysis report, and conduct result verification and optimization.
[0114] Further, input processing: Perform the aforementioned preprocessing work on the network security threat analysis report that needs to extract knowledge as the input text. Model inference: Use the fine-tuned model to perform inference on the input text, identify and extract key information. The entity extraction sub-model outputs the BIO tags of the token sequence, and the relationship extraction sub-model outputs the relationship type distribution between entity pairs. Post-processing: Post-process the extracted information, obtain the type of the entity according to the BIO tags, and then obtain the entity relationship triple according to the relationship type. After deduplication, merging, formatting, etc., to generate a structured knowledge representation.
[0115] Further, the result verification and optimization include:
[0116] Manual review: Manually review the extraction results to ensure the accuracy and integrity of the information;
[0117] Feedback loop: According to the manual review results, further fine-tune the model to improve the accuracy of knowledge extraction.
[0118] Continuous update: As the network security threats continue to evolve, regularly update the labeled dataset and the model to maintain the timeliness and accuracy of the method.
[0119] Embodiment 2:
[0120] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the knowledge extraction method for network security intention understanding described in Embodiment 1 are implemented.
[0121] The computer device of the present invention may be a device including a processor and a memory, such as a single-chip microcomputer including a central processing unit. And, when the processor is used to execute the computer program stored in the memory, the steps of the above-mentioned knowledge extraction method for network security intention understanding are implemented.
[0122] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0123] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0124] Embodiment 3:
[0125] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the knowledge extraction method for network security intent understanding described in Embodiment 1.
[0126] The computer-readable storage medium of the present invention may be any form of storage medium readable by the processor of a computer device, including but not limited to non-volatile memory, volatile memory, ferroelectric memory, etc. A computer program is stored on the computer-readable storage medium. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of the above-mentioned knowledge extraction method for network security intent understanding can be implemented.
[0127] The computer program includes computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, removable hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0128] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0129] Although the present application has been described above with reference to specific embodiments, various improvements can be made to it and components therein can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application can be combined with each other in any way, and the exhaustive description of these combinations is not given in this specification only for the sake of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A knowledge extraction method for network security intention understanding, characterized in that: The steps include: S1. Collect the original text of the network security threat analysis report and divide it into training data and prediction data; S2. Perform data preprocessing on the training data of the collected original text of the network security analysis report, including IOC word unit replacement and structural word unit insertion, and then use a word segmenter to divide the processed original text to obtain a processed word unit sequence as a training set; S3. Select a pre-training model, input the processed word sequence obtained in step S2 into the selected pre-training model for masked language model MLM task training, and obtain the model weight after continued pre-training; S4. Based on the pre-trained model selected in step S3, the model is adjusted, and an entity extraction sub-model and a relationship extraction sub-model are added downstream of the pre-trained model. After the model weights obtained in step S3 are used to continue pre-training, the model is trained to obtain a trained knowledge extraction model for network security threat analysis report; S5. Use the prediction data of the original text of the network security threat analysis report obtained in step S1 to perform model prediction on the trained knowledge extraction model for the network security threat analysis report, and perform result verification and optimization.
2. According to claim 1, a knowledge extraction method for network security intention understanding is characterized in that: The specific implementation method of step S2 includes the following steps: S2.
1. Perform IOC word replacement on the training data of the original text of the collected network security analysis report, and replace the hash values, URLs, domain names, and IP addresses mentioned in the text with the corresponding words [HASH], [URL], [DOMAIN], and [IP] through regular replacement; S2.
2. Then insert a structural word element, insert a word element representing the structure of the line in the text at the beginning of each line of the text, and set [h1], [h2], [p], and [tr] to represent the first-level title, second-level title, body paragraph, and table row respectively; S2.
3. Use a tokenizer to divide the processed text into a token sequence T = {w1, w2, w3, ..., w n }, where w n is the nth word, and the sequence length does not exceed the maximum length of the pre-trained model.
3. A knowledge extraction method for network security intention understanding according to claim 2, characterized in that: In step S2.1, the original text of the network security analysis report that exceeds the maximum length of the pre-trained model is first divided into multiple segments that meet the length requirements.
4. A knowledge extraction method for network security intention understanding according to claim 3, characterized in that: The specific implementation method of step S3 includes the following steps: S3.
1. Select BERT as the pre-trained model and change the tokenizer of the pre-trained model to divide the token sequence; S3.
2. Set the input text to the processed word sequence T obtained in step S2, perform masking operation, define the masking function as M(T,p), accept the processed word sequence T and masking probability p as input, and the masking function replaces the word in the text with the verification mark [MASK] with the masking probability p. The masked sequence T' is expressed as: T’={w’1,w’2,w’3,…,w’ n } Among them, w' n is the nth masked word; S3.
3. Use the pre-trained model selected in step S3.1, accept the masked sequence as input, perform masked language model MLM task training, and output the predicted distribution of the masked words, expressed as: P(w i ∣T’)=f(T’)[i] Where f(T')[i] represents the predicted distribution of the output of the pre-trained model at position i, P(w i |T') is the predicted distribution of the i-th masked word when the input is T'; For each masked token, the cross entropy loss between the distribution predicted by the pre-trained model and the distribution of the real word is calculated, expressed as: L i =-log P(w i ∣T’) Among them, L i is the cross entropy loss at position i; Then sum the losses of all masked positions to get the loss function L, which is expressed as: L=∑L i ; Then, according to the loss function L, the parameters of the pre-trained model are updated through the back-propagation algorithm to obtain the model weights after further pre-training.
5. A knowledge extraction method for network security intention understanding according to claim 4, characterized in that: The specific implementation method of step S4 includes the following steps: S4.
1. Add an entity extraction sub-model downstream of the pre-trained model and use the conditional random field CRF as the entity extraction sub-model. The expression is: Where x is the input sequence, P(y|x) is the conditional probability distribution of the output sequence y given the input sequence x; f k (y i-1 ,y i ,x,i) is the characteristic function, λ k is the weight of the characteristic function; Z(x) is the normalization factor; Set the input sequence x=BERT(T) as the output feature vector sequence of the pre-trained model, y as the BIO label corresponding to each word, and the normalization factor is obtained by summing all possible output sequences, expressed as: Among them, y′ represents all possible values of the output label sequence, y′ i is the label of the i-th word; S4.
2. Use the feedforward neural network FFNN as the relation extraction sub-model, obtain the location of each entity according to the output of the entity extraction sub-model, and calculate the average value x of the feature vectors of all word units of each entity. ei , and combine them in pairs with the sentence-initial element x [CLS] The corresponding feature vectors are concatenated to obtain the input x of the relation extraction sub-model, which is expressed as: x=x [CLS] ||x ei ||x ej ; Then the feedforward neural network outputs the relationship type distribution between entity pairs, expressed as: F(x)=max(0,xW1+b1)W2+b2 Among them, W1 and W2 are the weight matrix of the first hidden layer and the weight matrix of the second hidden layer of the feedforward neural network model, respectively; b1 and b2 are the bias term of the first hidden layer and the bias term of the second hidden layer of the feedforward neural network model, respectively; S4.
3. Use the labeled dataset to train the entity extraction sub-model and set the CRF loss function L CRF , the expression is: L CRF =-logP(y∣x); Set the FFNN loss function L FFNN , the expression is: L FFNN =-∑ i and i ×log(y' i ); Among them, y i is the real relationship type, y' i is the type of relationship predicted by FFNN; Finally, based on the CRF loss function L CRF And the FFNN loss function L FFNN Calculate the weighted average L', the expression is: L'=λ'L CRF +(1-λ')L FFNN Among them, λ' is the weight corresponding to the weighted average.
6. An electronic device, characterized in that: It includes a memory and a processor, the memory stores a computer program, and the processor implements the steps of a knowledge extraction method for network security intention understanding as described in any one of claims 1 to 5 when executing the computer program.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements a knowledge extraction method for network security intention understanding as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Training method and device for network security threat knowledge extraction model
CN116579426A
Network space knowledge extraction method and device based on rule enhancement prompt learning
CN117391083A
Knowledge extraction method and system fusing pre-training language model
CN117521802A
Text encoder event extraction pre-training method and system and storage medium
CN118485078A
BERT model-based event atlas intelligent construction and analysis method and device
CN118747222A
Cited By
Network traffic early warning method, electronic equipment and storage medium
CN121037054A