Network security named entity recognition model construction method, electronic device and storage medium based on prompt learning idea
Through a prompt learning method, data is collected and marked for network security named entity recognition tasks, and data is augmented, which solves the problem that the model is difficult to learn the characteristics of network security entity in the pre-training stage, and improves the accuracy and generalization ability of recognition.
Patent Information
- Application Number
- CN202411190954.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-08-28
AI Technical Summary
In the network security named entity recognition task, it is difficult for the model to learn specific network security entity characteristics in the pre-training stage, resulting in poor application results when facing specific tasks.
Using a prompt learning method, by collecting and annotating network security data, generating annotation sequences, and data augmentation, we guide the pre-training model to learn more targeted network security features during the process of continuing pre-training and fine-tuning.
It improves the accuracy and generalization ability of network security named entity recognition, so that the model can more effectively identify entity when facing different network security scenarios.
Smart Images

Figure CN119167935B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security named entity recognition, and specifically relates to a network security named entity recognition model construction method based on prompt learning ideas, an electronic device and a storage medium. Background Art
[0002] With the development of deep learning technology, the pre-training + fine-tuning paradigm has become mainstream. The pre-training + fine-tuning paradigm is a widely used method in the training process of deep learning models. This method is usually divided into two main stages: pre-training and fine-tuning.
[0003] The pre-training phase refers to training a general model on a large-scale dataset. In this phase, the model learns some common features and patterns. These datasets are usually large-scale datasets in general fields, such as ImageNet (for image processing) or Wikipedia (for natural language processing). The purpose of pre-training is to allow the model to obtain preliminary feature representations on a wide range of data, which can be reused in various tasks.
[0004] The fine-tuning stage is to further train the model on the dataset of a specific task based on pre-training. In this stage, the model adjusts its weights and parameters by training on the specific task data to better adapt to the needs of the specific task. Fine-tuning can be regarded as a customization of the pre-trained model so that it can perform better in a specific application scenario.
[0005] The pre-training + fine-tuning paradigm has been successful in many fields. For example, the BERT model has performed well in various NLP tasks such as question-answering systems, sentiment analysis, and named entity recognition. In the field of computer vision, pre-trained ResNet, VGG and other models have also been widely used in tasks such as image classification and object detection. However, this approach also has many problems. Taking the network security named entity recognition (NER) task as an example, we can see the significant differences and challenges between the pre-training and fine-tuning stages.
[0006] In the pre-training stage, the data used by the model is usually unlabeled network security data. This data may include various forms of network logs, communication records, threat intelligence reports, etc. Although this data is rich, due to the lack of clear annotations, the model cannot directly learn the fine-grained information required for specific tasks. To address this challenge, a common method is to use a masked language model (MLM) or an auto-regressive task. For example, when using the MLM method, the model randomly masks some words and then predicts these masked words to learn contextual information and the relationship between words. Although these methods allow the model to be trained on large-scale unlabeled data, they mainly focus on the statistical characteristics and grammatical structure of the language itself, rather than specific entities or relationships. At the same time, this method also has a problem that the masked tokens in the traditional MLM process are random, and the model may mask out key network security entities, making it difficult for the model to capture the features of the entities required to be identified in downstream tasks during the pre-training process. This problem can be supplemented by massive data in general fields to make up for the missing knowledge, but network security data is often small in scale, so this problem is particularly critical in network security data. Summary of the invention
[0007] The problem to be solved by the present invention is that the general language and context information learned by the model in the pre-training stage is directly applied when facing specific network security entity recognition tasks. A network security named entity recognition model construction method, electronic device and storage medium based on the prompt learning idea are proposed.
[0008] To achieve the above object, the present invention is implemented through the following technical solutions:
[0009] A method for constructing a network security named entity recognition model based on prompt learning ideas includes the following steps:
[0010] S1. Collect network security data and obtain the network space security data sequence X = (x1, x2, ... x i …x n ) where x i is the i-th word, to be used;
[0011] S2. Setting a labeling set and generating a labeling sequence based on labeling rules, wherein the labeling rules include the entity type of the labeled data and the labeled data does not belong to any entity;
[0012] S3. Based on the annotation rule of step S2, the cyberspace security data sequence obtained in step S1 is segmented, and then corresponding annotation subsequences are generated to obtain processed cyberspace security data;
[0013] S4. Based on the annotation sequence obtained in step S2, define data augmentation rules, perform data augmentation on the processed cyberspace security data obtained in step S3, and obtain a data augmented cyberspace security data set;
[0014] S5. Use the data-augmented cyberspace security dataset obtained in step S4 to continue pre-training and fine-tuning the pre-trained model to obtain a network security named entity recognition data extraction model.
[0015] Furthermore, the specific implementation method of step S2 includes the following steps:
[0016] S2.1. Set the entity types of the labeled data, including 11 types of network security data, AttackPattern, the method used by the actor to attack the target, Campaign, the measures used to prevent or respond to the attack, Identity, Indicator, a pattern that can be used to detect malicious network behavior, Intrusion Set, a grouping set of hostile behaviors and resources with common attributes, Organization, illegal software used by the threat actor to carry out the attack, Malware, legal software used by the threat actor to carry out the attack, Vulnerability, and operating system System that is considered to be related to malicious operations;
[0017] S2.2. Set the annotation set T = (AttackPattern, CampaignMalware, Organization, System, O…), where O represents that the annotated data does not belong to any entity;
[0018] Generate a labeling sequence Y=(y1,y2,…y i …y n ), where y i ∈T.
[0019] Furthermore, in step S3, the cyberspace security data sequence X is processed into sentences according to natural sentences to obtain subsequences X1, X2, X3...X n , the corresponding generated annotation subsequence Y1, Y2, Y3...Y n .
[0020] Furthermore, step S4 defines data augmentation rules. If any X n If multiple entities of different types appear in the annotation set T, then the label y n and the original data x nCombined with augmented data, the specific steps are as follows:
[0021] S4.1. Assume that the rule set R is as follows:
[0022]
[0023] Among them, m represents when x m The corresponding position when the mark is Malware, o represents when x o The mark is Organization, and s1 and s2 represent the positions corresponding to the first and second appearances of System in the label sequence.
[0024] S4.2. Generate augmented dataset based on rule set R If the subsequence of the original statement set X cannot satisfy any rule R, then the corresponding Leave it empty, and then Concatenate with X to get the data-augmented cyberspace security dataset
[0025] Furthermore, the specific implementation method of step S5 includes the following steps:
[0026] S5.1. For X final The part belonging to the original sequence X is determined by the labeled sequence Y, and the mask probability p mask Depends on the corresponding label y i Is it equal to that the labeled data does not belong to any entity O? The expression is:
[0027]
[0028] Where Mask(x i ,y i ) is to mask the sequence X, that is, to take the probability p mask Decide whether to i Perform mask replacement;
[0029] S5.2. For X final The augmented sequence Part, combined with downstream tasks to predict the model The y in n Part, set the mask probability p mask =0.5 pairs All the y n Perform probability replacement to complete the mask operation on the augmented sequence;
[0030] S5.3. Combine the data sets obtained in step S5.1 and step S5.2 to obtain the pre-training data set Input the pre-trained dataset into the BERT model to generate three embeddings, namely word embedding w i , position embedding p i and segment embedding d i , and get the input representation r i =w i +p i +d i ;
[0031] r i Passed as input to the BERT model to get the output e i Then, the bidirectional long short-term memory neural network BiLSTM is combined with CRF for downstream network security named entity recognition, and the expression is:
[0032] e i =BERT(r i )
[0033] The BiLSTM layer embeds the word into the sequence E = {e1, e2, …, e n} is passed into the bidirectional long short-term memory neural network to obtain the context representation h of each word i , the expression is:
[0034]
[0035] in, is the hidden state of the corresponding position of the forward LSTM, is the hidden state of the corresponding position of the reverse LSTM, [;] represents the concatenation operation of the vector;
[0036] S5.4. The linear layer represents the context of each word in the bidirectional long short-term memory neural network h i Pass it into the linear layer, map it to the label space, and get the score s i , the expression is:
[0037] s i =Wh i +b
[0038] Among them, W and b are the learned parameters;
[0039] S5.5. The CRF layer is used to model the dependency between label sequences and define the transfer matrix A, where A y',y represents the score transferred from label y' to y, and the expression is:
[0040]
[0041] in, Indicates that at time step i, the label is y iscore;
[0042] S5.6. Set the normalization factor Z(X) to be the exponential sum of the scores of all possible label sequences, expressed as:
[0043]
[0044] The conditional probability of calculating the label sequence Y is:
[0045]
[0046] S5.7. Set the training objective to maximize the conditional probability, use the Viterbi algorithm to find the label sequence with the highest score, and minimize the loss expression as:
[0047]
[0048] Then save the model parameters when minimizing the loss, complete the downstream tasks, and obtain the network security named entity recognition data extraction model.
[0049] An electronic device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a method for constructing a network security named entity recognition model based on a prompt learning idea are implemented.
[0050] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a method for constructing a network security named entity recognition model based on a prompt learning idea.
[0051] Beneficial effects of the present invention:
[0052] The method for constructing a network security named entity recognition model based on the prompt learning idea described in the present invention first uses downstream tasks to guide data augmentation to perform prompt augmentation operations, and then masks the augmented data to prevent the model from randomly masking out key information during continued pre-training. Then, the augmented data is used together with the original data to continue pre-training and fine-tuning the pre-trained model, so that the language and context information learned by the model in the pre-training stage can be directly applied when facing specific network security entity recognition tasks.
[0053] The method for constructing a network security named entity recognition model based on the idea of prompt learning described in the present invention can improve the accuracy of network security named entity recognition: by collecting and processing network security data, generating a labeling sequence based on labeling rules, and performing data augmentation to ensure the diversity and coverage of the data set, thereby improving the model's recognition accuracy for different types of network security entities.
[0054] The method for constructing a network security named entity recognition model based on the idea of prompt learning described in the present invention can enhance the generalization ability of the model: by defining data augmentation rules and performing data augmentation on the processed cyberspace security data, the diversity of training data is increased, and the generalization ability of the model for unseen data is enhanced, so that the model can more effectively perform entity recognition when facing different network security scenarios.
[0055] The method for constructing a network security named entity recognition model based on the idea of prompt learning described in the present invention can improve the adaptability and robustness of the model: continuing pre-training and fine-tuning operations are performed based on the pre-trained model, utilizing the general characteristics of the pre-trained model and the special characteristics after fine-tuning while combining data augmentation rules customized according to downstream tasks, so that the model has better adaptability and robustness in the field of network security and can cope with different types of network security entity recognition tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 The present invention provides a flowchart of a method for constructing a network security named entity recognition model based on the idea of prompt learning. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solution and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only part of the embodiments of the present invention, rather than all of the specific embodiments. The components of the specific embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.
[0058] Therefore, the following detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected specific embodiments of the present invention. Based on the specific embodiments of the present invention, all other specific embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0059] In order to further understand the content, features and effects of the present invention, the following specific implementation methods are given as examples, and the attached Figure 1 The detailed instructions are as follows:
[0060] Embodiment 1:
[0061] A method for constructing a network security named entity recognition model based on prompt learning ideas includes the following steps:
[0062] S1. Collect network security data and obtain the network space security data sequence X = (x1, x2, ... x i …x n ) where x i is the i-th word, to be used;
[0063] S2. Setting a labeling set and generating a labeling sequence based on labeling rules, wherein the labeling rules include the entity type of the labeled data and the labeled data does not belong to any entity;
[0064] Furthermore, the specific implementation method of step S2 includes the following steps:
[0065] S2.1. Set the entity types of the labeled data, including 11 types of network security data, AttackPattern, the method used by the actor to attack the target, Campaign, the measures used to prevent or respond to the attack, Identity, Indicator, a pattern that can be used to detect malicious network behavior, Intrusion Set, a grouping set of hostile behaviors and resources with common attributes, Organization, illegal software used by the threat actor to carry out the attack, Malware, legal software used by the threat actor to carry out the attack, Vulnerability, and operating system System that is considered to be related to malicious operations;
[0066] S2.2. Set the annotation set T = (AttackPattern, CampaignMalware, Organization, System, O…), where O represents that the annotated data does not belong to any entity;
[0067] Generate a labeling sequence Y=(y1,y2,…y i …y n ), where y i ∈T.
[0068] S3. Based on the annotation rule of step S2, the cyberspace security data sequence obtained in step S1 is segmented, and then corresponding annotation subsequences are generated to obtain processed cyberspace security data;
[0069] Furthermore, in step S3, the cyberspace security data sequence X is processed into sentences according to natural sentences to obtain subsequences X1, X2, X3...X n , the corresponding generated annotation subsequence Y1, Y2, Y3...Y n ;
[0070] S4. Based on the annotation sequence obtained in step S2, define data augmentation rules, perform data augmentation on the processed cyberspace security data obtained in step S3, and obtain a data augmented cyberspace security data set;
[0071] Furthermore, step S4 defines data augmentation rules. If any X n If multiple entities of different types appear in the annotation set T, then the label y n and the original data x n Combined with augmented data, the specific steps are as follows:
[0072] S4.1. Assume that the rule set R is as follows:
[0073]
[0074] Among them, m represents when x m The corresponding position when the mark is Malware, o represents when x o The mark is Organization, and s1 and s2 represent the positions corresponding to the first and second appearances of System in the label sequence.
[0075] S4.2. Generate augmented dataset based on rule set R If the subsequence of the original statement set X cannot satisfy any rule R, then the corresponding Leave it empty, and then Concatenate with X to get the data-augmented cyberspace security dataset
[0076] S5. Use the cyberspace security dataset augmented with data obtained in step S4 to continue pre-training and fine-tuning the pre-trained model to obtain a cybersecurity named entity recognition data extraction model;
[0077] Furthermore, the specific implementation method of step S5 includes the following steps:
[0078] S5.1. For X final The part belonging to the original sequence X is determined by the labeled sequence Y, and the mask probability p mask Depends on the corresponding label y i Is it equal to that the labeled data does not belong to any entity O? The expression is:
[0079]
[0080] Where Mask(x i ,y i ) is to mask the sequence X, that is, to take the probability pmask Decide whether to i Perform mask replacement;
[0081] S5.2. For X final The augmented sequence Part, combined with downstream tasks to predict the model The y in n Part, set the mask probability p mask =0.5 pairs All the y n Perform probability replacement to complete the mask operation on the augmented sequence;
[0082] S5.3. Combine the data sets obtained in step S5.1 and step S5.2 to obtain the pre-training data set Input the pre-trained dataset into the BERT model to generate three embeddings, namely word embedding w i , position embedding p i and segment embedding d i , and get the input representation r i =w i +p i +d i ;
[0083] r i Passed as input to the BERT model to get the output e i Then, the bidirectional long short-term memory neural network BiLSTM is combined with CRF for downstream network security named entity recognition, and the expression is:
[0084] e i =BERT(r i )
[0085] The BiLSTM layer embeds the word into the sequence E = {e1, e2, …, e n} is passed into the bidirectional long short-term memory neural network to obtain the context representation h of each word i , the expression is:
[0086]
[0087] in, is the hidden state of the corresponding position of the forward LSTM, is the hidden state of the corresponding position of the reverse LSTM,
[0088] [;] represents the concatenation operation of vectors;
[0089] S5.4. The linear layer represents the context of each word in the bidirectional long short-term memory neural network h i Pass it into the linear layer, map it to the label space, and get the score si , the expression is:
[0090] s i =Wh i +b
[0091] Among them, W and b are the learned parameters;
[0092] S5.5. The CRF layer is used to model the dependency between label sequences and define the transfer matrix A, where A y',y represents the score transferred from label y' to y, and the expression is:
[0093]
[0094] in, Indicates that at time step i, the label is y i score;
[0095] S5.6. Set the normalization factor Z(X) to be the exponential sum of the scores of all possible label sequences, expressed as:
[0096]
[0097] The conditional probability of calculating the label sequence Y is:
[0098]
[0099] S5.7. Set the training objective to maximize the conditional probability, use the Viterbi algorithm to find the label sequence with the highest score, and minimize the loss expression as:
[0100]
[0101] Then save the model parameters when minimizing the loss, complete the downstream tasks, and obtain the network security named entity recognition data extraction model.
[0102] Embodiment 2:
[0103] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a method for constructing a network security named entity recognition model based on a prompt learning idea described in Example 1 are implemented.
[0104] The computer device of the present invention may be a device including a processor and a memory, such as a single chip microcomputer including a central processing unit, etc. Furthermore, the processor is used to implement the steps of the method for constructing a network security named entity recognition model based on the prompt learning idea when executing the computer program stored in the memory.
[0105] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0106] The memory may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0107] Embodiment 3:
[0108] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a method for constructing a network security named entity recognition model based on a prompt learning idea as described in Example 1.
[0109] The computer-readable storage medium of the present invention can be any form of storage medium that can be read by a processor of a computer device, including but not limited to non-volatile memory, volatile memory, ferroelectric memory, etc. A computer program is stored on the computer-readable storage medium. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of the above-mentioned method for constructing a network security named entity recognition model based on the idea of prompt learning can be implemented.
[0110] The computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. The computer readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer readable media do not include electric carrier signals and telecommunication signals.
[0111] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0112] Although the present application has been described above with reference to specific embodiments, various modifications may be made thereto and parts thereof may be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application may be used in combination with each other in any manner, and the fact that these combinations are not exhaustively described in this specification is only for the sake of omitting space and saving resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A method for constructing a network security named entity recognition model based on the idea of prompt learning, characterized in that: The steps include: S1. Collect network security data and obtain the network space security data sequence X = (x1, x 2, …x i …x n ) where x i is the i-th word, to be used; S2. Setting a labeling set and generating a labeling sequence based on labeling rules, wherein the labeling rules include the entity type of the labeled data and the labeled data does not belong to any entity; The specific implementation method of step S2 includes the following steps: S2.
1. Set the entity types of the labeled data, including 11 types of network security data, AttackPattern, the method used by the actor to attack the target, Campaign, the measures used to prevent or respond to the attack, Identity, Indicator, the pattern used to detect malicious network behavior, Intrusion Set, a grouping set of hostile behaviors and resources with common attributes, Organization, the illegal software used by the threat actor to carry out the attack, Malware, the legal software used by the threat actor to carry out the attack, Vulnerability, and the operating system System that is considered to be related to malicious operations; S2.
2. Set the annotation set T = (AttackPattern, Campaign, Malware, Organization, System, O…), where O represents that the annotated data does not belong to any entity; Generate a labeling sequence Y=(y1,y 2, …y i …y n ), where y i ∈T; S3. Based on the annotation rule of step S2, the cyberspace security data sequence obtained in step S1 is segmented, and then corresponding annotation subsequences are generated to obtain processed cyberspace security data; In step S3, the cyberspace security data sequence X is processed into sentences according to natural sentences to obtain subsequences of X, X1, X2, X3...X n , the corresponding generated annotation subsequence Y1, Y2, Y3...Y n ; S4. Based on the annotation sequence obtained in step S2, define data augmentation rules, perform data augmentation on the processed cyberspace security data obtained in step S3, and obtain a data augmented cyberspace security data set; Step S4 defines data augmentation rules. If any X n If multiple entities of different types appear in the annotation set T, the annotation y n and the original data x n Combined with augmented data, the specific steps are as follows: S4.
1. Assume that the data augmentation rule set R is as follows: Among them, m represents when x m The corresponding position when the mark is Malware, o represents when x o The mark is Organization, and s1 and s2 represent the positions corresponding to the first and second appearances of System in the label sequence. S4.
2. Generate augmented dataset based on data augmentation rule set R If the subsequence of the cyberspace security data sequence X cannot satisfy any rule of the data augmentation rule set R, then the corresponding Leave it empty, and then Concatenate with X to get the data-augmented cyberspace security dataset S5. Use the cyberspace security dataset with data augmentation obtained in step S4 to continue pre-training and fine-tuning the pre-trained model to obtain a cybersecurity named entity recognition data extraction model; The specific implementation method of step S5 includes the following steps: S5.
1. For X final The part of the cyberspace security data sequence X in the mask is determined according to the label sequence Y. The mask probability p mask Depends on the corresponding label y i Is it equal to that the labeled data does not belong to any entity O? The expression is: Where Mask(x i ,y i ) is to mask the cyberspace security data sequence X, that is, to take the probability p mask Decide whether to i Perform mask replacement; S5.
2. For X final The augmented dataset part, combined The y in n , set the mask probability p mask =0.5 pairs All the y n Perform probability replacement to complete the mask operation on the augmented data set; S5.
3. Combine the data sets obtained in step S5.1 and step S5.2 to obtain the pre-training data set Input the pre-trained dataset into the BERT model to generate three embeddings, namely word embedding w i , position embedding p i and segment embedding d i , and get the input representation r i =w i +p i +d i ; r i Passed as input to the model BERT to get the output e i Then, the bidirectional long short-term memory neural network BiLSTM is combined with CRF for downstream network security named entity recognition, and the expression is: e i =BERT(r i ) The BiLSTM layer embeds the word into the sequence E = {e1, e2, …, e n } is passed into the bidirectional long short-term memory neural network to obtain the context representation h of each word i , the expression is: in, is the hidden state of the corresponding position of the forward LSTM, is the hidden state of the corresponding position of the reverse LSTM, [;] represents the concatenation operation of the vector; S5.
4. The linear layer represents the context of each word in the bidirectional long short-term memory neural network h i Pass it into the linear layer, map it to the annotation space, and get the score s i , the expression is: Among them, W and b are the learned parameters; S5.
5. The CRF layer is used to model the dependencies between annotation sequences and define the transfer matrix A, where A y',y represents the score transferred from the label y' to y, and the expression is: in, Indicates that at time step i, labeled y i score; S5.
6. Set the normalization factor Z(X) to be the exponential sum of the scores of all possible labeled sequences, expressed as: The conditional probability of the labeled sequence Y is calculated as: S5.
7. Set the training objective to maximize the conditional probability, use the Viterbi algorithm to find the highest-scoring annotation sequence, and minimize the loss expression as: Then save the model parameters when minimizing the loss, complete the downstream tasks, and obtain the network security named entity recognition data extraction model.
2. An electronic device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a method for constructing a network security named entity recognition model based on a prompt learning idea as described in claim 1 when executing the computer program.
3. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing a network security named entity recognition model based on the prompt learning concept described in claim 1 is implemented.
Citation Information
Patent Citations
Network security named entity identification method and device, equipment and storage medium
CN116545779A
Network security named entity identification method based on threat intelligence
CN116611436A