Network security entity identification method and system based on multilayer channel attention
By adopting the multi-layer channel attention mechanism of BERT, BiLSTM, MSCA and CRF in the network security entity recognition method, the shortcomings of the existing methods in recognition accuracy and computing efficiency are solved, and high-precision and low resource consumption are achieved.
Patent Information
- Application Number
- CN202510068949.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
Smart Images

Figure CN119990129A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security entity recognition, and more specifically, relates to a network security entity recognition method and system based on multi-layer channel attention. Background Art
[0002] At present, the number and complexity of network security incidents are increasing, and security threats such as abnormal activities, malware, and intrusion behaviors in the network environment are becoming more and more hidden, resulting in the increasing difficulty of threat detection and security incident analysis. As a key technology for dealing with security threats, Network Security Entity Recognition (NSER) aims to automatically extract key information from a large amount of unstructured text data such as security logs, attack reports, and event alerts, and identify the types of entities involved, such as IP addresses, domain names, malware names, attack methods, etc. In order to improve the accuracy and adaptability of entity recognition, researchers have gradually explored combining BERT (Bidirectional Encoder Representations from Transformers, pre-trained language model) with other network modules. However, existing methods usually find it difficult to achieve lightweight design when integrating these modules, and the synergy between different modules is limited, failing to strike a balance between network complexity, recognition accuracy, and computational efficiency. Traditional named entity recognition methods may ignore the correlation between labels, resulting in illogical entity label output,
[0003] Prior art document 1 (CN118982025A) discloses a network security entity recognition method based on BERT-BiLSTM-CCA-CRF. Its shortcomings are that the CCA module only uses a single-scale convolution operation, fails to effectively capture multi-scale features, has limited collaboration between modules, and has lost feature information transmission. In addition, the label dependency of the CRF layer is not flexible enough in dynamic scenarios. Summary of the invention
[0004] In order to solve the deficiencies in the prior art, the present invention provides a network security entity recognition method and system based on multi-layer channel attention. The method adopts the BERT model in the embedding layer to comprehensively capture context information, introduces a BiLSTM (Bi-directional Long Short-Term Memory) model to enhance feature representation, and introduces MSCA (Multi-Scale Channel Attention) to fully capture important features. A CRF (Conditional Random Field) model is added to the output layer to capture the dependency between entity labels. Finally, a variety of entity types in network security texts can be accurately identified, and the monitoring and analysis capabilities of network security events can be improved.
[0005] The present invention adopts the following technical solution.
[0006] A first aspect of the present invention provides a network security entity recognition method based on multi-layer channel attention, comprising:
[0007] Perform data preprocessing on real cybersecurity entity text datasets;
[0008] Construct a named entity recognition model based on multi-layer channel attention, input the preprocessed network security entity text dataset into the named entity recognition model for model training, and obtain a trained named entity recognition model;
[0009] Inputting the actual network security entity text dataset into the trained named entity recognition model and outputting a label sequence;
[0010] According to the output label sequence, the entity category of each word in the actual network security entity text dataset is identified to realize network security entity recognition based on multi-layer channel attention.
[0011] Preferably, data preprocessing is performed on the real network security entity text dataset, specifically including:
[0012] Use the word segmentation tool to segment the original text in the real cybersecurity entity text dataset and annotate each word with a corresponding label;
[0013] Remove stop words from the segmented and annotated text;
[0014] Normalize the text after removing stop words;
[0015] Enhance the standardized text data;
[0016] The association rule algorithm is used to identify potential dangerous events in the data-enhanced text, and the preprocessed network security entity text dataset is obtained.
[0017] Preferably, a named entity recognition model based on multi-layer channel attention is constructed, and the preprocessed network security entity text dataset is input into the named entity recognition model for model training to obtain a trained named entity recognition model, which specifically includes:
[0018] A named entity recognition model based on multi-layer channel attention is constructed based on the BERT model, BiLSTM model, MSCA and CRF model;
[0019] The preprocessed network text dataset is divided into a test set and a training set;
[0020] The named entity recognition model is trained using 5-fold cross validation, the training data set is divided into 5 parts, and 5 rounds of cross validation are performed. In each round of cross validation, 1 part of the data is selected as the validation set, and the remaining 4 parts of the data are used as the training set. Each round of cross validation obtains an optimal named entity recognition model;
[0021] After 5-fold cross-validation, the test set is used to verify the optimal named entity recognition model obtained in each round of cross-validation to obtain the optimal named entity recognition model of 5-fold cross-validation.
[0022] Preferably, the named entity recognition model is trained using 5-fold cross validation, the training data set is divided into 5 parts, 5 rounds of cross validation are performed, 1 part of the data is selected as the validation set in each round of cross validation, and the remaining 4 parts of the data are used as the training set, and each round of cross validation obtains an optimal named entity recognition model, specifically including:
[0023] Inputting a single training set into the BERT model of the named entity recognition model to generate a context-sensitive embedding vector;
[0024] Inputting the generated embedding vector into the BiLSTM model of the named entity recognition model to perform bidirectional feature extraction of the text sequence to obtain a bidirectional integrated feature;
[0025] The bidirectional integration features extracted by the BiLSTM model are passed into the MSCA of the named entity recognition model to extract multi-scale features to obtain multi-scale features;
[0026] Input the generated multi-scale features into the CRF model of the named entity recognition model, decode the label of each position in the multi-scale features, and output the current optimal label sequence;
[0027] Repeatedly inputting the remaining single training sets in the training set into the named entity recognition model, and obtaining named entity recognition models with four different output optimal label sequences after four trainings;
[0028] A learning rate decay strategy is adopted to dynamically adjust the learning rate of the model according to the performance of the named entity recognition model with the optimal label sequence of different outputs on the validation set, and the named entity recognition model with the best performance among the four models is determined as the optimal named entity recognition model for this round of cross-validation.
[0029] Preferably, a single training set is input into the BERT model of the named entity recognition model to generate a context-sensitive embedding vector, and the generated embedding vector is input into the BiLSTM model of the named entity recognition model to perform bidirectional feature extraction of the text sequence, specifically including:
[0030] The BERT model captures the semantic information of the text through a multi-layer Transformer structure, generates context-sensitive word embeddings, and converts word embeddings into context-sensitive embedding vectors;
[0031] The BiLSTM model processes the forward and reverse sequences of the text in the embedded vector at the same time, captures the dependency between the previous and next sentences in the text, extracts bidirectional features, and integrates the extracted bidirectional features to obtain bidirectional integrated features.
[0032] Preferably, the bidirectional integration features extracted by the BiLSTM model are passed into the MSCA of the named entity recognition model to extract multi-scale features to obtain multi-scale features, specifically including:
[0033] The bidirectional integrated features extracted by the BiLSTM model are passed into the MSCA of the named entity recognition model, and the multi-scale convolution combined with ordinary convolution and dilated convolution is extracted and fused to obtain the preliminary fused features. Then the channel attention mechanism is used to assign weights to the preliminary fused features of different channels to obtain the channel attention weighted features. Then, the preliminary fused features and the channel attention weighted features are fused to obtain the final multi-scale features.
[0034] Preferably, the bidirectional features extracted by the BiLSTM model are passed into the MSCA of the named entity recognition model, and the multi-scale convolution extraction of ordinary convolution and dilated convolution is combined to obtain preliminary fusion features, which specifically includes:
[0035] MSCA sets the convolution kernel of ordinary convolution to 1×1 and 3×3, and extracts bidirectional features according to different convolution kernels of ordinary convolution;
[0036] By inserting a fixed hole interval into the convolution kernel of the dilated convolution, the features with different receptive fields from the ordinary convolution in the bidirectional features are extracted based on the convolution kernel of the dilated convolution;
[0037] The features extracted by the ordinary convolution and the dilated convolution are standardized by batch normalization, and the features extracted by the ordinary convolution and the dilated convolution are fused by addition operation to obtain preliminary fused features.
[0038] Preferably, a channel attention mechanism is used to assign weights to the context-dependent information features of different channels to obtain channel attention weighted features, specifically including:
[0039] The initial fusion features are respectively extracted by average pooling for the overall mean of all features in each channel and by maximum pooling for the strongest response value in each channel;
[0040] After the features after average pooling and maximum pooling are fused, they are linearly transformed through a shared convolutional layer;
[0041] The linearly transformed features are nonlinearly mapped through the activation function ReLU;
[0042] The features after nonlinear mapping are used to generate channel weight matrix through Sigmoid function;
[0043] Then, the channel weight matrix is multiplied element-by-element by the preliminary fusion feature, and the weight is applied to the feature channel of the preliminary fusion feature to obtain the channel attention weighted feature.
[0044] Preferably, the generated multi-scale features are input into the CRF model of the named entity recognition model, the label of each position in the multi-scale features is decoded, and the current optimal label sequence is output, which specifically includes:
[0045] The CRF model establishes a dependency model between labels through contextual dependencies in multi-scale features, calculates the optimal label sequence of each position label and its previous and subsequent words in the dependency model between labels through dynamic programming decoding, and outputs the current optimal label sequence.
[0046] The second aspect of the present invention proposes a network security entity recognition system based on multi-layer channel attention, and the network security entity recognition method based on multi-layer channel attention is executed, comprising:
[0047] Data preprocessing module, used to preprocess the real cybersecurity entity text dataset;
[0048] Model building module: used to build a named entity recognition model based on multi-layer channel attention, input the preprocessed network security entity text dataset into the named entity recognition model for model training, and obtain a trained named entity recognition model;
[0049] Label sequence acquisition module: used to input the actual network security entity text dataset into the trained named entity recognition model and output the label sequence;
[0050] Category identification module: used to identify the entity category of each word in the actual network security entity text dataset according to the output label sequence, and realize network security entity recognition based on multi-layer channel attention.
[0051] The third aspect of the present invention discloses an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the network security entity recognition method based on multi-layer channel attention when loaded into the processor.
[0052] The fourth aspect of the present invention discloses a computer-readable storage medium, which stores a computer program, and is characterized in that when the computer program is executed by a processor, it implements the network security entity identification method based on multi-layer channel attention.
[0053] Compared with the prior art, the beneficial effects of the present invention include at least:
[0054] (1) In the data preprocessing of the present invention, the word segmentation tool is specially customized to better identify the terms and abbreviations specific to the network security field and reduce word segmentation errors; stop words that have no practical significance for the task are removed to reduce the model calculation burden and improve processing efficiency; standardized processing ensures the feasibility and consistency of batch input;
[0055] (2) The named entity recognition model based on the multi-scale convolutional attention mechanism constructed by the present invention has the advantages of high-precision recognition, efficient calculation and multi-module deep fusion. Among them, the named entity recognition model based on the multi-scale convolutional attention mechanism is a BERT-BiLSTM-MSCA-CRF structure. By combining contextual semantics, serialized dependencies, multi-scale attention and label association decoding, the recognition accuracy of network security entities is significantly improved, and the shortcomings of existing methods in recognition accuracy are effectively solved; the model optimizes the combined design of BERT embedding, channel attention and CRF layers, so that the model has low computing resource consumption while maintaining a high recognition accuracy, and is suitable for actual resource-constrained environments, such as embedded systems or mobile devices; the multi-scale channel attention mechanism enables the model to adapt to different information features in network security texts, adapt to dynamic and changeable entity expressions and context changes, and is suitable for various security threat detection scenarios; the model realizes multi-module deep fusion in structure, innovatively combines pre-trained language models with multi-channel attention mechanisms, enhances the accurate recognition and effective extraction of network security entities, and provides more efficient and stable technical support for threat detection and intelligence analysis in the field of network security, which can greatly improve the response speed and monitoring capabilities of network security systems;
[0056] (3) This paper proposes a multi-scale channel attention mechanism, which combines ordinary convolution and dilated convolution, and dynamically adjusts the channel weights, thereby improving the model's ability to capture multi-scale features. In addition, deep fusion between modules optimizes the integrity of feature transfer, significantly reduces computational complexity, significantly improves the accuracy and efficiency of network security entity recognition, enhances the model's ability to adapt to dynamic scenarios, and is suitable for real-time threat analysis scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a schematic diagram of the process of a high-precision network security entity recognition method based on multi-layer channel attention provided in accordance with an embodiment of the present invention;
[0058] Figure 2 It is a schematic diagram of the overall network structure of a named entity recognition model based on a multi-scale convolutional attention mechanism provided in accordance with an embodiment of the present invention;
[0059] Figure 3 It is a schematic diagram of a specific process of inputting a preprocessed network text dataset into a named entity recognition model based on a multi-scale convolutional attention mechanism and outputting a label sequence according to an embodiment of the present invention;
[0060] Figure 4 is a schematic diagram of the structure of a BERT model provided according to an embodiment of the present invention;
[0061] Figure 5is a schematic diagram of the structure of a multi-scale convolutional attention mechanism provided according to an embodiment of the present invention;
[0062] Figure 6 is a schematic diagram of a high-precision network security entity recognition system based on multi-layer channel attention provided in accordance with an embodiment of the present invention. DETAILED DESCRIPTION
[0063] In order to make the purpose, technical scheme and advantages of the present invention clearer, the technical scheme of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. The described embodiments are only embodiments of a part of the present invention, not all embodiments. Based on the spirit of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention.
[0064] like Figure 1 As shown, embodiment 1 of the present invention provides a high-precision network security entity recognition method based on multi-layer channel attention, comprising the following steps:
[0065] Step 1: Preprocess the real cybersecurity entity text dataset.
[0066] In a preferred but non-limiting embodiment of the present invention, the text data in the real network security entity text data set includes network security entities such as intrusion detection logs, malware descriptions, IP addresses and domain names.
[0067] Further preferably, in order to ensure that the model can accurately process professional terms and abbreviations in the field of network security, the data set is first cleaned and preprocessed, and the data preprocessing includes the following steps:
[0068] Step 1.1: Word segmentation and labeling: Use the word segmentation tool to segment the original text and label each word with a corresponding label so that the subsequent named entity recognition model can perform named entity recognition.
[0069] Step 1.2: Remove stop words that have no practical significance for the task.
[0070] Step 1.3: Standardize all sentences into a fixed-length format. Fill in sentences that are not long enough and truncate overlong sentences to a length that fits the model.
[0071] Step 1.4: Data enhancement, using data enhancement techniques such as random deletion and replacement to improve data diversity and the generalization ability of the model, so as to better adapt to the variable data in real network security scenarios.
[0072] Step 1.5: The association rule algorithm processes the data to obtain potential dangerous events.
[0073] After data enhancement, the data is further processed using association rule algorithms to identify potential dangerous events. Association rule algorithms help discover patterns and events that may indicate potential dangers by analyzing the correlations between items in the network security data set. For example, by analyzing the correlations between data items such as intrusion detection logs, malware descriptions, and domain names / IP addresses, the algorithm can discover the connection between specific behavior patterns and potential attack events. The key to this step is to find possible dangerous events through rule mining and provide basic data for subsequent processing.
[0074] The potential dangerous events identified by the association rule algorithm are the preprocessed network security entity text dataset, which will be passed as input to the subsequent named entity recognition (NER) model. The task of the NER model is to further extract and classify key information from these potential dangerous events, such as attack source IP, attack target system, attack time, etc. The NER model can help better understand the specific details of potential dangerous events by identifying specific entities from text data, and provide information support for subsequent prevention and response.
[0075] Step 2: Build a named entity recognition model based on multi-layer channel attention.
[0076] like Figure 2 As shown in the figure, a named entity recognition model based on multi-layer channel attention is constructed based on the BERT model, BiLSTM model, MSCA and CRF model. The structure is BERT-BiLSTM-MSCA-CRF, which specifically includes:
[0077] Step 2.1, the named entity recognition model uses a BERT model in the embedding layer to capture the context information of the input data;
[0078] The embedding layer of the named entity recognition model uses the BERT model for word vector pre-training. The BERT model captures context information and processes the global dependencies of the text through the self-attention mechanism, so that the semantic connection between words can be fully understood. It provides a richer and more dynamic contextual semantic representation, especially for professional terms and abbreviations that appear in complex network security texts. The pre-training of the BERT model can effectively improve the processing capabilities of these words.
[0079] Step 2.2, the named entity recognition model uses a BiLSTM model in the feature extraction layer to extract features from the output data of the embedding layer;
[0080] In the feature extraction layer, the BiLSTM model is combined to learn the temporal features of the input sequence. BiLSTM can extract information from both the forward and reverse directions, thereby comprehensively modeling the grammar and context of the input text.
[0081] Step 2.3, the named entity recognition model uses MSCA in the attention layer to focus on the discriminant features of different dimensions in the output data of the feature extraction layer through different channels;
[0082] In addition, at the attention layer, a multi-layer channel attention mechanism is introduced, which can focus on key features of different dimensions in the input data through different channels. MSCA not only helps the model focus on the most discriminative features, but also effectively improves the extraction accuracy of important information through multi-level attention processing and avoids the interference of redundant data.
[0083] Step 2.4: The named entity recognition model uses a CRF model in the output layer to capture the serialized dependency relationship between entity tags based on the output data of MSCA.
[0084] In the output layer, a CRF model is added to capture the serialized dependencies between entity tags. CRF ensures that the output entity tags are coherent and consistent by learning the dependency structure between tags. The introduction of the CRF layer optimizes this dependency, making the final entity recognition result more accurate and in line with the actual grammatical structure. Through the collaborative work of the above layers, the model of the present invention can not only effectively identify multiple entity types in network security texts, but also maintain efficient performance in changing network security scenarios. Finally, the model can identify multiple entities including intrusion detection logs, malware descriptions, IP addresses, domain names, etc., providing strong technical support for network security analysis and defense.
[0085] Step 3: Input the preprocessed network text dataset into the named entity recognition model based on the multi-scale convolutional attention mechanism for model training to obtain a trained named entity recognition model.
[0086] In a preferred but non-limiting embodiment of the present invention, Figure 3 As shown in the figure, the preprocessed network text dataset is input into the named entity recognition model based on the multi-scale convolutional attention mechanism and the label sequence is output, including:
[0087] Step 3.1: Divide the preprocessed web text dataset into a test set and a dataset to be trained, and divide the dataset to be trained into 5 parts. Use 1 of the 5 parts as a validation set each time, and the other 4 parts as training sets. Specifically, in order to ensure the generalization ability of the model, 5-fold cross-validation is used to prevent the model from overfitting. Therefore, the dataset is divided into 5 parts, and a different part of the data will be selected as the validation set in each round of cross-validation, and the remaining 4 parts of the data will be used as training sets. Each training refers to the use of different validation sets and training sets for model training in different rounds of each cross-validation. In this way, during the cross-validation process, the model will undergo 5 rounds of training and validation, using a different validation set in each round to ensure the stability and reliability of the training results.
[0088] Step 3.2: Use the training set to train the named entity recognition model. Each round of cross-validation training process includes:
[0089] Step 3.2.1: Input a single training set into the BERT embedding layer to generate a context-sensitive embedding vector.
[0090] Specifically, the embedding layer of the named entity recognition model is based on the BERT pre-trained model. The BERT model can generate the semantic representation of each word in context, has two-way context understanding capabilities, and is particularly suitable for complex entity relationships in the field of network security. The specific implementation steps are as follows:
[0091] Step 3.2.1.1: Semantic information capture. The BERT model generates context-sensitive word embeddings through a multi-layer Transformer structure, which can deeply capture the semantic information of the text. The BERT model not only obtains the meaning of each word, but also captures its contextual relationship, thereby ensuring that the model accurately understands professional terms and entity relationships, such as Figure 4 shown.
[0092] Step 3.2.1.2: Embedding vector generation, convert word embedding into context-sensitive embedding vector. Each word embedding is represented as a high-dimensional vector, which contains rich semantic information and contextual dependencies. This embedding vector is used as the input of subsequent layers to provide an accurate semantic basis for the model.
[0093] Step 3.2.2: Input the embedding vector generated by each BERT model into the BiLSTM model for bidirectional feature extraction of the text sequence;
[0094] Specifically, in order to further enhance the understanding of sequence features, the model of the present invention adds a BiLSTM model after the BERT model is embedded. The BiLSTM model can better model the dependency relationship between the context and provide support for subsequent entity recognition. The specific implementation steps are as follows:
[0095] Step 3.2.2.1: Bidirectional feature extraction, the BiLSTM model can simultaneously process the forward and reverse sequences of the text in the embedding vector, capture the dependencies between the previous and next contexts within the text sentence, and extract bidirectional features. Through bidirectional feature extraction, the BiLSTM model generates rich context representations and enhances the model's ability to understand complex sequences.
[0096] Step 3.2.2.2: Integrate the extracted bidirectional features to obtain bidirectional integrated features. The features output by BiLSTM will contain the integrated information of the context, which can more accurately represent the semantics and positional relationship of the entity and improve the recognition accuracy.
[0097] Step 3.2.3: Pass the output features of the BiLSTM layer into MSCA, perform attention allocation on the channel dimension, and generate the final multi-scale feature representation;
[0098] Specifically, Figure 5 Based on the BiLSTM output, the MSCA proposed in the embodiment of the present application can effectively improve the model's ability to identify network security entities. The specific implementation steps of the multi-scale convolutional attention module are as follows:
[0099] Step 3.2.3.1: Multi-scale convolution extraction, MSCA combines ordinary convolution (Conv1, Conv3) and dilated convolution (Dilate) to achieve comprehensive capture of features of different scales. Ordinary convolution extracts local features through 1×1 and 3×3 convolution kernels. The 1×1 convolution kernel focuses on point-by-point feature transformation, while the 3×3 convolution kernel can capture local patterns in a larger range. These convolution operations provide basic local feature information for subsequent processing. At the same time, dilated convolution expands the receptive field by inserting a fixed dilated interval in the convolution kernel, which can enhance the ability to capture long-distance contextual information without increasing computational complexity, thereby supplementing the lack of ordinary convolution in capturing global features.
[0100] The features after the convolution operation are standardized by BN (Batch Normalization) to ensure the stability of data distribution and accelerate the convergence process of the model. Subsequently, the features extracted by different convolution kernels are integrated and feature fusion is performed using the addition operation (Add). This fusion process not only retains the detailed information of local features, but also introduces global context features, making the final generated preliminary multi-scale comprehensive features have stronger expressive power.
[0101] Through this multi-scale convolution design, MSCA can focus on local details and global structures at the same time, making the model more adaptable and robust when processing complex network security text data, and providing diversified feature inputs for subsequent attention weighting.
[0102] Step 3.2.3.2: Channel attention weighting, the fused multi-scale features are respectively subjected to average pooling (AvgPooling) and maximum pooling (Max Pooling) operations. Average pooling is used to extract the overall mean of all features in each channel, which can reflect the global distribution information within the channel; maximum pooling extracts the strongest response value from each channel to highlight significant features. These two pooling operations generate two global feature representations, which summarize channel information from different perspectives. Next, the features obtained by average pooling and maximum pooling are fused and linearly transformed through a shared convolutional layer (Conv1) to extract more discriminative feature representations. In order to introduce nonlinear characteristics, the features after convolution transformation are further nonlinearly mapped through the activation function ReLU to capture the complex relationship between channels. This process makes the features more adaptable and can effectively describe important task-related information. The features after convolution and activation are used to generate a channel weight matrix through the Sigmoid function. Each value of the weight matrix ranges from 0 to 1 and is used to measure the importance of each feature channel. The higher the weight, the more critical the channel is in the current task; the lower the weight, the less important the channel is. Finally, the channel weight matrix is element-by-element multiplied with the preliminary fusion features, and the weights are applied to the feature channels of the preliminary fusion features. The features of high-weight channels are amplified, thereby enhancing their ability to express key information; the features of low-weight channels are suppressed, reducing the interference of irrelevant information. This dynamic weighting process enables the model to focus more on important features related to the task, greatly improving the effectiveness and discrimination ability of feature representation.
[0103] The introduction of channel attention weighting enables the MSCA module to dynamically adjust its focus according to the features of the current input, ensuring that the model can make full use of the most valuable information while suppressing the interference of noise and redundant features, providing optimized input for subsequent feature fusion and decoding steps.
[0104] Step 3.2.4: Input the multi-scale features generated by MSCA into the CRF model, decode the label at each position, and output the label sequence;
[0105] Specifically, to ensure the logical consistency of sequence annotation, the model adds a CRF decoding layer to the output layer. CRF captures the dependencies between labels to ensure that the generated label sequence conforms to the logical rules of named entity recognition. The specific steps are as follows:
[0106] Step 3.2.4.1: Label dependency modeling, CRF ensures the rationality of the label sequence by modeling the dependency between labels (for example, the start label "B" of a named entity should be followed by the internal label "I"). Through this modeling, the CRF layer can effectively avoid the label conflict problem that may occur during independent annotation.
[0107] Step 3.2.4.2: Dynamic programming decoding,CRF decoding uses a dynamic programming algorithm to calculate the optimal label sequence.,Dynamic programming ensures the efficiency of the decoding process, and can output the optimal label path in a reasonable time,,thus improving the practicality of the model.
[0108] Step 3.2.4.3: Label output, the CRF layer finally outputs a label sequence to determine the entity category of each word, thereby achieving accurate entity recognition of network security text.
[0109] Step 3.2.5: Repeat steps 3.2.1 to 3.2.5, and obtain 4 different models after 4 trainings;
[0110] Step 3.2.6: Use the learning rate decay strategy to dynamically adjust the model's learning rate according to the performance of different models on the validation set, determine the model with the best performance among the four models, and set it as the model with the best performance in this cross-validation.
[0111] Specifically, the learning rate decay strategy is adopted during the training process, and the learning rate is dynamically adjusted according to the performance on the validation set to ensure smooth training and achieve the best results. At the same time, evaluation indicators are evaluated: by calculating the precision rate (Precision), recall rate (Recall), F1 score (F1 Score) and other indicators on the validation set, the model performance is evaluated to ensure that it achieves the best in terms of recognition accuracy and efficiency. Experiments show that a learning rate of 5e-5 can achieve the best results for the model.
[0112] Step 3.3: After 5-fold cross-validation, the test set is used to verify the optimal named entity recognition model obtained in each round of 5 cross-validations to obtain the optimal named entity recognition model of 5-fold cross-validation.
[0113] Step 4: Input the actual network security entity text dataset into the trained named entity recognition model and output a label sequence;
[0114] After training, the model is applied to actual cybersecurity text data to test its performance in real scenarios. By comparing with manually annotated labels, the accuracy of the model in complex cybersecurity text is further verified. The specific application test steps are as follows:
[0115] Step 4.1: Actual data testing: input actual cybersecurity text data into the model and compare and analyze the entity labels output by the model.
[0116] Step 4.2: Effect verification, by calculating indicators such as precision, recall and F1 score, to evaluate the performance of the model in the actual environment.
[0117] Step 4.3: Error analysis: Analyze the cases with recognition errors to find out the reasons for model misjudgment and provide a basis for subsequent model optimization.
[0118] Step 5: Identify the entity category of each word in the actual network security entity text dataset based on the output label sequence to achieve entity recognition of network security text.
[0119] The network security entity recognition method based on multi-scale convolutional attention mechanism proposed in this paper successfully realizes efficient and accurate entity recognition, and verifies the practicality and effectiveness of the model in complex domain tasks.
[0120] Compared with the prior art, the beneficial effects of the present invention include at least:
[0121] (1) In the data preprocessing of the present invention, the word segmentation tool is specially customized to better identify the terms and abbreviations specific to the network security field and reduce word segmentation errors; stop words that have no practical significance for the task are removed to reduce the model calculation burden and improve processing efficiency; standardized processing ensures the feasibility and consistency of batch input;
[0122] (2) The named entity recognition model based on the multi-scale convolutional attention mechanism constructed by the present invention has the advantages of high-precision recognition, efficient calculation and multi-module deep fusion. Among them, the named entity recognition model based on the multi-scale convolutional attention mechanism is a BERT-BiLSTM-MSCA-CRF structure. By combining contextual semantics, serialized dependencies, multi-scale attention and label association decoding, the recognition accuracy of network security entities is significantly improved, and the shortcomings of existing methods in recognition accuracy are effectively solved; the model optimizes the combined design of BERT embedding, channel attention and CRF layers, so that the model has low computing resource consumption while maintaining a high recognition accuracy, and is suitable for actual resource-constrained environments, such as embedded systems or mobile devices; the multi-scale channel attention mechanism enables the model to adapt to different information features in network security texts, adapt to dynamic and changeable entity expressions and context changes, and is suitable for various security threat detection scenarios; the model realizes multi-module deep fusion in structure, innovatively combines pre-trained language models with multi-channel attention mechanisms, enhances the accurate recognition and effective extraction of network security entities, and provides more efficient and stable technical support for threat detection and intelligence analysis in the field of network security, which can greatly improve the response speed and monitoring capabilities of network security systems;
[0123] (3) This paper proposes a multi-scale channel attention mechanism, which combines ordinary convolution and dilated convolution, and dynamically adjusts the channel weights, thereby improving the model's ability to capture multi-scale features. In addition, deep fusion between modules optimizes the integrity of feature transfer, significantly reduces computational complexity, significantly improves the accuracy and efficiency of network security entity recognition, enhances the model's ability to adapt to dynamic scenarios, and is suitable for real-time threat analysis scenarios.
[0124] like Figure 6 As shown, embodiment 2 of the present invention provides a network security entity recognition system based on multi-layer channel attention, and runs the network security entity recognition method based on multi-layer channel attention described in embodiment 1, including:
[0125] Data preprocessing module, used to preprocess the real cybersecurity entity text dataset;
[0126] Model building module: used to build a named entity recognition model based on multi-layer channel attention, input the preprocessed network security entity text dataset into the named entity recognition model for model training, and obtain a trained named entity recognition model;
[0127] Label sequence acquisition module: used to input the actual network security entity text dataset into the trained named entity recognition model and output the label sequence;
[0128] Category identification module: used to identify the entity category of each word according to the output label sequence, and realize network security entity recognition based on multi-layer channel attention.
[0129] Example 3 of the present invention further illustrates the network security entity recognition method based on the multi-scale convolutional attention mechanism provided by the present invention through experiments.
[0130] The experiment uses a self-constructed dataset to evaluate the model performance. This dataset contains a large number of professional terms and abbreviations in the field of network security and is suitable for the verification of named entity recognition tasks. The focus of the experiment is to evaluate the performance of the model in terms of precision, recall, and F1 score.
[0131] (1) Experimental setup
[0132] The experiment was conducted on a computer equipped with an NVIDIA RTX 3090 graphics card and 32GB of memory. The software environment included Python 3.8 and PyTorch deep learning framework. The BERT pre-trained model was loaded through the transformers library.
[0133] (2) Dataset characteristics
[0134] The experimental dataset contains text data from real network security events, covering network security entities such as intrusion detection logs, malware descriptions, IP addresses and domain names. To enhance the generalization ability of the model, the dataset has been cleaned and enhanced, and divided into training and test sets.
[0135] (3) Experimental comparison
[0136] The experiment set up multiple sets of comparative tests, including different model configurations, to verify the performance of the model proposed in this paper in the NER task. The comparative models include: the combination model without BERT integration, the combination model with BERT integration, and the multi-scale convolutional attention model proposed in this paper. The specific experimental results are shown in Table 1.
[0137] Table 1 Comparative experiment
[0138]
[0139] (4) Experiments with different learning rates
[0140] The experiment also examined the effects of different learning rates on the model proposed in this paper, including learning rates 1e-4, 5e-5, 3e-5, 1e-5 and 5e-6. The experimental results are shown in Table 2, revealing the performance differences of the model under different learning rates.
[0141] Table 2 Comparison of different learning rates
[0142]
[0143]
[0144] (5) Experimental results analysis
[0145] The proposed model outperforms other comparison models in F1 score, especially in precision and recall. The introduction of multi-scale convolutional attention mechanism enables the model to focus on local and global features at the same time and adaptively adjust channel weights. The learning rate of 5e-5 performs best in all experimental groups, indicating that a moderate learning rate can strike a balance between training speed and stability.
[0146] Embodiment 4 of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is loaded into the processor, a network security entity recognition method based on multi-layer channel attention according to embodiment 1 is implemented.
[0147] Embodiment 5 of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements a network security entity recognition method based on multi-layer channel attention according to embodiment 1.
[0148] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A network security entity recognition method based on multi-layer channel attention, characterized in that: include: Perform data preprocessing on real cybersecurity entity text datasets; Construct a named entity recognition model based on multi-layer channel attention, input the preprocessed network security entity text dataset into the named entity recognition model for model training, and obtain a trained named entity recognition model; Inputting the actual network security entity text dataset into the trained named entity recognition model and outputting a label sequence; According to the output label sequence, the entity category of each word in the actual network security entity text dataset is identified to realize network security entity recognition based on multi-layer channel attention.
2. According to claim 1, a network security entity recognition method based on multi-layer channel attention is characterized in that: Data preprocessing is performed on real cybersecurity entity text datasets, including: Use the word segmentation tool to segment the original text in the real cybersecurity entity text dataset and annotate each word with a corresponding label; Remove stop words from the segmented and annotated text; Normalize the text after removing stop words; Enhance the standardized text data; The association rule algorithm is used to identify potential dangerous events in the data-enhanced text, and the preprocessed network security entity text dataset is obtained.
3. According to claim 1, a network security entity recognition method based on multi-layer channel attention is characterized in that: Construct a named entity recognition model based on multi-layer channel attention, input the preprocessed network security entity text dataset into the named entity recognition model for model training, and obtain a trained named entity recognition model, specifically including: A named entity recognition model based on multi-layer channel attention is constructed based on the BERT model, BiLSTM model, MSCA and CRF model; The preprocessed network text dataset is divided into a test set and a training set; The named entity recognition model is trained using 5-fold cross validation, the training data set is divided into 5 parts, and 5 rounds of cross validation are performed. In each round of cross validation, 1 part of the data is selected as the validation set, and the remaining 4 parts of the data are used as the training set. Each round of cross validation obtains an optimal named entity recognition model; After 5-fold cross-validation, the test set is used to verify the optimal named entity recognition model obtained in each round of cross-validation to obtain the optimal named entity recognition model of 5-fold cross-validation.
4. According to claim 3, a network security entity recognition method based on multi-layer channel attention is characterized in that: The named entity recognition model is trained using 5-fold cross validation. The training data set is divided into 5 parts and 5 rounds of cross validation are performed. In each round of cross validation, 1 part of the data is selected as the validation set and the remaining 4 parts of the data are used as the training set. Each round of cross validation obtains an optimal named entity recognition model, which specifically includes: Inputting a single training set into the BERT model of the named entity recognition model to generate a context-sensitive embedding vector; Inputting the generated embedding vector into the BiLSTM model of the named entity recognition model to perform bidirectional feature extraction of the text sequence to obtain a bidirectional integrated feature; The bidirectional integration features extracted by the BiLSTM model are passed into the MSCA of the named entity recognition model to extract multi-scale features to obtain multi-scale features; Input the generated multi-scale features into the CRF model of the named entity recognition model, decode the label of each position in the multi-scale features, and output the current optimal label sequence; Repeatedly inputting the remaining single training sets in the training set into the named entity recognition model, and obtaining named entity recognition models with four different output optimal label sequences after four trainings; A learning rate decay strategy is adopted to dynamically adjust the learning rate of the model according to the performance of the named entity recognition model with the optimal label sequence of different outputs on the validation set, and the named entity recognition model with the best performance among the four models is determined as the optimal named entity recognition model for this round of cross-validation.
5. According to claim 4, a network security entity recognition method based on multi-layer channel attention is characterized in that: Input a single training set into the BERT model of the named entity recognition model to generate a context-sensitive embedding vector, and input the generated embedding vector into the BiLSTM model of the named entity recognition model to perform bidirectional feature extraction of the text sequence, specifically including: The BERT model captures the semantic information of the text through a multi-layer Transformer structure, generates context-sensitive word embeddings, and converts word embeddings into context-sensitive embedding vectors; The BiLSTM model processes the forward and reverse sequences of the text in the embedded vector at the same time, captures the dependency between the previous and next sentences in the text, extracts bidirectional features, and integrates the extracted bidirectional features to obtain bidirectional integrated features.
6. A network security entity recognition method based on multi-layer channel attention according to claim 4, characterized in that: The bidirectional integration features extracted by the BiLSTM model are passed into the MSCA of the named entity recognition model to extract multi-scale features, and obtain multi-scale features, specifically including: The bidirectional integrated features extracted by the BiLSTM model are passed into the MSCA of the named entity recognition model, and the multi-scale convolution combined with ordinary convolution and dilated convolution is extracted and fused to obtain the preliminary fused features. Then the channel attention mechanism is used to assign weights to the preliminary fused features of different channels to obtain the channel attention weighted features. Then, the preliminary fused features and the channel attention weighted features are fused to obtain the final multi-scale features.
7. A network security entity recognition method based on multi-layer channel attention according to claim 6, characterized in that: The bidirectional features extracted by the BiLSTM model are passed into the MSCA of the named entity recognition model, and the multi-scale convolution extraction of ordinary convolution and dilated convolution is combined to obtain preliminary fusion features, including: MSCA sets the convolution kernel of ordinary convolution to 1×1 and 3×3, and extracts bidirectional features according to different convolution kernels of ordinary convolution; By inserting a fixed hole interval into the convolution kernel of the dilated convolution, the features with different receptive fields from the ordinary convolution in the bidirectional features are extracted based on the convolution kernel of the dilated convolution; The features extracted by the ordinary convolution and the dilated convolution are standardized by batch normalization, and the features extracted by the ordinary convolution and the dilated convolution are fused by addition operation to obtain preliminary fused features.
8. The network security entity recognition method based on multi-layer channel attention according to claim 6 is characterized by: The channel attention mechanism is used to assign weights to the context-dependent information features of different channels and obtain the channel attention weighted features, including: The initial fusion features are respectively extracted by average pooling for the overall mean of all features in each channel and by maximum pooling for the strongest response value in each channel; After the features after average pooling and maximum pooling are fused, they are linearly transformed through a shared convolutional layer; The linearly transformed features are nonlinearly mapped through the activation function ReLU; The features after nonlinear mapping are used to generate channel weight matrix through Sigmoid function; Then, the channel weight matrix is multiplied element-by-element by the preliminary fusion feature, and the weight is applied to the feature channel of the preliminary fusion feature to obtain the channel attention weighted feature.
9. A network security entity recognition method based on multi-layer channel attention according to claim 4, characterized in that: The generated multi-scale features are input into the CRF model of the named entity recognition model, the labels at each position in the multi-scale features are decoded, and the current optimal label sequence is output, which specifically includes: The CRF model establishes a dependency model between labels through contextual dependencies in multi-scale features, calculates the optimal label sequence of each position label and its previous and subsequent words in the dependency model between labels through dynamic programming decoding, and outputs the current optimal label sequence.
10. A network security entity recognition system based on multi-layer channel attention, running a network security entity recognition method based on multi-layer channel attention according to any one of claims 1 to 9, characterized in that: Data preprocessing module, used to preprocess the real cybersecurity entity text dataset; Model building module: used to build a named entity recognition model based on multi-layer channel attention, input the preprocessed network security entity text dataset into the named entity recognition model for model training, and obtain a trained named entity recognition model; Label sequence acquisition module: used to input the actual network security entity text dataset into the trained named entity recognition model and output the label sequence; Category identification module: used to identify the entity category of each word in the actual network security entity text dataset according to the output label sequence, and realize network security entity recognition based on multi-layer channel attention.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, a network security entity recognition method based on multi-layer channel attention is implemented as described in any one of claims 1 to 9.
12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a network security entity recognition method based on multi-layer channel attention as described in any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Named entity identification method based on label attention mechanism
CN111199152A
Game-oriented named entity identification method and device
CN117010393A
Network security entity identification method based on BERT-BiLSTM-CCA-CRF
CN118982025A
Counterfactual context-aware texture learning for camouflaged object detection
US20240312194A1
Cited By
Document-level saliency detection method and related device
CN121542806A