Entity identification method, apparatus and device, and medium

By combining the BERT model, CNN and dependent syntax analysis methods, the feature vectors of deep semantic information, local features and syntactic structure information are integrated, and the problem of poor effect and low accuracy of existing entity recognition methods in security-related text processing is solved, which significantly improves the effect and accuracy of entity recognition.

CN120218071APending Publication Date: 2025-06-27CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510330405.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing entity recognition methods have poor results and low accuracy when processing security-related texts, especially when faced with strong professionalism, diverse entity categories and high data noise.

Method used

Using a method combining trained BERT model, convolutional neural network CNN and dependent syntax analysis, the feature vectors of deep semantic information, local features and syntactic structure information are determined through three independent feature extraction channels, and integrated to enhance the flexibility and expression ability of feature fusion.

Benefits of technology

It significantly improves the understanding and entity recognition capabilities of the named entity recognition model for complex texts, and improves the effectiveness and accuracy of entity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218071A_ABST
    Figure CN120218071A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to an entity recognition method and device, equipment and a medium. In the embodiment of the invention, three independent feature extraction channels are analyzed through a BERT model, a convolutional neural network and a dependency syntax, a first feature vector, a second feature vector and a third feature vector are determined, the obtained three feature vectors are integrated, and a finally obtained target feature vector comprises deep semantic information and context information. The method has the advantages that feature fusion is realized, local features and syntactic structure information are also included, flexibility and expression ability of feature fusion are enhanced, complex text understanding and entity recognition ability of a named entity recognition model are remarkably improved in the entity recognition process, and the entity recognition effect and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to an entity recognition method, apparatus, device and medium. Background Art

[0002] Named Entity Recognition (NER) is a core task in Natural Language Processing (NLP), aiming to automatically identify and classify entities with specific meanings from text, such as person names, place names, organization names, time, quantity, etc. NER technology plays an important role in multiple fields such as information retrieval, question answering systems, knowledge graph construction, text summarization, etc. Through effective entity recognition, computers can better understand and process natural language text, improving the intelligence level and user experience of various application systems.

[0003] In the field of information security, named entity recognition has broad and important application value. Named entity recognition technology can automatically extract key entities from a large number of security-related texts. However, due to the strong professionalism of security-related texts, diverse entity categories, and high data noise, the existing entity recognition has poor effects and low accuracy. Summary of the Invention

[0004] This application provides an intention recognition method, apparatus, device and medium to solve the problems of poor effects and low accuracy in existing entity recognition of security-related texts.

[0005] In a first aspect, an embodiment of this application provides an entity recognition method, the method includes: Using a trained BERT model, perform embedding representation on the target text to be processed, and determine the first feature vector of the target text; Using a trained Convolutional Neural Network (CNN), perform feature extraction on the first feature vector to determine the second feature vector; According to the character position corresponding to each character in the target text and the saved dependency relation label corresponding to each character, determine the dependency relation label sequence, and perform embedding representation on the dependency relation label sequence to obtain the third feature vector; According to the first feature vector, the second feature vector and the third feature vector, determine the target feature vector of the target text; Input the target feature vector into the named entity recognition model, so that the named entity recognition model performs named entity recognition on the target text according to the target feature vector, and determine the target entities included in the target text.

[0006] Second aspect, an embodiment of the present application further provides an entity recognition device, which includes: A processing module, configured to use a trained BERT model to perform embedding representation on a target text to be processed, and determine a first feature vector of the target text; use a trained convolutional neural network CNN to perform feature extraction on the first feature vector to determine a second feature vector; according to the character position corresponding to each character in the target text and the saved dependency relation label corresponding to each character, determine a dependency relation label sequence, and perform embedding representation on the dependency relation label sequence to obtain a third feature vector; determine a target feature vector of the target text according to the first feature vector, the second feature vector, and the third feature vector; An entity recognition module, configured to input the target feature vector into a named entity recognition model, so that the named entity recognition model performs named entity recognition on the target text according to the target feature vector, and determine a target entity included in the target text.

[0007] Third aspect, an embodiment of the present application further provides an electronic device, which at least includes a processor and a memory. When the processor executes a computer program stored in the memory, the steps of the entity recognition method as described in any one of the above are implemented.

[0008] Fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the entity recognition method as described in any one of the above are implemented.

[0009] In the embodiment of the present application, through three independent feature extraction channels of the BERT model, the convolutional neural network, and the dependency syntax analysis, the first feature vector, the second feature vector, and the third feature vector are determined, and the three obtained feature vectors are integrated. The finally obtained target feature vector contains both deep semantic information and context information, and also contains local features and syntactic structure information, enhancing the flexibility and expression ability of feature fusion, and significantly improving the understanding and entity recognition ability of the named entity recognition model for complex texts in the entity recognition process, and improving the effect and accuracy of entity recognition. Description of the Drawings

[0010] To more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0011] Figure 1Schematic diagram of an entity recognition process provided by an embodiment of the present application; Figure 2 System architecture diagram provided by an embodiment of the present application; Figure 3 Schematic diagram of the structure of an entity recognition device provided by an embodiment of the present application; Figure 4 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0012] To make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0013] In the field of information security, named entity recognition has extensive and important application values. Named entity recognition technology can automatically extract key entities from a large number of security-related texts, such as vulnerability numbers (CVE), malware names, IP addresses, attack techniques, etc., which is crucial for enhancing network protection capabilities and response speeds. For example, when analyzing network security event logs and security reports, named entity recognition technology can quickly identify potential threat sources and affected system components, helping security analysts take corresponding protection measures in a timely manner. In addition, named entity recognition technology also plays a key role in malware analysis. By automatically extracting and classifying the names, version information, and propagation methods of malware, it assists security researchers in conducting in-depth threat analysis and comparative research.

[0014] At the same time, when building and maintaining a security knowledge base, named entity recognition promotes the structuring and retrievability of knowledge by automatically extracting key information from security documents, technical manuals, and standards, supporting the continuous update and expansion of the knowledge base. Although named entity recognition shows great potential in information security, the existing technologies still face many challenges when dealing with security texts with strong professionalism, diverse entity categories, and large data noise, such as the accurate recognition of professional terms and the requirements for character-level language processing.

[0015] Therefore, designing an efficient and accurate entity recognition method for the field of information security has become an important research direction for enhancing network security protection capabilities and intelligent levels.

[0016] Based on this, in order to improve the effect and accuracy of entity recognition, the embodiments of the present application provide an entity recognition method, device, equipment, and medium.

[0017] In the embodiment of the present application, a trained BERT model is used to perform an embedding representation on the target text to be processed, and a first feature vector of the target text is determined; a trained convolutional neural network (CNN) is used to extract features from the first feature vector to determine a second feature vector; according to the character positions corresponding to each character in the target text and the saved dependency relation labels corresponding to each character, a dependency relation label sequence is determined, and an embedding representation is performed on the dependency relation label sequence to obtain a third feature vector; according to the first feature vector, the second feature vector, and the third feature vector, a target feature vector of the target text is determined; the target feature vector is input into a named entity recognition model, so that the named entity recognition model performs named entity recognition on the target text according to the target feature vector to determine the target entities included in the target text.

[0018] Embodiment 1:

[0019] Figure 1 FIG. 1 is a schematic diagram of an entity recognition process provided by an embodiment of the present application, and the process includes: S101: Use a trained BERT model to perform an embedding representation on the target text to be processed, and determine a first feature vector of the target text; use a trained convolutional neural network (CNN) to extract features from the first feature vector to determine a second feature vector.

[0020] An entity recognition method provided by an embodiment of the present application is applied to an electronic device, and the electronic device may be a PC or a server.

[0021] In a security text, the semantic relationship and local features between characters are crucial for accurately identifying entities. Based on this, in the embodiment of the present application, the electronic device captures the deep semantic representation of the characters in the target text by introducing a pre-trained BERT model and leveraging the advantages of the pre-trained language model.

[0022] Specifically, a trained BERT model is pre-configured in the electronic device. The electronic device inputs the target text to be processed into the BERT model, so that the BERT model performs a deep semantic analysis on the target text and determines and outputs a first feature vector corresponding to the target text.

[0023] In a possible implementation manner, the electronic device does not modify the structure of the BERT model, but uses a large number of texts in the field of information security as samples to train the BERT model, so that the BERT model can well understand the deep semantic information in the texts in the field of information security, and finally the first feature vector of the target text output by the trained BERT model can be more accurate.

[0024] Among them, the process of the electronic device training the BERT model based on a large number of texts in the field of information security is similar to the existing process of training the BERT model, and will not be elaborated here.

[0025] In addition, in the embodiment of the present application, the electronic device will also capture the n-gram features in the first feature vector through a Convolutional Neural Network (CNN), so as to extract the local dependency relationship between the characters of the target text.

[0026] Specifically, a trained CNN is also pre-configured in the electronic device. The electronic device inputs the first feature vector output by the BERT model into the CNN, so that the CNN performs local feature extraction on the first feature vector, determines and outputs a second feature vector. The electronic device obtains the second feature vector output by the CNN.

[0027] In a possible implementation manner, the CNN includes a convolutional layer, an activation function, and a max pooling layer; wherein the convolutional layer contains 1D convolutional kernels of n-gram size, which can perform local n-gram feature extraction on the first feature vector; the activation function and the max pooling layer can process the extracted intermediate feature vector to obtain a second feature vector of a preset dimension.

[0028] It should be noted that in the embodiment of the present application, the length of the first feature vector is the same as the length of the second feature vector. Generally, the first feature vector and the second feature vector are vectors of 768 dimensions.

[0029] In addition, in the embodiment of the present application, the electronic device can also use multiple CNNS containing convolutional kernels of different sizes to extract the first feature vector, so that features can be captured from different lengths, improving the effectiveness of local feature extraction.

[0030] If the electronic device uses multiple CNNS containing convolutional kernels of different sizes to extract the first feature vector, the feature vectors output by each CNN are concatenated, and the concatenated feature vector is used as the second feature vector. In this scenario, the length of the concatenated second feature vector is the same as the length of the first feature vector, and the length of the feature vector output by each CNN is the ratio of the length of the first feature vector to the number of CNNS. For example, there are 3 CNNS, and the length of the first feature vector is 768 dimensions, then the length of the feature vector output by each CNN is 768 / 3 = 256.

[0031] S102: Determine a dependency label sequence according to the character position corresponding to each character in the target text and the saved dependency label corresponding to each character, and perform embedding representation on the dependency label sequence to obtain a third feature vector.

[0032] In the embodiments of the present application, structured information such as dependency syntactic analysis can provide more syntactic and context clues for entity recognition. Based on this, in the embodiments of the present application, the electronic device will also perform dependency syntactic analysis on the target text to determine and save the dependency relationship tags corresponding to each character of the target text. The electronic device determines the third feature vector corresponding to the target text according to the dependency relationship tags corresponding to each character in the target text, so that during subsequent entity recognition, it can be recognized according to the third feature vector, improving the understanding ability of the named entity recognition model for specific named entities in the security field. The named entity recognition model can make full use of the syntax and the relationship between entities, thereby improving the recognition accuracy.

[0033] Specifically, the electronic device obtains the dependency relationship tags corresponding to each character in the saved target text, and sorts the dependency relationship tags corresponding to each character according to the character position of each character in the target text to obtain a dependency relationship tag sequence. The electronic device performs embedding representation on the dependency relationship tag sequence according to a preset embedding representation algorithm to obtain a third feature vector.

[0034] Among them, the dependency relationship tag is a tag used to label the dependency relationship between words in a sentence. Dependency relationship analysis is a technique in natural language processing for parsing the syntactic relationship between words in a sentence, usually represented by a tree structure. The dependency tag defines the syntactic role of a word in a sentence, such as subject, predicate, object, etc. There are various types of the dependency relationship tags, including but not limited to amod: possessive adjective, such as "my", "his"; tmod: time adverbial, such as "yesterday", "this year", etc.; nsubj: nominal subject, such as "I", "you", etc.; csubj: clause subject; dobj: direct object, etc., which are not limited here.

[0035] In addition, in the embodiments of the present application, the electronic device can perform embedding representation on the dependency relationship tag sequence through an embedding layer such as nn.Embedding to obtain a third feature vector. The electronic device can also use other methods to perform embedding representation on the dependency relationship tag sequence, which is not limited here.

[0036] In the embodiments of the present application, the length of the third feature vector is the same as that of the first feature vector and the second feature vector.

[0037] S103: Determine the target feature vector of the target text according to the first feature vector, the second feature vector, and the third feature vector; input the target feature vector into a named entity recognition model, so that the named entity recognition model performs named entity recognition on the target text according to the target feature vector to determine the target entity included in the target text.

[0038] In an embodiment of the present application, after an electronic device determines a first feature vector including deep semantic information and context information, a second feature vector including local features, and a third feature vector including syntactic structure information, the electronic device integrates the first feature vector, the second feature vector, and the third feature vector to obtain a target feature vector. Among them, the target feature vector includes both deep semantic information and context information, and also includes local information and syntactic structure information.

[0039] Specifically, the electronic device can directly splice the first feature vector, the second feature vector, and the third feature vector to obtain a target feature vector; the electronic device can also perform weighted summation on the first feature vector, the second feature vector, and the third feature vector to obtain a target feature vector; the electronic device can also perform dimensionality reduction on the first feature vector, the second feature vector, and the third feature vector, and then splice them to obtain a target feature vector, etc.

[0040] In an embodiment of the present application, after the electronic device determines a target feature vector including deep semantic information, context information, local information, and syntactic structure information, the electronic device inputs the target feature vector into a named entity recognition model, so that the named entity recognition model performs named entity recognition on the target text according to the target feature vector to determine the target entity included in the target text.

[0041] In an embodiment of the present application, through three independent feature extraction channels of a BERT model, a convolutional neural network, and dependency syntactic analysis, the first feature vector, the second feature vector, and the third feature vector are determined, and the three obtained feature vectors are integrated. The finally obtained target feature vector includes both deep semantic information and context information, and also includes local features and syntactic structure information, enhancing the flexibility and expression ability of feature fusion, and significantly improving the understanding and entity recognition ability of the named entity recognition model for complex texts during the entity recognition process, improving the effect and accuracy of entity recognition.

[0042] Embodiment 2: In order to improve the effect and accuracy of entity recognition, on the basis of the above embodiments, in an embodiment of the present application, the CNN includes a preset number of sub-CNNs, and the convolutional kernels included in each sub-CNN are different; Using the trained convolutional neural network (CNN) to extract features from the first feature vector and determine the second feature vector includes: Input the first feature vector into each sub-CNN respectively, so that each sub-CNN extracts features from the first feature vector, and obtain each sub-feature vector output by each sub-CNN; Concatenate the sub-feature vectors in a set order, and determine the obtained concatenated feature vector as the second feature vector.

[0043] In the embodiments of the present application, the CNN may be composed of a preset number of sub-CNNs. Each sub-CNN is independent, and the convolutional kernels included in each sub-CNN are different. The electronic device can call each sub-CNN to perform local feature extraction on the first feature vector respectively, and each sub-CNN outputs a sub-feature vector. The electronic device concatenates the sub-feature vectors output by each sub-CNN, and determines the obtained concatenated feature vector as the second feature vector.

[0044] Specifically, the electronic device inputs the first feature vector into each sub-CNN respectively, so that each sub-CNN extracts features from the first feature vector, and obtain each sub-feature vector output by each sub-CNN; concatenate the sub-feature vectors in a set order, and determine the obtained concatenated feature vector as the second feature vector.

[0045] In a possible implementation manner, the electronic device calls three sub-CNNs to perform local feature extraction on the first feature vector respectively. The convolutional kernels deployed in the three sub-CNNs are 1D convolutional kernels with 3-gram size, 1D convolutional kernels with 4-gram size, and 1D convolutional kernels with 5-gram size respectively. The electronic device calls the 3 sub-CNNs respectively, so that each sub-CNN outputs sub-feature vectors with the same dimension. The electronic device performs horizontal concatenation on each sub-feature vector to obtain the second feature vector. The dimension of the second feature vector is the same as that of the first feature vector.

[0046] Embodiment 3: In order to improve the effect and accuracy of entity recognition, on the basis of the above embodiments, in the embodiments of the present application, the sub-CNN includes a convolutional layer, an activation function, and a max-pooling layer; Each sub-CNN extracting features from the first feature vector includes: For each sub-CNN, the convolutional layer of the sub-CNN performs convolutional processing on the first feature vector using a preset convolutional kernel to obtain an intermediate feature vector; the activation function and max-pooling layer of the sub-CNN perform dimension compression processing on the intermediate feature vector to obtain a sub-feature vector.

[0047] In the embodiments of the present application, each sub-CNN includes a convolutional layer, an activation function, and a max-pooling layer, and the convolutional layer further includes a convolutional kernel of a preset size.

[0048] Based on this, after each sub-CNN receives the first feature vector input by the electronic device, the convolutional layer of the sub-CNN performs a convolution process on the first feature vector using a preset convolutional kernel to obtain an intermediate feature vector; the activation function and max-pooling layer of the sub-CNN perform a dimensionality compression process on the intermediate feature vector to obtain a sub-feature vector.

[0049] Among them, the activation function can be a ReLU activation function.

[0050] In a possible implementation manner, there are 3 sub-CNNs, namely the first sub-CNN, the second sub-CNN, and the third sub-CNN. The first sub-CNN includes a 1D convolutional kernel of 3-gram size, the second sub-CNN includes a 1D convolutional kernel of 4-gram size, and the third sub-CNN includes a 1D convolutional kernel of 5-gram size.

[0051] On this basis, the convolutional layer of the first sub-CNN performs a convolution operation on the first feature vector using a 1D convolutional kernel of 3-gram size to extract local n-gram features and obtain an intermediate feature vector; the activation function and max-pooling layer of the first sub-CNN compress the feature dimension of the intermediate feature vector to obtain a sub-feature vector with a fixed length of 256 dimensions. The convolutional layer of the second sub-CNN performs a convolution operation on the first feature vector using a 1D convolutional kernel of 4-gram size to extract local n-gram features and obtain an intermediate feature vector; the activation function and max-pooling layer of the second sub-CNN compress the feature dimension of the intermediate feature vector to obtain a sub-feature vector with a fixed length of 256 dimensions. The convolutional layer of the third sub-CNN performs a convolution operation on the first feature vector using a 1D convolutional kernel of 5-gram size to extract local n-gram features and obtain an intermediate feature vector; the activation function and max-pooling layer of the third sub-CNN compress the feature dimension of the intermediate feature vector to obtain a sub-feature vector with a fixed length of 256 dimensions. The electronic device horizontally splices the three sub-feature vectors and determines the spliced intermediate feature vector as the second feature vector.

[0052] Embodiment 4: In order to improve the effect and accuracy of entity recognition, on the basis of the above embodiments, in the embodiments of the present application, the process of determining the dependency relationship label corresponding to each character includes: Using a preset dependency relationship label determination tool, perform dependency syntactic analysis on the target text to determine each word and its corresponding dependency relationship label in the target text; For each character in the target text, determine the target word to which the character belongs, and determine the dependency relation label corresponding to the target word as the dependency relation label corresponding to the character.

[0053] In the embodiments of the present application, structured information such as dependency syntactic analysis can provide more syntactic and context clues for entity recognition. Based on this, in the embodiments of the present application, the electronic device can use a preset dependency relation label determination tool to perform dependency syntactic analysis on the target text, determine each word and its corresponding dependency relation label in the target text, and further determine the dependency relation label corresponding to each character.

[0054] Specifically, the electronic device uses the dependency relation label determination tool to segment the target text, and uses the dependency relation label determination tool to perform dependency syntactic analysis on the target text according to each word to determine the dependency relation label corresponding to each word. The electronic device maps the word-level dependency relation labels to the character level through a full-character assignment strategy, that is, assigns the dependency label of a word to all characters under the word to ensure that each character can obtain corresponding syntactic information.

[0055] In a possible implementation manner, for each character in the target text, the electronic device determines the target word to which the character belongs, and determines the dependency relation label corresponding to the target word as the dependency relation label corresponding to the character.

[0056] In the embodiments of the present application, by converting these structured features into character-level embeddings, the model can make full use of the syntactic and entity relationships, thereby improving the recognition accuracy.

[0057] Among them, in the embodiments of the present application, the dependency relation label determination tool can be SpaCy.

[0058] Embodiment 5: In order to improve the effect and accuracy of entity recognition, on the basis of the above embodiments, in the embodiments of the present application, the determining the target feature vector of the target text according to the first feature vector, the second feature vector, and the third feature vector includes: Construct a feature vector matrix according to the first feature vector, the second feature vector, and the third feature vector; Obtain a pre-configured weight matrix; Determine the target feature vector according to the weight matrix and the feature vector matrix.

[0059] In the embodiments of the present application, the electronic device fuses the first feature vector, the second feature vector, and the third feature vector into a target feature vector through an adaptive fusion attention mechanism. Each attention head can learn different feature relationships, enhancing the flexibility and expressive ability of feature fusion. This enables the model to adaptively select the most useful features according to the requirements of the task, thereby improving performance.

[0060] Specifically, the electronic device uses an adaptive self-attention mechanism to assign weights to each of the three ways of determining feature vectors. By training to learn the weights, each weight can reflect the importance of the corresponding way of determining feature vectors in a specific task. The electronic device performs a weighted sum of each feature vector according to the weight corresponding to each feature vector, generating a fused comprehensive target feature vector, where the dimension of the target feature vector is the same as that of the first feature vector, the second feature vector, and the third feature vector. Generally, it is 768 dimensions.

[0061] In a possible implementation manner, the electronic device stores the weight corresponding to each feature vector in the form of a weight matrix. When performing feature fusion, the electronic device can construct a feature vector matrix based on the first feature vector, the second feature vector, and the third feature vector; then obtain the pre-configured weight matrix, and the electronic device determines the target feature vector according to the weight matrix and the feature vector matrix.

[0062] Among them, when the electronic device constructs a feature vector matrix based on the first feature vector, the second feature vector, and the third feature vector, the first feature vector, the second feature vector, and the third feature vector are constructed as row vectors, that is, the first feature vector, the second feature vector, and the third feature vector are aligned in dimension.

[0063] Embodiment 6: To improve the effect and accuracy of entity recognition, based on the above embodiments, in the embodiments of the present application, the named entity recognition model includes a fully connected layer and a conditional random field CRF; The named entity recognition model performs named entity recognition on the target text according to the target feature vector, including: The fully connected layer of the named entity recognition model determines a score vector for each character in the target text belonging to each entity according to the target feature vector; The CRF of the named entity recognition model determines the target entity according to the score vector corresponding to each character and the pre-configured constraint conditions.

[0064] In the embodiments of the present application, the named entity recognition model includes a fully connected layer and a conditional random field (CRF). The electronic device can input the target feature vector into the named entity recognition model, so that the named entity recognition model performs entity recognition on the target text according to the target feature vector.

[0065] Specifically, the fully connected layer of the named entity recognition model determines a score vector for each character in the target text belonging to each entity according to the target feature vector; the CRF of the named entity recognition model determines the target entity according to the score vector corresponding to each character and the pre-configured constraint conditions.

[0066] In a possible implementation manner, the target feature vector can be converted into scores of predicted categories through the fully connected layer, and is ready to be input into the CRF layer for sequence labeling; the CRF performs sequence decoding on the output of the fully connected layer, considers the dependencies between labels, and outputs the final entity label sequence; the electronic device determines the entities included in the target text according to the output entity label sequence.

[0067] Embodiment 7: In order to improve the effect and accuracy of entity recognition, on the basis of the above embodiments, in the embodiments of the present application, the process of determining the target text includes: Obtain a security text; Segment the security text according to a preset segmentation rule; Determine each segmented sub-text as the target text.

[0068] In the embodiments of the present application, the electronic device can obtain the security text from various channels, such as obtaining from a database, receiving user input, obtaining from a security system, etc. The electronic device segments the security text according to a preset segmentation rule to obtain each sub-text, and the electronic device determines each sub-text as the target text to be subjected to entity recognition.

[0069] In a possible implementation manner, the electronic device divides the collected security text by characters and at the same time divides it by sentence size, and divides it by punctuation marks nearby if it exceeds 512 characters.

[0070] In the embodiments of the present application, first, the electronic device respectively extracts deep semantic representations at the character level and local n-gram features by combining the BERT and BERT-CNN models; second, the electronic device uses dependency syntactic analysis to extract dependency relationship features to enhance the model's understanding of syntactic structures and entity relationships; third, the electronic device adopts an adaptive attention mechanism to dynamically adjust the weights of features in different channels to achieve flexible feature fusion. Through these innovations, the recognition accuracy and robustness are significantly improved.

[0071] The present application has the following advantages over the prior art: 1. In the current security entity named entity recognition technology, common methods mainly include models based on BiLSTM-CRF, BERT-CRF models, etc. Although these methods have improved the accuracy of entity recognition to a certain extent, they still have significant limitations when dealing with complex and fine-grained threat intelligence texts. The embodiments of the present application significantly improve the understanding and entity recognition capabilities of complex threat intelligence texts by innovatively combining three independent feature extraction channels: a pre-trained language model, a convolutional neural network, and dependency syntactic analysis. First, BERT is responsible for capturing deep semantic and context information, CNN is responsible for extracting local n-gram features, and dependency syntactic analysis is responsible for obtaining syntactic structure information. Second, a multi-head self-attention mechanism is used to dynamically weight and fuse features from different channels, enabling the weights of each channel's features to be adaptively adjusted according to the context, enhancing the flexibility and expressive power of feature fusion.

[0072] 2. The embodiments of the present application uniquely map the results of dependency syntactic analysis into character-level embeddings through a full-character assignment strategy, ensuring that each character can obtain corresponding syntactic information. The character-level integration of dependency syntactic analysis enhances the model's ability to understand syntactic structures and entity relationships, and is particularly suitable for dealing with complex entity categories in security entities (such as CVE numbers, malware names, IP addresses, etc.).

[0073] In the actual application process, the entity recognition method provided by the embodiments of the present application can be implemented and applied in multiple scenarios. Specific embodiments are listed as follows: 1. Threat intelligence analysis: The entity recognition method provided by the embodiments of the present application can accurately identify and extract security entities involved (such as vulnerability numbers, malware names, attacker organizations, CVE numbers, etc.) by automatically analyzing unstructured texts such as network threat intelligence reports, vulnerability announcements, and security blogs. The NER ability that integrates structured information can help enterprises quickly collect and process threat intelligence, build a more comprehensive threat intelligence library, and thus improve threat perception and response capabilities.

[0074] 2. Intrusion Detection System (IDS): The entity recognition method provided by the embodiments of the present application can analyze security entity information (such as IP addresses, malware hash values, etc.) in data such as network traffic logs and abnormal behavior reports, helping the intrusion detection system to more intelligently detect potential attack behaviors. Combining multi-channel feature extraction can enhance the ability to recognize complex and polymorphic attack patterns and effectively detect Advanced Persistent Threat (APT) attacks.

[0075] 3. Malware Analysis: The entity recognition method provided in the embodiments of this application can be used to automatically analyze the behavior logs and research reports of malware, and accurately extract malicious behavior patterns and corresponding security entities (such as malicious code snippets, C2 server addresses, etc.). The multi-channel feature fusion technology can enhance the understanding of diverse malware variants, help security analysts quickly generate threat summaries and provide countermeasures. 4. Automated Vulnerability Management: Among a large number of vulnerability reports and update announcements, the entity recognition method provided in the embodiments of this application can automatically extract information such as vulnerability numbers, affected software, and vulnerability descriptions, helping the security team quickly analyze vulnerabilities and determine the priority patching order. Through dependency syntactic analysis and an adaptive fusion mechanism, the model can more accurately process the complex technical descriptions and vulnerability impact assessments in vulnerability reports.

[0076] Figure 2 FIG. is a system architecture diagram provided for the embodiments of this application. As shown in Figure 2 the figure, the system at least includes an input layer (Input) for inputting the target text; a BERT encoder (BERT Encoder) for extracting word features to obtain a first feature vector; a dependency parsing encoder (Dependency Parsing Encoder) for performing dependency relationship extraction and determining a third feature vector; a convolutional neural network (Convolutional Neural Network) for performing local feature extraction on the first feature vector to obtain a second feature vector; a BERT-CNN encoder (BERT CNN Encoder) for constructing a feature vector matrix based on the first feature vector, the second feature vector, and the third feature vector; an adaptive attention (AdaptiveAttention) for performing adaptive attention fusion (Adaptive Attention Fusion) on the feature vector matrix to obtain a target feature vector; and a CRF for determining entities.

[0077] Embodiment 8: Based on the above embodiments, the embodiments of this application further provide an entity recognition device. Figure 3 FIG. is a schematic structural diagram of an entity recognition device provided for the embodiments of this application. The device includes: The processing module 301 is configured to use the trained BERT model to perform embedding representation on the target text to be processed, and determine the first feature vector of the target text; use the trained convolutional neural network CNN to perform feature extraction on the first feature vector to determine the second feature vector; according to the character positions corresponding to each character in the target text and the saved dependency relation labels corresponding to each character, determine the dependency relation label sequence, and perform embedding representation on the dependency relation label sequence to obtain the third feature vector; determine the target feature vector of the target text according to the first feature vector, the second feature vector, and the third feature vector. The entity recognition module 302 is configured to input the target feature vector into the named entity recognition model, so that the named entity recognition model performs named entity recognition on the target text according to the target feature vector, and determine the target entities included in the target text.

[0078] In a possible implementation manner, the CNN includes a preset number of sub-CNNs, and the convolutional kernels included in each sub-CNN are different; The processing module 301 is specifically configured to input the first feature vector into each sub-CNN respectively, so that each sub-CNN performs feature extraction on the first feature vector, and obtain each sub-feature vector output by each sub-CNN; splice the each sub-feature vector in a set order, and determine the spliced feature vector as the second feature vector.

[0079] In a possible implementation manner, the sub-CNN includes a convolutional layer, an activation function, and a max pooling layer; The processing module 301 is specifically configured to, for each sub-CNN, the convolutional layer of the sub-CNN performs convolutional processing on the first feature vector by using a preset convolutional kernel to obtain an intermediate feature vector; the activation function and the max pooling layer of the sub-CNN perform dimension compression processing on the intermediate feature vector to obtain a sub-feature vector.

[0080] In a possible implementation manner, the processing module 301 is further configured to use a preset dependency relation label determination tool to perform dependency syntactic analysis on the target text, and determine each word and the corresponding dependency relation label in the target text; for each character in the target text, determine the target word to which the character belongs, and determine the dependency relation label corresponding to the target word as the dependency relation label corresponding to the character.

[0081] In a possible implementation manner, the processing module 301 is specifically configured to construct a feature vector matrix according to the first feature vector, the second feature vector, and the third feature vector; obtain a pre-configured weight matrix; and determine the target feature vector according to the weight matrix and the feature vector matrix.

[0082] In a possible implementation manner, the named entity recognition model includes a fully connected layer and a conditional random field CRF; The processing module 301 is specifically configured to: specifically, the fully connected layer of the named entity recognition model determines a score vector of each character in the target text belonging to each entity according to the target feature vector; and the CRF of the named entity recognition model determines the target entity according to the score vector corresponding to each character and a pre-configured constraint condition.

[0083] In a possible implementation manner, the processing module 301 is further configured to obtain a security text; segment the security text according to a preset segmentation rule; and respectively determine each segmented sub-text as the target text.

[0084] Embodiment 9: Based on the above embodiments, an embodiment of the present application further provides an electronic device, Figure 4 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application, as Figure 4 shown, including: a processor 401, a communication interface 402, a memory 403, and a communication bus 404. Among them, the processor 401, the communication interface 402, and the memory 403 complete communication with each other through the communication bus 404; A computer program is stored in the memory 403. When the program is executed by the processor 401, the processor 401 is caused to execute the following steps: Adopt a trained BERT model to perform embedding representation on a target text to be processed, and determine a first feature vector of the target text; Adopt a trained convolutional neural network CNN to perform feature extraction on the first feature vector, and determine a second feature vector; According to the character position corresponding to each character in the target text and the saved dependency relationship label corresponding to each character, determine a dependency relationship label sequence, and perform embedding representation on the dependency relationship label sequence to obtain a third feature vector; Determine a target feature vector of the target text according to the first feature vector, the second feature vector, and the third feature vector; Input the target feature vector into a named entity recognition model, so that the named entity recognition model performs named entity recognition on the target text according to the target feature vector, and determines the target entities included in the target text.

[0085] In a possible implementation manner, the CNN includes a preset number of sub-CNNs, and the convolutional kernels included in each sub-CNN are different; The step of using the trained convolutional neural network CNN to perform feature extraction on the first feature vector to determine the second feature vector includes: Input the first feature vector into each sub-CNN respectively, so that each sub-CNN performs feature extraction on the first feature vector, and obtains each sub-feature vector output by each sub-CNN; Concatenate the each sub-feature vector in a set order, and determine the obtained concatenated feature vector as the second feature vector.

[0086] In a possible implementation manner, the sub-CNN includes a convolutional layer, an activation function, and a max pooling layer; The step of each sub-CNN performing feature extraction on the first feature vector includes: For each sub-CNN, the convolutional layer of the sub-CNN performs convolutional processing on the first feature vector using a preset convolutional kernel to obtain an intermediate feature vector; the activation function and max pooling layer of the sub-CNN perform dimensionality compression processing on the intermediate feature vector to obtain a sub-feature vector.

[0087] In a possible implementation manner, the process of determining the dependency relation label corresponding to each character includes: Use a preset dependency relation label determination tool to perform dependency syntactic analysis on the target text, and determine each word in the target text and the corresponding dependency relation label; For each character in the target text, determine the target word to which the character belongs, and determine the dependency relation label corresponding to the target word as the dependency relation label corresponding to the character.

[0088] In a possible implementation manner, the step of determining the target feature vector of the target text according to the first feature vector, the second feature vector, and the third feature vector includes: Construct a feature vector matrix according to the first feature vector, the second feature vector, and the third feature vector; Obtain a pre-configured weight matrix; Determine the target feature vector according to the weight matrix and the feature vector matrix.

[0089] In a possible implementation manner, the named entity recognition model includes a fully connected layer and a conditional random field (CRF); The named entity recognition model performing named entity recognition on the target text according to the target feature vector includes: The fully connected layer of the named entity recognition model determines a score vector of each character in the target text belonging to each entity according to the target feature vector; The CRF of the named entity recognition model determines the target entity according to the score vector corresponding to each character and a preconfigured constraint condition.

[0090] In a possible implementation manner, the process of determining the target text includes: Obtain a security text; Segment the security text according to a preset segmentation rule; Determine each segmented sub-text as the target text respectively.

[0091] Since the principle of the above electronic device for solving problems is similar to that of the entity recognition method, the implementation of the above electronic device can refer to the embodiments of the method, and the repeated parts will not be described again.

[0092] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface 402 is used for communication between the above electronic device and other devices. The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0093] The above processor may be a general-purpose processor, including a central processor, a Network Processor (NP), etc.; it may also be a Digital Signal Processing (DSP), an application-specific integrated circuit, a field-programmable gate array or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0094] Example 10: Based on the above embodiments, an embodiment of the present invention further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the processor is caused to perform the following steps when executing: Adopt a trained BERT model to perform embedding representation on the target text to be processed, and determine the first feature vector of the target text; Adopt a trained convolutional neural network CNN to perform feature extraction on the first feature vector, and determine the second feature vector; According to the character position corresponding to each character in the target text and the saved dependency relation label corresponding to each character, determine the dependency relation label sequence, and perform embedding representation on the dependency relation label sequence to obtain the third feature vector; According to the first feature vector, the second feature vector and the third feature vector, determine the target feature vector of the target text; Input the target feature vector into the named entity recognition model, so that the named entity recognition model performs named entity recognition on the target text according to the target feature vector, and determines the target entity included in the target text.

[0095] In a possible implementation manner, the CNN includes a preset number of sub-CNNs, and the convolutional kernels included in each sub-CNN are different; The step of adopting a trained convolutional neural network CNN to perform feature extraction on the first feature vector and determine the second feature vector includes: Input the first feature vector into each sub-CNN respectively, so that each sub-CNN performs feature extraction on the first feature vector, and obtain each sub-feature vector output by each sub-CNN; Concatenate the each sub-feature vector in a set order, and determine the concatenated feature vector obtained as the second feature vector.

[0096] In a possible implementation manner, the sub-CNN includes a convolutional layer, an activation function and a max pooling layer; The step of each sub-CNN performing feature extraction on the first feature vector includes: For each sub-CNN, the convolutional layer of the sub-CNN performs convolutional processing on the first feature vector by using a preset convolutional kernel to obtain an intermediate feature vector; the activation function and max pooling layer of the sub-CNN perform dimensionality compression processing on the intermediate feature vector to obtain a sub-feature vector.

[0097] In a possible implementation manner, the process of determining the dependency relation label corresponding to each character includes: Using a preset dependency relation label determination tool, perform dependency syntactic analysis on the target text to determine each word and its corresponding dependency relation label in the target text; For each character in the target text, determine the target word to which the character belongs, and determine the dependency relation label corresponding to the target word as the dependency relation label corresponding to the character.

[0098] In a possible implementation manner, the determining the target feature vector of the target text according to the first feature vector, the second feature vector, and the third feature vector includes: Construct a feature vector matrix according to the first feature vector, the second feature vector, and the third feature vector; Obtain a pre-configured weight matrix; Determine the target feature vector according to the weight matrix and the feature vector matrix.

[0099] In a possible implementation manner, the named entity recognition model includes a fully connected layer and a conditional random field CRF; The named entity recognition model performing named entity recognition on the target text according to the target feature vector includes: The fully connected layer of the named entity recognition model determines a score vector for each character in the target text belonging to each entity according to the target feature vector; The CRF of the named entity recognition model determines the target entity according to the score vector corresponding to each character and a pre-configured constraint condition.

[0100] In a possible implementation manner, the process of determining the target text includes: Obtain a security text; Segment the security text according to a preset segmentation rule; Determine each segmented sub-text as the target text respectively.

[0101] Since the principle of the above computer program product for solving problems is similar to the entity recognition method, the implementation of the above computer program product can refer to the implementation of the method, and the repeated parts will not be elaborated.

[0102] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0103] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0104] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0106] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.

Claims

1. An entity recognition method, characterized in that: The method comprises: Using the trained BERT model, embedding the target text to be processed, and determining the first feature vector of the target text; Using a trained convolutional neural network (CNN) to extract features from the first feature vector, and determine a second feature vector; Determine a dependency label sequence according to the character position corresponding to each character in the target text and the saved dependency label corresponding to each character, and embed the dependency label sequence to obtain a third feature vector; Determining a target feature vector of the target text according to the first feature vector, the second feature vector and the third feature vector; The target feature vector is input into a named entity recognition model, so that the named entity recognition model performs named entity recognition on the target text according to the target feature vector to determine the target entity contained in the target text.

2. The method according to claim 1, characterized in that The CNN includes a preset number of sub-CNNs, and each sub-CNN includes a different convolution kernel; The using a trained convolutional neural network (CNN) to extract features from the first feature vector to determine the second feature vector comprises: Inputting the first feature vector into each sub-CNN respectively, so that each sub-CNN performs feature extraction on the first feature vector, and obtaining each sub-feature vector output by each sub-CNN; Each of the sub-feature vectors is concatenated in a set order, and the obtained concatenated feature vector is determined as the second feature vector.

3. The method according to claim 2, characterized in that The sub-CNN includes a convolutional layer, an activation function and a maximum pooling layer; Each sub-CNN extracts features from the first feature vector, including: For each sub-CNN, the convolution layer of the sub-CNN uses a preset convolution kernel to perform convolution processing on the first feature vector to obtain an intermediate feature vector; the activation function and the maximum pooling layer of the sub-CNN perform dimensional compression processing on the intermediate feature vector to obtain a sub-feature vector.

4. The method according to claim 1, characterized in that: The process of determining the dependency label corresponding to each character includes: Using a preset dependency label determination tool to perform dependency syntactic analysis on the target text to determine each word in the target text and the corresponding dependency label; For each character in the target text, the target word to which the character belongs is determined, and the dependency label corresponding to the target word is determined as the dependency label corresponding to the character.

5. The method according to claim 1, characterized in that Determining a target feature vector of the target text according to the first feature vector, the second feature vector, and the third feature vector comprises: Constructing an eigenvector matrix according to the first eigenvector, the second eigenvector and the third eigenvector; Get a pre-configured weight matrix; The target eigenvector is determined according to the weight matrix and the eigenvector matrix.

6. The method according to claim 1, characterized in that The named entity recognition model includes a fully connected layer and a conditional random field CRF; The named entity recognition model performs named entity recognition on the target text according to the target feature vector, comprising: The fully connected layer of the named entity recognition model determines the score vector of each character in the target text belonging to each entity according to the target feature vector; The CRF of the named entity recognition model determines the target entity according to the score vector corresponding to each character and the pre-configured constraints.

7. The method according to claim 1, characterized in that The process of determining the target text includes: Get secure text; Segment the security text according to a preset segmentation rule; Each sub-text obtained by segmentation is determined as the target text.

8. An entity recognition device, characterized in that: The device comprises: A processing module is used to use the trained BERT model to embed the target text to be processed and determine the first feature vector of the target text; use the trained convolutional neural network CNN to extract features from the first feature vector and determine the second feature vector; determine the dependency label sequence according to the character position corresponding to each character in the target text and the dependency label corresponding to each character that is saved, and embed the dependency label sequence to obtain a third feature vector; determine the target feature vector of the target text according to the first feature vector, the second feature vector and the third feature vector; The entity recognition module is used to input the target feature vector into a named entity recognition model, so that the named entity recognition model performs named entity recognition on the target text according to the target feature vector to determine the target entity contained in the target text.

9. An electronic device, characterized in that: The electronic device comprises at least a processor and a memory, and the processor is used to implement the steps of the entity recognition method according to any one of claims 1 to 7 when executing a computer program stored in the memory.

10. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements the steps of the entity recognition method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Text entity recognition model construction method and equipment based on large model data enhancement

    CN120995985A

  • Text entity recognition model construction method and device based on large model data augmentation

    CN120995985B