Domain knowledge graph entity recognition method based on multi-feature fusion learning
By adopting a multi-feature fusion learning method in the domain knowledge graph, combined with the attention mechanism of BERT and BiLSTM, the accuracy of named entity recognition in the case of data scarcity is solved, and higher knowledge graph coverage and accuracy are achieved.
Patent Information
- Application Number
- CN202510204332.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-13
AI Technical Summary
The recognition of named entity in the domain knowledge graph is limited in the coverage and accuracy of the professional field due to the scarcity or difficulty in obtaining data.
Using a multi-feature fusion learning method, BERT's self-attention and BiLSTM's sequence attention are combined, and global optimization is performed through multi-task learning and conditional random field (CRF) layer to extract and fuse multiple features to improve the accuracy of entity recognition.
It significantly improves the training effect and generalization performance of the model under limited data conditions, improves the accuracy and coverage of the domain knowledge graph, and enhances the expression ability and prediction accuracy of the model.
Smart Images

Figure CN120146049A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of entity recognition in domain knowledge graphs, and particularly relates to a method for entity recognition in domain knowledge graphs based on multi-feature fusion learning. Background Art
[0002] As a structured knowledge representation method, domain knowledge graphs have demonstrated their unique value in many specific fields. For example, in professional fields such as artificial intelligence, medicine, law, and education, there are extremely high requirements for the accuracy and depth of knowledge. In professional fields, named entity recognition, as a key technology and link in constructing knowledge graphs, due to the scarcity or difficulty of obtaining data, there are limitations in the coverage and accuracy of the generated knowledge graphs in professional fields. The following are several common technical solutions:
[0003] 1. A cascaded Transformer-based Chinese medical entity recognition method based on multi-semantics. By matching sentences with an external medical dictionary, a lattice sequence containing Chinese characters and matching words is constructed. A syntactic parser is used to parse the input sentence to obtain the syntactic tree of the sentence, and syntactic relationship conversion rules are designed to convert the word-based syntactic tree result into a character-based syntactic relationship matrix. The lattice sequence is passed into the Lattice Transformer module, and the vocabulary-aware self-attention mechanism is used to model the lattice features. The obtained lattice features and syntactic relationship matrix are passed into the Syntax Transformer module to further extract syntactic features. Through this cascaded modeling method, multi-semantic features integrating word and syntactic information are obtained. Finally, the multi-semantic features are passed into the conditional random field module to predict labels. It has the following problems: First, it highly depends on external libraries. It depends on an external medical dictionary to construct the lattice sequence. If the dictionary is incomplete or not updated in a timely manner, it will affect the accuracy of recognition. Second, the correctness of the parser. Using a syntactic parser to parse sentences, any parsing error may lead to error propagation and affect the model performance. Third, high computational complexity. Due to the involvement of complex cascaded Transformer models and conditional random field modules, it may face high computational costs and resource consumption when processing large-scale medical text data. Deep learning models usually require a large amount of labeled data for training. If the training data is insufficient or of low quality, it may affect the performance and accuracy of the model. Fourth, the generalization ability of the model is limited to a specific field. Since it is mainly trained for the medical field, its application in other fields may require further adjustment and optimization.
[0004] 2. The joint extraction method of entity relations based on semantic enhancement and multi-feature fusion regards the joint extraction task of entity relations as a function that maps an object through a subject with a relation as a condition, and adopts the extraction idea of first identifying the head entity and then identifying the tail entity under each relation. First, it identifies the head entity information. It uses a pointer network to identify the start and end positions of the head entity by using the features that enhance the sequence dependence information after being encoded by the RNN, and takes the head entity and its entity type as prior information. It fuses multi-features with enhanced features to obtain a fusion vector with enhanced semantic expression ability, reducing the model's attention to semantically irrelevant entities. It has the following problems: First, this method depends on a complex joint decoding algorithm, and the efficiency may not be high when extracting triples. Second, the existing pre-trained encoders may not have strong enough semantic breadth when dealing with specific problems, which limits the semantic expression ability of the model. Third, since the model is mainly designed for specific tasks, its generalization ability may be limited, and corresponding adjustments and optimizations are required for different domains.
[0005] 3. The multi-feature fusion named entity recognition method for police situation texts. First, it constructs a dataset for police situation named entity recognition, defines the entity types to be recognized, and divides them into a training set, a validation set, and a test set. Second, it uses pre-trained word vectors to obtain the character features of the text, performs text matching based on rules and dictionaries to obtain pre-recognized label features, and converts the text into pinyin features. Finally, it fuses the above three features and sends them into a bidirectional long short-term memory network conditional random field model for named entity recognition. However, it also has the following disadvantages: First, this method depends on a large number of pre-trained word vectors and feature extraction, which may lead to high computational resources and time costs when training and deploying the model. Second, the performance of the model depends to a large extent on a high-quality labeled dataset, and the labeled data in the police situation field is often scarce and expensive, which may limit the training effect and generalization ability of the model. Third, although this method fuses character features, pre-recognized label features, and pinyin features, there may still be challenges in dealing with colloquial expressions and typos in police situation texts because these features may not be sufficient to capture all language variations. Summary of the Invention
[0006] The purpose of the present invention is to solve the problem that when performing text recognition, the named entity recognition in the domain knowledge graph has limitations in the coverage and accuracy in the professional field due to the scarcity or difficulty of obtaining data, and a domain knowledge graph entity recognition method based on multi-feature fusion learning is proposed.
[0007] The technical solution of the present invention is: A domain knowledge graph entity recognition method based on multi-feature fusion learning, including the following steps:
[0008] S1. Input the sentence sequence into the BERT model to convert it into high-dimensional vectors, extract global semantic features, and output word vectors;
[0009] S2. Input the word vectors into the BiLSTM+Fused attention network for bidirectional sequence modeling, obtain enhanced features, and perform multi-feature fusion on the enhanced features to output a label sequence;
[0010] S3. Input the label sequence into the CRF layer for global optimization to obtain an output sequence that conforms to the syntax rules and context logic of named entities, and complete the entity recognition of the domain knowledge graph.
[0011] Preferably, step S1 specifically includes the following sub-steps:
[0012] S11. Convert the sentence sequence in text form into a sentence set S, and use each sentence in the sentence set S as a basic processing unit;
[0013] S12. Input each sentence in the sentence set into the tokenizer tool for processing, and output a character ID sequence (x 1 ,x 2 ,...,x i ,...,x n ), where x i represents the ID of the i-th character, and i ∈ [1, 2,..., n];
[0014] S13. Input the character ID sequence (x 1 ,x 2 ,...,x i ,...,x n ) into the BERT model for feature extraction, and output the embedding features of each character;
[0015] S14. According to the embedding features, output a vector with a dimension of [n l , 768] for each sentence, and then obtain the word vector of each sentence, where n l represents the sentence length.
[0016] Preferably, the sentence set S in step S11 is:
[0017] S = {s 1 , s 2 ,..., s l ,... s m}
[0018] where s l represents the l-th sentence, and l ∈ [1, 2,..., m], and m represents the total number of sentences in the sentence set S;
[0019] The word vector is as follows:
[0020] H l = BERT(s l )
[0021] Wherein, H l represents the word vector, and the word vector H l is a high-dimensional vector that captures global semantic features, and BERT represents the BERT model.
[0022] Preferably, the BiLSTM + Fused attention network in step S2 includes a BiLSTM module and a Fused attention module;
[0023] The BiLSTM module is used to receive the word vector and perform bidirectional sequence modeling on the word vector to obtain enhanced features;
[0024] The Fused attention module is used to receive the enhanced features and perform multi-feature fusion on the enhanced features to output a label sequence.
[0025] Preferably, the enhanced feature B l is expressed by the formula:
[0026] B l = BiLSTM(H l )
[0027] Wherein, BiLSTM represents the bidirectional sequence modeling operation of BiLSTM, and H l represents the word vector.
[0028] Preferably, the BiLSTM module includes a forward LSTM link and a backward LSTM link, and both the forward LSTM link and the backward LSTM link are composed of multiple LSTM units; each LSTM unit includes three gate structures: an input gate, a forget gate, and an output gate;
[0029] The forward LSTM link is used to process the input word vector forward;
[0030] The backward LSTM link is used to process the input word vector backward;
[0031] The Fused attention module includes multiple parallel Attention units, and each Attention unit receives the output of each LSTM unit respectively.
[0032] Preferably, the multi-feature fusion of the enhanced features in step S2 includes the following steps:
[0033] Generate queries, keys, and values based on the enhanced features; where the query represents the target, that is, the focus that needs to be concerned currently, the key represents the enhanced features after processing each word of the word vector, and the value represents the word vector;
[0034] Calculate the similarity between the query and the key using dot product or cosine similarity;
[0035] Normalize the similarity to obtain the attention weights;
[0036] Perform weighted summation on the values according to the attention weights to complete multi-feature fusion.
[0037] Preferably, the step S3 specifically includes the following sub-steps:
[0038] Initialize the transition matrix of the CRF layer; the transition matrix represents the probability score of transferring from one label to another at each character position i in the sentence sequence;
[0039] Convert the label sequence into a scoring vector, each element of the scoring vector corresponds to the score of a potential label, and calculate the transition score between each pair of consecutive labels in the sentence sequence according to the transition matrix and the scoring vector;
[0040] Calculate the comprehensive score of the path based on the transition score;
[0041] Use the dynamic programming algorithm to select the path with the highest score as the final prediction result.
[0042] Preferably, the expression formula of the transition matrix is:
[0043] M i (y i-1 ,y i |x) = exp(W i (y i-1 ,y i |x))
[0044] where x represents the input feature vector, y i represents the label of the i-th character in the sequence, M i (y i-1 ,y i |x) represents the probability score of transferring from label y i-1 to label y i under the condition of the given input feature vector x, that is, the transition matrix from label y i-1 to label y i , W i (y i-1 ,y i |x) represents from label yi-1 Transfer to label y i The original score, where exp represents the exponential function with the natural constant e as the base, is used to convert the original score into a positive number.
[0045] Preferably, the comprehensive score includes an emission score and a transfer score;
[0046] The emission score is the element in the score vector obtained by multiplying the feature vector of each character or word in the sequence by the emission matrix;
[0047] The calculation formula for the emission score is:
[0048]
[0049] where, represents the score of predicting the i-th character as the y-th i label, x i is the feature vector of the i-th character, is the row vector of the label y corresponding to the i-th character in the emission matrix i in the sequence, and x T represents the transpose of x i ; the emission matrix is a weight matrix that maps the feature space to the label space, with each row corresponding to a label and each column corresponding to a dimension of the feature vector;
[0050] The transfer score is the transfer score between all adjacent labels;
[0051] The calculation formula for the comprehensive score is:
[0052]
[0053] where, score(x, y) represents the comprehensive score, represents the transfer score from label y i to label y i+1 , n represents the total number of characters, and y represents the label sequence of the entire observed sequence.
[0054] The beneficial effects of the present invention are:
[0055] 1. The present invention proposes a method based on multi-feature fusion learning, which combines the self-attention of BERT and the sequence attention of BiLSTM, and innovatively constructs a more comprehensive attention mechanism. This mechanism can utilize context information to strengthen the extraction of key features, which is particularly significant in the process of entity recognition.
[0056] 2. The present invention proposes a multi-task learning method, which shares and optimizes the underlying feature representations of the model by simultaneously training multiple related tasks, effectively reducing the risk of overfitting, improving the training effect and generalization performance of the model under limited data conditions, and significantly improving the performance of entity recognition through training and experiments in the artificial intelligence field dataset, providing new theoretical support and practical guidance for research applications in professional fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 The flowchart of the domain knowledge graph entity recognition method based on multi-feature fusion learning provided in Embodiment 1 of the present invention is shown.
[0058] Figure 2 The framework diagram of the BERT-BiLSTM-Fused attention-CRT model provided in Embodiment 2 of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] Now, the exemplary embodiments of the present invention will be described in detail with reference to the drawings. It should be understood that the embodiments shown and described in the drawings are merely exemplary, intended to illustrate the principles and spirit of the present invention, and not to limit the scope of the present invention.
[0060] Embodiment 1:
[0061] As Figure 1 shown, a domain knowledge graph entity recognition method based on multi-feature fusion learning includes the following steps:
[0062] S1. Input the sentence sequence into the BERT model to convert it into high-dimensional vectors to extract global semantic features, and output word vectors;
[0063] S2. Input the word vectors into the BiLSTM+Fused attention network for bidirectional sequence modeling to obtain enhanced features, and perform multi-feature fusion on the enhanced features to output a label sequence;
[0064] S3. Input the label sequence into the CRF layer for global optimization to obtain an output sequence that conforms to the syntactic rules and context logic of named entities, and complete the domain knowledge graph entity recognition.
[0065] In this embodiment, the BERT model includes multiple Transformer layers; each Transformer layer includes a feed-forward neural network for performing non-linear transformation on the output of the self-attention layer to further enhance feature representation; the BERT model uses multiple Transformer layers to deeply analyze and process the input word ID sequence. In this process, the self-attention mechanism enables the BERT model to capture the complex relationships between words, while multi-head attention allows the BERT model to learn multiple different representations simultaneously. To process sequence position information, BERT introduces position encoding to ensure that the model can understand the order of words. After being processed by multiple layers of encoders, BERT outputs the embedding representation of each word, and these embedding vectors are rich in semantic and context information, providing strong feature support for downstream tasks such as named entity recognition.
[0066] In this embodiment, the BERT model deeply explores the deep meaning of Chinese text by taking each Chinese character in the sentence as an independent unit. Using the self-attention mechanism and the feed-forward neural network, BERT can finely capture the complex relationships between words. The output vector not only contains the information of a single Chinese character but also integrates the context information of the Chinese character in the sentence, thus achieving a comprehensive understanding and in-depth representation of the text semantics. This method effectively avoids the complex word segmentation steps in English text processing and improves the efficiency of text processing. The step S1 specifically includes the following sub-steps:
[0067] S11. Convert the sentence sequence in text form into a sentence set S, and each sentence in the sentence set S is used as a basic processing unit;
[0068] S12. Input each sentence in the sentence set into the tokenizer tool for processing, and the output is a word ID sequence (x 1 ,x 2 ,...,x i ,...,x n ), where x i represents the ID of the i-th character, and i ∈ [1, 2,..., n];
[0069] S13. Input the word ID sequence (x 1 ,x 2 ,...,x i ,...,x n ) into the BERT model for feature extraction, and the output is the embedding feature of each character;
[0070] S14. According to the embedding feature, output a vector with a dimension of [n l , 768] for each sentence, and then obtain the word vector of each sentence, where n lIndicates the sentence length.
[0071] In this embodiment, the sentence set S in step S11 is:
[0072] S = {s 1 , s 2 ,..., s l ,... s m}
[0073] where s l represents the l-th sentence, and l ∈ [1, 2,..., m], where m represents the total number of sentences in the sentence set S;
[0074] The word vector is:
[0075] H l = BERT(s l )
[0076] where H l represents the word vector, and the word vector H l is a high-dimensional vector that captures the global semantic features. H l captures the global semantic features, and BERT represents the BERT model.
[0077] In this embodiment, the BiLSTM + Fused attention network in step S2 includes a BiLSTM module and a Fused attention module;
[0078] The BiLSTM module is used to receive the word vector and perform bidirectional sequence modeling on the word vector to enhance the sequential modeling ability of the sentence, so as to strengthen the sequential correlation of the output vector, and then obtain the enhanced feature B l , and its expression formula is:
[0079] B l = BiLSTM(H l )
[0080] where BiLSTM represents the bidirectional sequence modeling operation of BiLSTM, and H l represents the word vector;
[0081] The Fused attention module is used to receive the enhanced feature and perform multi-feature fusion on the enhanced feature to output the label sequence.
[0082] In this embodiment, the BiLSTM module includes a forward LSTM link and a backward LSTM link, which can process forward and backward text information simultaneously. When the BiLSTM module processes sequential data, it can take into account the information before and after the current word, effectively capture long-distance dependencies, and improve the accuracy of the BiLSTM+Fused attention network. Both the forward LSTM link and the backward LSTM link are composed of multiple LSTM units. Each LSTM unit includes three gate structures: an input gate, a forget gate, and an output gate, which can achieve the long-term memory ability of the BiLSTM+Fused attention network.
[0083] The forward LSTM link is used to process the input word vectors in the forward direction.
[0084] The backward LSTM link is used to process the input word vectors in the reverse direction.
[0085] The Fused attention module includes multiple parallel Attention units, which are used to receive the outputs of multiple LSTM units.
[0086] In this embodiment, the goal of adding the fused attention mechanism is to allow the BiLSTM+Fused attention network to focus on the most important parts of the input sequence when processing sequential data. The multi-feature fusion of the enhanced features in step S2 includes the following steps:
[0087] Generate Query, Key, and Value based on the enhanced features. Among them, the Query represents the target, that is, the focus that needs to be concerned currently. The Key represents the enhanced features of each word in the word vector after processing, and the Value represents the word vector.
[0088] Use a compatibility function such as dot product or cosine similarity to calculate the similarity or matching degree between the Query and the Key.
[0089] Normalize the similarity to obtain attention weights, which determine the degree to which each part should be "attended to" when processing the input. The normalization process can be implemented using the softmax function.
[0090] Perform weighted summation on the Value according to the attention weights. The BiLSTM+Fused attention network will focus more on the parts that are considered more important or relevant, thereby completing multi-feature fusion.
[0091] In this embodiment, the Conditional Random Field (CRF) is used in the final stage of the sequence labeling task to ensure that the label sequence output by the model is globally optimal. The specific steps of step S3 include the following sub-steps:
[0092] Initialize the transition matrix of the CRF layer; the transition matrix represents the probability scores of transitioning from one label to another at each character position i in the sentence sequence; these weights reflect the possible dependencies or sequential constraints between different labels. The initial state uses random values or sets a reasonable prior distribution based on domain knowledge, but as the training process progresses, these weights will be automatically adjusted according to the data to better adapt to the needs of the specific task; the transition matrix is used to describe the transition probabilities between labels, and the expression formula is:
[0093] M i (y i ,y i+1 |x) = exp(W i (y i ,y i+1 |x))
[0094] where x represents the input feature vector, M i (y i ,y i+1 |x) represents the probability score of transitioning from label y i to label y i+1 , that is, the transition matrix from label y i to label y i+1 , W i (y i ,y i+1 |x) represents the raw score of transitioning from label y i to label y i+1 , exp represents the exponential function with the natural constant e as the base, which is used to convert the raw score into a positive number, and can be normalized so that the sum of the scores of all possible transitions is 1, obtaining a valid probability distribution. These scores represent the probability of transitioning from the previous label to the next label in a specific context. The CRF can take into account the overall structure of the entire label sequence, rather than just isolated decisions at a single time point. Such a global perspective helps to avoid generating illogical label combinations and improves the consistency and accuracy of the prediction results.
[0095] Convert the label sequence into a scoring vector, where each element of the scoring vector corresponds to the score of a potential label, and calculate the transition scores between each pair of consecutive labels in the sentence sequence according to the transition matrix and the scoring vector; this step can quantitatively evaluate the likelihood of all possible labels at each position under the given input sequence;
[0096] Calculate the comprehensive score of the path based on the transition scores; the comprehensive score includes the emission score and the transition score;
[0097] The emission score is a score vector obtained by multiplying the feature vector of each character or word in the sequence by the emission matrix. Each element of this score vector corresponds to the score of a label. In the CRF layer, the emission matrix is a weight matrix that maps the feature space to the label space. Each row of it corresponds to a label, and each column corresponds to a dimension of the feature vector. The parameters of the emission matrix are learned through the training process. The calculation formula of the emission score is:
[0098] where, represents the score of predicting the i-th character as the y i -th label, x i is the feature vector of the i-th character, is the row vector corresponding to the label y i of the i-th character in the sequence in the emission matrix, and x T represents the transpose of x i ;
[0099] The transition score is the transition score between all adjacent labels;
[0100] The calculation formula of the comprehensive score is:
[0101]
[0102] where, score(x, y) represents the comprehensive score, represents the transition score from label y i to label y i+1 , n represents the total number of characters, and y represents the label sequence of the entire observed sequence;
[0103] Using the Viterbi algorithm, find the path with the highest score in polynomial time as the final prediction result.
[0104] The present invention proposes a method based on multi-feature fusion learning, which combines the self-attention of BERT and the sequence attention of BiLSTM, and innovatively constructs a more comprehensive attention mechanism. For example, "Xiaomi" may be food or a device brand in different language environments. By highlighting the features closely related to entity boundaries and category discrimination, it helps to more accurately identify entities in the text and more precisely determine the boundaries and categories of entities.
[0105] As an effective decoding tool, CRF integrates the advantages of local feature recognition and global constraints, and can achieve excellent performance in sequence labeling tasks, thus solving the problem that BiLSTM cannot correctly recognize the dependencies between adjacent labels. It ensures that when the model processes sequence labeling problems, it not only focuses on local information but also globally considers the mutual dependencies between labels, so as to identify the overall optimal label sequence.
[0106] In the present invention, the multi-feature fusion mechanism can be replaced by feature enhancement techniques. By generating additional synthetic features, such as polynomial features, interaction features, etc., these new features can capture more complex non-linear relationships between the original features, thereby improving the expressive power and prediction accuracy of the BERT-BiLSTM-Fused attention-CRT model. This method helps to mine the implicit correlation structure in the data. For example, in a specific field, certain symptom combinations in medical data may strongly indicate a specific disease. By appropriately selecting the number and type of newly added features, the generalization ability of the model to new data can be enhanced while avoiding overfitting. In addition, feature enhancement can also simplify the optimization process and accelerate the training speed. For relatively simple models, by introducing reasonable feature engineering, the algorithm can be guided to converge to the optimal solution faster, reducing the dependence on complex models.
[0107] The method for entity recognition in a domain knowledge graph based on multi-feature fusion learning proposed by the present invention. First, in the pre-training stage, by leveraging the advantages of the BERT (Bidirectional Encoder Representations from Transformers) model, large-scale domain text data is processed diversely, such as synonym replacement, sentence pattern transformation, etc., effectively expanding the training data set and enhancing the generalization ability of the model and its adaptability to unseen data. Next, in the functional training stage, a data set for a specific task such as named entity recognition or relation extraction is selected to finely train the model, and a template-based method is adopted to embed task-related features and prior knowledge into the input data, which not only enhances the prediction ability of the model under domain data training conditions but also improves the task adaptability. In addition, for the named entity recognition model, a multi-feature fusion learning mechanism is introduced. By fusing the self-attention of BERT and the sequence attention of the bidirectional long short-term memory network BiLSTM (Bi-directional LSTM), key feature extraction is strengthened using context information, effectively combining general knowledge in the general knowledge graph with data in the professional domain to improve the accuracy and reliability of the domain knowledge graph. At the same time, to further improve the prediction accuracy of the model, a multi-task learning method is adopted. By simultaneously training multiple related tasks to share and optimize the underlying feature representations of the model, the risk of overfitting is effectively reduced, and the training effect and generalization performance of the model under limited data conditions are improved. Finally, the conditional random field CRF (Conditional Random Field) is used for label prediction to obtain the optimal sequence between labels and automatically learn the constraint relationships between labels, significantly enhancing the expressive ability of the model.
[0108] Embodiment 2:
[0109] On the basis of Embodiment 1, a reference comparison experiment is conducted on a data set in the field of artificial intelligence in this embodiment of the present invention to illustrate the technical effects of the present invention.
[0110] In this embodiment, a BERT-BiLSTM-Fused attention-CRT model based on the BERT model, the BiLSTM+Fused attention network, and the CRF layer is constructed to implement the method for entity recognition in a domain knowledge graph based on multi-feature fusion learning as described in Embodiment 1. The structure of the BERT-BiLSTM-Fused attention-CRT model is as Figure 2 shown.
[0111] To ensure the accuracy of experimental data, multiple sets of iterative experiments were carried out. The following are the specific performance indicators of each group of experimental models: Precision, representing the proportion of correctly predicted instances among those predicted as entities by the model; Recall, indicating the proportion of all labeled entities successfully predicted as entities by the model; and the F1 value, which is an indicator that comprehensively considers Precision and Recall. When both are at a relatively high level, the F1 value will also be relatively high.
[0112] In this embodiment, in order to verify the effectiveness of the domain knowledge graph entity recognition method based on multi-feature fusion learning proposed by the present invention, an in-house dataset in the field of artificial intelligence was used for experimental testing. The in-house dataset in the field of artificial intelligence was formed by collecting artificial intelligence-related course resources and network resources in cooperation with the Open University of China, and after repeatedly performing manual and semi-automatic annotation on the data, a dataset with a data scale of 11,000 nodes was formed. This dataset covers a wide range of terms, proper nouns, and complex contexts in the field of artificial intelligence, ensuring the high quality and diversity of the data. When constructing the validation set and the training set, the embodiments of the present invention shuffled the order of the 11,000 training sets and constructed the training set and the validation set in the way that one positive sample corresponds to two negative samples. The experimental results are shown in Table 1.
[0113] Table 1 Experimental Results of the Dataset in the Field of Artificial Intelligence
[0114]
[0115] As can be seen from Table 1, the performance of the dataset in the field of artificial intelligence on the CNN-LSTM-CRF model is not good, and none of the three indicators exceeds 50%. Since the data quality of the in-house dataset is related to data collection, cleaning, and annotation, the accuracy is often not high during experiments. However, through the experimental results of the dataset in the field of artificial intelligence, it can still be seen that after adding BERT, all three indicators are significantly improved. The accuracy has increased by 24%, and the recall rate and the F1 value have increased by more than 10%. This is because BERT can provide rich context-based word embeddings and capture more in-depth semantic information, while the BiLSTM layer can effectively process sequence information and obtain more refined feature representations on this basis. At the same time, after adding the multi-feature fusion learning mechanism, the accuracy and the F1 value are still continuously improved. This method of comprehensively utilizing the advantages of various technologies enables the BERT-BiLSTM-Fused attention-CRT model to greatly exceed the traditional combined model in terms of recognition accuracy and robustness. Through the above experimental verification, the application based on multi-feature fusion learning can more accurately capture and focus on information highly relevant to the task, thus significantly improving the accuracy and efficiency of entity recognition.
[0116] Those of ordinary skill in the art will realize that the embodiments described herein are for the purpose of assisting the reader in understanding the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the present invention.
Claims
1. A domain knowledge graph entity recognition method based on multi-feature fusion learning, characterized in that: The following steps are involved: S1. Input the sentence sequence into the BERT model and convert it into a high-dimensional vector to extract global semantic features, and output the word vector; S2. Input the word vector into the BiLSTM+Fused attention network for bidirectional sequence modeling to obtain enhanced features, and then perform multi-feature fusion on the enhanced features to output the label sequence; S3. Input the label sequence into the CRF layer for global optimization to obtain an output sequence that conforms to the grammatical rules and contextual logic of the named entity, completing the entity recognition of the domain knowledge graph.
2. According to claim 1, the domain knowledge graph entity recognition method based on multi-feature fusion learning is characterized in that: The step S1 specifically includes the following sub-steps: S11. Convert the sentence sequence in text form into a sentence set S, and take each sentence in the sentence set S as a basic processing unit; S12. Input each sentence in the sentence set into the tokenizer tool for processing, and output the word ID sequence (x1, x2, ..., x i ,...,x n ), where x i represents the ID of the i-th character, i∈[1,2,...,n]; S13. Substitute the word ID sequence (x1, x2, ..., x i ,...,x n ) is input into the BERT model for feature extraction, and the embedded features of each character are output; S14. Based on the embedded features, output a dimension of [n l ,768], and then get the word vector of each sentence, where n l Indicates the length of the sentence.
3. The domain knowledge graph entity recognition method based on multi-feature fusion learning according to claim 2 is characterized in that: The sentence set S in step S11 is: S={s1,s2,...,s l ,...s m } Among them, s l represents the lth sentence, and l∈[1,2,...,m], where m represents the total number of sentences in the sentence set S; The word vector is: H l =BERT(s l ) Among them, H l Represents word vector, word vector H l BERT represents the BERT model, which captures the high-dimensional vector of global semantic features.
4. The domain knowledge graph entity recognition method based on multi-feature fusion learning according to claim 1 is characterized in that: The BiLSTM+Fused attention network in step S2 includes a BiLSTM module and a Fused attention module; The BiLSTM module is used to receive word vectors and perform bidirectional sequence modeling on the word vectors to obtain enhanced features; The Fused attention module is used to receive the enhanced features, perform multi-feature fusion on the enhanced features, and output a label sequence.
5. The domain knowledge graph entity recognition method based on multi-feature fusion learning according to claim 4 is characterized in that: The enhanced feature B l The expression formula is: B l =BiLSTM(H l ) Among them, BiLSTM represents the bidirectional sequence modeling operation of BiLSTM, H l Represents word vector.
6. The domain knowledge graph entity recognition method based on multi-feature fusion learning according to claim 4 is characterized in that: The BiLSTM module includes a forward LSTM link and a reverse LSTM link, and the forward LSTM link and the reverse LSTM link are both composed of multiple LSTM units; each LSTM unit includes three gate structures: an input gate, a forget gate, and an output gate; The forward LSTM link is used to forward process the input word vector; The reverse LSTM link is used to reversely process the input word vector; The Fused attention module includes multiple parallel Attention units, each Attention unit receives the output of each LSTM unit respectively.
7. The domain knowledge graph entity recognition method based on multi-feature fusion learning according to claim 6 is characterized in that: The multi-feature fusion of the enhanced features in step S2 includes the following steps: Generate query, key and value based on the enhanced features; query represents the target, that is, the current focus, key represents the enhanced features of each word in the word vector after processing, and value represents the word vector; Calculate the similarity between query and key using dot product or cosine similarity; Normalize the similarity to get the attention weight; The values are weighted and summed according to the attention weights to complete multi-feature fusion.
8. The domain knowledge graph entity recognition method based on multi-feature fusion learning according to claim 1 is characterized in that: The step S3 specifically includes the following sub-steps: Initialize the transfer matrix of the CRF layer; the transfer matrix represents the probability score of transferring from one label to another label at each character position i in the sentence sequence; Convert the label sequence into a score vector, where each element of the score vector corresponds to the score of a potential label. Calculate the transfer score between each pair of consecutive labels in the sentence sequence based on the transfer matrix and the score vector. Calculate a composite score for the pathway based on the transfer score; Using the dynamic programming algorithm, the path with the highest score is selected as the final prediction result.
9. The domain knowledge graph entity recognition method based on multi-feature fusion learning according to claim 8 is characterized in that: The transfer matrix is expressed as follows: M i (y i-1 ,y i |x)=exp(W i (y i-1 ,y i |x)) Among them, x represents the input feature vector, y i Indicates the label of the i-th character in the sequence, M i (y i-1 ,y i |x) means that given the input feature vector x, from label y i-1 Move to label y i The probability score from the label y i-1 Move to label y i The transfer matrix, W i (y i-1 ,y i |x) indicates that the label y i-1 Move to label y i The original fraction, exp represents the exponential function with the natural constant e as the base, which is used to convert the original fraction into a positive number.
10. The domain knowledge graph entity recognition method based on multi-feature fusion learning according to claim 8 is characterized in that: The comprehensive score includes a launch score and a transfer score; The emission score is an element in a score vector obtained by multiplying the feature vector of each character or word in the sequence by the emission matrix; The calculation formula of the emission score is: in, Indicates that the i-th character is predicted to be the y-th character i The score of labels, x i is the feature vector of the ith character, is the label y of the i-th character in the corresponding sequence in the emission matrix i The row vector, x T Represents x i The emission matrix is a weight matrix that maps the feature space to the label space, each row of which corresponds to a label and each column corresponds to a dimension of the feature vector; The transfer score is the transfer score between all adjacent labels; The calculation formula of the comprehensive score is: Among them, score(x,y) represents the comprehensive score, Indicates that from label y i Move to label y i+1 The transfer score of , n represents the total number of characters, and y represents the label sequence of the entire observation sequence.
Citation Information
Patent Citations
Knowledge graph intelligent question-answer method fusing pointer generation network
CN113010693A
BERT-BiLSTM-CRF-based ship named entity identification method
CN117744658A
Training a Joint Many-Task Neural Network Model using Successive Regularization
US20180121799A1