Knowledge graph construction method based on long and short term memory network introducing attention mechanism

By introducing a self-attention mechanism in LSTM and combining the BERT model, the shortcomings of the existing knowledge graph construction methods in processing unstructured data are solved, the expression ability and completeness of the knowledge graph are improved, and the ability to capture long-distance dependencies is enhanced.

CN120012893APending Publication Date: 2025-05-16ZHEJIANG UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510122563.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing knowledge graph construction methods are difficult to remove noise when processing unstructured data, process long entity annotations in an inaccurate manner, and traditional word embedding methods cannot fully consider context information, resulting in poor performance of models when understanding complex semantics and polysynonyms.

Method used

A knowledge graph method is used to construct a long and short-term memory network (LSTM) based on the introduction of attention mechanism. By introducing a self-attention mechanism in LSTM, the similarity between each time step and the previous time step is calculated, the context vector of weighted representation is generated, and deep semantic understanding is captured in combination with the BERT model.

Benefits of technology

It improves the expression ability and completeness of the knowledge graph, enhances the model's ability to capture long-distance dependencies, improves the information extraction effect of long entities, and solves the problem of unstructured data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012893A_ABST
    Figure CN120012893A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph construction method based on a long and short term memory network introducing an attention mechanism, and the method specifically comprises the steps: S1, carrying out the sentence segmentation and word segmentation preprocessing of a data set, including the removal of noise, special characters and stop words; s2, performing BIOES label labeling on the data after word segmentation, and introducing a hierarchical relationship; s3, word vector representation is established, a BERT model is introduced to capture context information, and deep semantic understanding is provided for word vectors; s4, establishing a Bi-LSTM neural network layer, obtaining a score probability of each word corresponding to each tag by using the Bi-LSTM neural network layer, adding an attention mechanism behind the Bi-LSTM layer, and capturing a long-distance dependency relationship; s5, establishing a CRF-GNN layer, and obtaining an output labeling sequence with the maximum probability; and S6, carrying out post-processing on the extracted information. According to the method, the Bi-LSTM model and the CRF model are fused, corresponding constraints can be given to the tag sequence, and the problem that information extraction output logic is disordered is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a knowledge graph construction method based on a long short-term memory network introducing an attention mechanism. Background Art

[0002] With the rapid development of information technology, unstructured data (such as text, images, audio, etc.) has exploded. Among them, text data, as an important carrier of knowledge, contains rich information, but its unstructured characteristics make information extraction and knowledge management face huge challenges.

[0003] Existing technical solutions lack data processing capabilities, and when processing large-scale unstructured data, it is often difficult to effectively remove noise, special characters, and stop words. Traditional BIO tagging methods have obvious deficiencies when processing long entities, and it is difficult to accurately mark the boundaries of long entities, resulting in reduced accuracy in entity recognition. In terms of word vector representation, existing technologies mostly rely on traditional word embedding methods, such as Word2Vec or GloVe. Although these methods can capture the basic semantics of words, they have defects in deep semantic understanding and cannot fully consider the impact of contextual information on word meaning, resulting in poor performance of the model when understanding complex semantics and polysemous words. The traditional long short-term memory network Bi-LSTM treats all time steps equally, without considering the differences in importance of different time steps, resulting in the model having difficulty effectively distinguishing key information from secondary information when processing complex texts.

[0004] In summary, existing knowledge graph construction methods focus on entity recognition and relationship extraction, but they are insufficient in the expressiveness and completeness of knowledge graphs. Summary of the invention

[0005] In order to solve the technical problems existing in the construction of existing knowledge graphs, such as being unfavorable for calculation and being unable to make explanations, the present invention provides a knowledge graph construction method based on a long short-term memory network introducing an attention mechanism, which is helpful to improve the expressiveness of the knowledge graph and the completeness of the knowledge graph, and solve the technical problems that the current unstructured data is difficult to process and the BIO tag is not applicable to long entity features. The present invention uses an attention mechanism in LSTM to improve the prediction accuracy of the model, and the self-attention mechanism can help the model pay attention to different positions in the input data. In LSTM, a self-attention mechanism can be introduced at each time step to calculate the similarity between the current time step and the previous time step, thereby obtaining a weighted context vector, which can be summed with the input of the current time step to obtain a new input vector. The knowledge graph can be described in a structured manner. After adding the attention mechanism, the attention weight of each neighbor can be used to explain the prediction results of the model.

[0006] The technical solution adopted by the present invention is:

[0007] A knowledge graph construction method based on a long short-term memory network introducing an attention mechanism is characterized in that the construction method specifically includes the following steps:

[0008] S1: Sentence and word segmentation preprocessing of the dataset, including removing noise, special characters and stop words;

[0009] S2: Label the segmented data with BIOES tags and introduce hierarchical relationships;

[0010] S3: Build word vector representation and introduce BERT model to capture context information and provide deep semantic understanding for word vectors;

[0011] S4: Establish a Bi-LSTM neural network layer, and use the Bi-LSTM neural network layer to obtain the score probability of each word corresponding to each label, and add an attention mechanism after the Bi-LSTM layer to capture long-distance dependencies;

[0012] S5: Establish the CRF-GNN layer to obtain the output annotation sequence with the maximum probability;

[0013] S6: Post-process the extracted information.

[0014] Furthermore, in step S1, sentence and word segmentation preprocessing of the data set specifically includes: removing irrelevant symbols, numbers, URLs, and email addresses through regular expressions; using NLTK and pySBD libraries to perform text sentence segmentation; using jieba to perform Chinese word segmentation; and filtering special characters and stop words to reduce the computational burden.

[0015] Furthermore, in the step S2, the segmented data is annotated with BIOES tags, and a hierarchical relationship is introduced, specifically including: splitting the long text of the same attribute in the segmented data into sentences and clauses, and annotating them; using the BIOES annotation method, and adding a hierarchical relationship to the annotation.

[0016] Furthermore, in step S3, word vector representation is established, and the BERT model is introduced to capture context information, so as to provide deep semantic understanding for word vectors. Specifically, the following steps are performed:

[0017] S31: Building a multilingual dictionary;

[0018] The dictionary covers all characters that appear in the dataset. This multilingual dictionary will serve as the basis for character-to-digital index mapping in subsequent steps.

[0019] S32: pre-trained word embedding, encoding word vector;

[0020] The numeric index of each word is read through the input layer and mapped into a word vector using the word2vec CBOW model in the lookup layer;

[0021] S33: training vector matrix, domain knowledge graph integration;

[0022] The Word2Vec model is used to fine-tune word vectors to make them more suitable for the semantic and grammatical features of specific fields; the trained word vectors are combined with the domain knowledge graph to enhance entity recognition.

[0023] Furthermore, the step S4 specifically includes:

[0024] S41: input sequence processing;

[0025] Take a sequence of sentences x {×1, x2, …, xn} as input at each time step, where each xi represents a word or token in the sequence;

[0026] S42: Forward and Backward Propagation of Bi-LSTM Layers

[0027] Use forward LSTM to process the sequence and obtain the hidden layer output sequence (h1, h2, …, hn);

[0028] Use reverse LSTM to process the sequence and obtain another hidden layer output sequence (h′1, h′2, …, h′n);

[0029] S43: Output fusion of Bi-LSTM layer

[0030] Concatenate the outputs of the forward and reverse LSTM to get:

[0031]

[0032] in, represents the concatenation operation, j represents the time step, 1≤j≤n; S44: introduces the attention mechanism

[0033] Calculate the attention weight oαj at each time step, which can be done by performing a dot product between a trainable weight vector W and the context vector v and the output of the Bi-LSTM, and then generating the weight through a softmax function;

[0034] S45: Weighted sum and context vector

[0035] The output vectors of all time steps are weighted and summed according to the calculated attention weights to obtain a comprehensive context vector C, as shown below:

[0036]

[0037] This context vector C will be used in subsequent classification or regression tasks to improve the model's ability to capture long-distance dependencies;

[0038] S46: Calculation of attention weights

[0039] The calculation of the attention weight αj can be achieved by the following formula:

[0040]

[0041] Among them, e j is the attention score, which can be calculated as follows:

[0042]

[0043] Where v is the context vector and W is the trainable weight matrix;

[0044] S47: Score calculation of output layer

[0045] Using the context vector C and the attention-weighted Bi-LSTM output, a fully connected layer is used to calculate the score probability of each word corresponding to each label.

[0046] Furthermore, the step S5 specifically includes:

[0047] S51: Output processing of CRF layer

[0048] On the basis of the Bi-LSTM layer, the score probability of each word corresponding to each label is passed to the CRF layer to consider the transition probability between labels; S52: Construction of graph structure

[0049] Based on the output of the CRF layer, a graph structure is constructed, where nodes represent words in the text and edges represent potential dependencies between words;

[0050] S53: Introducing Graph Neural Networks (GNN)

[0051] A GNN layer is introduced after the CRF layer to further process sequence data; the GNN layer can capture the complex relationships between entities and incorporate these relationships into the final labeled sequence;

[0052] S54: Training of GNN layers

[0053] Train the GNN layer to learn the representation of node features and graph structure through supervised learning;

[0054] S55: Fusion of CRF and GNN outputs

[0055] The output of the CRF layer is combined with the graph structure features learned by the GNN layer to enhance the model’s ability to judge sequence annotations.

[0056] S56: Decoding of the maximum probability sequence

[0057] Use the Viterbi algorithm or the Beam Search algorithm to decode the output label sequence with the highest probability from the fused output.

[0058] Furthermore, the step S6 specifically includes: performing unstructured information extraction according to the output annotation sequence with the maximum probability; formatting and integrating the extracted information to facilitate the construction of a knowledge graph.

[0059] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0060] 1. The present invention can solve the problem of chaotic information extraction output logic by fusing the Bi-LSTM and CRF models and giving corresponding constraints to the label sequence.

[0061] 2. In view of the problem that traditional BIO tag annotation is not applicable to long entity features, the present invention proposes a BIOES annotation method, which adds a hierarchical relationship to the annotation to improve the information extraction effect of long texts, and solves the current technical problems of difficult processing of unstructured data, single-word entity annotation, and unclear entity boundary clarity. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 This is a flowchart of constructing a Bi-LSTM knowledge graph based on the attention mechanism of the present invention.

[0063] Figure 2 It is the flow chart of word vector representation of the present invention.

[0064] Figure 3 It is a schematic diagram of the integrated knowledge graph in the field of the present invention.

[0065] Figure 4 It is a schematic flow chart of step S4 of the present invention. DETAILED DESCRIPTION

[0066] The specific implementation of the embodiment of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the embodiment of the present invention, and is not used to limit the embodiment of the present invention.

[0067] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0068] The present invention will be described in detail below with reference to the accompanying drawings and in combination with exemplary embodiments. This embodiment relates to a patient information extraction method using a knowledge graph construction method based on Bi-LSTM of the present invention. Taking the extraction of information from medical records and the construction of a knowledge graph as an example, it is assumed that we have a medical record text from which we need to extract the patient's personal information, symptoms, diagnosis results and treatment recommendations, and construct a knowledge graph. We will use the technologies mentioned in the present invention (Bi-LSTM, CRF, GNN, etc.) to complete this task.

[0069] Input Data

[0070] The following is an example of a simplified medical record text:

[0071] Patient Zhang San, male, 35 years old, complained of chest tightness, shortness of breath, and intermittent palpitations for two weeks. After preliminary examination, he was diagnosed with coronary heart disease and was advised to take aspirin.

[0072] A patient information extraction method based on a knowledge graph construction method of a long short-term memory network Bi-LSTM introducing an attention mechanism is applied to the present invention, such as Figure 1 As shown, the specific steps include:

[0073] Step S1: Preprocess the data set by sentence and word segmentation

[0074] Noise removal: Remove irrelevant symbols, numbers, URLs, etc.

[0075] Sentence and word segmentation: divide the text into sentences and words.

[0076] Filter stop words and special characters: Remove common stop words and punctuation marks.

[0077] Processing results:

[0078] ·Input text: Patient Zhang San, male, 35 years old, complained of chest tightness, shortness of breath, and intermittent palpitations, which have lasted for two weeks. After preliminary examination, he was diagnosed with coronary heart disease and was recommended to take aspirin. ·Word segmentation results: ['patient', 'Zhang San', 'male', '35 years old', 'main complaint', 'chest tightness', ', 'shortness of breath', ', ', 'accompanied by', 'intermittent', 'palpitations', ', ', 'has lasted', 'two weeks', '. ', 'After', 'preliminary', 'examination', ', ', 'diagnosed as', 'coronary heart disease', ', ', 'recommended', 'to take', 'aspirin', '. ']

[0079] Step S2: Label the segmented data with BIOES tags and introduce hierarchical relationships

[0080] 1. Split the long text of the same attribute in the word segmented data into sentences and clauses, and mark them.

[0081] 2. Use the BIOES annotation method to clearly mark the start (B), inside (1), end (E) and single word entities (S) of entities.

[0082] 3. Introduce hierarchical relationships in annotations, such as self-reference tables, closure tables, path enumerations, etc., to handle complex entity relationships.

[0083] Actual Cases

[0084] Original text excerpt: "Patient Zhang San, male, 35 years old, complained of chest tightness, shortness of breath, and intermittent palpitations, which have lasted for two weeks. After preliminary examination, he was diagnosed with coronary heart disease and was advised to take aspirin."

[0085] Text fragment after word segmentation: ["patient","Zhang San",","male",","35 years old",","chief complaint","chest tightness","shortness of breath","accompanied by","intermittent","palpitations","has lasted for","two weeks",".","after","preliminary","examination","diagnosed as","coronary heart disease",","recommendation","take","aspirin","."]

[0086] Annotated text snippet:

[0087] [“Patient” (B-Patient), “Zhang San” (E-Patient), “,” (O), “Male” (B-Gender), “,” (O), “35 years old” (B-Age), “,” (O), “Child complaint” (O), “Chest tightness” (B-SympTom), “,” (O), “Shortness of breath” (I-Symptom), “,” (O), “Accompanied by” (O), “Intermittent” (O), “Palpitations” (E-Symptom), “,” (O), “Continued” (O), “Two weeks” (B-Time), “.” (O), “After” (O), “Preliminary” (O), “Examination” (O), “,” (O), “Diagnosed as” (O), “Coronary heart disease” (B-Disease), “,” (O), “Recommendation” (O), “Take” (O), “Aspirin” (S-Drug), “.” (O)]

[0088] · Annotation sequence: [

[0089] ·'B-Patient', 'E-Patient', 'B-Gender', 'B-Age', 'O', 'B-Symptom', 'O', 'I-Symptom', 'O', 'O', 'O ′, ′E-SympTom′, ′O′, ′O′, ′B-Time′, ′O′, ′O′, ′O′, ′O′, ′O′, ′B-Disease′, ′O′, ′O′, ′O′, ′S-Drug′, ′O′

[0090] ·]

[0091] Step S3: Create word vector representation

[0092] 1. Build a multilingual dictionary covering all characters appearing in the dataset, providing a basis for character-to-digital index mapping.

[0093] 2. Read the numeric index of each word through the input layer and map it to a word vector using the word2vec CBOW model.

[0094] 3. Fine-tune the word vectors through the gradient descent algorithm to make them more in line with the semantic and grammatical features of specific fields, and combine the trained word vectors with the domain knowledge graph to enhance entity recognition capabilities.

[0095] Actual Cases

[0096] Suppose we are processing text data in the medical field, and use the Word2Vec model to fine-tune the word vector to make it more suitable for the semantic and grammatical features of the medical field. For example: Input text: 'Patient Zhang San, male, 35 years old, complained of chest tightness and shortness of breath'

[0097] After word segmentation: ['patient', 'Zhang San', ', ', 'male', ', ', '35 years old', ', ', 'chief complaint', 'chest tightness', ', ', '

[0098] Shortness of breath

[0099] Numeric index: [1, 2, 3, 4, 3, 5, 3, 6, 7, 8, 9]

[0100] Word vector (before fine-tuning):

[0101]

[0102] Word vector (after fine-tuning):

[0103]

[0104] The fine-tuned word vectors are combined with the domain knowledge graph to enhance entity recognition capabilities. For example, word vectors such as "patient", "chest tightness", and "shortness of breath" are matched with entities in the medical knowledge graph to improve the accuracy of entity recognition.

[0105] Suppose we have a medical knowledge graph that contains the following entities and relations:

[0106]

[0107]

[0108] Match the fine-tuned word vectors with entities in the knowledge graph to enhance entity recognition capabilities.

[0109] Matching results:

[0110]

[0111] Step S4: Establish a Bi-LSTM neural network layer, and use the Bi-LSTM neural network layer to obtain the score probability of each word corresponding to each label, and add an attention mechanism after the Bi-LSTM layer to capture long-distance dependencies

[0112] 1. Take the sentence sequence {×1, x2, …, xn} as the input of each time step.

[0113] Actual case:

[0114] These word vectors are fed into the Bi-LSTM model as input sequences:

[0115]

[0116]

[0117] 2. Use forward LSTM and backward LSTM to process the sequence and obtain the hidden layer output sequence.

[0118] Actual Cases

[0119] Suppose we use a Bi-LSTM model with 128 hidden units in each LSTM layer. The forward LSTM and the backward LSTM process the input sequence respectively and obtain the hidden layer output sequence:

[0120] The forward LSTM hidden layer outputs the sequence (h1, h2, ..., hn):

[0121]

[0122]

[0123] The reverse LSTM hidden layer outputs a sequence (h′1, h′2, ..., h′n):

[0124]

[0125] 3. Concatenate the outputs of the forward and reverse LSTM to get the output of the Bi-LSTM.

[0126] Actual Cases

[0127] Concatenate the outputs of the forward and reverse LSTM to get the output vector for each time step:

[0128] Bi-LSTM Output [

[0130] [0.1, 0.2, ..., 0.128, 0.136, 0.135, ..., 0.128], # patients

[0131] [0.2, 0.3, ..., 0.129, 0.135, 0.134, ..., 0.127], # Zhang San

[0132] [0.3, 0.4, ..., 0.130, 0.134, 0.133, ..., 0.126], #,

[0133] …

[0134] [0.9, 0.10, ..., 0.136, 0.1, 0.2, ..., 0.128] #. ]

[0136] 4. Introduce the attention mechanism, calculate the attention weight of each time step, and generate the context vector.

[0137] Actual Cases

[0138] Calculate the attention weight aj for each time step, perform a dot product with the output of Bi-LSTM through a trainable weight vector W and the context vector v, and then generate the weight through the softmax function:

[0139] Attention weights α = [α1, α2, ..., αn]:

[0140]

[0141]

[0142] Generate the context vector C, and sum the output vectors of all time steps according to the calculated attention weights:

[0143] Context vector C = α1*P1+α2*P2+...+αn*Pn: [

[0145] 0.15, 0.25, ..., 0.132 # Comprehensive context vector ]

[0147] 5. Using the context vector and the attention-weighted Bi-LSTM output, the score probability of each word corresponding to each tag is calculated through a fully connected layer.

[0148] Actual Cases

[0149] Suppose we have 5 tags (B-Patient, E-Patient, B-Symptom, E-Symptom, O), and calculate the score probability of each word corresponding to each tag through the fully connected layer:

[0150] Score probability matrix S = [s1, s2, ..., sn]: [

[0152] [0.8, 0.1, 0.05, 0.03, 0.02], #patient (B-Patient)

[0153] [0.01, 0.9, 0.02, 0.05, 0.02], # Zhang San (E-Patient)

[0154] [0.05, 0.05, 0.05, 0.05, 0.8], #, (O)

[0155] …

[0156] [0.05, 0.05, 0.05, 0.05, 0.8]#. (O) ]

[0158] Step S5: Establish the CRF-GNN layer to obtain the output annotation sequence with the maximum probability

[0159] 1. Pass the output of the Bi-LSTM layer to the CRF layer, considering the transition probability between labels.

[0160] 2. Build a graph structure based on the output of the CRF layer, where nodes represent words and edges represent potential dependencies.

[0161] 3. The GNN layer is introduced after the CRF layer to capture the complex relationships between entities.

[0162] Actual Cases

[0163] Use the GNN layer to process the graph structure and learn the representation of node features and graph structure. For example, the GNN layer may update the feature vector of each node, considering the influence of its neighboring nodes:

[0164] Updated node feature vector:

[0165]

[0166]

[0167] 4. Train the GNN layer to learn the representation of node features and graph structure.

[0168] 5. Fusion the output of the CRF layer with the graph structure features learned by the GNN layer.

[0169] Actual Cases

[0170] Use the labeled data to train the GNN layer and optimize the model parameters. Assuming we have a set of labeled data, we calculate the loss through the cross entropy loss function and perform back propagation to update the model parameters:

[0171] Annotated data:

[0172]

[0173]

[0174] Loss function: cross entropy loss

[0175] Optimizer: Adam

[0176] Final annotation sequence Y_final:

[0177]

[0178]

[0179] 6. Use the Viterbi algorithm or the Beam Search algorithm to decode the output label sequence with the highest probability from the fused output.

[0180] Example

[0181] Time step t: 0, 1, 2, 3

[0182] Tag set tag_set: [B-Patient, E-Patient, O, B-Symptom, I-Symptom, E-Symptom]

[0183] Score matrix S:

[0184]

[0185]

[0186] Transition probability matrix transition_matrix:

[0187]

[0188] Diagram:

[0189] copy

[0190]

[0191] Final path: B-Patient->E-Patient->O->E-Symptom

[0192] Final annotation sequence:

[0193]

[0194] Step S6: post-process the extracted information.

[0195] 1. Perform unstructured information extraction based on the output annotation sequence with the highest probability.

[0196] 2. Format and integrate the extracted information to facilitate the construction of the knowledge graph.

[0197] Step S7: Apply the constructed knowledge graph to extract patient information.

[0198] Actual example: Integrate the extracted entities and relations into triples of the knowledge graph:

[0199]

[0200]

[0201] These triplets can be further integrated into the structure of the final medical knowledge graph, for example:

[0202]

Claims

1. A knowledge graph construction method based on a long short-term memory network with an attention mechanism, characterized in that: The construction method specifically includes the following steps: S1: Preprocess the data set by sentence and word segmentation, including removing noise, special characters and stop words; S2: Label the segmented data with BIOES tags and introduce hierarchical relationships; S3: Build word vector representation and introduce BERT model to capture context information and provide deep semantic understanding for word vectors; S4: Establish a Bi-LSTM neural network layer, and use the Bi-LSTM neural network layer to obtain the score probability of each word corresponding to each label, and add an attention mechanism after the Bi-LSTM layer to capture long-distance dependencies; S5: Establish the CRF-GNN layer to obtain the output annotation sequence with the maximum probability; S6: Post-process the extracted information.

2. The method for constructing a knowledge graph based on a long short-term memory network with an attention mechanism as claimed in claim 1, characterized in that: In step S1, the sentence and word segmentation preprocessing of the data set specifically includes: removing irrelevant symbols, numbers, URLs, and email addresses through regular expressions; using NLTK and pySBD libraries to perform text sentence segmentation; using jieba to perform Chinese word segmentation; and filtering special characters and stop words to reduce the computational burden.

3. The method for constructing a knowledge graph based on a long short-term memory network with an attention mechanism as claimed in claim 1, characterized in that: In the step S2, the segmented data is annotated with BIOES tags, and a hierarchical relationship is introduced, which specifically includes: splitting the long text of the same attribute in the segmented data into sentences and clauses, and annotating them; using the BIOES annotation method, and adding a hierarchical relationship to the annotation.

4. The method for constructing a knowledge graph based on a long short-term memory network with an attention mechanism as claimed in claim 1, characterized in that: In step S3, word vector representation is established, and the BERT model is introduced to capture context information to provide deep semantic understanding for word vectors. Specifically, the following steps are performed: S31: Building a multilingual dictionary; The dictionary covers all characters that appear in the dataset. This multilingual dictionary will serve as the basis for character-to-digital index mapping in subsequent steps. S32: pre-trained word embedding, encoding word vector; The numeric index of each word is read through the input layer and mapped into a word vector using the word2vec CBOW model in the lookup layer; S33: training vector matrix, domain knowledge graph integration; The Word2Vec model is used to fine-tune word vectors to make them more suitable for the semantic and grammatical features of specific fields. The trained word vectors are combined with the domain knowledge graph to enhance entity recognition.

5. The method for constructing a knowledge graph based on a long short-term memory network with an attention mechanism as claimed in claim 1, characterized in that: The step S4 specifically includes: S41: input sequence processing; Take a sequence of sentences X {x1,x2,…,xn} as input for each time step, where each xi represents a word or token in the sequence; S42: Forward and Backward Propagation of Bi-LSTM Layers Use forward LSTM to process the sequence and obtain the hidden layer output sequence (h1,h2,…,hn); Use reverse LSTM to process the sequence and obtain another hidden layer output sequence (h'1,h'2,…,h'n); S43: Output fusion of Bi-LSTM layer Concatenate the outputs of the forward and reverse LSTM to get: in, represents the concatenation operation, j represents the time step, 1≤j≤n; S44: Introduce the attention mechanism and calculate the attention weight αj of each time step. This can be done by performing a dot product between a trainable weight vector W and the context vector v and the output of Bi-LSTM, and then generating the weight through the softmax function; S45: Weighted sum and context vector The output vectors of all time steps are weighted and summed according to the calculated attention weights to obtain a comprehensive context vector C, as shown below: This context vector C will be used in subsequent classification or regression tasks to improve the model's ability to capture long-distance dependencies; S46: Calculation of attention weights The calculation of attention weight αj can be achieved by the following formula: Among them, e j is the attention score, which can be calculated as follows: Among them, v is the context vector and W is the trainable weight matrix; S47: Score calculation of output layer Using the context vector C and the attention-weighted Bi-LSTM output, a fully connected layer is used to calculate the score probability of each word corresponding to each label.

6. The method for constructing a knowledge graph based on a long short-term memory network with an attention mechanism as claimed in claim 1, characterized in that: The step S5 specifically includes: S51: Output processing of CRF layer Based on the Bi-LSTM layer, the score probability of each word corresponding to each label is passed to the CRF layer to consider the transition probability between labels; S52: Construction of graph structure Based on the output of the CRF layer, a graph structure is constructed, in which nodes represent words in the text and edges represent potential dependencies between words; S53: Introducing Graph Neural Networks (GNN) A GNN layer is introduced after the CRF layer to further process sequence data; the GNN layer can capture the complex relationships between entities and incorporate these relationships into the final labeled sequence; S54: Training of GNN layers Train the GNN layer to learn the representation of node features and graph structure through supervised learning; S55: Fusion of CRF and GNN outputs The output of the CRF layer is combined with the graph structure features learned by the GNN layer to enhance the model’s ability to judge sequence annotations. S56: Decoding of the maximum probability sequence Use the Viterbi algorithm or the Beam Search algorithm to decode the output label sequence with the highest probability from the fused output.

7. The method for constructing a knowledge graph based on a long short-term memory network with an attention mechanism as claimed in claim 1, characterized in that: The step S6 specifically includes: performing unstructured information extraction according to the output annotation sequence with the maximum probability; and formatting and integrating the extracted information to facilitate the construction of a knowledge graph.

Citation Information

Cited By

  • Data visualization method and system based on artificial intelligence

    CN120852561A

  • Method and device for querying path based on associated knowledge graph, medium and equipment

    CN120873171A