Dynamic fusion word embedding and graph structure information query method and system based on attention mechanism

By introducing a dynamic fusion word embedding and graph structure information query method with attention mechanism in natural language processing technology, the problems of in-depth semantic understanding and inaccurate search results in the existing technology are solved, and more efficient and accurate natural language query is achieved.

CN119961428APending Publication Date: 2025-05-09TSINGHUA UNIVERSITY +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411729152.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing natural language processing technology has problems such as in-depth semantic understanding, inaccurate search results, and ineffective use of graph structure information.

Method used

A dynamic fusion word embedding and graph structure information query method based on attention mechanism is adopted, and a word embedding vector database is constructed by obtaining knowledge graphs and performing vector representations; a pre-trained large language model is used to convert natural language queries into a query language of knowledge graphs; entities and relationships are extracted and mapped to word embedding vector database; entity and relationship information are fused through attention mechanism, alternative languages ​​are generated and statement queries are performed.

Benefits of technology

It enhances the semantic recognition and information retrieval capabilities of artificial intelligence systems when processing complex queries, realizes semantic understanding and accurate retrieval of natural language queries, and improves the efficiency and accuracy of queries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961428A_ABST
    Figure CN119961428A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic fusion word embedding and graph structure information query method and system based on an attention mechanism, and the method comprises the steps: obtaining a knowledge graph, carrying out the vectorization expression of the knowledge graph, and constructing a word embedding vector database; converting the pre-acquired natural language query of the user into a query language of the knowledge graph through a pre-trained large language model; extracting entities and relationships in a query language, and mapping the entities and the relationships to the word embedding vector database; fusing and reconstructing entities and relationships mapped to the word embedding vector database through an attention mechanism, and replacing the query language to generate a replacement language; and performing statement query through a built-in search engine of the knowledge graph based on the replacement language to generate a query result. The method solves the problems that existing natural language processing does not deeply understand semantics, retrieval results are inaccurate, and graph structure information cannot be effectively utilized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a method and system for dynamically fusion word embedding and graph structure information query based on an attention mechanism. Background Art

[0002] With the rapid development of artificial intelligence technology, especially in the field of natural language processing, the combination of knowledge graphs and large models has become the key to improving semantic understanding capabilities. As a structured knowledge representation method, knowledge graphs are widely used in information retrieval, recommendation systems and other fields. However, traditional knowledge graph query methods often rely on keyword matching or fixed query languages, lacking a deep understanding of natural language queries. In addition, existing technologies rarely consider the structural position and neighbor information of entities in the knowledge graph when processing queries, resulting in the inability to fully utilize the semantic richness of the graph. Summary of the invention

[0003] The present invention provides a method and system for querying graph structure information by dynamically fusion of word embedding and graph structure information based on an attention mechanism, so as to solve the problems that the existing natural language processing has a shallow understanding of semantics, inaccurate retrieval results and cannot effectively utilize graph structure information.

[0004] The present invention provides a method for querying dynamic fusion word embedding and graph structure information based on an attention mechanism, comprising: Obtain the knowledge graph and vectorize it to build a word embedding vector database; The pre-acquired natural language queries of users are converted into the query language of the knowledge graph through the pre-trained large language model; Extracting entities and relations in the query language, and mapping the entities and relations to the word embedding vector database; Reconstructing entities and relations mapped to the word embedding vector database through attention fusion, replacing the query language, and generating a replacement language; Based on the replacement language, a statement query is performed through a search engine built into the knowledge graph to generate a query result.

[0005] According to a method for querying word embedding and graph structure information based on a dynamic fusion of an attention mechanism, the present invention provides a method for obtaining a knowledge graph and vectorizing the knowledge graph to construct a word embedding vector database, which specifically includes: Select a base point from the knowledge graph to traverse the entire knowledge graph to obtain entities and relationships in the knowledge graph; Use word embedding technology to map entities and relationships in the knowledge graph into low-dimensional dense vectors; Insert the low-dimensional dense vector into the preset vector database, assign a unique ID to each low-dimensional dense vector, and build a word embedding vector database.

[0006] According to a dynamic fusion word embedding and graph structure information query method based on attention mechanism provided by the present invention, the training process of the large language model is: Obtain a base model and a dataset containing natural query input and corresponding query language; Fine-tune the base model using the fine-tuned data set to learn how to generate correct query language based on input; The cross entropy loss function is used to measure the gap between the output of the base model and the actual query language, and the parameters of the base model are updated through back propagation and optimizer to obtain the trained large language model.

[0007] According to a method for querying word embedding and graph structure information based on a dynamic fusion of an attention mechanism, the method extracts entities and relationships in a query language and maps the entities and relationships to the word embedding vector database, including: Obtain entities and relations in the query language, and perform word embedding on the entities and relations in the query language; Entities and relationships are mapped to the word embedding vector database to make the word embedding space of the query language consistent with the word embedding space of the knowledge graph.

[0008] According to a method for querying word embedding and graph structure information based on a dynamic fusion of attention mechanism provided by the present invention, the entities and relationships mapped to the word embedding vector database are reconstructed by fusing the attention mechanism, the query language is replaced, and the replacement language is generated, which specifically includes: Query the word embedding vector database for the vector with the highest similarity to the entities and relations in the query language, and extract the neighbor information in the knowledge graph for each entity in the query language; Performing a dot product operation on the entity word embedding in the query language and the word embedding of the neighbor entity and relationship in the neighbor information through the attention mechanism to obtain the attention weight and normalize the attention weight to obtain the attention distribution; The entity words in the query language are embedded in the weighted neighbor information and aggregated to obtain the entity representation that integrates the graph structure; Based on the text represented by the entity representation in the fused query statement in the knowledge graph, the original query language is replaced to generate a replacement language.

[0009] According to a method for querying dynamic fusion word embedding and graph structure information based on an attention mechanism provided by the present invention, the sentence query is performed through a search engine built into the knowledge graph based on the replacement language to generate query results, specifically including: The knowledge graph is saved in the form of a graph database, and the search engine in the graph database is used to perform the search of query statements, generate query results, and feed back the query results to the user who inputs the natural language query.

[0010] The present invention also provides a dynamic fusion word embedding and graph structure information query system based on attention mechanism, the system comprising: The word embedding vector database construction module is used to obtain the knowledge graph and vectorize the knowledge graph to build a word embedding vector database; A language conversion module is used to convert the pre-acquired natural language queries of users into the query language of the knowledge graph through a pre-trained large language model; A mapping module, used to extract entities and relations in the query language, and map the entities and relations to the word embedding vector database; A language reconstruction module, used for fusing and reconstructing entities and relations mapped to the word embedding vector database through an attention mechanism, replacing the query language, and generating a replacement language; The query feedback module is used to perform statement query based on the replacement language through the search engine built into the knowledge graph to generate query results.

[0011] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for dynamically fusion word embedding and graph structure information query based on the attention mechanism as described above is implemented.

[0012] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for dynamically fusion word embedding and graph structure information query based on the attention mechanism.

[0013] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned methods for dynamically fusion word embedding and graph structure information query based on the attention mechanism.

[0014] The present invention provides a method and system for querying word embedding and graph structure information by dynamic fusion based on attention mechanism. By introducing the attention mechanism, the dynamic fusion of word embedding and graph structure information is realized, and the semantic recognition and information retrieval capabilities of the artificial intelligence system when processing complex queries are enhanced. The advantages of knowledge graphs, word embedding, large language models and attention mechanisms are combined to realize semantic understanding and precise retrieval of natural language queries, and the efficiency and accuracy of queries are improved. The method can not only enhance the intelligence level of search engines, but also provide more precise semantic services for intelligent assistants, personalized recommendation systems, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0016] Figure 1 It is a flow chart of the method for dynamically fusion word embedding and graph structure information query based on attention mechanism provided by the present invention.

[0017] Figure 2 It is a schematic diagram of the module connection of the dynamic fusion word embedding and graph structure information query system based on the attention mechanism provided by the present invention.

[0018] Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention.

[0019] Figure numerals: 110: word embedding vector database construction module; 120: language conversion module; 130: mapping module; 140: language reconstruction module; 150: query feedback module; 310: processor; 320: communication interface; 330: memory; 340: communication bus. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0021] Combine the following Figure 1The present invention describes a method for querying word embedding and graph structure information by dynamic fusion based on attention mechanism, including: step 100, obtaining a knowledge graph and vectorizing the knowledge graph to construct a word embedding vector database.

[0022] Specifically, a base point is selected from the knowledge graph to traverse the entire knowledge graph to obtain entities and relationships in the knowledge graph; Use word embedding technology to map entities and relationships in the knowledge graph into low-dimensional dense vectors; Insert the low-dimensional dense vector into the preset vector database, assign a unique ID to each low-dimensional dense vector, and build a word embedding vector database.

[0023] As a structured knowledge representation method, knowledge graphs are widely used in information retrieval and recommendation systems. However, traditional knowledge graph query methods often rely on keyword matching or fixed query languages, lacking a deep understanding of natural language queries. In addition, the existing technology rarely considers the structural position and neighbor information of entities in the knowledge graph when processing queries, resulting in the inability to fully utilize the semantic richness of the graph. The present invention dynamically aggregates word embedding and graph structure information by introducing an attention mechanism.

[0024] In the present invention, all entities, their attributes and relationships are extracted from the knowledge graph, and all entities, their attributes and relationships are obtained by selecting a node and traversing the entire knowledge graph; all the above entities, attributes and relationships are mapped into low-dimensional dense vectors using word embedding technology (Word2Vec, GloVe, FastText or BERT); a vector database of word embedding is created. Select a vector database (such as Faiss, Milvus, Annoy or Elastic search), configure and create a vector database instance according to the relevant documents. Insert the vectors of entities, attributes and relationships into the database. When inserting vectors, a unique ID needs to be specified for each vector. This ID can be the name of an entity, attribute or relationship, or it can be a unique digital ID.

[0025] Step 200: Convert the pre-acquired natural language query of the user into the query language of the knowledge graph through the pre-trained large language model. In the present invention, Cypher QL is used to refer to the query language of the knowledge graph.

[0026] In the present invention, the large language model training process includes: obtaining a base model and a data set including a natural query sentence input and a corresponding query language; Fine-tune the base model using the fine-tuned data set to learn how to generate correct query language based on input; The cross entropy loss function is used to measure the gap between the output of the base model and the actual query language, and the parameters of the base model are updated through back propagation and optimizer to obtain the trained large language model.

[0027] Specifically, choose a pre-trained base model, such as GLM4, GPT3, T5, and other open source models.

[0028] Prepare a dataset containing natural query input and corresponding Cypher QL statements. For example, if the input is "find all actors born in New York", the corresponding Cypher QL statement may be "MATCH (a:Actor) WHEREa.BirthPlace = 'New York' RETURN a".

[0029] Set the fine-tuning parameters, including learning rate, batch size, and number of training rounds. These parameters need to be adjusted according to the aforementioned dataset size and computing resources. Use the constructed dataset to train and fine-tune the base model. During the fine-tuning process, the base model will learn how to generate correct Cypher QL statements based on the input. Use the cross entropy loss function to measure the gap between the model's output and the actual Cypher QL statement, and update the model's parameters through backpropagation and optimizer.

[0030] During the training process, it is necessary to perform prompt engineering on the Cypher QL statement generation task of the large language model; design an initial prompt to explain the definition of the Cypher QL statement generation task to the large model; construct a thinking chain for the Cypher QL statement generation task, break the generation task into several steps, and guide the large model to reason according to the thinking chain. First, let the large model extract all the important entities and their attributes in the user's query, then extract the relationship between the entities, and finally guide the large model to use Cypher QL statements to represent the aforementioned entities and their attributes and relationships; design a few-shot context in combination with the above-mentioned thinking chain, and add 10-100 standard answers and thinking chain processes of "input-Cypher QL statement" conversion to the prompt of the large model to activate the context learning ability of the model during reasoning.

[0031] To evaluate the performance of a fine-tuned large language model, run the model on a separate test set and calculate some evaluation metric such as accuracy or F1-Score.

[0032] Step 300: extract entities and relations in the query language, and map the entities and relations to the word embedding vector database.

[0033] Specifically, entities and relations in the query language are obtained, and word embedding is performed on the entities and relations in the query language; Entities and relationships are mapped to the word embedding vector database to make the word embedding space of the query language consistent with the word embedding space of the knowledge graph.

[0034] Extract the entities, their attributes, and relationships in the Cypher QL statements generated by the large model. Since Cypher QL statements are relatively structured languages, they are implemented through rule-based matching methods. The extracted entities, their attributes, and relationships are embedded in words to ensure that the word embedding space of the user's query is consistent with the word embedding space of the knowledge graph.

[0035] Step 400: Reconstruct the entities and relationships mapped to the word embedding vector database through attention mechanism fusion, replace the query language, and generate a replacement language.

[0036] Specifically, the word embedding vector database is searched for the vector with the highest similarity to the entities and relations in the query language, and the neighbor information in the knowledge graph is extracted for each entity in the query language; Performing a dot product operation on the entity word embedding in the query language and the word embedding of the neighbor entity and relationship in the neighbor information through the attention mechanism to obtain the attention weight and normalize the attention weight to obtain the attention distribution; The entity words in the query language are embedded in the weighted neighbor information and aggregated to obtain the entity representation that integrates the graph structure; Based on the text represented by the entity representation in the fused query statement in the knowledge graph, the original query language is replaced to generate a replacement language.

[0037] In the present invention, the created vector database is queried for the vectors with the highest similarity to the entities and their attributes and relationships in the Cypher QL statements generated by the large model; for each entity in the query language, its neighbor information in the knowledge graph is extracted, that is, other entities and relationships directly connected to the entity.

[0038] The attention mechanism is used to adaptively assign weights to different neighbor entities and relations. Specifically, the entity word embedding is dot-producted with the word embedding of its neighbor entities and relations to obtain the attention weight. The weight is normalized by softmax to obtain the final attention distribution.

[0039] Aggregate entity word embeddings and weighted neighbor information to obtain entity representations that integrate graph structures. Fusion can be performed using weighted averaging or concatenation; use the text represented in the knowledge graph by the fused entity representation to replace the text in the Cypher QL statement generated by the large language model.

[0040] Step 500: Perform a statement query based on the replacement language through a search engine built into the knowledge graph to generate a query result.

[0041] In the present invention, the knowledge graph is saved in the form of a graph database, and the search engine in the graph database is used to perform the search of query statements, generate query results, and feed back the query results to the user who inputs the natural language query.

[0042] In another embodiment, a visual representation of the knowledge graph is used to allow users to query through an interactive graphical interface. Users can browse the knowledge graph by clicking on nodes and edges, and construct queries by selecting specific nodes and edges. The system converts the user's operations into corresponding query statements and returns the results. This method is intuitive and easy to use, but for complex queries, users may need to perform a lot of interactive operations.

[0043] In another embodiment, the meta information in the knowledge graph, such as entity type, relationship type, etc., is used to enhance the semantic understanding ability of the query. Specifically, when converting the user query into a Cypher QL statement, not only the keywords in the query are considered, but also the type information of these keywords. For example, if the user queries "Who directed Forrest Gump", the system will not only include the entity "Forrest Gump" in the Cypher QL statement, but also specify the type of this entity as "movie". This method can improve the accuracy of the query, especially for ambiguous queries.

[0044] The present invention maps the natural language text of entities, relationships and attributes in the knowledge graph into a low-dimensional dense vector representation, i.e., word embedding, to realize the vectorized representation of the knowledge graph. Then, the user's natural language query is converted into the query language of the knowledge graph, such as Cypher QL, using a pre-trained large language model. The system extracts the entities and relationships in the generated Cypher QL statement and maps them to the pre-built knowledge graph word embedding space. In order to better capture the semantic information in the query statement, the system introduces an attention mechanism to dynamically aggregate entity word embedding information and the structural information of the knowledge graph. Specifically, after identifying the entity and relationship from the query statement, not only the corresponding word embedding vector is retrieved, but also the neighbor information of the entity in the knowledge graph is considered, i.e., other entities and relationships directly connected to the entity. Using the attention mechanism, weights are adaptively assigned to different neighbor entities and relationships, highlighting the important information related to the query and suppressing noise. The entity word embedding is aggregated with the neighbor information to obtain an entity representation that integrates the graph structure. Finally, the most relevant entities and relationships in the knowledge graph are retrieved through vector similarity, and re-entered into the Cypher QL statement to perform a true semantic query on the knowledge graph and return the final result to the user. This method combines the advantages of knowledge graph, word embedding, large language model and attention mechanism, realizes semantic understanding and precise retrieval of natural language queries, and greatly improves the efficiency and accuracy of queries.

[0045] refer to Figure 2 The present invention also discloses a dynamic fusion word embedding and graph structure information query system based on attention mechanism, the system comprising: A word embedding vector database construction module 110 is used to obtain a knowledge graph and vectorize the knowledge graph to construct a word embedding vector database; A language conversion module 120, used to convert the pre-acquired natural language query of the user into the query language of the knowledge graph through a pre-trained large language model; A mapping module 130, for extracting entities and relations in the query language, and mapping the entities and relations to the word embedding vector database; A language reconstruction module 140 is used to reconstruct entities and relationships mapped to the word embedding vector database through attention mechanism fusion, replace the query language, and generate a replacement language; The query feedback module 150 is used to perform statement query based on the replacement language through the search engine built into the knowledge graph to generate query results.

[0046] The step of obtaining a knowledge graph and vectorizing the knowledge graph to construct a word embedding vector database specifically includes: Select a base point from the knowledge graph to traverse the entire knowledge graph to obtain entities and relationships in the knowledge graph; Use word embedding technology to map entities and relationships in the knowledge graph into low-dimensional dense vectors; Insert the low-dimensional dense vector into the preset vector database, assign a unique ID to each low-dimensional dense vector, and build a word embedding vector database.

[0047] The training process of the large language model is: Obtain a base model and a dataset containing natural query input and corresponding query language; Fine-tune the base model using the fine-tuned data set to learn how to generate correct query language based on input; The cross entropy loss function is used to measure the gap between the output of the base model and the actual query language, and the parameters of the base model are updated through back propagation and optimizer to obtain the trained large language model.

[0048] During the training process, it is necessary to perform prompt engineering on the Cypher QL statement generation task of the large language model; design an initial prompt to explain the definition of the Cypher QL statement generation task to the large model; construct a thinking chain for the Cypher QL statement generation task, break the generation task into several steps, and guide the large model to reason according to the thinking chain. First, let the large model extract all the important entities and their attributes in the user's query, then extract the relationship between the entities, and finally guide the large model to use Cypher QL statements to represent the aforementioned entities and their attributes and relationships; design a few-shot context in combination with the above-mentioned thinking chain, and add 10-100 standard answers and thinking chain processes of "input-Cypher QL statement" conversion to the prompt of the large model to activate the contextual learning ability of the model during reasoning.

[0049] Extracting entities and relations in the query language and mapping the entities and relations to the word embedding vector database, including: Obtain entities and relations in the query language, and perform word embedding on the entities and relations in the query language; Entities and relationships are mapped to the word embedding vector database to make the word embedding space of the query language consistent with the word embedding space of the knowledge graph.

[0050] Reconstructing entities and relationships mapped to the word embedding vector database through attention fusion, replacing the query language, and generating a replacement language, specifically including: Query the word embedding vector database for the vector with the highest similarity to the entities and relations in the query language, and extract the neighbor information in the knowledge graph for each entity in the query language; Performing a dot product operation on the entity word embedding in the query language and the word embedding of the neighbor entity and relationship in the neighbor information through the attention mechanism to obtain the attention weight and normalize the attention weight to obtain the attention distribution; The entity words in the query language are embedded in the weighted neighbor information and aggregated to obtain the entity representation that integrates the graph structure; Based on the text represented by the entity representation in the fused query statement in the knowledge graph, the original query language is replaced to generate a replacement language.

[0051] Based on the replacement language, a sentence query is performed through a search engine built into the knowledge graph to generate query results, specifically including: The knowledge graph is saved in the form of a graph database, and the search engine in the graph database is used to perform the search of query statements, generate query results, and feed back the query results to the user who inputs the natural language query.

[0052] A dynamic fusion word embedding and graph structure information query system based on an attention mechanism provided by the present invention realizes the dynamic fusion of word embedding and graph structure information by introducing the attention mechanism, thereby enhancing the semantic recognition and information retrieval capabilities of the artificial intelligence system when processing complex queries; combining the advantages of knowledge graphs, word embeddings, large language models and attention mechanisms, realizes semantic understanding and precise retrieval of natural language queries, and improves the efficiency and accuracy of queries; it can not only improve the intelligence level of search engines, but also provide more accurate semantic services for smart assistants, personalized recommendation systems, etc.

[0053] Figure 3 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 3As shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330 and a communication bus 340, wherein the processor 310, the communication interface 320 and the memory 330 communicate with each other through the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute a method for querying information based on a dynamic fusion of word embedding and graph structure information based on an attention mechanism, the method comprising: obtaining a knowledge graph and vectorizing the knowledge graph to construct a word embedding vector database; converting the pre-acquired user's natural language query into a query language of the knowledge graph through a pre-trained large language model; extracting entities and relations in the query language, and mapping the entities and relations to the word embedding vector database; fusing and reconstructing the entities and relations mapped to the word embedding vector database through an attention mechanism, replacing the query language, and generating a replacement language; performing a sentence query based on the replacement language through a search engine built into the knowledge graph to generate a query result.

[0054] In addition, the logic instructions in the above-mentioned memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0055] On the other hand, the present invention also provides a computer program product, which includes a computer program, and the computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a dynamic fusion word embedding and graph structure information query method based on an attention mechanism provided by the above methods, the method including: obtaining a knowledge graph and vectorizing the knowledge graph to construct a word embedding vector database; converting the pre-acquired user's natural language query into a query language of the knowledge graph through a pre-trained large language model; extracting entities and relationships in the query language, and mapping the entities and relationships to the word embedding vector database; fusing and reconstructing the entities and relationships mapped to the word embedding vector database through an attention mechanism, replacing the query language, and generating a replacement language; based on the replacement language, performing statement query through a search engine built into the knowledge graph to generate query results.

[0056] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute a method for querying word embedding and graph structure information based on a dynamic fusion of an attention mechanism provided by the above-mentioned methods, the method comprising: obtaining a knowledge graph and vectorizing the knowledge graph to construct a word embedding vector database; converting a pre-acquired natural language query of a user into a query language of the knowledge graph through a pre-trained large language model; extracting entities and relationships in the query language, and mapping the entities and relationships to the word embedding vector database; fusing and reconstructing the entities and relationships mapped to the word embedding vector database through an attention mechanism, replacing the query language, and generating a replacement language; performing statement query based on the replacement language through a search engine built into the knowledge graph to generate query results.

[0057] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0058] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0059] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A dynamic fusion word embedding and graph structure information query method based on attention mechanism, characterized in that: include: Obtain the knowledge graph and vectorize it to build a word embedding vector database; The pre-acquired natural language queries of users are converted into the query language of the knowledge graph through the pre-trained large language model; Extracting entities and relations in the query language, and mapping the entities and relations to the word embedding vector database; Reconstructing entities and relations mapped to the word embedding vector database through attention fusion, replacing the query language, and generating a replacement language; Based on the replacement language, a statement query is performed through a search engine built into the knowledge graph to generate a query result.

2. According to the attention mechanism-based dynamic fusion word embedding and graph structure information query method of claim 1, it is characterized in that: The step of obtaining a knowledge graph and vectorizing the knowledge graph to construct a word embedding vector database specifically includes: Select a base point from the knowledge graph to traverse the entire knowledge graph to obtain entities and relationships in the knowledge graph; Use word embedding technology to map entities and relationships in the knowledge graph into low-dimensional dense vectors; Insert the low-dimensional dense vector into the preset vector database, assign a unique ID to each low-dimensional dense vector, and build a word embedding vector database.

3. The method for querying dynamic word embedding and graph structure information based on attention mechanism according to claim 1, characterized in that: The training process of the large language model is as follows: Obtain a base model and a dataset containing natural query sentence input and corresponding query language; Fine-tune the base model using the fine-tuned data set to learn how to generate correct query language based on input; The cross entropy loss function is used to measure the gap between the output of the base model and the actual query language, and the parameters of the base model are updated through back propagation and optimizer to obtain the trained large language model.

4. The method for querying dynamic word embedding and graph structure information based on attention mechanism according to claim 1, characterized in that: The extracting entities and relations in the query language and mapping the entities and relations to the word embedding vector database includes: Obtain entities and relations in the query language, and perform word embedding on the entities and relations in the query language; Entities and relationships are mapped to the word embedding vector database to make the word embedding space of the query language consistent with the word embedding space of the knowledge graph.

5. The method for querying dynamic word embedding and graph structure information based on attention mechanism according to claim 1, characterized in that: The fusion and reconstruction of entities and relationships mapped to the word embedding vector database through the attention mechanism, replacing the query language, and generating the replacement language specifically include: Query the word embedding vector database for the vector with the highest similarity to the entities and relations in the query language, and extract the neighbor information in the knowledge graph for each entity in the query language; Performing a dot product operation on the entity word embedding in the query language and the word embedding of the neighbor entity and relationship in the neighbor information through the attention mechanism to obtain the attention weight and normalize the attention weight to obtain the attention distribution; The entity words in the query language are embedded in the weighted neighbor information and aggregated to obtain the entity representation that integrates the graph structure; Based on the text represented by the entity representation in the fused query statement in the knowledge graph, the original query language is replaced to generate a replacement language.

6. The method for querying dynamic word embedding and graph structure information based on attention mechanism according to claim 1, characterized in that: The sentence query based on the replacement language is performed through the built-in search engine of the knowledge graph to generate the query result, specifically including: The knowledge graph is saved in the form of a graph database, and the search engine in the graph database is used to perform the search of query statements, generate query results, and feed back the query results to the user who inputs the natural language query.

7. A dynamic fusion word embedding and graph structure information query system based on attention mechanism, characterized in that: The system comprises: The word embedding vector database construction module is used to obtain the knowledge graph and vectorize the knowledge graph to build a word embedding vector database; A language conversion module is used to convert the pre-acquired natural language queries of users into the query language of the knowledge graph through a pre-trained large language model; A mapping module, used to extract entities and relations in the query language, and map the entities and relations to the word embedding vector database; A language reconstruction module, used for fusing and reconstructing entities and relations mapped to the word embedding vector database through an attention mechanism, replacing the query language, and generating a replacement language; The query feedback module is used to perform statement query based on the replacement language through the search engine built into the knowledge graph to generate query results.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for dynamically fusion word embedding and graph structure information query based on the attention mechanism as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the dynamic fusion word embedding and graph structure information query method based on the attention mechanism as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the dynamic fusion word embedding and graph structure information query method based on the attention mechanism as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Address search method and system based on spatial relation knowledge graph enhanced LLM

    CN120407606A