Knowledge query processing method and apparatus
By combining knowledge graphs and large language models, and utilizing entity extraction and alignment models, the problem of biased understanding of professional knowledge in vertical domain intelligent question answering systems has been solved, resulting in more accurate knowledge retrieval and improved user experience.
Patent Information
- Application Number
- PCT/CN2025/076662
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-07
- Filing Date
- 2025-02-10
- Publication Date
- 2025-12-11
AI Technical Summary
Existing intelligent question-answering systems have biases in their understanding of professional knowledge in vertical fields, resulting in inaccurate answers.
By combining knowledge graphs and large language models, and through pre-trained entity extraction and alignment models, intent recognition, entity extraction, and entity alignment are performed to form standard graph database query statements, thereby obtaining knowledge query results.
It improved the accuracy of knowledge retrieval in vertical domain question-and-answer systems and enhanced the user experience.
Smart Images

Figure CN2025076662_11122025_PF_FP_ABST
Abstract
Description
Knowledge query processing method and device
[0001] Cross-reference to related applications
[0002] The present disclosure is based on and claims priority from Chinese patent application 2024107401798 filed on June 7, 2024, the disclosure of which is incorporated herein in its entirety by reference. TECHNICAL FIELD
[0003] Embodiments of the present disclosure relate to the technical field of wireless communication, in particular to a knowledge query processing method and device. BACKGROUND
[0004] Currently, there are mainly two implementation methods for intelligent question answering systems: intelligent question answering systems based on knowledge graphs and intelligent question answering systems based on general large models. The question answering system based on the knowledge graph has the advantages of fast knowledge iteration and accurate answers, but has poor natural language understanding capability. The question answering system based on the general large model has good natural language understanding capability, but has illusion in question answering and cannot understand professional knowledge. Based on this, the existing intelligent question answering system scheme combines the general large model and the knowledge graph, and the general large model reduces the illusion of the general large model and improves the professional nature of the answers by integrating the professional knowledge searched by the knowledge graph. However, for the intelligent question answering system in the vertical field (such as medical diagnosis, smart power grid, fault diagnosis, etc.) with strict knowledge accuracy requirements and fast knowledge iteration speed, the illusion problem cannot be overcome from the principle of probability maximization of the general model, and the general model has certain deviation in understanding professional knowledge.
[0005] For the problem that the intelligent question answering system in the related art has certain deviation in understanding professional knowledge in the vertical field and the answers are not accurate enough, no solution has been proposed. SUMMARY
[0006] Embodiments of the present disclosure provide a knowledge query processing method and device to at least solve the problem that the intelligent question answering system in the related art has certain deviation in understanding professional knowledge in the vertical field and the answers are not accurate enough.
[0007] According to an embodiment of the present disclosure, a knowledge query processing method is provided, the method comprising:
[0008] obtaining a language to be queried;
[0009] performing intent recognition and entity extraction on the language to be queried based on a pre-trained target entity extraction model to form a first query sentence;
[0010] aligning, based on a pre-trained target entity alignment model, entities in the first query sentence with the graph database to form a second query sentence;
[0011] extracting, based on the second query sentence, knowledge from the graph database to obtain a knowledge query result.
[0012] According to another embodiment of the present disclosure, a knowledge query processing apparatus is provided, and the apparatus comprises:
[0013] an acquisition module configured to acquire a language to be queried;
[0014] an extraction module configured to perform intent recognition and entity extraction on the language to be queried based on a pre-trained target entity extraction model to form a first query sentence;
[0015] an alignment module configured to align entities in the first query sentence with a graph database based on a pre-trained target entity alignment model to form a second query sentence;
[0016] an extraction module configured to extract knowledge from the graph database based on the second query sentence to obtain a knowledge query result.
[0017] According to still another embodiment of the present disclosure, a computer program product is also provided, comprising computer program instructions, wherein the computer program instructions cause a computer to implement the steps in any of the above method embodiments.
[0018] According to still another embodiment of the present disclosure, a computer readable storage medium is also provided, wherein the storage medium stores a computer program, and the computer program is configured to execute the steps in any of the above method embodiments when running.
[0019] According to still another embodiment of the present disclosure, an electronic device is also provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the computer program to perform the steps in any of the above method embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0020] FIG. 1 is a hardware structure block diagram of a computer device of a knowledge query processing method according to an embodiment of the present disclosure;
[0021] FIG. 2 is a flowchart of a knowledge query processing method according to an embodiment of the present disclosure;
[0022] FIG. 3 is a flowchart of a knowledge query processing method according to an optional embodiment of the present disclosure;
[0023] FIG. 4 is a flowchart of a corpus generation process according to an embodiment of the present disclosure;
[0024] FIG. 5 is a block diagram of a knowledge query processing apparatus according to an embodiment of the present disclosure;
[0025] FIG. 6 is a block diagram of a knowledge query processing apparatus according to an optional embodiment of the present disclosure;
[0026] FIG. 7 is a structural schematic diagram of an intelligent question-answering system based on a large language model + knowledge graph according to the present embodiment. DETAILED DESCRIPTION
[0027] Hereinafter, the embodiments of the present disclosure will be described in detail with reference to the accompanying drawings and in conjunction with embodiments.
[0028] It should be noted that the terms "first", "second", and the like in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence.
[0029] The method embodiments provided in the embodiments of the present disclosure can be executed in a computer device or a similar computing device. Taking an example of running on a computer device, FIG. 1 is a hardware structure block diagram of a computer device of a knowledge query processing method according to an embodiment of the present disclosure, as shown in FIG. 1, the computer device can include one or more (only one is shown in FIG. 1) processors 102 (the processor 102 can include but not limited to a processing device such as a microprocessor MCU or programmable logic device) and a memory 104 for storing data, wherein the above-mentioned computer device can further include a transmission device 106 for communication function and an input and output device 108. Those skilled in the art can understand that the structure shown in FIG. 1 is only schematic, which does not limit the structure of the above-mentioned computer device. For example, the computer device can further include more or less components than those shown in FIG. 1, or have a different configuration from that shown in FIG. 1.
[0030] The memory 104 can be used to store computer programs, for example, software programs of application software and modules, such as the computer program corresponding to the knowledge query processing method in the embodiments of the present disclosure. The processor 102 executes various functions and applications and single board matching by running the computer program stored in the memory 104, that is, implements the above-mentioned method. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, which can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0031] The transmission device 106 is configured to receive or send data via a network. The network can include a wireless network provided by a communication provider of the computer device. In an example, the transmission device 106 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In an example, the transmission device 106 can be a radio frequency (RF) module configured to communicate with the Internet in a wireless manner.
[0032] In the embodiment, a knowledge query processing method running on the computer device is provided. FIG. 2 is a flowchart of the knowledge query processing method according to the embodiment of the present disclosure. As shown in FIG. 2, the flowchart includes the following steps:
[0033] In step S202, the language to be queried is obtained.
[0034] In step S204, an intent recognition and entity extraction are performed on the language to be queried based on a pre-trained target entity extraction model to form a first query sentence.
[0035] In step S206, an entity alignment is performed between the entities in the first query sentence and the graph database based on a pre-trained target entity alignment model to form a second query sentence.
[0036] In step S208, knowledge is extracted from the graph database based on the second query sentence to obtain a knowledge query result.
[0037] Through the above steps S202 to S208, the problem that the professional knowledge understanding of the intelligent question and answer system in the related art in the vertical field has a certain deviation and the answer is not accurate can be solved. The question and answer system in the vertical field can directly query the knowledge of the graph database in the form of natural language, the accuracy of the knowledge query result is improved, and the user experience is improved.
[0038] In the embodiment of the present disclosure, step S204 can specifically include: forming a query sentence structure according to the recognized intent based on the target entity extraction model, and filling the entities in the first query sentence into the corresponding positions of the query sentence structure to form the first query sentence.
[0039] In the embodiment of the present disclosure, step S206 can specifically include: obtaining one or more entities from the graph database based on the target entity alignment model, the similarity of the one or more entities to the entities in the first query sentence being greater than a preset threshold, determining a target entity with the highest similarity from the one or more entities, and replacing the entities in the first query sentence with the target entity to obtain the second query sentence.
[0040] FIG. 3 is a flowchart of a knowledge query processing method according to an optional embodiment of the present disclosure. As shown in FIG. 3, the method further includes the following steps:
[0041] Step S302, constructing a first preset number of corpora from the knowledge graph, the corpora including natural language and corresponding labels;
[0042] Step S304, obtaining a plurality of similar corpora corresponding to each of the first preset number of corpora;
[0043] Step S306, training an initial entity extraction model according to the first preset number of corpora to obtain a target entity extraction model;
[0044] Step S308, training an initial target entity alignment model according to the first preset number of corpora and the corresponding plurality of similar corpora to obtain a target entity alignment model.
[0045] In the embodiments of the present disclosure, the above step S302 can specifically include: constructing the first preset number of corpora from the graph database according to entities, attributes, and relationships based on a preset template.
[0046] In the embodiments of the present disclosure, the above step S306 can specifically include: taking the natural language of the first preset number of corpora as the input of the initial entity extraction model, taking the labels corresponding to the first preset number of corpora as the output of the initial entity extraction model, training the initial entity extraction model, obtaining the trained target entity extraction model, and satisfying a first preset condition between the labels corresponding to the natural language of the first preset number of corpora output by the target entity model and the labels actually corresponding to the natural language of the first preset number of corpora.
[0047] In the embodiments of the present disclosure, the above step S308 can specifically include: setting a similarity value for each of the first preset number of corpora and the corresponding similar corpora; selecting a second preset number of corpora from the corpora other than the target corpora in the first preset number of corpora, and determining the similarity value between the target corpora and each of the second preset number of corpora, the target corpora being any of the first preset number of corpora, and the second preset number being less than the first preset number; taking the target corpora, the plurality of similar corpora corresponding to the target corpora, and the second preset number of corpora corresponding to the target corpora as the input of the initial entity alignment model, taking the corresponding similarity as the output of the initial entity alignment model, training the initial entity alignment model, obtaining the trained target entity alignment model, and satisfying a second preset condition between the similarity value of the target corpora and the corresponding corpora output by the target entity alignment model and the actual similarity value of the target corpora and the corresponding corpora.
[0048] The embodiments of the present disclosure are aimed at a scene with strict knowledge question and answer accuracy requirements and fast knowledge iteration. A large amount of high-quality corpus is generated by a knowledge graph and a large model to train an entity extraction model and an entity alignment model. The large model is used for natural language understanding. The knowledge graph is used for knowledge retrieval and extraction. A question and answer system suitable for vertical field knowledge is constructed. Knowledge query of a graph database can be directly performed in a natural language form. User experience is improved. The embodiments of the present disclosure combine a large model and a knowledge graph. A general large model is trained by using professional corpus. The general large model can understand professional knowledge in a vertical field. The knowledge graph is used as the only output of an answer to solve the illusion problem caused by the integration of knowledge of the large model. A method for constructing an intelligent question and answer system based on a large language model (LLM) + knowledge graph. The method is based on a constructed knowledge graph. According to an expert question and answer template, corpus is produced by using a tool and a large model for entity extraction model (intention recognition) training and entity alignment training. The converged model is deployed on a question and answer system. During intelligent question and answer, the trained entity extraction model first performs intention recognition and entity extraction on natural language input by a user to obtain corresponding entities and labels. Then, the entity alignment model is used to perform entity alignment with a graph database. A standard query statement is formed to perform knowledge query. User experience is improved.
[0049] The knowledge in the knowledge graph is stored in a structured manner, and the corpus can be directly constructed according to entities, attributes and relationships. The expert template (corresponding to the preset template) is mainly to ensure that the generated corpus conforms to human language expression as much as possible, so that the model can better understand human language. Therefore, the expert generates a template for a certain problem or a certain scene, and then uses a tool to instantiate the expert template. For example, if the entity is XX function, and the attribute is principle, then according to the expert template, the generated corpus is what is the principle of XX function? The corpus generated based on the above expert template can comprehensively cover the corresponding knowledge, but these corpora are all based on tool expansion, and the corpora lack diversity. To solve this problem, a large model can be used to generalize the already generated corpus to obtain similar corpora, generate corpora in different syntax formats, and improve the diversity of the corpora. FIG. 4 is a flowchart of a corpus generation process according to an embodiment of the present disclosure. As shown in FIG. 4, the entity with the label function is searched from the knowledge graph, the expert template is XX function, and the principle is what is the principle of XX function? Based on the expert template, the corpus is generated, and the principle of MIMO (Multiple-Input Multiple-Output) function is what is the principle of MIMO function? Then the generated corpus is generalized to obtain the following different corpora: 1. How does MIMO technology work? 2. What is the working principle of the MIMO system? 3. How to understand the basic principle of MIMO technology? Then store in the corpus database. These corpora can be used for training of entity extraction and entity alignment models after being labeled differently.
[0050] After the corpus is collected, the model is trained, the entity extraction corpus is used to train the entity extraction model, and the entity alignment corpus is used to train the entity alignment model. In each training period, the training data is traversed, forward and backward propagation is performed, and the model parameters are updated according to the loss. At the same time, according to the training result, the learning rate, the number of training periods and other parameters are adjusted to obtain the converged entity extraction model and entity alignment model.
[0051] The converged entity extraction model and entity alignment model are deployed in the question and answer system. The user first inputs a natural language question, the entity extraction module uses the entity extraction model to extract keywords and identify intent, then the entity alignment module uses the entity alignment model to align the extracted keywords with the entities in the graph database, and combines the aligned keywords and the query logic obtained by identifying the intent to generate a graph database query statement, and then uses the generated graph database query statement to extract knowledge from the graph database and present the extracted knowledge to the user.
[0052] First, according to the entity type in the knowledge graph, generate vertical domain-specific word groups, and set up question and answer templates according to expert templates, such as What is the principle of XX function? Here, the graph database used by the knowledge graph is NebulaGraph. Based on each template, use tools to extract the corresponding entity name, attribute name, and relationship from the graph database for filling, for example, What is the principle of MIMO function? After generating the question and answer question, the question needs to be tagged, and the tags mainly include entity extraction and intent recognition, such as { 'MIMO': 'function entity', 'principle': 'function entity. attribute name'}, this type of tag already contains the query logic (intent recognition) of the graph database, and the model can directly generate standard graph database query statements after generating the question tag.
[0053] In order to ensure the diversity of the corpus, the generated corpus is generalized to obtain multiple similar pre-requisitions corresponding to the pre-requisition, such as,
[0054] {
[0055] You are an excellent sentence constructor. Now, please construct sentences with the given keywords. I will provide the keywords and the sentence I have constructed. Please construct 3 more sentences based on the given keywords and the sentence I have constructed.
[0056] Note: The keywords must appear in the sentence.
[0057] Note: The output format is based on the example.
[0058] Here is an example:
[0059] Keywords: MIMO function, principle
[0060] Sentence I constructed: What is the principle of MIMO function?
[0061] Your 3 constructed sentences:
[0062] 1. Can you introduce the principle of MIMO?
[0063] 2. Principle of MIMO?
[0064] 3. MIMO, principle?
[0065] Of course, the constructed sentences do not have to be the same as the example. Please answer:
[0066] {question}
[0067] Your 3 constructed sentences:
[0068] }
[0069] Finally, two formats of corpora are formed, as follows:
[0070] {"Question": "What is the principle of MIMO?", label: {"MIMO": "entity", "principle": "attribute"}, This format of corpus is used for entity extraction model training;
[0071] {"Q1": "What is the principle of MIMO?", "Q2": "What is the principle of MIMO?", "label": "0.9"}, This format of corpus is used for embedding model training.
[0072] After the corpus is generated, the Chinese-Llama2 7B model is fine-tuned using format 1 corpus, and the Sentence Transformers model is fine-tuned using format 2 corpus. The generated corpus is divided into training set, validation set and test set according to 8:1:1. The training set corpus is input into the model. In each training cycle, the corpus set is traversed, and forward and backward propagation is performed. The model parameters are updated through loss. At the same time, during the learning process, according to the performance of the validation set, the learning rate, the number of training cycles and other parameters are adjusted to find the best performance of the model through multiple iterations, and the converged entity extraction model and entity alignment model are obtained.
[0073] The converged entity extraction model and entity alignment model are deployed in the question and answer system. The user first inputs a natural language question. The entity extraction module uses the entity extraction model to extract keywords and identify intent. The entity alignment module uses the entity alignment model Sentence Transformers model to align the extracted keywords with the entities in the graph database, and combines the aligned entities and the query logic obtained by intent recognition to generate a standard nGQL query statement of the NebulaGraph graph database. Then, the query statement is used to extract knowledge from the graph database, and finally returned to the user.
[0074] The embodiment of the present disclosure also provides a knowledge query processing device. FIG. 5 is a block diagram of a knowledge query processing device according to an embodiment of the present disclosure. As shown in FIG. 5, the device comprises:
[0075] The acquisition module 52 is configured to acquire a language to be queried;
[0076] The extraction module 54 is configured to perform intent recognition and entity extraction on the language to be queried based on a pre-trained target entity extraction model, to form a first query statement;
[0077] The alignment module 56 is configured to perform entity alignment between the entities in the first query statement and a graph database based on a pre-trained target entity alignment model, to form a second query statement;
[0078] The extraction module 58 is configured to perform knowledge extraction from the graph database based on the second query statement to obtain a knowledge query result.
[0079] In the embodiments of the present disclosure, the extraction module 54 is further configured to form a query statement structure according to the identified intent based on the target entity extraction model, and fill the entity in the first query statement into a corresponding position of the query statement structure to form the first query statement.
[0080] In an embodiment, the alignment module 56 is further configured to obtain one or more entities from the graph database based on the target entity alignment model, the one or more entities having a similarity greater than a preset threshold to the entity in the first query statement, determine a target entity having a highest similarity from the one or more entities, and replace the entity in the first query statement with the target entity to obtain the second query statement.
[0081] FIG. 6 is a block diagram of a knowledge query processing apparatus according to an optional embodiment of the present disclosure. As shown in FIG. 6, the apparatus further includes:
[0082] The construction module 62 is configured to construct a first preset number of corpora from a knowledge graph, wherein the corpora include natural languages and corresponding labels.
[0083] The generalization module 64 is configured to perform generalization processing on the first preset number of corpora to obtain a plurality of similar corpora corresponding to each corpus.
[0084] The first training module 66 is configured to train an initial entity extraction model based on the first preset number of corpora to obtain the target entity extraction model.
[0085] The second training module 68 is configured to train an initial target entity alignment model based on the first preset number of corpora and the corresponding plurality of similar corpora to obtain the target entity alignment model.
[0086] In an embodiment, the construction module 62 is further configured to construct the first preset number of corpora from the graph database according to entities, attributes, and relationships based on a preset template.
[0087] In an embodiment, the first training module 66 is further configured to train the initial entity extraction model by taking the natural language of the first preset number of corpora as an input of the initial entity extraction model and taking the label corresponding to the first preset number of corpora as an output of the initial entity extraction model, to obtain the trained target entity extraction model, wherein the label output by the target entity model corresponding to the natural language of the first preset number of corpora satisfies a first preset condition with the label actually corresponding to the natural language of the first preset number of corpora.
[0088] In an embodiment, the second training module 68 is further configured to set the first preset number of corpora and the similarity values of the corresponding each similar corpus respectively, select a second preset number of corpora from the corpora other than the target corpus in the first preset number of corpora, and determine the similarity values of the target corpus and each corpus in the second preset number of corpora, wherein the target corpus is any corpus in the first preset number of corpora, and the second preset number is less than the first preset number, take the target corpus, the plurality of similar corpora corresponding to the target corpus, and the second preset number of corpora corresponding to the target corpus as inputs of the initial entity alignment model, take the corresponding similarity values as outputs of the initial entity alignment model, train the initial entity alignment model to obtain the trained target entity alignment model, and wherein the similarity values of the target corpus and the corresponding corpora output by the target entity alignment model satisfy a second preset condition.
[0089] FIG. 7 is a structural schematic diagram of an intelligent question answering system based on a large language model + knowledge graph according to the present embodiment. As shown in FIG. 7, it mainly includes a corpus generation module, an entity extraction module, an entity alignment module, a knowledge graph module, and some or all of them are combined to realize part or all of the functions of the above modules.
[0090] The corpus generation module extracts relevant information from the knowledge graph according to the expert template to construct training corpora, and then generalizes these corpora through a large model to improve the diversity of the corpora.
[0091] The entity extraction module trains the general large language model using the generated corpora, so that the general large language model has the ability of vertical domain entity extraction. This module mainly includes two functions of keyword extraction and intent recognition. Keyword extraction is mainly to extract vertical domain keywords from user input natural language; intent recognition is mainly to obtain the query logic of user input natural language.
[0092] The entity alignment module trains the general large language model using the generated corpora, so that the general large language model has the ability of vertical domain entity alignment. This module mainly includes two functions of keyword alignment and query sentence generation. Keyword alignment is mainly to align the keywords extracted from the user input natural language with the entities in the knowledge graph; query sentence generation is to combine the aligned entities and the query logic obtained in the entity extraction module to obtain a complete graph database query sentence.
[0093] The knowledge graph module mainly has two functions: knowledge storage and knowledge extraction. The knowledge storage is to store the knowledge in a vertical field into a graph database, and is performed under the premise that the knowledge graph is constructed, so it does not involve knowledge storage. The main function of the knowledge extraction is to acquire knowledge according to an input graph database query statement, to help generate training corpus and output question answers.
[0094] The embodiments of the present disclosure also provide a computer program product, comprising computer program instructions, wherein the computer program instructions enable a computer to implement the steps in any of the above method embodiments.
[0095] The embodiments of the present disclosure also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.
[0096] In an example embodiment, the above computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0097] The embodiments of the present disclosure also provide an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the computer program to perform the steps in any of the above method embodiments.
[0098] In an example embodiment, the above electronic device can further comprise a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0099] The specific examples in the embodiments can refer to the examples described in the above embodiments and example embodiments, which will not be described herein again.
[0100] Obviously, those skilled in the art should understand that the above modules or steps of the present disclosure can be realized by a general computing device, which can be concentrated on a single computing device or distributed on a network composed of multiple computing devices, and can be realized by program codes executable by a computing device, so that they can be stored in a storage device and executed by a computing device, and in some cases, the steps shown or described can be executed in different order, or they can be manufactured into individual integrated circuit modules, or multiple modules or steps can be manufactured into a single integrated circuit module. Therefore, the present disclosure is not limited to any specific hardware and software combination.
[0101] The above merely provides exemplary embodiments of the present disclosure, but is not for limiting the present disclosure. For those skilled in the art, the present disclosure can have various modifications and changes. Any modified, equivalent replaced, improved and other technical solutions made within the principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A knowledge query processing method, the method comprising: acquiring a language to be queried; performing intent recognition and entity extraction on the language to be queried based on a pre-trained target entity extraction model to form a first query sentence; performing entity alignment between entities in the first query sentence and a graph database based on a pre-trained target entity alignment model to form a second query sentence; performing knowledge extraction from the graph database based on the second query sentence to obtain a knowledge query result.
2. The method of claim 1, wherein, Performing intent recognition and entity extraction on the language to be queried based on a pre-trained target entity extraction model to form a first query sentence comprises: forming a query sentence structure according to the recognized intent based on the target entity extraction model, and filling the entities in the first query sentence into the corresponding positions of the query sentence structure to form the first query sentence.
3. The method of claim 1, wherein, Performing entity alignment between entities in the first query sentence and a graph database based on a pre-trained target entity alignment model to form a second query sentence comprises: acquiring one or more entities from the graph database that have a similarity greater than a preset threshold to the entities in the first query sentence based on the target entity alignment model, determining a target entity with the highest similarity from the one or more entities, and replacing the entities in the first query sentence with the target entity to obtain the second query sentence.
4. The method of claim 1, wherein, The method further comprises: constructing a first preset number of corpora from a knowledge graph, wherein the corpora comprise natural language and corresponding labels; acquiring a plurality of similar corpora corresponding to each corpus in the first preset number of corpora; training an initial entity extraction model based on the first preset number of corpora to obtain the target entity extraction model; training an initial target entity alignment model based on the first preset number of corpora and the corresponding plurality of similar corpora to obtain the target entity alignment model.
5. The method of claim 4, wherein, Constructing a first preset number of corpora from a knowledge graph comprises: constructing the first preset number of corpora from the graph database according to entities, attributes, and relationships based on a preset template.
6. The method of claim 4, wherein, Training an initial entity extraction model based on the first preset number of corpora to obtain the target entity extraction model comprises: training the initial entity extraction model by taking the natural language of the first preset number of corpora as input and the labels corresponding to the first preset number of corpora as output, to obtain the trained target entity extraction model, wherein the labels corresponding to the natural language of the first preset number of corpora output by the target entity model satisfy a first preset condition with the actual labels corresponding to the natural language of the first preset number of corpora.
7. The method of claim 4, wherein, Training an initial target entity alignment model based on the first preset number of corpora and the corresponding plurality of similar corpora to obtain the target entity alignment model comprises: respectively setting similarity values for the first preset number of corpora and each corresponding similar corpus. selecting a second preset number of corpora from the corpora other than the target corpus in the first preset number of corpora, and determining a similarity value of the target corpus and each of the second preset number of corpora, wherein the target corpus is any corpus in the first preset number of corpora, and the second preset number is less than the first preset number; training the initial entity alignment model by taking the target corpus, the plurality of similar corpora corresponding to the target corpus, and the second preset number of corpora corresponding to the target corpus as inputs, and taking the corresponding similarity values as outputs, to obtain the trained target entity alignment model, wherein the similarity value of the target corpus and the corresponding corpus output by the target entity alignment model satisfies a second preset condition. 8.A knowledge query processing apparatus, comprising: an acquisition module configured to acquire a language to be queried; an extraction module configured to perform intent recognition and entity extraction on the language to be queried based on a pre-trained target entity extraction model, to form a first query sentence; an alignment module configured to perform entity alignment between entities in the first query sentence and a graph database based on a pre-trained target entity alignment model, to form a second query sentence; an extraction module configured to perform knowledge extraction from the graph database based on the second query sentence, to obtain a knowledge query result.
9. A computer-readable storage medium having stored therein a computer program, wherein, The computer program is configured to execute the method described in any one of claims 1 to 7 when running. 10.A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Semantic matching method and device for power transformer knowledge questions and answers
CN113919366A
Auxiliary retrieval method fusing knowledge graph and large language model
CN117633252A
Query statement generation method and device, electronic equipment and storage medium
CN117708297A
Generating question templates in a knowledge-graph based question and answer system
US20210201174A1