Data query method and device, electronic equipment and storage medium

By identifying and combining entities and intention information in the knowledge graph within a single module, and fusing them with the graph query grammar to generate target query statements, the problems caused by identification of entities and intentions in different modules in the question-and-answer system are solved, the recognition efficiency and accuracy are improved, and the system's generalization and recall capabilities are enhanced.

CN120162399APending Publication Date: 2025-06-17BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311707999.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The knowledge graph-based question-and-answer system has low problem identification efficiency due to the individual identification of entities and intentions in different modules, and the information of entities and intention pairing is lost, and the matching is inaccurate, which affects the recognition accuracy.

Method used

Through a pre-established generative model, the graph information in the problem statement, including entity information and intention information, is identified in a single module, and combined it in a preset format to generate graph information represented by structured text. Then, it is fused with the graph query syntax, generate a target query statement, and obtain the matching target answer through the graph database query.

Benefits of technology

It improves the efficiency and accuracy of the Q&A system to identify problems, ensures that the pairing information between entities and intentions is retained, reduces the problem of inaccurate matching, and enhances the system's generalization and recall capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162399A_ABST
    Figure CN120162399A_ABST
Patent Text Reader

Abstract

The invention provides a data query method and device, electronic equipment and a storage medium. The method comprises the following steps: obtaining a question statement; identifying map information in the question statement according to a pre-established generative model, and combining the identified map information according to a preset format to obtain map information represented by a structured text; the map information at least comprises entity information and intention information; the generative model comprises a preset format of map information combination; fusing the graph information represented by the structured text with the graph query grammar to generate a target query statement corresponding to the question statement; and querying a graph database according to the target query statement so as to obtain a target answer matched with the question statement. Through the generative model and the preset format, the entity and the intention are identified in the same module, the length of an identification link is reduced, and pairing information between the entity and the intention is reserved, so that the identification accuracy and the generalization ability of the question and answer system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of knowledge graphs, and in particular, to a data query method, apparatus, electronic device, and storage medium. Background Art

[0002] Currently, the use based on human-computer interaction is becoming more and more frequent. The question-answering system is the main information exchange tool for human-computer interaction and is a tool for generating a target answer according to an input question statement (such as a user query); the current question-answering system generally includes a question-answering system based on a knowledge graph and a question-answering system based on a retrieval system.

[0003] However, for a question-answering system based on a knowledge graph, since entity recognition and intent recognition are respectively recognized and extracted through a dictionary and a learning model in different modules, the recognition of questions in the question-answering system needs to be switched between different modules, resulting in a low efficiency of recognizing questions in the question-answering system; and since entities and intents are separately recognized in different modules, even if an entity and an intent are recognized from a question statement, the pairing information between them is not retained, and it is necessary to query an intent library to match the existing intents in the intent library and re-establish the pairing information of the entity and the intent. When matching the intent in the intent library, as long as the intent matching degree meets a certain threshold similarity, when the matching threshold similarity is met but the similarity is low, there may be a problem of inaccurate matching, resulting in insufficient accuracy of recognizing questions in the question-answering system. Summary of the Invention

[0004] The present disclosure provides a data query method, apparatus, electronic device, and storage medium.

[0005] According to a first aspect of the present disclosure, there is provided a data query method, including: obtaining a question statement; recognizing graph information in the question statement according to a pre-established generative model, and combining the recognized graph information in a preset format to obtain graph information represented by structured text; the graph information includes at least entity information and intent information; the generative model includes a preset format for combining graph information; fusing the graph information represented by structured text with a graph query syntax to generate a target query statement corresponding to the question statement; querying a graph database according to the target query statement so as to obtain a target answer matching the question statement.

[0006] In some embodiments of the present disclosure, the graph information in the question statement is identified according to a pre-established generative model, and the identified graph information is combined in a preset format to obtain the graph information represented by structured text, including: identifying the graph information in the question statement according to the pre-established generative model; pairing the identified graph information to form a data group; and filling the graph information in the data group with label information corresponding to different dimensions in the preset format to obtain the graph information represented by structured text.

[0007] In some embodiments of the present disclosure, the graph information represented by structured text is fused with the graph query syntax to generate a target query statement corresponding to the question statement, including: determining the number of entities in the graph information represented by structured text; if the number of entities is 1, fusing the graph information represented by structured text containing one said entity with the graph query syntax to generate at least one target query statement corresponding to the question statement; if the number of entities is at least 2, fusing the graph information represented by structured text containing at least two entities with the graph query syntax to generate at least two target query statements corresponding to the question statement, and each target query statement in the at least two target query statements contains at least two entities, and the positions of the at least two entities in each target query statement are different.

[0008] In some embodiments of the present disclosure, after querying the graph database according to the target query statement, it further includes: when there is an answer in the answers corresponding to the entities included in the target query statement that matches the intent information of the question statement, outputting the answer that matches the intent information of the question statement as the target answer; when there is no answer in the answers corresponding to the entities included in the target query statement that matches the intent information of the question statement, determining whether there is a parent node for the entities included in the target query statement; if there is a parent node, querying whether there is an answer in the answers included in the parent node that matches the intent information of the question statement; if there is an answer in the answers included in the parent node that matches the intent information of the question statement, outputting the answer that matches the intent information of the question statement as the target answer; if there is no parent node, returning a prompt message indicating that no answer matching the intent information of the question statement is found in the query result.

[0009] In some embodiments of the present disclosure, querying whether there is an answer in the answers included in the parent node that matches the intent information of the question statement includes: fusing the information of having a parent node into the target query statement based on the graph query syntax to generate a new target query statement; querying the graph database according to the new target query statement to determine whether there is an answer in the answers included in the parent node that matches the intent information of the question statement.

[0010] In some embodiments of the present disclosure, after querying the graph database according to the target query statement, the following steps are further included: If the entity included in the target query statement is not found, it is determined whether the target query statement includes attribute information; if it includes attribute information, the attribute information is deleted and the graph database is queried again.

[0011] In some embodiments of the present disclosure, deleting the attribute information and continuing to query the graph database includes: deleting the content corresponding to the attribute in the target query statement to construct a target query statement that does not include the attribute; executing the target query statement that does not include the attribute to continue querying the graph database.

[0012] According to a second aspect of the present disclosure, there is provided a data query device, including:

[0013] An acquisition module, configured to acquire a problem statement;

[0014] An identification module, configured to identify the graph information in the problem statement according to a pre-established generative model, and combine the identified graph information in a preset format to obtain graph information represented by structured text; the graph information includes at least entity information and intention information; the generative model includes a preset format for combining graph information;

[0015] A statement generation module, configured to fuse the graph information represented by structured text with graph query syntax to generate a target query statement corresponding to the problem statement;

[0016] A query module, configured to query the graph database according to the target query statement to obtain a target answer that matches the problem statement.

[0017] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0018] At least one processor; and

[0019] A memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the foregoing first aspect.

[0021] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the foregoing first aspect.

[0022] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the method described in the foregoing first aspect.

[0023] The data query method, device, electronic device, and storage medium provided by the present disclosure obtain a question statement; identify the graph information in the question statement according to a pre-established generative model, and combine the identified graph information in a preset format to obtain graph information represented by structured text; the graph information includes at least entity information and intent information; the generative model includes a preset format for combining graph information; fuse the graph information represented by structured text with graph query syntax to generate a target query statement corresponding to the question statement; query a graph database according to the target query statement to obtain a target answer matching the question statement. In this application, entities and intents are identified within the same pre-established generative model, realizing the extraction of entities and intents in the same module. Since entities and intents are identified and extracted in the same module, it is avoided that entities and intents are identified in different modules, so that the identification of questions in the question-answering system does not need to switch between different modules, improving the efficiency of identifying questions in the question-answering system; and since entities and intents are extracted and identified from the same question statement in the same module, the pairing information between entities and intents is retained. At the same time, the identified graph information is combined using a preset format, and the combined content is the entities and intents identified in the question statement, rather than the intent matched based on similarity in the prior art, improving the accuracy of entity recognition and intent recognition, and thus ensuring the accuracy of identifying questions in the question-answering system.

[0024] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understandable through the following description. Brief Description of the Drawings

[0025] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0026] Figure 1 is a schematic flowchart of a data query method provided by an embodiment of the present disclosure;

[0027] Figure 2 is a schematic flowchart of another data query method provided by an embodiment of the present disclosure;

[0028] Figure 2a is a schematic diagram of obtaining graph information provided by an embodiment of the present disclosure;

[0029] Figure 2b is a schematic diagram of a knowledge graph provided by an embodiment of the present disclosure;

[0030] Figure 2c is a schematic diagram of querying a graph database provided by an embodiment of the present disclosure;

[0031] Figure 3 A schematic diagram of an architecture for implementing a data query method based on an end-to-end generative model provided by an embodiment of the present disclosure;

[0032] Figure 4 A schematic diagram of the structure of a data query device provided by an embodiment of the present disclosure;

[0033] Figure 5 A schematic block diagram of an exemplary electronic device 500 provided by an embodiment of the present disclosure. Detailed implementation manners

[0034] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0035] To solve the problems in the related art, the present invention can identify entities and intents through a single module by presetting a graph structure and the expandable multi-tasking ability of an end-to-end model, retain the pairing information between the two, and improve the recognition accuracy and generalization ability. At the same time, entities are used to recall questions, narrowing the recall range and improving the recall ability for candidate questions.

[0036] The following describes a data query method, device, electronic device, and storage medium according to embodiments of the present disclosure with reference to the accompanying drawings.

[0037] Figure 1 A flowchart of a data query method provided by an embodiment of the present disclosure. This method can be executed by an electronic device, specifically, by an electronic device such as a personal computer (PC), a server, etc. As Figure 1 shown, the method includes:

[0038] Step 101: Obtain a question statement.

[0039] In some embodiments, a question statement generally refers to a query statement input by a user or a system, also known as a user query.

[0040] Specifically, the terminal can generate a question statement based on information such as voice and text input by the user or the system, and send the question statement to the electronic device; where the terminal can be a mobile phone, a laptop computer, a tablet computer, a wearable intelligent device, etc.

[0041] Step 102: Identify the graph information in the question statement according to the pre-established generative model, and combine the identified graph information in a preset format to obtain the graph information represented by structured text; the graph information includes at least entity information and intent information; the generative model includes the preset format for graph information combination.

[0042] In some embodiments, through the pre-established generative model, extract the graph information in the question statement, obtain the entity information and intent information corresponding to the question statement, and combine the extracted graph information according to the preset format for graph information combination included in the generative model to obtain the graph information represented by structured text.

[0043] The pre-established generative model refers to the generative model obtained through pre-training. The generative model can specifically be an end-to-end generative model. The end-to-end generative model can first tokenize the user query, then load the tokenization result into the encoder of the end-to-end generative model through the vocabulary vector to generate the sentence vector of the text, and then output the extracted text, that is, the graph information, through the decoder of the end-to-end generative model. After that, the extracted graph information is combined through the preset format in the end-to-end generative model to obtain the graph information represented by structured text.

[0044] The preset format is a pre-set structure or pattern for combining the identified graph information. In a data query system or a question-and-answer system, the preset format can be a pre-constructed knowledge graph used to represent the connection structure of various entities, intents, etc. The preset format can include label information in different dimensions. Combining the identified graph information according to the preset format specifically means filling the information of the identified graph information according to the label information in different dimensions in the preset format. Among them, before combining the identified graph information, it is necessary to pair the identified graph information to form an array value so as to combine the graph information in the data group according to the preset format. The graph information represented by structured text is the text display of the graph information according to the graph structure and is the text expression of the structured graph information.

[0045] Among them, in the present disclosure, the question statement can be identified and extracted through the pre-established generative model, or the graph information corresponding to the question statement can be obtained by parsing the question statement based on the preset graph structure; or the graph information corresponding to the question statement can be obtained by annotating the question statement based on the preset graph structure.

[0046] Step 103: Integrate the graph information represented by structured text with the graph query syntax to generate the target query statement corresponding to the question statement.

[0047] In some embodiments, the graph information represented by structured text and the graph query syntax can be fused according to the graph query syntax format in the graph query syntax, to generate multiple target query statements. Among them, the target query statements can be expressed by graph query statements in the present disclosure. When constructing the target query statements, it is necessary to fuse the graph information represented by structured text with the graph query syntax, so as to generate the target query statements corresponding to the question statements. Different graph query syntaxes can generate different graph query statements. The graph query syntax is determined according to the actual business and is not limited in the embodiments of the present disclosure.

[0048] Step 104: Query the graph database according to the target query statement, so as to obtain the target answer that matches the question statement.

[0049] In some embodiments, the query requests encapsulated by each target query statement can be used in sequence to query the graph database, to determine whether there is a target answer in the graph database that matches the currently queried target query statement. If there is, stop the query and obtain the target answer. If not, continue to query using the next target query statement. Among them, the graph database is a database that represents the relationships between data in a graphical manner.

[0050] In summary, the technical solution provided by the present disclosure realizes the recognition and extraction of entities and intents in the same module through the pre-established generative model, reducing the length of the recognition link. Since the entities and intents are recognized and extracted in the same module, it is avoided that the entities and intents are recognized in different modules, so that the recognition of questions in the question-and-answer system does not need to be switched between different modules, improving the efficiency of the question-and-answer system in recognizing questions; and since the entities and intents are extracted and recognized from the same question statement in the same module, the pairing information between the entities and intents is retained, and at the same time, the recognized graph information is combined in a preset format. The combined content is the entities and intents recognized in the question statement, rather than the intents matched based on similarity in the prior art, improving the accuracy of entity recognition and intent recognition, and thus ensuring the accuracy of the question-and-answer system in recognizing questions.

[0051] Figure 2 FIG. is a schematic flowchart of another data query method provided by an embodiment of the present disclosure. This method can be executed by an electronic device, specifically, it can be executed by an electronic device such as a PC or a server. Figure 2 Based on Figure 1 the embodiments shown, steps 103 and 104 are further defined. In Figure 2 the embodiments shown, step 103 includes step 203 and step 204, and step 104 includes step 205. As Figure 2 shown, this method may include:

[0052] Step 201: Obtain the problem statement.

[0053] In some embodiments, the terminal can generate a problem statement based on information such as voice or text input by the user or the system, and send the problem statement to the electronic device; wherein, the terminal can be a mobile phone, a laptop computer, a tablet computer, a wearable intelligent device, etc.

[0054] Step 202: Identify the graph information in the problem statement according to the pre-established generative model, and combine the identified graph information in a preset format to obtain the graph information represented by structured text; the graph information includes at least entity information and intent information; the generative model includes a preset format for combining graph information.

[0055] In some embodiments, before identifying the graph information in the problem statement according to the pre-established generative model, it is necessary to train the generative model so that the trained generative model can identify and extract the graph information contained in the problem statement to obtain the graph information corresponding to the problem statement. Identifying the graph information in the problem statement according to the pre-established generative model and combining the identified graph information in a preset format to obtain the graph information represented by structured text specifically includes: identifying the graph information in the problem statement according to the pre-established generative model; pairing the identified graph information to form a data group; filling the graph information in the data group with the label information corresponding to different dimensions in the preset format to obtain the graph information represented by structured text.

[0056] The pre-established generative model is a model trained based on the problem statement and the graph information represented by structured text, wherein the input of the generative model is the problem statement, and the output is the graph information represented by structured text. The graph information represented by structured text is the graph information in the data group after being filled with the label information corresponding to different dimensions in the preset format.

[0057] The graph information can include one or more of: entity information, intent information, attribute information, location information, vehicle model information, etc.; wherein, the attribute information can include temperature information, length information, etc. When identifying the graph information, if there are entities and intents in the problem statement, at least the entities and intents are identified to obtain specific entity information and intent information; if there are no entities and intents in the problem statement, the entities and intents are also identified, and the entity information is empty and the intent information is empty.

[0058] Among them, the present disclosure can also parse the problem statement based on a preset graph structure to obtain the graph information corresponding to the problem statement; or annotate the problem statement based on a preset graph structure to obtain the graph information corresponding to the problem statement.

[0059] For example, assume the question statement is: "What should I do if the windshield wiper won't turn on?" Based on a pre-established generative model, the entity information corresponding to this question statement can be extracted as [device] windshield wiper, and the intent information as [intent] fault handling. The entity information and intent information are combined according to the preset format corresponding to the graph information represented by structured text, that is, [device] windshield wiper and [intent] fault handling are combined according to the positional relationship between [device] and [intent], to obtain the graph information represented by structured text, that is, "[device] windshield wiper [intent] fault handling".

[0060] Among them, the generative model in this disclosure can be an end-to-end model. The end-to-end model can first tokenize the question statement, then load the tokenization result into the encoder of the end-to-end model through the vocabulary vector to generate the sentence vector of the text, and then output the extracted text through the decoder of the end-to-end model, that is, the graph information corresponding to the question statement. Then, the graph information is combined according to the preset format to obtain the graph information represented by structured text.

[0061] Exemplarily, as Figure 2a shown in the schematic diagram of obtaining graph information, the end-to-end model is a seq2seq model, the question statement is "How to turn on the massage function of the second-row seat", use the tokenizer to tokenize the question statement, and then use the vocabulary ID vector to load the tokenization result into the seq2seq model encoder to generate the sentence vector of the text to be matched. Then, load the text vector to be matched into the seq2seq model decoder to output the graph information. Then, the graph information is combined through the preset format to output the graph information represented by structured text: [device] seat [function] massage [position] second row [intent] opening method, where [device] seat and [function] massage are entities, [position] second row is an attribute, and [intent] opening method is an intent.

[0062] In some embodiments, the end-to-end model can be trained using the question statement and the graph information represented by structured text, so that when the end-to-end model inputs the question statement, it extracts graph information such as entities, intents, and attributes through the learned information, and uses the graph information output format (i.e., the preset format) in the end-to-end model to output the graph information represented by structured text. The graph information represented by this structured text is mapped in the preset graph structure to generate a knowledge graph as shown in Figure 2b shown, that is, the graph information represented by structured text is represented in the form of a knowledge graph.

[0063] Among them, for the acquisition of graph information, in addition to the above method, it is also possible to directly obtain the graph information corresponding to the question statement by parsing the question statement based on the preset graph structure or by annotating the question statement based on the preset graph structure.

[0064] Exemplarily, on the left side of Table 1 are problem statements, and on the right side are graph information represented by structured text generated through a preset format of training samples in the training sample set:

[0065]

[0066] Among them, for some of the graph information in Table 1, for the graph information that has been defined in the knowledge graph, the problem statement can be directly parsed to obtain it by means of a preset graph structure, that is, graph information such as vehicle model information or other graph information without generalization problems, such as device information, function information, intention information, etc. For graph information closely related to actual business, text pairs of problems and graph information can be generated through manual annotation with the help of an entity list and the extracted graph information, that is, a new problem is obtained by splicing the vehicle model with the generalization problem, such as vehicle model information, location information, etc.

[0067] Step 203: Determine the number of entities in the graph information represented by the structured text.

[0068] In some embodiments, the problem statement is recognized and extracted through a pre-established generative model to determine the entity information and intention information in the problem statement, so as to determine the number of entities in the graph information represented by the structured text.

[0069] Step 204, according to the number of entities, fuse the graph information represented by the structured text with the graph query syntax to generate a target query statement corresponding to the problem statement.

[0070] In actual application, when the graph information represented by the structured text is obtained, the graph inference module can be used to perform inference on the target query statement. The embodiments of the present disclosure perform queries through entities, that is, recall the problem through entities, which can narrow the recall range of the problem, thereby improving the recall ability for the problem; by first combining entities with entity attributes for querying and then using entities for querying when the query fails, the recall range of the problem can be further narrowed, thereby further improving the recall ability for the problem.

[0071] In some embodiments, if the number of entities is 1, the graph information represented by the structured text containing one entity is fused with the graph query syntax to generate at least one target query statement corresponding to the question statement; if the number of entities is at least 2, the graph information represented by the structured text containing at least two entities is fused with the graph query syntax to generate at least two target query statements corresponding to the question statement. Each of the at least two target query statements contains at least two entities, and the positions of the at least two entities are different in each target query statement. In other words, when the number of entities is 1, only the edge relationship between the entity and other information such as attributes needs to be considered to generate at least one target query statement. When the number of entities is at least 2, since the positions (i.e., pointing relationships) of the entities in the extracted graph information represented by the structured text are different, in addition to considering the edge relationship between the entity and other information such as attributes, the directivity between entities, that is, the subordination relationship between different entities, also needs to be considered to generate at least two target query statements.

[0072] In an alternative embodiment of the present disclosure, taking the Cypher graph query syntax as an example. When the number of entities is 1, the graph information represented by the structured text containing one entity can be fused with the graph query syntax to generate at least one target query statement corresponding to the question statement. Exemplarily, the graph information represented by the extracted structured text is [device] windshield wiper [position] front [intention] fault handling. Using the Cypher graph query syntax, the graph query syntax is fused (i.e., combined) with the graph information represented by the structured text to obtain "MATCH (v1: device {name: \"windshield wiper\"}) - [e1: Q&A] -> (v2: answer) WHERE v1.device.position == \"front\" AND e1.intention == \"fault handling\" RETURN v2.answer AS Answer;".

[0073] When the number of entities is 2, for example, when the graph information represented by the extracted structured text is [Device] Seat [Function] Massage [Location] Second row [Intention] Opening method, the graph information represented by the structured text contains two entities, [Device] Seat and [Function] Massage. Since the directivity of the seat and the massage is unknown, therefore, for each pointing relationship, the graph information represented by the structured text and the graph query syntax can be combined to generate a target query statement. Specifically, for the seat pointing to the massage, using the Cypher graph query syntax, the graph query syntax is combined with the graph information represented by the structured text to obtain "MATCH (v1:Device{name:\"Seat\"})-[e1:Has]->(v2:Function{name:\"Massage\"})-[e2:Q&A]->(v2:Answer) WHERE v1.Device.Location == \"Second row\" AND e2.Intention == \"Opening method\" RETURN v3.Answer AS Answer;". For the massage pointing to the seat, using the Cypher graph query syntax, the graph query syntax is combined with the graph information represented by the structured text to obtain "MATCH (v1:FUNCTION{name:\"Massage\"})-[e1:Has]->(v2:Device{name:\"Seat\"})-[e2:Q&A]->(v2:Answer) WHERE v1.Device.Location == \"Second row\" AND e2.Intention == \"Opening method\" RETURN v3.Answer AS Answer;".

[0074] Step 205, encapsulate the target query statement into a query request message and send it to the graph database for querying.

[0075] In some embodiments, as Figure 2c shown in the schematic diagram of querying the graph database, when one or more target query statements are generated, the retrieval steps can be executed by the graph retrieval module. Specifically, for one or more target query statements generated by the graph reasoning module, a target query statement can be directly encapsulated into a request message, or multiple target query statements can be combined into a target query statement list and encapsulated into a request message, and the graph database is requested for querying in sequence according to the order of the target query statement list. Among them, when there are multiple target query statements, combining them into a target query statement list and encapsulating them into a request message together can improve the data query efficiency.

[0076] In some embodiments, when querying the graph database using different target query statements in sequence, when the target answer is obtained using the current target query statement, the query is stopped and the obtained target answer is sent to the terminal; when the target answer is not obtained using the current target query statement, the next target query statement is continued to be used for querying.

[0077] In some embodiments, after querying the graph database according to the target query statement, the following steps are further included: when it is queried that there is an answer in the answers corresponding to the entity included in the target query statement that matches the intent information of the question statement, the answer that matches the intent information of the question statement is output as the target answer; when it is queried that there is no answer in the answers corresponding to the entity included in the target query statement that corresponds to the intent information of the question statement, it is determined whether the entity included in the target query statement has a parent node; if there is a parent node, it is queried whether there is an answer in the answers included in the parent node that matches the intent information of the question statement; if there is an answer in the answers included in the parent node that matches the intent information of the question statement, the answer that matches the intent information of the question statement is output as the target answer; if there is no parent node, a prompt message indicating that no answer matching the intent information of the question statement is queried is returned.

[0078] Among them, querying whether there is an answer in the answers included in the parent node that matches the intent information of the question statement includes: integrating the information of the parent node into the target query statement based on the graph query syntax to generate a new target query statement; querying the graph database according to the new target query statement to determine whether there is an answer in the answers included in the parent node that matches the intent information of the question statement.

[0079] In an alternative embodiment of the present disclosure, when the number of entities is 1, if no answer corresponding to the intent is queried (i.e., retrieved) under the entity included in the target query statement, that is, no question-and-answer edge with the intent of "fault handling" is retrieved, then it is queried whether there is a relevant edge for its parent node. If there is a parent node, a target query statement including the parent node is constructed through the target query statement and the relevant information existing for the parent node, that is, "MATCH(v1:Device{name:\"Wiper\"})-[e1:has parent node]->()-[e2:question and answer]->(v2:Answer) WHERE v1.Device.location == \"front\" AND e2.intent == \"fault handling\" RETURN v2.Answer AS Answer;"; the target query statement including the parent node is executed to determine whether there is an answer corresponding to the intent.

[0080] When the number of entities is 2, two target query statements for a seat with a massage function or a seat with a massage function can be generated. Since the edge relationship between the seat and the massage is one-way, only one target query statement can take effect when the graph retrieval module retrieves in the graph database, and a reply answer about the way to turn on the seat massage will be returned. If no answer is retrieved, a graph query statement for the edge of the entity's parent node can also be automatically generated as when the number of entities is 1, and the query for the edge of the parent node is selected to be performed n times according to the actual business logic, which will not be elaborated here.

[0081] In addition, after querying the graph database according to the target query statement, the following steps are further included: if the entity included in the target query statement is not found, it is determined whether the target query statement includes attribute information; if it includes attribute information, the attribute information is deleted and the graph database is queried continuously. Among them, deleting the attribute information and continuing to query the graph database includes: deleting the content corresponding to the attribute in the target query statement to construct a target query statement that does not include the attribute; executing the target query statement that does not include the attribute to continue querying the graph database.

[0082] Specifically, when the entity included in the target query statement is not found in the graph database, it is further checked whether the target query statement includes attribute information. If it includes attribute information, the graph database can be continuously queried by gradually deleting some attribute information. Among them, by deleting the attribute information to narrow the query scope, it is easier to find the matching entity. Because overly specific attribute conditions may lead to too small a query result set and it is difficult to find the matching entity. By gradually deleting the attribute information, the query scope can be expanded and the possibility of finding the matching entity can be increased.

[0083] In summary, for the technical solution provided by the present disclosure, by using the pre-established generative model to identify the question statement, the entity and the intent can be identified and extracted within the same module, reducing the length of the identification link, avoiding the identification of the entity and the intent in different modules, enabling the identification of questions in the question-answering system without switching between different modules, and improving the efficiency of identifying questions in the question-answering system; and because the entity and the intent are extracted and identified from the same question statement in the same module, the pairing information between the entity and the intent is retained. At the same time, the identified graph information is combined using a preset format to obtain graph information represented by structured text. The graph information represented by structured text includes the entity and the intent identified in the question statement, rather than the intent matched based on similarity in the prior art, improving the accuracy, recall ability of entity recognition and intent recognition, and the generalization ability of the data query system or the question-answering system, thereby ensuring the accuracy of identifying questions in the question-answering system and retaining the pairing information between the entity and the intent.

[0084] The following further illustrates the embodiments of the present disclosure in conjunction with specific application embodiments.

[0085] The application embodiment of the present disclosure provides a data query method based on an end-to-end generative model. The data query method based on the end-to-end generative model can be implemented using the Figure 3 architecture shown. This method may include:

[0086] Step S1: After the question statement (user query) is input into the system, the graph information extraction module first identifies and extracts the entities and intents in the question statement through a pre-established generative model (end-to-end model), and combines the identified graph information in a preset format to obtain graph information represented by structured text.

[0087] Specifically, as Figure 2a shown, when the user query is "How to turn on the massage function of the second-row seat", the graph information extraction module uses a tokenizer to tokenize the user query, and then loads the tokenization result into the seq2seq model encoder using the vocabulary ID vector to generate a text sentence vector to be matched. Then, the text vector to be matched is loaded into the seq2seq model decoder, and combined with the preset format, the graph information represented by structured text is output as [Device] Seat [Function] Massage [Location] Second row [Intent] Opening method, where [Device] Seat and [Function] Massage are entities, [Location] Second row is an attribute, and [Intent] Opening method is an intent.

[0088] Step S2: Combining the graph information represented by the extracted structured text, the graph reasoning module executes a predetermined reasoning link according to the graph query syntax and generates multiple candidate target query statements corresponding to the question statement for the graph query module to call.

[0089] Specifically, as Figure 2b shown, if an entity is retrieved, the method of step S21 or S22 is executed; if no entity is retrieved, the filtering attributes (such as location, etc.) are deleted, and the method of step S21 or S22 is executed.

[0090] Step S21: When the number of entities is 1, for example, the graph information represented by the extracted structured text is: [Device] Wiper [Location] Front [Intent] Fault handling. According to the extracted information, a graph query statement can be automatically generated. Taking the Cypher graph query syntax as an example: "MATCH(v1:Device{name:\"Wiper\"})-[e1:Question and Answer]->(v2:Answer) WHERE v1.Device.Location == \"Front\" AND e1.Intent == \"Fault handling\" RETURN v2.Answer AS Answer;"; if no question-and-answer edge with the intent of "Fault handling" is retrieved under this entity, check whether there is a relevant edge in its parent node. The graph query statement is as follows: "MATCH(v1:Device{name:\"Wiper\"})-[e1:Has a parent node]->()->[e2:Question and Answer]->(v2:Answer) WHERE v1.Device.Location == \"Front\" AND e2.Intent == \"Fault handling\" RETURN v2.Answer AS Answer;".

[0091] Among them, the query of the edges connected to the parent node can be performed n times, and the value of n can be selected according to the actual business logic.

[0092] Step S22: When the number of entities is 2, for example, the graph information represented by the extracted structured text is: [Device] Seat [Function] Massage [Location] Second row [Intention] Opening method. According to the extracted information, since the directivity of the seat and the massage is unknown, two target query statements can be generated by combination. Taking the Cypher graph query syntax as an example: "MATCH (v1:Device{name:\"Seat\"})-[e1:Has]->(v2:Function{name:\"Massage\"})-[e2:Question and Answer]->(v2:Answer) WHERE v1.Device.Location == \"Second row\" AND e2.Intention == \"Opening method\" RETURN v3.Answer AS Answer;" and "MATCH (v1:Function{name:\"Massage\"})-[e1:Has]->(v2:Device{name:\"Seat\"})-[e2:Question and Answer]->(v2:Answer) WHERE v1.Device.Location == \"Second row\" AND e2.Intention == \"Opening method\" RETURN v3.Answer AS Answer;". Since the edge relationship between the seat and the massage is unidirectional, that is, the seat has the massage function. Therefore, only the first target query statement can take effect when the graph retrieval module retrieves in the graph database, and the reply answer for the opening method of the seat massage will be returned.

[0093] It should be noted that when no answer is retrieved after executing step S22, the graph query statement for the edges connected to the entity parent node can be automatically generated according to the method of step S21.

[0094] Step S3: For the candidate target query statements, the graph query module will encapsulate the candidate target query statements into message requests to sequentially query the graph database, return the target answer if the target answer is found, and continue to execute the next statement if the target answer is not found.

[0095] Specifically, as Figure 2c shown, after multiple target query statements are generated, the retrieval steps can be executed by the graph retrieval module. Specifically, for the list of candidate target query statements generated by the graph reasoning module, that is, multiple target query statements, the graph retrieval module encapsulates them into message requests and sequentially queries the graph database.

[0096] Step S4: Return the retrieved target answer reply to the user.

[0097] It should be noted that when the corresponding target answer is not queried, it is possible not to reply to the user, that is, not to send a reply message to the user terminal, or to reply to the user with a preset answer, that is, to send a preset answer to the user terminal.

[0098] The application embodiments of the present disclosure have the following advantages:

[0099] (1) Through the pre-established generative model, the recognition of entities and intents is achieved in a single module, which can improve the accuracy generalization ability of the recognition of both;

[0100] (2) A large number of training samples can be generated with the help of the multi-dimensional structured data of the graph, with a low dependence on manual annotation, improving the effect of the graph information generation model;

[0101] (3) By using entities to recall questions, the recall range is narrowed, and the recall ability for candidate questions is improved.

[0102] Corresponding to the above data query method, the present invention also proposes a data query device. Since the device embodiments of the present invention correspond to the above method embodiments, for the details not disclosed in the device embodiments, reference may be made to the above method embodiments, and no further elaboration will be made in the present invention.

[0103] Figure 4 FIG. is a schematic structural diagram of a data query device provided by an embodiment of the present disclosure, as Figure 4 shown, including:

[0104] An acquisition module 410, configured to acquire a question statement;

[0105] An identification module 420, configured to identify the graph information in the question statement according to a pre-established generative model, and combine the identified graph information in a preset format to obtain graph information represented by structured text; the graph information includes at least entity information and intent information; the generative model includes a preset format for graph information combination;

[0106] A statement generation module 430, configured to fuse the graph information represented by structured text with graph query syntax to generate a target query statement corresponding to the question statement;

[0107] A query module 440, configured to query a graph database according to the target query statement to obtain a target answer matching the question statement.

[0108] In some embodiments of the present disclosure, the identification module 420 is configured to: identify the graph information in the question statement according to a pre-established generative model; pair the identified graph information to form a data group; fill the graph information in the data group with label information corresponding to different dimensions in the preset format to obtain graph information represented by structured text

[0109] In some embodiments of the present disclosure, the statement generation module 430 is configured to: determine the number of entities in the graph information of the structured text representation; if the number of entities is 1, fuse the graph information of the structured text representation containing one entity with the graph query syntax to generate at least one target query statement corresponding to the question statement; if the number of entities is at least 2, fuse the graph information of the structured text representation containing at least two entities with the graph query syntax to generate at least two target query statements corresponding to the question statement, and each target query statement in the at least two target query statements contains at least two entities, and the positions of the at least two entities in each target query statement are different.

[0110] In some embodiments of the present disclosure, after querying the graph database according to the target query statement, the query module 440 is configured to: when there is an answer matching the intent information of the question statement in the answers corresponding to the entities included in the target query statement, output the answer matching the intent information of the question statement as the target answer; when there is no answer corresponding to the intent information of the question statement in the answers corresponding to the entities included in the target query statement, determine whether there is a parent node for the entities included in the target query statement; if there is a parent node, query whether there is an answer matching the intent information of the question statement in the answers included in the parent node; if there is an answer matching the intent information of the question statement in the answers included in the parent node, output the answer matching the intent information of the question statement as the target answer; if there is no parent node, return a prompt message indicating that no answer matching the intent information of the question statement is found in the query result.

[0111] In some embodiments of the present disclosure, the query module 440 is configured to: fuse the information of the parent node into the target query statement based on the graph query syntax to generate a new target query statement; query the graph database according to the new target query statement to determine whether there is an answer matching the intent information of the question statement in the answers included in the parent node.

[0112] In some embodiments of the present disclosure, after querying the graph database according to the target query statement, the query module 440 is configured to: if the entities included in the target query statement are not found, determine whether the target query statement includes attribute information; if it includes attribute information, delete the attribute information and continue to query the graph database.

[0113] In some embodiments of the present disclosure, the query module 440 is configured to: delete the content corresponding to the attribute in the target query statement to construct a target query statement without attributes; execute the target query statement without attributes to continue querying the graph database.

[0114] It should be noted that the foregoing explanation of the method embodiments also applies to the device in this embodiment, with the same principle, and will not be further limited in this embodiment.

[0115] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0116] Figure 5 A schematic block diagram of an exemplary electronic device 500 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0117] As Figure 5 shown, the device 500 includes a computing unit 501 that can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded from a storage unit 508 into a RAM (Random Access Memory) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.

[0118] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0119] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, CPU (Central Processing Unit), GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the data query method. For example, in some embodiments, the data query method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the aforementioned data query method in any other suitable manner (e.g., by means of firmware).

[0120] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, FPGA (Field Programmable Gate Array), ASIC (Application-Specific Integrated Circuit), ASSP (Application Specific Standard Product), SOC (System On Chip), CPLD (Complex Programmable Logic Device), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0121] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes may be executed entirely on the machine, partially on the machine, executed partially on the machine as a stand-alone software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0122] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0123] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0124] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0125] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.

[0126] Among them, it should be noted that artificial intelligence is a discipline that studies enabling a computer to simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and there are both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0127] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0128] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A data query method, characterized in that, The method includes: Obtain a problem statement; Identify the graph information in the problem statement according to a pre-established generative model, and combine the identified graph information in a preset format to obtain graph information represented by structured text; the graph information at least includes entity information and intent information; the generative model includes a preset format for graph information combination; Fuse the graph information represented by the structured text with graph query syntax to generate a target query statement corresponding to the problem statement; Query a graph database according to the target query statement to obtain a target answer that matches the problem statement.

2. The method according to claim 1, characterized in that, The step of identifying the graph information in the problem statement according to a pre-established generative model and combining the identified graph information in a preset format to obtain graph information represented by structured text includes: Identify the graph information in the problem statement according to a pre-established generative model; Pair the identified graph information to form data groups; Fill the graph information in the data groups with information according to the label information corresponding to different dimensions in the preset format to obtain graph information represented by structured text.

3. The method according to claim 1, characterized in that, The step of fusing the graph information represented by the structured text with graph query syntax to generate a target query statement corresponding to the problem statement includes: Determine the number of entities in the graph information represented by the structured text; If the number of entities is 1, fuse the graph information represented by the structured text containing one such entity with graph query syntax to generate at least one target query statement corresponding to the problem statement; If the number of entities is at least 2, fuse the graph information represented by the structured text containing at least two such entities with graph query syntax to generate at least two target query statements corresponding to the problem statement, and each target query statement in the at least two target query statements contains the at least two entities, and the positions of the at least two entities are different in each target query statement.

4. The method according to claim 1, characterized in that, After querying the graph database according to the target query statement, it further includes: When there is an answer that matches the intent information of the problem statement among the answers corresponding to the entities included in the target query statement, output the answer that matches the intent information of the problem statement as the target answer; When there is no answer corresponding to the intent information of the problem statement among the answers corresponding to the entities included in the target query statement, determine whether there is a parent node for the entity included in the target query statement; If there is a parent node, query whether there is an answer that matches the intent information of the problem statement among the answers included in the parent node; if there is an answer that matches the intent information of the problem statement among the answers included in the parent node, output the answer that matches the intent information of the problem statement as the target answer; If there is no parent node, return a prompt message indicating that no answer matching the intent information of the problem statement is found in the query result.

5. The method according to claim 4, characterized in that, The step of querying whether there is an answer that matches the intent information of the problem statement among the answers included in the parent node includes: Based on the graph query syntax, fuse the information of the parent node into the target query statement to generate a new target query statement; Query the graph database according to the new target query statement to determine whether there is an answer in the answers included in the parent node that matches the intent information of the question statement.

6. The method according to claim 4 or 5, characterized in that, After querying the graph database according to the target query statement, it further includes: If the entity included in the target query statement is not queried, determine whether the target query statement includes attribute information; If it includes attribute information, delete the attribute information and continue to query the graph database.

7. The method according to claim 6, characterized in that, The deleting the attribute information and continuing to query the graph database includes: Delete the content corresponding to the attribute in the target query statement to construct a target query statement without attributes; Execute the target query statement without attributes to continue querying the graph database.

8. A data query device, characterized in that, It includes: An acquisition module, configured to acquire a question statement; An identification module, configured to identify the graph information in the question statement according to a pre-established generative model, and combine the identified graph information in a preset format to obtain graph information represented by structured text; The graph information at least includes entity information and intent information; The generative model includes a preset format for graph information combination; A statement generation module, configured to fuse the graph information represented by the structured text with the graph query syntax to generate a target query statement corresponding to the question statement; A query module, configured to query the graph database according to the target query statement to obtain a target answer matching the question statement.

9. An electronic device, characterized in that, It includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Question and answer method and device, computer equipment and program product

    CN120429412A