Data query method and device, electronic equipment and computer readable medium

By using knowledge graphs and model output prompt technology in the data query system, identifying user query intentions and generating accurate query statements, the problem of low data query accuracy in the existing technology is solved, and higher query accuracy is achieved.

CN120216535APending Publication Date: 2025-06-27JINGDONG ALLIANZ GENERAL INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311803312.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, when processing user data query services, the data query accuracy is low, especially when the user does not bring in the relevant table description information, it is easy to lead to incorrect data recall.

Method used

In response to user query statements, intent recognition is performed to extract field name information, matched field entities are positioned from the pre-constructed knowledge graph, and corresponding table entities are found based on relationships, and model output prompts are constructed to generate more accurate query statements.

Benefits of technology

It improves the accuracy of data queries, ensures that when processing user data query services, the required data can be obtained more accurately and reduces error recalls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216535A_ABST
    Figure CN120216535A_ABST
Patent Text Reader

Abstract

The invention discloses a data query method and device, electronic equipment and a computer readable medium, and relates to the technical field of computers.The specific implementation mode comprises the steps that in response to an obtained user query statement, intention recognition is conducted on the user query statement to extract field name information; positioning a field entity matched with the field name information from a pre-constructed knowledge graph, searching a table entity corresponding to the field entity based on a relationship in the pre-constructed knowledge graph, and constructing a model output prompt based on the table entity; generating a query statement based on the model output prompt; and executing the query statement to obtain query result data. The more accurate query statement can be generated by constructing the model output prompt, and the accuracy of data query can be improved when the data query service of the user is processed by executing the generated query statement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a data query method, apparatus, electronic device, and computer-readable medium. Background Art

[0002] Currently, when constructing the direct input (Prompt) of a large model for converting natural language text (text) into structured query language SQL (i.e., constructing text2SQL), there are two existing approaches: One is to use all tables and their fields as the context, and then construct the Prompt in combination with the user's query requirements. Limited by the context length of the large model, this method is often not applicable in production, and moreover, too many irrelevant tables and fields will reduce the retrieval performance of the large model; the other is to use only the relevant tables and relevant fields as the context based on the user's query requirements, and then construct the Prompt in combination with the user's query requirements. However, when the user queries, they often do not bring in the description information of the relevant tables. For example: Please help me query what types of goods are in XX warehouse. When directly recalling the tables related to the user's query at this time, it is easy to make mistakes. When processing the user's data query service, the data query accuracy is low. Summary of the Invention

[0003] In view of this, embodiments of the present application provide a data query method, apparatus, electronic device, and computer-readable medium, which can solve the technical problem of low data query accuracy when processing the user's data query service.

[0004] To achieve the above object, according to one aspect of the embodiments of the present application, a data query method is provided, including:

[0005] In response to obtaining a user query statement, perform intent recognition on the user query statement to extract field name information;

[0006] Locate a field entity matching the field name information from a pre-constructed knowledge graph, and then based on the relationships in the pre-constructed knowledge graph, find a table entity corresponding to the field entity, and construct a model output prompt based on the table entity;

[0007] Generate a query statement based on the model output prompt;

[0008] Execute the query statement to obtain query result data.

[0009] Optionally, before locating a field entity matching the field name information from a pre-constructed knowledge graph, the data query method further includes:

[0010] Obtain a preset business database identifier, and then obtain corresponding entity relationship data based on the business database identifier;

[0011] Construct entity triples and relationship triples based on entity relationship data, and then pre-construct a knowledge graph based on the entity triples and relationship triples.

[0012] Optionally, constructing entity triples and relationship triples based on entity relationship data includes:

[0013] Extract the table entity, field entity, data type, and remarks from the entity relationship data, and then construct table entity triples based on the table entity and the corresponding remarks, and construct field entity triples based on the field entity, the corresponding data type, and the corresponding remarks;

[0014] Extract the relationship between the table entity and the field entity in the entity relationship data, and then construct relationship triples based on the field entity, the table entity, and the relationship.

[0015] Optionally, pre-constructing a knowledge graph includes:

[0016] Set the label of the field node to the field entity corresponding to the entity triple;

[0017] Set the label of the table node to the table entity corresponding to the entity triple;

[0018] Pre-construct a knowledge graph based on the field node, the table node, the label of the field node, the label of the table node, and the relationship in the relationship triple.

[0019] Optionally, locate the field entity matching the field name information in the pre-constructed knowledge graph, including:

[0020] Convert the remarks of the field entity in the pre-constructed knowledge graph into corresponding remark vectors;

[0021] Convert the field name information into a corresponding field name information vector;

[0022] Calculate the similarity between the field name information vector and the remark vector, determine the target remark vector according to the similarity, and then determine the field entity corresponding to the target remark vector as the field entity matching the field name information in the pre-constructed knowledge graph.

[0023] Optionally, construct a model output prompt based on the table entity, including:

[0024] Determine the number of different table entities in the table entity. If the number exceeds the preset threshold, obtain the corresponding foreign key fields;

[0025] If the field entity does not exactly correspond to the foreign key field, traverse the table corresponding to the table entity, and judge whether other tables in the table except the currently traversed table can cover the field entity. If so, delete the currently traversed table. If not, retain the currently traversed table, and then update the table entity;

[0026] Construct model output prompts based on the updated table entity.

[0027] Optionally, construct model output prompts based on the table entity, including:

[0028] If the field entity exactly corresponds to a foreign key field, generate confirmation information based on the table corresponding to the table entity and the corresponding remarks and display it to the user, and determine the target table entity based on the user's selection operation;

[0029] Construct model output prompts based on the target table entity.

[0030] Optionally, construct model output prompts based on the target table entity, including:

[0031] Generate context based on the target table entity and the field entity;

[0032] Construct model output prompts according to the preset instructions, context, user query statement, and preset output guidelines.

[0033] Optionally, generate query statements, including:

[0034] Determine the scenario identifier according to the model output prompt;

[0035] Generate query statements based on the scenario identifier, preset inference path, and deviation verification mechanism.

[0036] Optionally, execute the query statement to obtain query result data, including:

[0037] Execute each query statement to obtain each query result data;

[0038] Perform consistency verification on each query result data, and return the query result data with the highest consistency as the execution result data.

[0039] Optionally, after obtaining the query result data, the method further includes:

[0040] Obtain user requirement data, and determine the data display method based on the user requirement data;

[0041] Convert the execution result data into the corresponding data form according to the data display method and output it.

[0042] In addition, the present application further provides a data query device, including:

[0043] An information extraction unit configured to, in response to obtaining a user query statement, perform intent recognition on the user query statement to extract field name information;

[0044] A model output prompt construction unit, configured to locate a field entity matching the field name information from a pre-constructed knowledge graph, and then find a table entity corresponding to the field entity based on the relationships in the pre-constructed knowledge graph, and construct a model output prompt based on the table entity;

[0045] A query statement generation unit, configured to generate a query statement based on the model output prompt;

[0046] An execution unit, configured to execute the query statement to obtain query result data.

[0047] Optionally, the data query device further includes a knowledge graph construction unit, configured to:

[0048] Obtain a preset business database identifier, and then obtain corresponding entity relationship data based on the business database identifier;

[0049] According to the entity relationship data, construct entity triples and relationship triples, and then pre-construct a knowledge graph based on the entity triples and relationship triples.

[0050] The knowledge graph construction unit is further configured to:

[0051] Extract table entities, field entities, data types, and remarks from the entity relationship data, and then construct table entity triples based on the table entities and corresponding remarks, and construct field entity triples based on the field entities, corresponding data types, and corresponding remarks;

[0052] Extract the relationship between the table entity and the field entity in the entity relationship data, and then construct a relationship triple based on the field entity, table entity, and relationship.

[0053] The knowledge graph construction unit is further configured to:

[0054] Set the label of the field node to the field entity corresponding to the entity triple;

[0055] Set the label of the table node to the table entity corresponding to the entity triple;

[0056] Pre-construct a knowledge graph based on the field node, table node, label of the field node, label of the table node, and the relationship in the relationship triple.

[0057] The model output prompt construction unit is further configured to:

[0058] Convert the remarks of the field entity in the pre-constructed knowledge graph into corresponding remark vectors;

[0059] Convert the field name information into a corresponding field name information vector;

[0060] Calculate the similarity between the calculation field name information vector and the note vector, determine the target note vector according to the similarity, and then determine the field entity corresponding to the target note vector as the field entity in the pre-constructed knowledge graph that matches the field name information.

[0061] The model output prompt construction unit is further configured to:

[0062] Determine the number of different table entities in the table entity. If the number exceeds the preset threshold, obtain the corresponding foreign key field;

[0063] If the field entity does not exactly correspond to the foreign key field, traverse the table corresponding to the table entity, and determine whether other tables in the table except the currently traversed table can cover the field entity. If so, delete the currently traversed table; if not, retain the currently traversed table, and then update the table entity.

[0064] Construct a model output prompt based on the updated table entity.

[0065] The model output prompt construction unit is further configured to:

[0066] If the field entity exactly corresponds to the foreign key field, generate the information to be confirmed based on the table corresponding to the table entity and the corresponding note, display it to the user, and determine the target table entity based on the user's selection operation;

[0067] Construct a model output prompt based on the target table entity.

[0068] The model output prompt construction unit is further configured to:

[0069] Generate context based on the target table entity and the field entity;

[0070] Construct a model output prompt according to the preset instructions, context, user query statement, and preset output guidelines.

[0071] The query statement generation unit is further configured to:

[0072] Determine the scenario identifier according to the model output prompt;

[0073] Generate a query statement based on the scenario identifier, preset inference path, and deviation verification mechanism.

[0074] The execution unit is further configured to:

[0075] Execute each query statement to obtain each query result data;

[0076] Perform consistency verification on each query result data, and return the query result data with the highest consistency as the execution result data.

[0077] The data query device further includes a data display unit, which is configured to:

[0078] Obtain user requirement data and determine a data display mode based on the user requirement data;

[0079] Convert the execution result data into a corresponding data form according to the data display mode and output it.

[0080] In addition, the present application also provides a data analysis electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the data query method as described above.

[0081] In addition, the present application also provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the data query method as described above is implemented.

[0082] One embodiment of the above invention has the following advantages or beneficial effects: In the present application, in response to obtaining a user query statement, the user query statement is subjected to intention recognition to extract field name information; a field entity matching the field name information is located from a pre-constructed knowledge graph, and then a table entity corresponding to the field entity is found based on the relationship in the pre-constructed knowledge graph, and a prompt is output based on the table entity to construct a model; a query statement is generated based on the prompt output by the model; the query statement is executed to obtain query result data. By constructing a model to output a prompt, a more accurate query statement can be generated, and by executing the generated query statement, the accuracy of data query can be improved when processing the user's data query service.

[0083] The further effects of the above non-conventional optional manner will be described in combination with specific embodiments below. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] The drawings are used to better understand the present application and do not constitute an improper limitation to the present application. Among them:

[0085] Figure 1 is a schematic diagram of the main process of the data query method provided by an embodiment of the present application;

[0086] Figure 2 is a schematic diagram of the main process of constructing a knowledge graph required for the data query method provided by an embodiment of the present application;

[0087] Figure 3 is a schematic diagram of the main process of the data query method provided by an embodiment of the present application;

[0088] Figure 4It is a schematic diagram of a pre - constructed knowledge graph for the data query method provided by an embodiment of the present application;

[0089] Figure 5 It is a schematic diagram of the main units of the data query device according to an embodiment of the present application;

[0090] Figure 6 It is an exemplary system architecture diagram to which the embodiments of the present application can be applied;

[0091] Figure 7 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing the embodiments of the present application. Detailed implementation manners

[0092] The following makes an explanation of the exemplary embodiments of the present application with reference to the accompanying drawings. Various details of the embodiments of the present application are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, the description of well - known functions and structures is omitted below. It should be noted that in the technical solutions of the present disclosure, in terms of the collection, gathering, updating, analysis, processing, use, transmission, storage, etc. of user personal information, they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken for user personal information to prevent illegal access to user personal information data and to maintain the security of user personal information, network security, and national security.

[0093] Figure 1 It is a schematic diagram of the main process of the data query method provided by an embodiment of the present application. As Figure 1 shown, the data query method includes:

[0094] Step S101, in response to obtaining a user query statement, perform intent recognition on the user query statement to extract field name information.

[0095] In this embodiment, the execution subject of the data query method may be a large model-based data analysis system, abbreviated as Data assistant. When the execution subject detects a user input query statement, it can obtain the input user query statement and perform intent recognition on the user query statement. The user query statement can be, for example, "Please help me query what are the names of the goods with a purchase price less than 1000 and the buyer being XXX company". When performing intent recognition on the user query statement, the extracted field name information can be purchase price, buyer, and goods name. The execution subject can call the large model to perform intent recognition on the user query statement and extract the field name information therein. Among them, in the field of artificial intelligence, a large model refers to a deep neural network with more than 1 billion parameters, which can process massive amounts of data and complete various complex tasks, such as natural language processing, computer vision, speech recognition, etc.

[0096] Step S102, locate the field entity in the pre-constructed knowledge graph that matches the field name information, and then based on the relationships in the pre-constructed knowledge graph, find the table entity corresponding to the field entity, and construct a model output prompt based on the table entity.

[0097] The constructed model output prompt is the Prompt.

[0098] Specifically, locating the field entity in the pre-constructed knowledge graph that matches the field name information includes: converting the remarks of the field entities in the pre-constructed knowledge graph into corresponding remark vectors; converting the field name information into corresponding field name information vectors; calculating the similarity between the field name information vectors and the remark vectors, and determining the target remark vector according to the similarity, and then determining the field entity corresponding to the target remark vector as the field entity in the pre-constructed knowledge graph that matches the field name information.

[0099] Specifically, the remark vector corresponding to when the similarity is greater than the preset threshold can be determined as the target remark vector. The target remark vector is used to represent the remark vector with a relatively high similarity to the field name information corresponding to the user query statement in the pre-constructed knowledge graph. Find the corresponding field entity in the pre-constructed knowledge graph according to the target remark vector in the pre-constructed knowledge graph. Then, find the table entity corresponding to the field entity in the pre-constructed knowledge graph to construct the Prompt (that is, the direct input data of the large model) based on the table entity.

[0100] Exemplarily, in the pre-constructed knowledge graph, in addition to the two field entities of purchase price and goods name, there is also the field entity of "purchasing party name". "Purchasing party name" and "buyer" are synonyms, and the Euclidean distance between semantic vectors is the shortest. Through entity disambiguation, it can be determined that the field entity in the pre-constructed knowledge graph that matches the "buyer" in the field name information of the user query statement is "purchasing party name". Based on the field entities of "purchase price", "goods name", and "purchasing party name" found in the pre-constructed knowledge graph, find the corresponding table entities with subordinate relationships from the pre-constructed knowledge graph. And construct a model output prompt Prompt based on the found table entities.

[0101] Specifically, constructing a model output prompt based on table entities includes: determining the number of different table entities in the table entities. If the number exceeds the preset threshold, obtain the corresponding foreign key fields; if the field entities do not exactly correspond to the foreign key fields, traverse the table corresponding to the table entity, and judge whether other tables in the table except the currently traversed table can cover the field entities. If so, delete the currently traversed table, if not, retain the currently traversed table, and then update the table entities; construct a model output prompt based on the updated table entities.

[0102] Exemplarily, if there are multiple different table entities found according to the field entities corresponding to the target note vector (that is, the number exceeds the preset threshold, and the preset threshold can be 1, 2, 3, etc. The embodiments of the present application do not make specific limitations on the preset threshold), then the foreign keys between these multiple table entities need to be recalled additionally. If there are multiple different table entities found according to the field entities corresponding to the target note vector, the self-evaluation mechanism is triggered. Exemplarily, the self-evaluation mechanism may include:

[0103] The first step: Based on the pre-constructed knowledge graph, judge whether all the fields involved in the user query statement are foreign keys (the degree of the foreign key field entity is greater than 1). If they are not all foreign keys (that is, the common keywords of multiple tables), then jump to the second step. If they are all foreign keys, guide the user to confirm the table to be queried. For example, when the user queries all goods names, since the goods name field is a foreign key, it will be located to the goods information table, inventory information table, and purchase information table. At this time, the self-evaluation Self-Evaluate mechanism will give the relevant tables and note information to the user, and the user will assist in determining the table to be queried.

[0104] The second step: Loop through the tables corresponding to the table entities. If the remaining tables except the currently traversed table can cover all the field information included in the user query statement, then delete the currently traversed table, otherwise retain the currently traversed table.

[0105] Specifically, constructing a model output prompt based on the table entity includes: if the field entity exactly corresponds to the foreign key field, generating confirmation information based on the table corresponding to the table entity and the corresponding remarks and presenting it to the user, and determining the target table entity based on the user's selection operation (i.e., assisting the user to determine the table to be queried); constructing a model output prompt based on the target table entity.

[0106] The confirmation information can be a list of confirmation information composed of the table corresponding to the table entity and the corresponding remarks. Presenting this list of confirmation information to the user for the user to select and determine the target table entity. Constructing a model output prompt Prompt based on the target table entity.

[0107] Specifically, constructing a model output prompt based on the target table entity includes: generating context based on the target table entity and the field entity; constructing a model output prompt according to the preset instruction, context, user query statement, and preset output guidance.

[0108] Exemplarily, the model output prompt Prompt consists of four parts: instruction, context, user input, and output guidance. Among them, the instruction is used to indicate the role and task played by the large model. For example: You are a text-to-SQL generator, and your main goal is to assist users as much as possible in converting the input text into correct SQL statements. The context is composed of the recalled relevant tables and relevant fields. The user input corresponds to the user query statement. The output guidance (the data type and data format of the output) is used to indicate that the large model should not output any information other than the SQL statement.

[0109] Exemplarily, the constructed model output prompt Prompt is as follows:

[0110] """

[0111] Instruction: You are a professional text-to-SQL generator,and your task is toassist users in converting their input text into correct SQL statements asmuch as possible.

[0112] Context: Context begins

[0113] The table names and table fields come from the following tables:

[0114] Table 1: cargo_info (Cargo Information Table)

[0115] Field 1: cargo_name (string, name of goods), Field 2: price (float, purchase price)

[0116] Table 2: purchase_info (purchase information table)

[0117] Field 1: cargo_name (string, name of goods), Field 2: company_name (string, name of the purchaser)

[0118] Context ends

[0119] User input: input: Please help me query what are the names of the goods whose purchase price is less than 1000 and the name of the purchaser is XXX Company.

[0120] Output instruction: Please generate the correct SQL statement without outputting irrelevant information.

[0121] """

[0122] Step S103, generate a query statement based on the model output prompt.

[0123] Specifically, generating a query statement includes: determining a scenario identifier according to the model output prompt; generating a query statement based on the scenario identifier, a preset inference path, and a deviation verification mechanism.

[0124] The execution entity can determine the scenario identifier based on the context and user input in the model output prompt. The scenario identifier is used to represent the scenario to which the query statement applies. For example, the scenario identifier can correspond to only the sorting scenario. The execution entity generates multiple SQL statements based on different preset inference paths. Due to the inherent deviation awareness of the large model, when generating SQL statements, a deviation verification mechanism is introduced in the way of historical conversations. The deviation verification mechanism consists of two parts. One is not to provide additional quantity information unless necessary. For example, when the scenario identifier corresponds to only the sorting scenario, prompt the large model not to include COUNT(*) information in the SELECT statement. The other is to prompt the large model not to misuse keywords such as LEFT JOIN, OR, and IN. Then, multiple SQL query statements are generated based on different preset inference paths.

[0125] Step S104, execute the query statement to obtain the query result data.

[0126] Specifically, execute query statements to obtain query result data, including: execute each query statement to obtain each query result data; perform consistency verification based on each query result data, and return the query result data with the highest consistency as the execution result data.

[0127] The execution entity triggers the Self-Consistency mechanism for consistency verification, executes the SQL query statements generated by different inference paths, and from the correctly executed results, with the help of the deepdiff library in Python (used for comparing differences in tabular data), uses the voting principle to determine the result with the highest consistency as the final result. Consistency verification is to perform consistency verification on the results of the SQL query statements generated by different inference paths of the large model when executed in the database, and return the result with the highest consistency, which can improve the accuracy of the obtained query result data.

[0128] Specifically, after obtaining the query result data, the data query method further includes: obtaining user requirement data, determining the data display method based on the user requirement data; converting the execution result data into the corresponding data form according to the data display method and outputting it.

[0129] The execution entity can convert the table result into a bar chart, a pie chart, or a line chart for output based on the user's requirements. By constructing a model output prompt, more accurate query statements can be generated, and by executing the generated query statements, the accuracy of data query can be improved when processing the user's data query service.

[0130] In this embodiment, in response to obtaining a user query statement, the intent of the user query statement is recognized to extract field name information; the field entity matching the field name information is located from the pre-constructed knowledge graph, and then the table entity corresponding to the field entity is found based on the relationships in the pre-constructed knowledge graph, and a model output prompt is constructed based on the table entity; based on the model output prompt, a query statement is generated; the query statement is executed to obtain query result data. By constructing a model output prompt, more accurate query statements can be generated, and by executing the generated query statements, the accuracy of data query can be improved when processing the user's data query service.

[0131] Figure 2 It is a schematic diagram of the main process of constructing the knowledge graph required for the data query method provided by an embodiment of the present application. As Figure 2 shown, the method for constructing the knowledge graph required for the data query method of the embodiment of the present application includes:

[0132] Step S201, obtain the preset business database identifier, and then obtain the corresponding entity relationship data based on the business database identifier.

[0133] In the embodiment of the present application, text2SQl: Convert natural language into SQL query statements. Self-Evaluate mechanism: Self-evaluation mechanism. ER diagram: The full name is Entity Relationship Diagram, which provides a method to represent entity types, attributes, and relationships, and is used to describe the conceptual model of the real world. Prompt: The model output prompt, that is, the input of the large model. Self-Consistency mechanism: Perform consistency verification based on the voting principle. embedding model: An embedding model that converts text into vectors.

[0134] Before performing data query, the execution entity can obtain the preset business database identifier, that is, obtain the business database name or number during the design of the business database. And obtain the entity relationship data during the design of the business database through the business database identifier. Exemplarily, the entity relationship data can be the entity relationship table or entity relationship diagram (i.e., ER diagram) during the design of the business database. Exemplarily, the entity relationship diagram can include entities, attributes, and relationships. Entities, for example, can include fields and tables. Relationships are the relationships between entities, such as the relationship between fields and tables.

[0135] Step S202: Extract the table entities, field entities, data types, and remarks in the entity relationship data, and then construct table entity triples based on the table entities and the corresponding remarks, and construct field entity triples based on the field entities, the corresponding data types, and the corresponding remarks.

[0136] The execution entity can construct two types of triples according to the entity relationship data. The first represents entity attributes: There are two types of entities in total, fields and tables. For fields, there are data type and remark attributes, and for tables, there is a remark attribute. Exemplarily, the specific form of the entity triple is: (field1, data type, string), (field1, remark, name), (table1, remark, cargo information table); The second represents the relationship between tables and fields.

[0137] Exemplarily, the table entity triple is, for example, (table1, remark, cargo information table), and the field entity triples are, for example, (field1, data type, string), (field1, remark, name). The specific form of the relationship triple is: (field1, belongs to, table1).

[0138] Exemplarily, the entity relationship data is, for example:

[0139] Table 1: cargo_info (cargo information table) Field 1: cargo_name (string, cargo name), Field 2: orgin_cargo (float, cargo origin), Field 3: price (float, purchase price);

[0140] Table 2: check_info (Inventory Information Table) Field 1: cargo_name (string, name of goods), Field 2: Person_name (string, name of the inventory taker);

[0141] Table 3: purchase_info (Purchase Information Table) Field 1: cargo_name (string, name of goods), Field 2: company_name (string, name of the purchaser), Field 3: purchase_nums (int, quantity purchased).

[0142] Extract the table entities, field entities, data types, and remarks in the entity relationship data, and perform the following triple construction:

[0143] Construct table entity triples based on the table entities and their corresponding remarks:

[0144] (cargo_info, remarks, Goods Information Table), (check_info, remarks, Inventory Information Table), (purchase_info, remarks, Purchase Information Table)

[0145] Construct field entity triples based on the field entities, their corresponding data types, and their corresponding remarks:

[0146] (cargo_name, data type, string), (cargo_name, remarks, name of goods), (orgin_cargo, data type, float), (orgin_cargo, remarks, place of origin of goods), (price, data type, float), (price, remarks, purchase price), (Person_name, data type, string), (Person_name, remarks, name of the inventory taker), (company_name, data type, string), (company_name, remarks, name of the purchaser), (purchase_nums, data type, int), (purchase_nums, remarks, quantity purchased)

[0147] Step S203: Extract the relationships between the table entities and field entities in the entity relationship data, and then construct relationship triples based on the field entities, table entities, and relationships.

[0148] For example, for a field entity such as Field 1 and a table entity such as Table 1, the relationship between the table entity and the field entity can be that Field 1 belongs to Table 1, and thus the resulting relationship triple can be (Field 1, belongs to, Table 1).

[0149] Example, relationship triples constructed based on field entities, table entities, and relationships:

[0150] (cargo_name, belongs to, cargo_info), (orgin_cargo, belongs to, cargo_info), (price, belongs to, cargo_info), (cargo_name, belongs to, check_info), (Person_name, belongs to, check_info), (cargo_name, belongs to, purchase_info), (company_name, belongs to, purchase_info), (purchase_nums, belongs to, purchase_info)

[0151] Step S204, pre-construct a knowledge graph based on entity triples and relationship triples.

[0152] The executing entity constructs entity triples and relationship triples based on entity relationship data, and then pre-constructs a knowledge graph based on the entity triples and relationship triples.

[0153] Specifically, pre-constructing the knowledge graph includes: setting the label of the field node as the field entity corresponding to the entity triple; setting the label of the table node as the table entity corresponding to the entity triple; pre-constructing the knowledge graph based on the field node, table node, the label of the field node, the label of the table node, and the relationship in the relationship triple.

[0154] Example, as Figure 4 shown, pre-construct a knowledge graph based on the triples, set the label of all field nodes as field entities, and the label of the table nodes as table entities. Connect the field entities and table entities through the corresponding relationships (such as connecting with a line and marking "belongs to"), and add the corresponding remarks and data types to each field node, thereby generating a pre-constructed knowledge graph composed of fields, tables, remarks, and data types. By pre-constructing the knowledge graph, the subordination relationship between fields and tables can be made more explicit, facilitating accurate and rapid data analysis based on the constructed pre-constructed knowledge graph.

[0155] Example, as Figure 4 shown, the content included in each table node in the pre-constructed knowledge graph can be:

[0156] check_info Remarks: Inventory information table, cargo_info Remarks: Goods information table, purchase_info Remarks: Purchase information table.

[0157] In the pre - constructed knowledge graph, the field node "price" connected to the table node "cargo_info" Note: purchase price, data type: float, the field node "orgin_cargo" connected to "cargo_info" Note: place of origin of goods, data type: float, the field node "cargo_name" connected to "cargo_info" Note: name of goods, data type: string;

[0158] The field node "Person_name" connected to the table node "check_info" Note: name of the checker, data type: string, the field node "cargo_name" connected to the table node "check_info" Note: name of goods, data type: string;

[0159] The field node "company_name" connected to the table node "purchase_info" Note: name of the purchaser, data type: string, the field node "purchase_nums" connected to the table node "purchase_info" Note: purchase quantity, data type: int. Figure 4 In the figure, the ellipse: table entity, the rectangle: field entity.

[0160] By pre - constructing the knowledge graph, the subordination relationship between fields and tables can be made more explicit, facilitating accurate and rapid data query based on the pre - constructed knowledge graph.

[0161] Figure 3 It is the main process schematic diagram of the data query method provided by an embodiment of this application. The data query method of the embodiment of this application is applied to the data analysis scenario based on the large model. As Figure 3 shown:

[0162] Step A: The execution subject constructs the knowledge graph of the table schema.

[0163] The first step: Construct triples. Based on the entity - relationship diagram during the design of the business database, that is, the ER diagram, construct two types of triples. The first type represents entity attributes: There are two types of entities in total, fields and tables. For fields, there are data type and note attributes, and for tables, there is a note attribute. The specific form is, (field1, data type, string), (field1, note, name), (table1, note, goods information table); The second type represents the relationship between tables and fields: (field1, belongs to, table1). For example, if there are the following three tables in the database:

[0164] Table 1: cargo_info (Cargo Information Table) Field 1: cargo_name (string, Cargo Name), Field 2: orgin_cargo (float, Cargo Origin), Field 3: price (float, Purchase Price)

[0165] Table 2: check_info (Inventory Check Information Table) Field 1: cargo_name (string, Cargo Name), Field 2: Person_name (string, Inventory Checker's Name)

[0166] Table 3: purchase_info (Purchase Information Table) Field 1: cargo_name (string, Cargo Name), Field 2: company_name (string, Purchaser's Name), Field 3: purchase_nums (int, Purchase Quantity)

[0167] Based on the above table information, the following triples can be converted:

[0168] Table Entities: (cargo_info, Remarks, Cargo Information Table), (check_info, Remarks, Inventory Check Information Table), (purchase_info, Remarks, Purchase Information Table)

[0169] Field Entities: (cargo_name, Data Type, string), (cargo_name, Remarks, Cargo Name), (orgin_cargo, Data Type, float), (orgin_cargo, Remarks, Cargo Origin), (price, Data Type, float), (price, Remarks, Purchase Price), (Person_name, Data Type, string), (Person_name, Remarks, Inventory Checker's Name), (company_name, Data Type, string), (company_name, Remarks, Purchaser's Name), (purchase_nums, Data Type, int), (purchase_nums, Remarks, Purchase Quantity)

[0170] Relationships: (cargo_name, Belongs To, cargo_info), (orgin_cargo, Belongs To, cargo_info), (price, Belongs To, cargo_info),

[0171] (cargo_name, belongs to, check_info), (Person_name, belongs to, check_info), (cargo_name, belongs to, purchase_info), (company_name, belongs to, purchase_info), (purchase_nums, belongs to, purchase_info)

[0172] Step 2: Pre-construct a knowledge graph based on the triples, and set the label of all field nodes to field entities and the label of table nodes to table entities.

[0173] Step B: The execution entity uses a large model to perform intent recognition on the user's query statement and extracts the field name information therein. For example, for the query statement: Please help me query what are the names of the goods whose purchase price is less than 1000 and the buyer is Company XXX, the field name information obtained by intent recognition is: purchase price, buyer, and goods name.

[0174] Step C: The execution entity performs minimum hint recall based on the pre-constructed knowledge graph, that is, recalls relevant tables and relevant fields. First, entity disambiguation needs to be performed on the fields. The specific method is to convert the field information extracted from the user's query statement and the field entity note information in the pre-constructed knowledge graph into vectors. The vectors have semantic information, which can solve the problem of synonyms and align two entities with the same meaning but different expressions. Then, calculate the similarity between the vectors, locate the specific field entity in the pre-constructed knowledge graph, and then locate the corresponding table based on the relationship query. If multiple tables are located, the foreign keys between these multiple tables also need to be recalled additionally. For example, for the query statement in Step B:

[0175] First: Convert the field names obtained by intent recognition into vectors using the word embedding model.

[0176] Second: Retrieve the vector library storing the vectors of the field entity note information, and recall the corresponding field entity notes based on the Euclidean distance metric. The result is: purchase price, name of the purchasing party, and goods name. And standardize the field names in the user's query statement: Please help me query what are the names of the goods whose purchase price is less than 1000 and the name of the purchasing party is Company XXX. Note: Since the buyer and the name of the purchasing party are synonyms, the Euclidean distance between the semantic vectors is the shortest, achieving entity disambiguation.

[0177] Next: Based on the field entity note information recalled in the previous step, locate the corresponding field entities in the pre-constructed knowledge graph: [{"price": {"data type": float, "note": "purchase price"}}, {"company_name": {"data type": string, "note": "name of the purchasing party"}}, {"cargo_name": {"data type": string, "note": "name of the goods"}}]

[0178] Then: Based on the relationship query, locate the relevant table entities: [{"cargo_info": {"note": "goods information table", "field list": [{"price": {"data type": float, "note": "purchase price"}}, {"cargo_name": {"data type": string, "note": "name of the goods"}}]}}, {"check_info": {"note": "inventory check information table", "field list": [{"cargo_name": {"data type": string, "note": "name of the goods"}}]}}, {"purchase_info": {"note": "purchase information table", "field list": [{"company_name": {"data type": string, "note": "name of the purchasing party"}}, {"cargo_name": {"data type": string, "note": "name of the goods"}}]}}]

[0179] Finally: Since the foreign key between the multiple located tables is cargo_name, the final result is: [{"cargo_info": {"note": "goods information table", "field list": [{"price": {"data type": float, "note": "purchase price"}}, {"cargo_name": {"data type": string, "note": "name of the goods"}}]}}, {"check_info": {"note": "inventory check information table", "field list": [{"cargo_name": {"data type": string, "note": "name of the goods"}}]}}, {"purchase_info": {"note": "purchase information table", "field list": [{"company_name": {"data type": string, "note": "name of the purchasing party"}}, {"cargo_name": {"data type": string, "note": "name of the goods"}}]}}]

[0180] Step D: The execution entity triggers the self-evaluation mechanism, i.e., the Self-Evaluate mechanism. The self-evaluation mechanism has two functions. The first is self-correction, which is used to filter out redundant information, and the second is to guide troubleshooting. The specific approach is as follows:

[0181] The first step: Based on the pre-constructed knowledge graph, determine whether all the fields involved in the user's query statement are foreign keys (the degree of the foreign key field entity is greater than 1, and the foreign key is the common keyword between multiple tables). If not all are foreign keys, jump to the second step; if all are foreign keys, guide the user to confirm the table to be queried. For example, when the user queries all cargo names, since the cargo name field is a foreign key, it will be located to the cargo information table, inventory information table, and purchase information table. At this time, the Self-Evaluate mechanism will give the relevant tables and note information to the user, and the user will assist in determining the table to be queried.

[0182] The second step: Loop through the tables recalled in step C. If the remaining tables can cover all the field information included in the user's query statement, delete the table; otherwise, keep the table.

[0183] For example, for the user query statement in step B, since not all the fields involved are foreign keys, the second step in step D is executed. It is found that after deleting the inventory information table, the cargo information table and the purchase information table contain all the field information in the user's query statement. Therefore, the inventory information table is deleted. After executing the Self-Evaluate mechanism, the result is: [{"cargo_info": {"note": "Cargo Information Table", "field list": [{"price": {"data type": float, "note": "Purchase price"}}, {"cargo_name": {"data type": string, "note": "Cargo name"}}]}}, {"purchase_info": {"note": "Purchase Information Table", "field list": [{"company_name": {"data type": string, "note": "Purchasing party name"}}, {"cargo_name": {"data type": string, "note": "Cargo name"}}]}}].

[0184] Step E: The execution entity constructs a Prompt (the Prompt is the direct input to the large model. Constructing a reasonable Prompt can guide the large model to obtain the correct result desired by the user). The Prompt consists of four parts: instruction, context, user input, and output guidance. The instruction mainly tells the large model the role and task it plays. For example: You are a text-to-SQL generator, and your main goal is to assist users as much as possible in converting the input text into correct SQL statements. The context is composed of the recalled relevant tables and relevant fields. The user input corresponds to the query statement. The output guidance (the data type and data format of the output) mainly tells the large model not to output any information other than the SQL statement. For example, for the user query statement in Step B, after going through Step C and Step D, the constructed prompt is as follows:

[0185] """

[0186] You are a professional text-to-SQL generator,and your task is toassist users in converting their input text into correct SQL statements asmuch as possible.

[0187] Context begins

[0188] The table names and table fields come from the following tables:

[0189] Table 1: cargo_info (Cargo Information Table)

[0190] Field 1: cargo_name (string, Cargo Name), Field 2: price (float, Purchase Price)

[0191] Table 2: purchase_info (Purchase Information Table)

[0192] Field 1: cargo_name (string, Cargo Name), Field 2: company_name (string, Purchasing Party Name)

[0193] Context ends

[0194] input: Please help me query what the cargo names are for which the purchase price is less than 1000 and the purchasing party name is XXX Company.

[0195] Please generate the correct SQL statement without outputting irrelevant information.

[0196] """

[0197] Step F: The execution entity generates multiple SQL statements based on different preset inference paths. Due to the inherent bias perception of the large model, a bias verification mechanism is introduced in the way of historical conversations when generating SQL statements. The bias verification mechanism consists of two parts. One is not to provide extra quantity information unless necessary. For example, when the scenario identifier corresponds to the sorting-only scenario, prompt the large model not to include COUNT(*) information in the SELECT statement. The other is to prompt the large model not to misuse the LEFT JOIN, OR, and IN keywords. Then, multiple SQL query statements are generated based on different preset inference paths.

[0198] For example, for the user query statement in Step B, when the large model selects GPT3.5-turbo, the parameter temperature is set to 0.7, and the number of inference paths is set to 3. After executing Step F, the generated SQL statements are as follows:

[0199] SQL1:

[0200] SELECT cargo_name

[0201] FROM cargo_info

[0202] INNER JOIN purchase_info ON cargo_info.cargo_name = purchase_info.cargo_name

[0203] WHERE cargo_info.price < 1000

[0204] AND purchase_info.company_name = 'XXX Company'

[0205] SQL2:

[0206] SELECT cargo_name

[0207] FROM cargo_info

[0208] JOIN purchase_info ON cargo_info.cargo_name = purchase_info.cargo_name

[0209] WHERE price < 1000 AND company_name = 'XXX Company'

[0210] SQL3:

[0211] SELECT cargo_name

[0212] FROM cargo_info

[0213] WHERE price < 1000

[0214] AND cargo_name IN(

[0215] SELECT cargo_name

[0216] FROM purchase_info

[0217] WHERE company_name = 'XXX Company' )

[0219] In addition, there can be SQL4, SQL5, etc. The specific content of SQL in this application embodiment is not limited.

[0220] Step G: The execution entity triggers the Self-Consistency mechanism for consistency verification. Execute the SQL query statements generated by different inference paths. From the correctly executed results, with the help of the deepdiff library of python (used for comparing differences in tabular data), use the voting principle to determine the result with the highest consistency as the final result.

[0221] Step H: Result visualization. The execution entity converts the tabular result into a bar chart, pie chart, or line chart based on the user's needs. By constructing a model output prompt, more accurate query statements can be generated. By executing the generated query statements, the accuracy of data query can be improved when processing the user's data query service.

[0222] When constructing the model output prompt (Prompt) of text2SQL, the embodiments of the present application provide a minimum prompt recall method based on a knowledge graph. By using the field information in the user query statement to locate relevant tables in the graph through relationship queries, the problem that relevant tables cannot be recalled when there is no table description information in the query statement can be solved. In addition, a self-evaluation mechanism, namely the Self-Evaluate mechanism, is provided. On the one hand, through the self-correction function, useless information can be filtered out automatically. On the other hand, it can automatically judge whether the recalled relevant tables are ambiguous. If there is ambiguity, it guides the user to assist in solving it.

[0223] Figure 5 It is a schematic diagram of the main units of the data query device according to the embodiments of the present application. As Figure 5 shown, the data query device 500 includes an information extraction unit 501, a model output prompt construction unit 502, a query statement generation unit 503, and an execution unit 504.

[0224] The information extraction unit 501 is configured to, in response to obtaining a user query statement, perform intent recognition on the user query statement to extract field name information.

[0225] The model output prompt construction unit 502 is configured to locate field entities matching the field name information from a pre-constructed knowledge graph, and then based on the relationships in the pre-constructed knowledge graph, find table entities corresponding to the field entities, and construct a model output prompt based on the table entities.

[0226] The query statement generation unit 503 is configured to generate a query statement based on the model output prompt.

[0227] The execution unit 504 is configured to execute the query statement to obtain query result data.

[0228] In some embodiments, the data query device further includes Figure 5 a knowledge graph construction unit not shown in the figure, which is configured to: obtain a preset business database identifier, and then obtain corresponding entity relationship data based on the business database identifier; according to the entity relationship data, construct entity triples and relationship triples, and then based on the entity triples and relationship triples, pre-construct a knowledge graph.

[0229] In some embodiments, the knowledge graph construction unit is further configured to: extract table entities, field entities, data types, and remarks from the entity relationship data, and then based on the table entities and corresponding remarks, construct table entity triples, and based on the field entities, corresponding data types, and corresponding remarks, construct field entity triples; extract the relationships between the table entities and field entities in the entity relationship data, and then based on the field entities, table entities, and relationships, construct relationship triples.

[0230] In some embodiments, the knowledge graph construction unit is further configured to: set the label of the field node as the field entity corresponding to the entity triple; set the label of the table node as the table entity corresponding to the entity triple; and pre-construct a knowledge graph based on the field node, the table node, the label of the field node, the label of the table node, and the relationship in the relationship triple.

[0231] In some embodiments, the model output prompt construction unit 502 is further configured to: convert the remarks of the field entity in the pre-constructed knowledge graph into corresponding remark vectors; convert the field name information into corresponding field name information vectors; calculate the similarity between the field name information vector and the remark vector, determine the target remark vector according to the similarity, and then determine the field entity corresponding to the target remark vector as the field entity in the pre-constructed knowledge graph that matches the field name information.

[0232] In some embodiments, the model output prompt construction unit 502 is further configured to: determine the number of different table entities in the table entity, and if the number exceeds a preset threshold, obtain the corresponding foreign key field; if the field entity does not completely correspond to the foreign key field, traverse the table corresponding to the table entity, and determine whether other tables in the table except the currently traversed table can cover the field entity. If so, delete the currently traversed table, and if not, retain the currently traversed table, and then update the table entity; construct a model output prompt based on the updated table entity.

[0233] In some embodiments, the model output prompt construction unit 502 is further configured to: if the field entity completely corresponds to the foreign key field, generate a confirmation information based on the table corresponding to the table entity and the corresponding remarks and display it to the user, and determine the target table entity based on the user's selection operation; construct a model output prompt based on the target table entity.

[0234] In some embodiments, the model output prompt construction unit 502 is further configured to: generate a context based on the target table entity and the field entity; construct a model output prompt according to the preset instruction, context, user query statement, and preset output guidance.

[0235] In some embodiments, the query statement generation unit 503 is further configured to: determine a scenario identifier according to the model output prompt; generate a query statement based on the scenario identifier, a preset inference path, and a deviation verification mechanism.

[0236] In some embodiments, the execution unit 504 is further configured to: execute each query statement to obtain each query result data; perform a consistency check on each query result data, and return the query result data with the highest consistency as the execution result data.

[0237] In some embodiments, the data query device further includesFigure 5 A data display unit not shown in the figure is configured to: obtain user requirement data, determine a data display method based on the user requirement data; and convert the execution result data into a corresponding data form according to the data display method and output it.

[0238] It should be noted that the data query method and data query device of the present application have a corresponding relationship in the specific implementation content, so the repeated content will not be described again.

[0239] Figure 6 An exemplary system architecture 600 to which the data query method or data query device of the embodiments of the present application can be applied is shown.

[0240] As Figure 6 shown, the system architecture 600 may include terminal devices 601, 602, 603, a network 604, and a server 605. The network 604 is used to provide a medium for communication links between the terminal devices 601, 602, 603 and the server 605. The network 604 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0241] Users can use the terminal devices 601, 602, 603 to interact with the server 605 through the network 604 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 601, 602, 603, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0242] The terminal devices 601, 602, 603 may be various electronic devices with a data analysis and processing screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0243] The server 605 may be a server providing various services, such as a background management server (only as an example) that provides support for data analysis requests submitted by users using the terminal devices 601, 602, 603. The background management server may, in response to obtaining a user query statement, perform intent recognition on the user query statement to extract field name information; locate field entities matching the field name information from a pre-constructed knowledge graph, and then find table entities corresponding to the field entities based on the relationships in the pre-constructed knowledge graph, construct a model based on the table entities to output a prompt; generate a query statement based on the model output prompt; execute the query statement to obtain query result data. By constructing a model to output a prompt, a more accurate query statement can be generated, and by executing the generated query statement, the accuracy of data query can be improved when processing the user's data query service.

[0244] It should be noted that the data query method provided by the embodiments of the present application is generally executed by the server 605. Correspondingly, the data query device is generally disposed in the server 605.

[0245] It should be understood that Figure 6 the numbers of the terminal devices, networks, and servers in

[0246] Reference is made below to Figure 7 which shows a schematic structural diagram of a computer system 700 of a terminal device suitable for implementing the embodiments of the present application. Figure 7 The terminal device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0247] As Figure 7 shown, the computer system 700 includes a central processing unit (CPU) 701 which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the computer system 700 are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0248] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including such as a cathode ray tube (CRT), a liquid crystal credit authorization query processor (LCD), etc. and a speaker; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as required. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as required so that the computer program read from it can be installed into the storage section 708 as required.

[0249] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 709 and / or installed from the removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above functions defined in the system of the present application are performed.

[0250] It should be noted that the computer-readable medium shown in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present application, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0251] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0252] The units described in the embodiments of the present application can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an information extraction unit, a model output hint construction unit, a query statement generation unit, and an execution unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases.

[0253] On the other hand, the present application also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist separately without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device responds to obtaining a user query statement, performs intent recognition on the user query statement to extract field name information; locates a field entity matching the field name information from a pre-constructed knowledge graph, and then searches for a table entity corresponding to the field entity based on the relationships in the pre-constructed knowledge graph, constructs a model output hint based on the table entity; generates a query statement based on the model output hint; and executes the query statement to obtain query result data.

[0254] According to the technical solution of the embodiments of the present application, a more accurate query statement can be generated by constructing a model output hint, and the accuracy of data query can be improved when processing the user's data query service by executing the generated query statement.

[0255] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. A data query method, characterized in that, Including: In response to obtaining a user query statement, perform intent recognition on the user query statement to extract field name information; Locate a field entity matching the field name information in a pre-constructed knowledge graph, then based on the relationships in the pre-constructed knowledge graph, find a table entity corresponding to the field entity, and construct a model output prompt based on the table entity; Generate a query statement based on the model output prompt; Execute the query statement to obtain query result data.

2. The method according to claim 1, wherein Before locating the field entity matching the field name information in the pre-constructed knowledge graph, the method further includes: Obtain a preset business database identifier, and then obtain corresponding entity relationship data based on the business database identifier; According to the entity relationship data, construct entity triples and relationship triples, and then based on the entity triples and the relationship triples, pre-construct a knowledge graph.

3. The method according to claim 2, wherein The constructing entity triples and relationship triples according to the entity relationship data includes: Extract table entities, field entities, data types, and remarks from the entity relationship data, then construct table entity triples based on the table entities and corresponding remarks, and construct field entity triples based on the field entities, corresponding data types, and corresponding remarks; Extract the relationship between the table entity and the field entity in the entity relationship data, and then construct a relationship triple based on the field entity, the table entity, and the relationship.

4. The method according to claim 2, wherein The pre-constructing a knowledge graph includes: Set the label of the field node as the field entity corresponding to the entity triple; Set the label of the table node as the table entity corresponding to the entity triple; Pre-construct a knowledge graph based on the field node, the table node, the label of the field node, the label of the table node, and the relationship in the relationship triple.

5. The method according to claim 1, wherein The locating a field entity matching the field name information in the pre-constructed knowledge graph includes: Convert the remarks of the field entities in the pre-constructed knowledge graph into corresponding remark vectors; Convert the field name information into a corresponding field name information vector; Calculate the similarity between the field name information vector and the remark vector, determine a target remark vector according to the similarity, and then determine the field entity corresponding to the target remark vector as the field entity matching the field name information in the pre-constructed knowledge graph.

6. The method according to claim 1, characterized in that, The constructing a model output prompt based on the table entity includes: Determine the number of different table entities in the table entity. If the number exceeds a preset threshold, obtain corresponding foreign key fields; If the field entity does not completely correspond to the foreign key field, traverse the table corresponding to the table entity, and judge whether other tables in the table except the currently traversed table can cover the field entity. If so, delete the currently traversed table. If not, retain the currently traversed table, and then update the table entity; Construct a model output prompt based on the updated table entity.

7. The method according to claim 6, characterized in that, The constructing a model output prompt based on the table entity includes: If the field entity exactly corresponds to the foreign key field, generate confirmation information to be presented to the user based on the table corresponding to the table entity and the corresponding remarks, and determine the target table entity based on the user's selection operation; Construct a model output prompt based on the target table entity.

8. The method according to claim 7, characterized in that, The constructing a model output prompt based on the target table entity includes: Generate a context based on the target table entity and the field entity; Construct a model output prompt according to a preset instruction, the context, the user query statement, and a preset output guideline.

9. The method according to claim 1, wherein The generating a query statement includes: Determine a scenario identifier according to the model output prompt; Generate a query statement based on the scenario identifier, a preset inference path, and a deviation verification mechanism.

10. The method according to claim 1, characterized in that, The executing the query statement to obtain query result data includes: Execute each of the query statements to obtain respective query result data; Perform a consistency check on the respective query result data, and return the query result data with the highest consistency as the execution result data.

11. The method according to claim 1, wherein After obtaining the query result data, the method further includes: Obtain user requirement data, and determine a data display mode based on the user requirement data; Convert the execution result data into a corresponding data form according to the data display mode and output it.

12. A data query device, characterized in that, including: An information extraction unit configured to, in response to obtaining a user query statement, perform intent recognition on the user query statement to extract field name information; A model output prompt construction unit configured to locate a field entity matching the field name information from a pre-constructed knowledge graph, and then, based on the relationships in the pre-constructed knowledge graph, find a table entity corresponding to the field entity, and construct a model output prompt based on the table entity; A query statement generation unit configured to generate a query statement based on the model output prompt; An execution unit configured to execute the query statement to obtain query result data.

13. A data analysis electronic device, characterized in that, including: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-11.

14. A computer-readable medium having a computer program stored thereon, characterized in that, The program, when executed by the processor, implements the method according to any one of claims 1-11.