Query method based on query auxiliary information

By obtaining user information and natural query statements, combining query auxiliary information and constellation data model, the text processing model is used to determine the target query fields and data tables, and query instructions are generated, which solves the problem of difficult intention understanding in natural language query technology and achieves more accurate data query.

CN120045582AActive Publication Date: 2025-05-27GUANGZHOU SMART SOFTWARE CO LTD

Patent Information

Application Number
CN202510101376.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Existing natural language query technologies are difficult to accurately understand the user's query intentions, especially because of the ambiguity and diversity of natural language, which leads to difficulties in capturing the user's real needs.

Method used

By obtaining the user's natural query statements and user information, combining query auxiliary information, using preset constellation data model and text processing model, the target query field and target data table are determined, and query instructions are generated to achieve accurate data query.

Benefits of technology

It significantly improves the accuracy of natural language query, can meet the complex and diverse data query needs of users, and brings users a more accurate data query experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045582A_ABST
    Figure CN120045582A_ABST
Patent Text Reader

Abstract

The invention relates to a query method based on query auxiliary information. The method comprises the following steps: acquiring a natural query statement of a user and user information; obtaining query auxiliary information according to the user information; determining candidate query fields associated with the natural query statement by matching query fields in a preset constellation data model; the natural query statement, the candidate query fields, the query auxiliary information and a preset first task processing text are input into a first text processing model, the query intention of the natural query statement is understood through the first text processing model in combination with the query auxiliary information, and the candidate query fields conforming to the query intention are determined as target query fields. And further, determining at least one target data table only comprising the target query field from the constellation data model, and inputting the natural query statement and the target data table into a query instruction generation model to generate a query instruction. And finally, based on the query instruction, querying from the data warehouse to obtain a query result of the natural query statement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language queries, and particularly to a query method based on query auxiliary information. Background Art

[0002] With the rapid development of information technology, especially the rise of fields such as big data analysis and business intelligence, users' demands for data queries have become increasingly complex and diverse. Traditional database query technologies rely on structured query languages (such as SQL) to achieve data retrieval. However, for non-professional users, this query method with a relatively high technical threshold is particularly inconvenient. To improve the user experience, existing database query technologies have begun to attempt to convert natural language into database query language, but there are many challenges in this process.

[0003] Specifically, natural language has fuzziness and diversity. Different users or user groups have different identity backgrounds and often have their own expression habits and professional terms, which are usually not understood by natural language processing models, resulting in difficulties for natural language processing models to accurately capture users' true needs when understanding users' query intentions. For example, when a user mentions "our department", it actually refers to a specific department where they are located, but the model usually cannot accurately understand this expression. In addition, even the same vocabulary may have different meanings for different users or in different scenarios, which further increases the difficulty of natural language processing. Summary of the Invention

[0004] Based on this, the purpose of this application is to provide a query method based on query auxiliary information to improve the accuracy of natural language queries, so as to meet users' complex and diverse data query needs.

[0005] The query method based on query auxiliary information described in the embodiments of this application includes the following steps:

[0006] Obtain the natural query statement input by the user and the user information of the user;

[0007] According to the user information, obtain the query auxiliary information of the user; wherein, the query auxiliary information is additional information associated with the user for the natural query statement;

[0008] Obtain each query field of the preset constellation data model, and determine the query field associated with the natural query statement as the candidate query field; wherein, the constellation data model includes a structured data table and query fields in the data table;

[0009] Input the natural query statement, the candidate query fields, the query auxiliary information, and a preset first task processing text into a first text processing model to obtain a number of target query fields; wherein, the first task processing text is used to prompt the first text processing model to understand the query intention of the natural query statement in combination with the query auxiliary information, and determine the candidate query fields that conform to the query intention as target query fields;

[0010] Determine at least one target data table in the constellation data model that only includes the target query fields;

[0011] Input the natural query statement and the target data table into a query instruction generation model to obtain a query instruction; wherein, the query instruction generation model is used to understand the query intention of the natural query statement, determine a query task that conforms to the query intention, and generate a query instruction based on the target data table and the query task;

[0012] Query the query result of the natural query statement from the data warehouse corresponding to the constellation data model based on the query instruction.

[0013] In the embodiments of the present application, the natural query statement and user information of the user are obtained, the query auxiliary information is obtained according to the user information, and then, by matching the query fields in the preset constellation data model, the candidate query fields associated with the natural query statement are determined. The natural query statement, the candidate query fields, the query auxiliary information, and the preset first task processing text are input into the first text processing model. The first text processing model combines the query auxiliary information to understand the query intention of the natural query statement, and determines the candidate query fields that conform to the query intention as target query fields. Further, at least one target data table that only includes the target query fields is determined from the constellation data model, and the natural query statement and the target data table are input into the query instruction generation model to generate a query instruction. Finally, based on the query instruction, the query result of the natural query statement is queried from the data warehouse. In summary, in the embodiments of the present application, for different users, the query auxiliary information associated with the user is introduced to enhance the ability of the large model to understand the user's query intention, help the large model more accurately interpret the user's natural query statement, so as to accurately screen out the target query fields from the candidate query fields, and then accurately determine the target data table. Finally, the large model generates a query instruction in combination with the query intention and the target data table and queries to obtain the query result. It significantly improves the accuracy of natural language queries, can meet the complex and diverse data query needs of users, and brings a more accurate data query experience to users.

[0014] For better understanding and implementation, the present application will be described in detail below with reference to the accompanying drawings. Description of the Drawings

[0015] Figure 1 Schematic flowchart of the query method based on query auxiliary information according to an embodiment of the present application;

[0016] Figure 2 Schematic diagram of the steps of obtaining synonym information from the user's synonym information library in an embodiment of the present application;

[0017] Figure 3 Schematic diagram of the steps of adding synonym information to the user's synonym information library in an embodiment of the present application;

[0018] Figure 4 Schematic diagram of the steps of generating a query instruction by combining the user's query semantic enhancement information in an embodiment of the present application. Detailed implementation manners

[0019] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings. Among them, when the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements.

[0020] It should be clear that the embodiments described in the following embodiments do not represent all embodiments consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0021] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application are also intended to include the plural forms unless the context clearly indicates otherwise. In addition, in the description of the present application, unless otherwise stated, "a plurality" means two or more. It should also be understood that the term " / and / " used herein refers to and includes any or all possible combinations of one or more of the associated listed items. For example, A and / or B may represent three cases: A exists alone, A and B exist simultaneously, and B exists alone; the character " / " generally represents an "or" relationship between the associated objects before and after.

[0022] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. Moreover, these terms are only used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances. Depending on the context, the words "if" / "when" used in this application can be interpreted as "when...", "while...", or "in response to determining".

[0023] This application relates to the field of natural language query technology. Existing database query technologies convert natural language into database query language, but there are many challenges in this process. Specifically, natural language has ambiguity and diversity, and different users or user groups often have their own expression habits and professional terms. These differences cause existing natural language processing models to often encounter difficulties in understanding the user's query intent and are difficult to accurately capture the user's true needs. For example, when a user mentions "our department", it actually refers to the specific department where they are located, but the model usually cannot accurately understand this expression. In addition, even the same word may have different meanings for different users or in different scenarios, which further increases the difficulty of natural language processing.

[0024] In response to this, the embodiments of this application provide a query method based on query auxiliary information, introducing corresponding query auxiliary information for different users to improve the accuracy of natural language queries, so as to meet the complex and diverse data query needs of users.

[0025] Please refer to Figure 1 , the query method based on query auxiliary information described in the embodiments of this application includes the following steps:

[0026] S101: Obtain the natural query statement input by the user and the user information of the user;

[0027] S102: According to the user information, obtain the query auxiliary information of the user; wherein, the query auxiliary information is additional information associated with the user for the natural query statement;

[0028] S103: Obtain each query field of the preset constellation data model, and determine the query field associated with the natural query statement as the candidate query field; wherein, the constellation data model includes a structured data table and query fields in the data table;

[0029] S104: Input the natural query statement, the candidate query fields, the query auxiliary information, and a preset first task processing text into a first text processing model to obtain a number of target query fields. Among them, the first task processing text is used to prompt the first text processing model to understand the query intention of the natural query statement in combination with the query auxiliary information, and determine the candidate query fields that meet the query intention as the target query fields.

[0030] S105: Determine at least one target data table in the constellation data model that only includes the target query fields.

[0031] S106: Input the natural query statement and the target data table into a query instruction generation model to obtain a query instruction. Among them, the query instruction generation model is used to understand the query intention of the natural query statement, determine the query task that meets the query intention, and generate a query instruction based on the target data table and the query task.

[0032] S107: Query the query result of the natural query statement from the data warehouse corresponding to the constellation data model based on the query instruction.

[0033] In the embodiment of the present application, the natural query statement and user information of the user are obtained, the query auxiliary information is obtained according to the user information, and then, by matching the query fields in the preset constellation data model, the candidate query fields associated with the natural query statement are determined. The natural query statement, the candidate query fields, the query auxiliary information, and the preset first task processing text are input into the first text processing model. The first text processing model combines the query auxiliary information to understand the query intention of the natural query statement, and determines the candidate query fields that meet the query intention as the target query fields. Further, at least one target data table that only includes the target query fields is determined from the constellation data model, and the natural query statement and the target data table are input into the query instruction generation model to generate a query instruction. Finally, based on the query instruction, the query result of the natural query statement is queried from the data warehouse. In summary, in the embodiment of the present application, for different users, the query auxiliary information associated with the user is introduced to enhance the ability of the large model to understand the user's query intention, help the large model more accurately interpret the natural query statement of the user, so as to accurately screen out the target query fields from the candidate query fields, and then accurately determine the target data table. Finally, the large model combines the query intention and the target data table to generate a query instruction and query the query result. It significantly improves the accuracy of natural language queries, can meet the complex and diverse data query needs of users, and brings a more accurate data query experience to users.

[0034] The query method based on query auxiliary information described in the embodiments of the present application takes a computer as the execution subject. Hereinafter, first, the text processing model of the embodiments of the present application will be described, and then each step will be described in detail.

[0035] In the embodiments of the present application, the text processing model can be based on one or more pre-trained language models, such as BERT, GPT, etc. (These models have been trained on a large amount of text data and can capture the deep semantic and syntactic features of language). In order to adapt to specific tasks in the present application, such as query intention understanding, query instruction generation, etc., it is obtained through fine-tuning. Of course, the text processing model can also be an existing large language model, as long as it has the processing ability for the specific tasks of the present application, it can be applied in the technical solution of the present application. In this regard, the embodiments of the present application do not make restrictions.

[0036] In order to enable the text processing model to complete tasks as required, it is usually necessary to provide task processing text. The task processing text is used to require or prompt the task requirements of the text processing model, including specific requirements and format instructions for the model output, and can also include task-related scenario information or context information for the model to learn and understand the task to be processed, etc. In the embodiments of the present application, different sub-texts can be set in the task processing text according to the needs of the query, respectively used to require or prompt the corresponding task requirements of the text processing model. The task processing text provided to the model in different task processing requirements is also correspondingly different, such as the first task processing text, the second task processing text, etc.

[0037] For step S101, obtain the natural query statement input by the user and the user information of the user.

[0038] In this step, receive the natural query statement input by the user through a user interface (such as a web page, APP, etc.). The natural query statement is the data query requirement expressed by the user in natural language, such as "query the sales data of our department last month". At the same time, the user information related to the user is also obtained, including but not limited to the user's identity information, such as user name, user identity type, user attributes, etc.

[0039] For step S102, obtain the query auxiliary information of the user according to the user information; wherein, the query auxiliary information is additional information associated with the user for the natural query statement.

[0040] Among them, the query auxiliary information is additional information related to the user's natural query statement, which is used to help the model more accurately understand the user's query intention. For example, if the user belongs to the "Sales Department", the query auxiliary information may include business terms related to the "Sales Department", common query conditions, etc. This information can be extracted from the database of the user's department, the user profile, or historical query data. This step determines the type to which the user belongs based on the user information, and accordingly selects the corresponding synonym information library. This information is crucial for understanding the user's query intention, especially when the user uses specific industry terms or idiomatic expressions.

[0041] Please refer to Figure 2 , in one embodiment, the query auxiliary information described in step S102 includes synonym information; the step of obtaining the query auxiliary information of the user according to the user information includes:

[0042] Step S1021, perform entity word segmentation recognition on the natural query statement to obtain a number of entity word segments;

[0043] Step S1022, determine the synonym information library corresponding to the user information, and determine the synonyms corresponding to any one of the entity word segments from the synonym information library;

[0044] Step S1023, obtain synonym information according to the synonyms of each entity word segment.

[0045] In this embodiment, in order to more accurately understand and respond to the query needs of different types of users, especially considering the differences in the expression idioms, identity backgrounds, etc. of different users, customized synonym information libraries are set up for different types of users respectively. These synonym information libraries are specially designed according to the characteristics of the users (such as the department they belong to, the position, the industry, etc.), and contain synonym and near-synonym information closely related to the users.

[0046] Step S1021 uses word segmentation technology to split the query statement into meaningful words or phrases and perform entity recognition to obtain a number of entity word segments. Step S1022 selects the corresponding synonym information library according to the type of the user (such as department, position, etc.). In one embodiment, the user information includes user department information or position information, so as to select the corresponding synonym information library based on the user department information or position information. Check whether there are synonyms for any entity word segment in the synonym information library. If there are, in step S1023, add the entity word segment and its synonyms to the synonym information. The synonym information will be used as query auxiliary information for understanding the user's query intention and generating query instructions in subsequent steps.

[0047] In one example, for users in the sales department, the thesaurus may contain synonym relationships such as "sales data = sales amount", "sales data = sales quantity", "department = specific department name of the user (such as sales department, finance department)". When it is recognized that the user enters a query statement such as "Query the sales data of our department last month", the thesaurus of this user can be used to replace "department" with the "specific department name" of this user, and replace "sales data" with the specific sales data indicator related to this department, such as it may be "sales (amount)" or "sales (quantity)", etc., so as to more accurately understand the user's query intention. In one example, for users in department A, "sales data" refers to "sales amount", while for department B, "sales data" refers to "sales quantity".

[0048] In summary, in this embodiment, by setting a customized thesaurus for different types of users, the query intention of the user can be more accurately understood. Even if the user uses specific industry terms or idiomatic expressions, query instructions that meet the user's needs can be generated. This innovation improves the powerful ability to adapt to different user needs and expression habits.

[0049] Please refer to Figure 3 , in one embodiment, before the step of obtaining the query auxiliary information of the user according to the user information in step S102, the following steps are further included:

[0050] Step S201, in response to a thesaurus information creation instruction, parse the thesaurus information creation instruction to obtain user information and thesaurus information; the thesaurus information includes a standard vocabulary and the habitual expression vocabulary of the user corresponding to the user information for the standard vocabulary;

[0051] Step S202, add the thesaurus information to the thesaurus corresponding to the user information.

[0052] Among them, in step S201, the thesaurus information creation instruction is an instruction triggered by a system administrator or a user with corresponding permissions, used to indicate creating thesaurus information for a specific user or user group. After receiving the thesaurus information creation instruction, it is first parsed to extract the user information and thesaurus information included in the instruction. Among them, the user information may include the user's identity information, such as username, department, position, etc., used to determine which user or user group the thesaurus information will be applied to. The thesaurus information is the core content of the instruction, including a standard vocabulary and the habitual expression vocabulary of the user for the standard vocabulary. For example, for users in the sales department, the standard vocabulary may be "sales volume", and the habitual expression vocabulary may be "sales performance" or "total sales volume", etc.

[0053] In step S202, the synonym information database is a database specifically for storing synonym information, which is organized according to the user's identity information (such as department, position, etc.). After parsing the user information and synonym information, these information are added to the synonym information database corresponding to the user. In this way, when this user or user group conducts a query, they can utilize these synonym information to more accurately understand the query intention of the user.

[0054] In summary, in this embodiment, the system administrator or a user with corresponding permissions can create and update the synonym information database for different users or user groups according to actual needs, so as to be able to provide sufficient query assistance information for the user's natural query statement, enabling the text processing model to more accurately understand the query intention of the user and improving the accuracy of the query. At the same time, it also provides convenience for subsequent maintenance and optimization, and can continuously optimize the synonym information database as the user's needs and expression habits change.

[0055] For step S103, obtain each query field of the preset constellation data model, and determine the query field associated with the natural query statement as the candidate query field; wherein, the constellation data model includes a structured data table and the query fields in the data table.

[0056] Among them, the constellation data model is an abstract representation used to describe the database structure. It is a data logic layer established in advance based on each data table stored in the data warehouse, which records several data tables, the basic fields included in each data table, and the graph relationship of each data table. In addition, it may also record the binding relationship between at least some basic fields and multi-dimensional operation fields. That is to say, the constellation data model does not record the specific business data of the fields of each data table, but only records the information related to generating the query statement. These information mainly include the table name of the data table, the field information of the data table, and the multi-dimensional operation field information bound to the basic field. In addition, a graph relationship is established among each data table in the constellation data model, and the graph relationship includes at least a directed graph relationship.

[0057] Among them, the basic fields are generally the fields in the database that record the original business data. More specifically, the basic fields are generally divided into measurement fields and dimension fields. The measurement fields generally refer to the fields in the fact table, and the dimension fields generally refer to the fields in the dimension table. Among them, the fact table and the dimension table are concepts in data modeling in the field of data warehousing. Currently, the main data models involved in the field of data warehousing are the star data model, the snowflake data model generated based on the star data model (which can actually be considered a kind of star data model), and the constellation data model. In the star data model, it includes a central data table and other data tables connected to the central data table. Among them, the central data table is the fact table, and the other data tables connected to it are the dimension tables.

[0058] In this step, access the preset constellation data model to obtain all available query fields, and then through text analysis of the natural query statement (such as keyword matching, semantic analysis, etc.), determine the query fields associated with the natural query statement as candidate query fields. For example, for the natural query statement "Query the sales data of our department last month", query fields such as "department", "time", and "sales amount" may be determined as candidate query fields.

[0059] In one embodiment, the step of obtaining each query field of the preset constellation data model in step S103 and determining the query fields associated with the natural query statement as candidate query fields includes:

[0060] Step S1031, obtain each query field of the preset constellation data model;

[0061] Step S1032, determine the query fields whose semantic similarity with the natural query statement or any of the entity segmentations is greater than the preset similarity threshold as candidate query fields.

[0062] Step S1031 obtains each query field in the preset constellation data model. Then, in step S1032, using the semantic similarity calculation algorithm, calculate the semantic similarity between each query field and the natural query statement, and calculate the semantic similarity between each query field and the identified entity segmentations. The purpose of this step is to find the query fields that are semantically closest to the natural query statement or the entity segmentations. Finally, according to the preset similarity threshold, determine the query fields with a semantic similarity greater than the threshold as candidate query fields. These fields have a high degree of semantic matching with the natural query statement or the entity segmentations, so they are considered as fields that may meet the user's query requirements. In summary, in this embodiment, by introducing entity segmentation recognition and semantic similarity calculation, the query fields associated with the natural query statement are accurately selected as candidate query fields.

[0063] For step S104, input the natural query statement, the candidate query fields, the query auxiliary information, and a preset first task processing text into a first text processing model to obtain a number of target query fields. Among them, the first task processing text is used to prompt the first text processing model to understand the query intent of the natural query statement in combination with the query auxiliary information, and determine the candidate query fields that conform to the query intent as target query fields.

[0064] In this step, the natural query statement, the candidate query fields, the query auxiliary information, and the preset first task processing text are combined into an input set and input into the first text processing model. The first text processing model understands the query intent of the natural query statement in combination with the query auxiliary information, and filters out the target query fields that conform to the query intent from the candidate query fields. For example, for the natural query statement "Query the sales data of our department last month", the model may determine the target query fields as "Department A", "Time", and "Sales amount" according to the query auxiliary information "User's department = Department A".

[0065] In one embodiment, the first task processing text in step S104 is further used to prompt the first text processing model to understand the query intent of the natural query statement in combination with the input business knowledge.

[0066] After the step of performing entity word segmentation recognition on the natural query statement in step S1021 to obtain a number of entity words, the following steps are further included:

[0067] Step S10211, retrieve relevant business knowledge from a preset business knowledge base based on the natural query statement and the entity words.

[0068] The step of inputting the natural query statement, the candidate query fields, the query auxiliary information, and the preset first task processing text into the first text processing model in step S104 to obtain a number of target query fields includes:

[0069] Step S1041, input the natural query statement, the candidate query fields, the query auxiliary information, the business knowledge, and the preset first task processing text into the first text processing model to obtain a number of target query fields.

[0070] In this embodiment, the first task processing text guides the first text processing model to deeply understand the query intention of the natural query statement by combining the input business knowledge. After performing entity word segmentation recognition on the natural query statement in step S1021 and obtaining several entity word segments, in step S10211, based on the natural query statement and the recognized entity word segments, relevant business knowledge is retrieved from the preset business knowledge base. That is, the business knowledge related to the natural query statement or entity word segments is obtained to more accurately understand the query intention in subsequent steps. Then, in step S1041, the natural query statement, candidate query fields, query auxiliary information, retrieved business knowledge, and the preset first task processing text are input into the first text processing model together, thereby obtaining several target query fields. The improvement in this step is that business knowledge is introduced as an input, enabling the first text processing model to more comprehensively understand the query intention of the natural query statement and generate more accurate target query fields accordingly. In summary, by introducing business knowledge, the accuracy and efficiency of processing the natural query statement and obtaining the target query fields are improved.

[0071] For step S105, determine at least one target data table in the constellation data model that only includes the target query fields.

[0072] Based on the target query fields obtained in S104, in this step, the data tables in the constellation data model that only contain these target query fields are filtered out as the target data tables. These target data tables are the basis for subsequent query instruction generation and data query. For example, if the target query fields are "Department A", "Time", and "Sales Amount", then the sales data table of Department A may be filtered out as the target data table.

[0073] In one embodiment, the target query fields include dimension fields and measure fields; the step of determining at least one target data table in the constellation data model that only includes the target query fields in step S105 includes:

[0074] Step S1051, determine the data tables in the constellation data model that contain any one or more of the measure fields as the target data tables;

[0075] Step S1052, using the target data tables as the root nodes, according to the graph relationship, determine the other data tables in the constellation data model that can reach the target data tables. Respectively, using each of the target data tables as the fact table and the other data tables that can be reached corresponding to the fact table as the dimension tables, several target star data models are obtained;

[0076] Step S1053, remove the dimension tables at the edges of the target star data model that do not include the dimension fields, remove the base fields other than the metric fields from the fact table of any of the target star data models, and remove the base fields other than the dimension fields from the dimension tables of any of the target star data models;

[0077] Step S1054, obtain a target data table according to the target star data model.

[0078] In this embodiment, the target query fields include metric fields and dimension fields. First, use the data tables in the constellation data model that contain any one or more metric fields as the root data tables; use the fact tables centered on the root data tables and the other data tables that can reach the fact tables as dimension tables to determine several target star data models; remove the dimension tables irrelevant to the query in each target star data model, and remove the base fields (metric fields or dimension fields) irrelevant to the query, so that the finally obtained target star data model only includes the fact tables, dimension tables, and base fields related to the query. Through the method of this embodiment of the present application, it is realized to trim the pre-modeled constellation data model according to the target query fields, determine the target data table based on the trimmed minimum available star data model, avoid querying other irrelevant data tables and irrelevant query fields, reduce the complexity of the query instruction generation model's understanding, and make the generation of query instructions more efficient and accurate.

[0079] For step S106, input the natural query statement and the target data table into a query instruction generation model to obtain a query instruction; wherein, the query instruction generation model is used to understand the query intention of the natural query statement, determine a query task that conforms to the query intention, and generate a query instruction based on the target data table and the query task.

[0080] In this step, the natural query statement and the target data table are input into the query instruction generation model. The query instruction generation model is a trained machine learning model that can understand the query intention of the natural query statement and generate a query instruction that conforms to the query intention according to the target data table. The query instruction can be a structured query statement such as an SQL statement for querying data in a database; it can also be a set of generated codes such as a Python script, which is executed by a code executor to query data from the database and perform specific calculations or processes. In one example, the model may generate an SQL query statement according to the input, which filters out the "sales amount" data of "Department A" at "time = October (assuming the current month is November)" from the target data table.

[0081] In one embodiment, the query instruction generation model in step S106 is a second text processing model; the step of inputting the natural query statement and the target data table into the query instruction generation model to obtain a query instruction includes:

[0082] Step S1061, input the natural query statement, the target data table, and a preset second task processing text into the second text processing model to obtain a query instruction; wherein, the second task processing text is used to prompt the second text processing model to understand the query intention of the natural query statement, determine a query task that conforms to the query intention, and generate a query instruction based on the target data table and the query task.

[0083] In this embodiment, the query instruction generation model is a second text processing model, and in order to enable the second text processing model to complete the task of step S106, a second task processing text is also preset. The second task processing text not only prompts the model to understand the query intention of the natural query statement, but also guides the model on how to determine the corresponding query task according to this intention. Once the query task is determined, the model will generate the final query instruction based on the structure and content of the target data table and the specific requirements of the query task.

[0084] Please refer to Figure 4 , in one embodiment, the second task processing text in step S1061 is further used to prompt the second text processing model to understand the query intention of the natural query statement in combination with the input query semantic enhancement information;

[0085] The step of inputting the natural query statement, the target data table, and the preset second task processing text into the second text processing model in step S1061 to obtain a query instruction includes:

[0086] Step S10611, obtain the user's query semantic enhancement information from a preset query semantic enhancement information library according to the user information;

[0087] Step S10612, input the natural query statement, the target data table, the query semantic enhancement information, and the preset second task processing text into the second text processing model to obtain a query instruction.

[0088] In this embodiment, the second task processing text guides the second text processing model to more deeply understand the query intention of the natural query statement in combination with the input query semantic enhancement information.

[0089] Among them, the query semantic enhancement information includes business knowledge related to natural query statements, data table processing, and relevant guidance for generating query instructions. Specifically, it may include date and time conditions. For example, it is required that all queries must include date and time conditions. If the natural query statement does not explicitly specify the date and time, the date (year-month-day) equal to "2024-10-20" is default added as a filtering condition; if it has been specified, there is no need to add. It can also include data table processing guidance. For example, when using sql_json for query, data is only queried from one table, and it is necessary to check whether all fields are in this table. It can also include query field selection. For example, when words such as "each" or "every" are included in the user request, the subsequent fields need to be added to the SELECT part to ensure that the query results include all necessary fields. When the user request contains the description of TOPN or "the first N", the subsequent fields also need to be added to the SELECT part. It can also include set operations and result display requirements. For example, when the user request contains "list", the set needs to be completely retained, and records with no values should also be filled with blanks for display. It can also include special indicator calculations. For example, for indicators related to the number of customers (such as the number of investment consulting customers), year-on-year, same-period values, month-on-month, previous-period values, etc., cannot be directly obtained from the table, but need to be calculated using a specified plugin. By using the query semantic enhancement information as part of the input, the first text processing model generates more accurate query instructions that meet business requirements.

[0090] Furthermore, in this embodiment, considering the differences in the identity backgrounds and query permissions of different users, different query semantic enhancement information is set for different users respectively. This embodiment maintains a query semantic enhancement information library, which can adjust the query semantic enhancement information for different users or user groups. The adjustment can involve any aspect. For example, different settings can be made for different users in any aspect related to business knowledge related to natural query statements, data table processing, or generating query instructions. Thus, the second text processing model can more appropriately understand the user's natural query statement and generate query instructions according to the required specifications. Specifically, in step S10611, the identity information of the user is used to find the corresponding query semantic enhancement information in the preset query semantic enhancement information library. This information library may be a database, a configuration file, or a data structure in memory, which stores the mapping relationship between different user identity information and the corresponding query semantic enhancement information. In one embodiment, the query semantic enhancement information may include, but is not limited to: First, business term mapping: mapping informal or specific business terms that the user may use to the standard terms used inside the system. Second, data table preference: the data tables or fields that the user often queries. Third, query mode: the query structures or modes that the user commonly uses, such as sorting methods, filtering conditions, etc. Fourth, permission restrictions: determining the data tables and fields that the user can access according to the user's permission level. In step S10612, the query semantic enhancement information is also input into the second text processing model to obtain a query instruction.

[0091] In summary, this embodiment can obtain relevant enhancement information from the preset query semantic enhancement information library according to the user's identity information. These information will subsequently be used in step S106 and input into the second text processing model together with the natural query statement, the target data table, and the preset second task processing text, etc., to generate a more accurate and user-demand-compliant composite query instruction, which can better adapt to the needs and query habits of different users.

[0092] In one embodiment, the second task processing text in step S106 further includes a preset task prompt text, and the task prompt text includes the usage scenario information and call interfaces of each preset function plug-in; the task prompt text is used to prompt the second text processing model to understand the usage scenario corresponding to the query task, determine a matching function plug-in according to the usage scenario information of each function plug-in, obtain the call interface corresponding to the function plug-in, and generate an execution code (i.e., a query instruction) according to the call interface.

[0093] In this embodiment, the task prompt text details the usage scenario information of each preset function plug-in and the corresponding call interfaces. These information are directly related to whether the second text processing model can accurately understand the usage scenario of the query task and select the correct function plug-in accordingly to execute the query task. Specifically, the usage scenario information in the task prompt text describes key elements such as the type of query task applicable to each function plug-in, the input data format, and the expected output result. These information provide a clear framework for the second text processing model, enabling it to screen out the most suitable one or more from numerous preset function plug-ins according to the specific requirements of the query task to execute the query task. At the same time, the task prompt text also contains the call interface information of each function plug-in. The call interface is the bridge for the query instruction executor to interact with the function plug-in, which defines how the query instruction executor calls the function plug-in, passes parameters, and receives return results, etc. Based on the call interface information, the second text processing model can accurately generate execution code, enabling the query instruction executor to implement the call and execution of the function plug-in. Similarly, in other embodiments, when the second text processing model needs to process other tasks, the task prompt text can be set to prompt the second text processing model to process the corresponding tasks and prompt relevant processing requirements.

[0094] In summary, the task prompt text of this embodiment provides a clear and accurate guiding framework for the second text processing model, enabling it to more quickly understand the specific requirements and usage scenarios of the query task and select the most suitable function plug-in accordingly to generate query instructions.

[0095] In one embodiment, the task prompt text further includes usage prompt information for the call interfaces of each of the function plug-ins; the usage prompt information is used to prompt the second text processing model to understand the target parameter information of the query task and pass the target parameter information into the call interface when generating execution code according to the call interface; wherein, the target parameter information includes a base field, a dimension field, a qualification condition, and an analysis index; the target parameter information is used to instruct the corresponding function plug-in to query target data according to the base field and the dimension field and perform analysis of the analysis index on the target data under the qualification condition.

[0096] Among them, the usage prompt information is a key component of the task prompt text, which details the specific usage methods and related details of the call interfaces of each function plug-in. These information are related to whether the second text processing model can accurately pass the target parameter information in the query task into the call interface in the specified format and order, so as to generate effective execution code.

[0097] Specifically, the use of hint information not only defines the structure of the target parameter information, but also clarifies the specific meanings and uses of these parameter information in the query task.

[0098] The target parameter information includes multiple key elements such as basic fields, dimension fields, qualification conditions, and analysis indicators. Among them, the basic fields are the basic data elements involved in the query task, and they constitute the data basis of the query task. For example, when the query task is "analyze whether the sales amount in the Guangzhou area in the past 3 months is abnormal", "sales amount" is the basic field. The dimension field is used to refine and classify the basic field so as to more accurately locate and analyze the target data. In the above example, "monthly" (in the past 3 months) is the dimension field. The qualification conditions are used to filter and screen the query results to ensure that the query results meet the actual requirements. The qualification conditions can include time range, data range, etc. In the above example, "region = Guangzhou" restricts the data range to only the Guangzhou area. The analysis indicator is the core purpose of the query task, and it indicates what kind of analysis needs to be performed on the target data. In the above example, "data abnormality" is the analysis indicator, which requires the function plug-in to analyze the sales amount in the Guangzhou area in the past 3 months and give a predicted value of whether it is abnormal.

[0099] In summary, the use of hint information not only clearly describes the meanings and uses of these parameter information, but also guides how the second text processing model correctly passes them into the call interface. In this way, the function plug-in can accurately query the target data according to the incoming target parameter information and perform the analysis of the corresponding analysis indicators on the target data under the qualification conditions.

[0100] In summary, the use of hint information provides a clear, detailed and specific guidance framework for the second text processing model, enabling it to more accurately understand the target parameter information in the query task and correctly pass them into the call interface in the specified format and order. In addition, since the various elements of the target parameter information, the relationships between them, and their specific meanings and uses in the query task are defined in detail in the use hint information, the second text processing model can easily handle various complex and diverse query tasks. At the same time, since the use hint information of the call interface is preset in the task hint text, it is possible to easily update and expand the function plug-in without large-scale modification and adjustment of the second text processing model.

[0101] For step S107, based on the query instruction, query the query result of the natural query statement from the data warehouse corresponding to the constellation data model.

[0102] In this step, according to the query instruction generated in S106, a data query task is executed from the data warehouse corresponding to the constellation data model, and the query result is obtained. The query result contains the data required by the user. For example, the query result may be a table or chart containing the "sales volume of Department A last month" as the query result. In one embodiment, the query instruction is executed by a preset query instruction executor to query the query result of the natural query statement from the data warehouse corresponding to the constellation data model.

[0103] In one embodiment, the query task pointed to by the natural query statement is a composite query task; the second task processing text is used to prompt the second text processing model to determine a composite query task that conforms to the query intention of the natural query statement, and generate a composite query instruction based on the target data table and the composite query task; the composite query instruction includes a number of query instructions with a hierarchical dependency relationship.

[0104] Among them, the compound query task is different from the simple single-condition query task. The compound query task involves the combination and collaborative work of multiple sub-query tasks, and there is a hierarchical dependency relationship among these sub-query tasks. Specifically, a compound query task cannot be completed through a single query operation, but requires multiple sub-query operations to be executed sequentially or in parallel according to a certain logical order and mutual dependency relationship to ultimately meet the user's information needs. Specifically, a compound query task generally has the following characteristics: First, the combination of sub-query tasks: The compound query task contains two or more sub-query tasks, and these sub-query tasks together constitute a complete set of operations to achieve the final query goal. For example, in a business data analysis scenario, a user may input the following natural query statement: "Find companies with an annual sales volume exceeding 10 million, and calculate the average profit growth rate of these companies in the past three years." This query task contains two sub-tasks: one is to find companies with an annual sales volume exceeding 10 million (which can be called sub-task A), and the other is to calculate the average profit growth rate of the companies that meet the conditions of sub-task A in the past three years (which can be called sub-task B). Second, the hierarchical dependency relationship: There is a clear hierarchical dependency relationship among the sub-query tasks, that is, the execution of subsequent sub-tasks depends on the results of previous sub-tasks. In the above example, the execution of sub-task B depends on the result of sub-task A, because only by first finding companies with an annual sales volume exceeding 10 million can the average profit growth rate of these companies be calculated. This dependency relationship may be data dependency or logical dependency, and subsequent sub-tasks need to use the data or calculation results filtered by previous sub-tasks as input or operation basis. Third, logical complexity: Due to involving multiple sub-tasks and their dependency relationships, the compound query task usually has a high logical complexity. This not only requires an accurate understanding and processing of the individual execution logics of each sub-task, but also needs to effectively manage and coordinate the relationships between them to ensure the smooth completion of the entire query task. For example, in a database environment with multiple data tables and complex associations, different sub-tasks may involve different data tables, and complex association operations (such as multi-table joins) and data filtering operations (such as conditional filtering, grouping and aggregation, etc.) are required, and the execution order and association methods of these operations need to be arranged according to the specific dependency relationships to ensure data consistency and the correctness of query results.

[0105] In this embodiment, a natural query statement input by the user is obtained, and a natural query statement involving a compound query task is obtained. For example, the user may input "Please query the top 3 customers in terms of sales among the brands with the top 3 sales, and the proportion of these customers in the brand sales". First, the target query fields in the preset constellation data model that match the query intention are determined, and then, based on the target query fields, the relevant data tables are filtered out from the constellation data model. On this basis, the compound query task of the natural query statement is decomposed into subtasks by the second text processing model, and according to the user's query intention and the structural information of the data table, a compound query instruction with a logical hierarchy and dependency relationship is generated. Finally, according to the generated compound query instruction, the query result that meets the user's needs is queried from the data warehouse. In summary, the embodiment of the present application realizes accurate query of the user's compound query task, improves the accuracy and efficiency of the query, and has a wide range of application prospects.

[0106] In one embodiment, the second task processing text includes task decomposition prompt text; the task decomposition prompt text is used to prompt the second text processing model to split the natural query statement into several sub-query statements with a hierarchical dependency relationship according to the semantic structure characteristics of the natural query statement and the preset semantic structure splitting rules; understand each sub-query task corresponding to the sub-query statement from bottom to top according to the hierarchical dependency relationship, and generate sub-query instructions according to the sub-query task; obtain a compound query instruction according to the sub-query instructions of each sub-query task.

[0107] Among them, the semantic structure splitting rules are used to analyze and process natural language text and split the natural language text into units with clear semantic components and structures. Specifically, the semantic structure splitting rules involve multiple aspects of natural language processing, including syntactic analysis, semantic role annotation, entity recognition, relationship extraction, etc., which can be achieved by fine-tuning existing pre-trained models (such as BERT, GPT, etc.). That is to say, in this embodiment, the semantic structure splitting rules can be to train the pre-trained model with training text and annotated semantic units, so that they are built into the processing logic of the pre-trained model. This pre-trained model after training is the second text processing model.

[0108] In this embodiment, the task decomposition prompt text provides clear guidance to the second text processing model, enabling it to decompose the original natural query statement into multiple sub-query statements with hierarchical dependency relationships based on the semantic structure features of the natural query statement and the preset semantic structure splitting rules. This process is similar to dividing a large problem into multiple small problems, and each small problem (i.e., sub-query statement) is more specific and easier to handle. For example, for the query "Please query the per capita income of the top five cities in terms of population", the splitting results in "Query the top five cities in terms of population" and "Query the per capita income of the top five cities in terms of population".

[0109] After the splitting is completed, the model will understand each sub-query task corresponding to the sub-query statement one by one from bottom to top according to the hierarchical dependency relationships between these sub-query statements. This processing method from the specific to the abstract and from the local to the whole helps the model to more accurately grasp the user's query intention and generate corresponding sub-query instructions. Subsequently, the model will further integrate these instructions according to the sub-query instructions generated by each sub-query task to form a complete composite query instruction. This composite query instruction not only contains all the information of the user's original query, but also improves the accuracy and efficiency of the query through task decomposition and the generation of sub-query instructions.

[0110] In summary, after introducing the task decomposition prompt text in this embodiment, through task decomposition and the understanding of hierarchical dependency relationships, the model can more accurately grasp the user's query requirements and generate query instructions that better fit the user's intention. This not only improves the query accuracy rate, but also greatly shortens the query response time and enhances the user experience.

[0111] The above embodiments only illustrate several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and the present application also intends to cover these changes and modifications.

Claims

1. A query method based on query auxiliary information, characterized in that: The following steps are involved: Obtaining a natural query statement input by a user and user information of the user; Acquire query auxiliary information of the user according to the user information; wherein the query auxiliary information is additional information associated with the user for the natural query statement; Acquire various query fields of a preset constellation data model, and determine the query field associated with the natural query statement as a candidate query field; wherein the constellation data model includes a structured data table and query fields in the data table; The natural query statement, the candidate query fields, the query auxiliary information, and the preset first task processing text are input into a first text processing model to obtain a plurality of target query fields; wherein the first task processing text is used to prompt the first text processing model to understand the query intent of the natural query statement in combination with the query auxiliary information, and determine the candidate query fields that meet the query intent as target query fields; Determining at least one target data table including only the target query field from the constellation data model; Inputting the natural query statement and the target data table into a query instruction generation model to obtain a query instruction; wherein the query instruction generation model is used to understand the query intent of the natural query statement, determine a query task that meets the query intent, and generate a query instruction based on the target data table and the query task; Based on the query instruction, a query result of the natural query statement is obtained by querying from a data warehouse corresponding to the constellation data model.

2. The query method based on query auxiliary information according to claim 1, characterized in that: The query auxiliary information includes synonym information; The step of acquiring the user's query auxiliary information according to the user information comprises: Perform entity segmentation recognition on the natural query sentence to obtain a number of entity segmentations; Determine a synonym information database corresponding to the user information, and determine a synonym corresponding to any of the entity participles from the synonym information database; According to the synonyms of each of the entity participles, synonym information is obtained.

3. The query method based on query auxiliary information according to claim 2, characterized in that: Before the step of obtaining the query auxiliary information of the user according to the user information, the step further includes: In response to a synonym information creation instruction, the synonym information creation instruction is parsed to obtain user information and synonym information; the synonym information includes a standard vocabulary and a user's habitual expression vocabulary corresponding to the user information for the standard vocabulary; The synonym information is added to a synonym information library corresponding to the user information.

4. The query method based on query auxiliary information according to claim 2, characterized in that: The step of obtaining each query field of the preset constellation data model and determining the query field associated with the natural query statement as a candidate query field includes: Get various query fields of the preset constellation data model; A query field whose semantic similarity with the natural query statement or any of the entity segmentations is greater than a preset similarity threshold is determined as a candidate query field.

5. The query method based on query auxiliary information according to claim 2, characterized in that: The first task processing text is also used to prompt the first text processing model to understand the query intent of the natural query statement in combination with the input business knowledge; After the step of performing entity segmentation recognition on the natural query sentence to obtain a plurality of entity segmentations, the step further includes: Based on the natural query sentence and the entity segmentation, relevant business knowledge is retrieved from a preset business knowledge base; The step of inputting the natural query sentence, the candidate query fields, the query auxiliary information and the preset first task processing text into the first text processing model to obtain a plurality of target query fields includes: The natural query statement, the candidate query fields, the query auxiliary information, the business knowledge and the preset first task processing text are input into a first text processing model to obtain a plurality of target query fields.

6. The query method based on query auxiliary information according to claim 1, characterized in that: The query instruction generation model is a second text processing model; the step of inputting the natural query sentence and the target data table into the query instruction generation model to obtain the query instruction includes: The natural query statement, the target data table and the preset second task processing text are input into the second text processing model to obtain a query instruction; wherein the second task processing text is used to prompt the second text processing model to understand the query intent of the natural query statement, determine the query task that meets the query intent, and generate a query instruction based on the target data table and the query task.

7. The query method based on query auxiliary information according to claim 6, characterized in that: The second task processing text is also used to prompt the second text processing model to understand the query intent of the natural query statement in combination with the input query semantic enhancement information; The step of inputting the natural query statement, the target data table and the preset second task processing text into the second text processing model to obtain the query instruction includes: According to the user information, obtaining the query semantic enhancement information of the user from a preset query semantic enhancement information library; The natural query statement, the target data table, the query semantic enhancement information and the preset second task processing text are input into a second text processing model to obtain a query instruction.

8. The query method based on query auxiliary information according to claim 6, characterized in that: The query task pointed to by the natural query statement is a compound query task; the second task processing text is used to prompt the second text processing model to determine a compound query task that meets the query intent of the natural query statement, and generate a compound query instruction based on the target data table and the compound query task; the compound query instruction includes several query instructions with hierarchical dependencies.

9. The query method based on query auxiliary information according to claim 8, characterized in that: The second task processing text includes a task decomposition prompt text; the task decomposition prompt text is used to prompt the second text processing model to split the natural query statement into a plurality of sub-query statements with hierarchical dependencies according to the semantic structure characteristics of the natural query statement and a preset semantic structure splitting rule; understand the sub-query tasks corresponding to each of the sub-query statements from bottom to top according to the hierarchical dependencies, and generate sub-query instructions according to the sub-query tasks; and obtain a compound query instruction according to the sub-query instructions of each of the sub-query tasks.

10. The query method based on query auxiliary information according to claim 1, characterized in that: The target query field includes a dimension field and a metric field; the constellation data model records a number of data tables, basic fields included in each data table, and graph relationships of each data table; the graph relationship includes at least a directed graph relationship; The step of determining at least one target data table including only the target query field from the constellation data model comprises: Determine a data table in the constellation data model that contains any one or more of the metric fields as a target data table; Taking the target data table as the root node, determining other data tables in the constellation data model that can reach the target data table according to the graph relationship, taking each of the target data tables as a fact table and taking the other data tables that can be reached by the fact table as dimension tables, to obtain several target star data models; Remove the edge dimension tables that do not include the dimension fields in the target star data model, remove the basic fields other than the metric fields in any fact table of the target star data model, and remove the basic fields other than the dimension fields in any dimension table of the target star data model; A target data table is obtained according to the target star data model.

Citation Information

Patent Citations

  • Data query method and device based on text processing model

    CN118708704A

  • Large model training method and data query method based on large model

    CN118780398A

Cited By

  • Data query method and device based on interface generation, equipment and medium

    CN120873072A

  • Query method based on multi-agent collaboration

    CN121301386A

  • Query methods based on multi-agent collaboration

    CN121301386B