Processing method of composite query task

By determining the target query fields and data tables, and using the text processing model to generate composite query instructions, the problem of inaccurate processing of composite query tasks in the existing technology is solved, and more efficient and accurate query results are achieved.

CN120045698AActive Publication Date: 2025-05-27GUANGZHOU SMART SOFTWARE CO LTD

Patent Information

Application Number
CN202510084358.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-27
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

When handling composite query tasks, existing database query technology is difficult to accurately capture the true intention of user query, resulting in the query results that cannot accurately meet user needs.

Method used

By obtaining the natural query statement input by the user, determining the target query fields and data tables that match the query intent, subtask decomposition of the natural query statements using the text processing model, and generating compound query instructions with hierarchical dependencies.

Benefits of technology

It realizes accurate understanding and processing of user compound query tasks, improves the accuracy and efficiency of query, and can better meet users' compound query needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045698A_ABST
    Figure CN120045698A_ABST
Patent Text Reader

Abstract

The invention relates to a processing method of a composite query task, which comprises the following steps of: obtaining a natural query statement which is input by a user and relates to the composite query task, and firstly determining a target query field matched with a query intention in a preset constellation data model; then, based on the target query field, screening out a related data table from the constellation data model; and inputting the natural query statement, the screened data table and preset prompt information into a text processing model, and generating a composite query instruction with a logic level and a dependency relationship through the model. And finally, according to the generated query instruction, querying from the data warehouse to obtain a query result meeting user requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language query, and particularly to a method for processing compound query tasks. Background Art

[0002] Today, with the rapid development of information technology, users' demands for data query are becoming increasingly complex and diverse. Especially in the fields of big data analysis, business intelligence, etc., users often hope to quickly obtain data results that have been precisely calculated and filtered through query statements expressed in natural language. However, although existing database query technologies have, to a certain extent, realized the conversion between natural language and database query language, they still seem inadequate when faced with compound query tasks involving multiple data tables, multiple data types, and complex calculation logics.

[0003] Specifically, existing query processing technologies usually rely on keyword matching and simple semantic understanding of natural query statements to convert user input into corresponding SQL query statements. However, this conversion method often has difficulty accurately capturing the true intention of the user's query and generating accurate query statements when faced with compound query requirements, resulting in query results that do not accurately meet the user's needs. Summary of the Invention

[0004] Based on this, the purpose of this application is to provide a method for processing compound query tasks, aiming to accurately understand the complex query intention in the user's natural query statement, generate accurate query instructions, and thus obtain accurate query results.

[0005] The method for processing compound query tasks described in the embodiments of this application includes the following steps:

[0006] Obtain the natural query statement input by the user; the query task pointed to by the natural query statement is a compound query task;

[0007] Determine, from a preset constellation data model, the query fields that match the query intention of the natural query statement as target query fields; the constellation data model includes structured data tables and query fields in the data tables;

[0008] According to the target query fields, determine at least one target data table that only includes the target query fields from the constellation data model;

[0009] Input the natural query statement, the target data table, and a preset first task processing text into a first text processing model to obtain a composite query instruction; wherein, the first task processing text is used to prompt the first text processing model to understand the query intention of the natural query statement, determine a composite query task that conforms to the query intention, and generate a composite query instruction based on the target data table and the composite query task; the composite query instruction includes several query instructions with hierarchical dependencies.

[0010] Based on the composite query instruction, query the query result of the natural query statement from the data warehouse corresponding to the constellation data model.

[0011] In the embodiment of the present application, a natural query statement involving a composite query task input by a user is obtained. First, a target query field that matches the query intention in a preset constellation data model is determined. Then, based on the target query field, relevant data tables are filtered out from the constellation data model. On this basis, the natural query statement, the filtered data tables, and a preset prompt message are input into a text processing model. The model decomposes the composite query task of the natural query statement into subtasks, and generates a composite query instruction with a logical hierarchy and dependencies according to the query intention of the user and the structural information of the data table. Finally, according to the generated query instruction, a query result that meets the user's needs is queried from the data warehouse. In summary, the embodiment of the present application realizes accurate query of the user's composite query task, improves the accuracy and efficiency of the query, and has a wide range of application prospects.

[0012] For better understanding and implementation, the present application will be described in detail below with reference to the accompanying drawings. Description of the Drawings

[0013] Figure 1 It is a schematic flowchart of the processing method for the composite query task in the embodiment of the present application;

[0014] Figure 2 It is a schematic diagram of the steps for obtaining synonyms from the synonym information library corresponding to the user identity information in the embodiment of the present application. Detailed Embodiments

[0015] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings. Among them, when the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements.

[0016] It should be clear that the embodiments described in the following embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0017] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application are also intended to include the plural forms unless the context clearly indicates otherwise. In addition, in the description of the present application, unless otherwise specified, "a plurality of" means two or more. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone; the character " / " generally represents an "or" relationship between the associated objects before and after.

[0018] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. Moreover, these terms are only used to distinguish similar objects and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances. Depending on the context, the words "if" / "when" used in the present application can be interpreted as "when...", "when...", or "in response to a determination".

[0019] The present application relates to the field of natural language query technology. Although the existing database query technology has achieved the conversion between natural language and database query language to a certain extent, it still seems powerless when faced with complex query tasks involving multiple data tables, multiple data types, and complex calculation logics. Specifically, the existing query processing technology usually relies on keyword matching and simple semantic understanding of natural query statements to convert user input into corresponding SQL query statements. However, this conversion method often has difficulty accurately capturing the true intention of the user's query and generating accurate query statements when faced with complex query requirements, resulting in the query results not accurately meeting the user's needs. In addition, the existing technology also lacks an in-depth understanding of the complex relationships between data tables, making it easy to make mistakes when processing operations such as multi-table joins and data aggregations.

[0020] In view of the deficiencies of the natural query statement processing solution in the prior art when dealing with compound query tasks, the embodiments of the present application provide a method for processing compound query tasks, aiming to accurately understand the complex query intent in the user's natural query statement, generate accurate query instructions, and thus obtain accurate query results.

[0021] Please refer to Figure 1 , the method for processing compound query tasks described in the embodiments of the present application includes the following steps:

[0022] S101: Obtain the natural query statement input by the user; the query task pointed to by the natural query statement is a compound query task;

[0023] S102: Determine, from a preset constellation data model, the query field that matches the query intent of the natural query statement as the target query field; the constellation data model includes a structured data table and the query fields in the data table;

[0024] S103: According to the target query field, determine at least one target data table that only includes the target query field from the constellation data model;

[0025] S104: Input the natural query statement, the target data table, and a preset first task processing text into a first text processing model to obtain a compound query instruction; wherein, the first task processing text is used to prompt the first text processing model to understand the query intent of the natural query statement, determine the compound query task that conforms to the query intent, and generate a compound query instruction based on the target data table and the compound query task; the compound query instruction includes several query instructions with a hierarchical dependency relationship;

[0026] S105: Based on the compound query instruction, query the query result of the natural query statement from the data warehouse corresponding to the constellation data model.

[0027] The embodiments of the present application obtain the natural query statement input by the user involving a compound query task. First, the target query field that matches the query intent in the preset constellation data model is determined. Then, based on the target query field, the relevant data tables are screened out from the constellation data model. On this basis, the natural query statement, the screened data tables, and the preset prompt information are input into the text processing model. The model decomposes the compound query task of the natural query statement into subtasks, and generates a compound query instruction with a logical hierarchy and dependency relationship according to the user's query intent and the structure information of the data table. Finally, according to the generated query instruction, the query result that meets the user's needs is queried from the data warehouse. In summary, the embodiments of the present application achieve accurate query of the user's compound query task, improve the accuracy and efficiency of the query, and have a wide range of application prospects.

[0028] The processing method of the composite query task described in the embodiments of the present application takes a computer as the execution subject. Hereinafter, first, the text processing model of the embodiments of the present application will be described, and then each step will be described in detail.

[0029] In the embodiments of the present application, the text processing model can be based on one or more pre-trained language models, such as BERT, GPT, etc. (these models have been trained on a large amount of text data and can capture the deep semantic and syntactic features of language). In order to adapt to specific tasks in the present application, such as query intention understanding, query instruction generation, etc., it is obtained through fine-tuning. Of course, the text processing model can also be an existing large language model, as long as it has the processing ability for the specific tasks of the present application, it can be applied to the technical solution of the present application. In this regard, the embodiments of the present application do not make any restrictions.

[0030] In order to enable the text processing model to complete the task as required, it is usually necessary to provide task processing text. The task processing text is used to require or prompt the task requirements of the text processing model, including specific requirements and format descriptions for the model output, and can also include task-related scenario information or context information for the model to learn from and understand the task to be processed, etc. In the embodiments of the present application, different sub-texts can be set in the task processing text according to the needs of the query, respectively used to require or prompt the corresponding task requirements of the text processing model. The task processing text provided to the model in different task processing requirements is also correspondingly different, such as the first task processing text, the second task processing text, etc.

[0031] For step S101, obtain the natural query statement input by the user; the query task pointed to by the natural query statement is a composite query task.

[0032] Among them, the compound query task, different from the simple single-condition query task, involves the combination and collaborative work of multiple sub-query tasks, and there is a hierarchical dependency relationship between these sub-query tasks. Specifically, a compound query task cannot be completed through a single query operation, but requires multiple sub-query operations to be executed sequentially or in parallel according to a certain logical order and mutual dependency relationship to ultimately meet the user's information needs. Specifically, compound query tasks generally have the following characteristics: First, the combination of sub-query tasks: A compound query task contains two or more sub-query tasks, and these sub-query tasks together constitute a complete set of operations to achieve the final query goal. For example, in a business data analysis scenario, a user may enter the following natural query statement: "Find companies with an annual sales of more than 10 million, and calculate the average profit growth rate of these companies in the past three years." This query task contains two sub-tasks: one is to find companies with an annual sales of more than 10 million (which can be called sub-task A), and the other is to calculate the average profit growth rate of the companies that meet the conditions of sub-task A in the past three years (which can be called sub-task B). Second, hierarchical dependency relationship: There is a clear hierarchical dependency relationship between sub-query tasks, that is, the execution of subsequent sub-tasks depends on the results of previous sub-tasks. In the above example, the execution of sub-task B depends on the result of sub-task A, because only by first finding companies with an annual sales of more than 10 million can the average profit growth rate of these companies be calculated. This dependency relationship may be data dependency or logical dependency, and subsequent sub-tasks need to use the data or calculation results filtered by previous sub-tasks as input or operation basis. Third, logical complexity: Due to involving multiple sub-tasks and their dependency relationships, compound query tasks usually have high logical complexity. This not only requires accurate understanding and processing of the individual execution logics of each sub-task, but also needs to effectively manage and coordinate the relationships between them to ensure the smooth completion of the entire query task. For example, in a database environment with multiple data tables and complex associations, different sub-tasks may involve different data tables, and complex association operations (such as multi-table joins) and data filtering operations (such as conditional filtering, grouping and aggregation, etc.) are required, and the execution order and association methods of these operations need to be arranged according to the specific dependency relationship to ensure data consistency and the correctness of query results.

[0033] In this step, first, receive the natural query statement input by the user through the interface or other means. For example, the user may enter "Please query the top 3 customers in terms of sales among the top 3 brands in terms of sales, and the proportion of these customers in the brand sales", and this query statement clearly indicates the user's query requirement, that is, to obtain the top 3 customers in terms of sales among the top 3 brands in terms of sales and their proportion in the sales of each brand.

[0034] For step S102, determine, from a preset constellation data model, a query field that matches the query intent of the natural query statement as the target query field; the constellation data model includes a structured data table and query fields in the data table.

[0035] Among them, the constellation data model is an abstract representation used to describe the database structure. It is a data logic layer established in advance based on each data table stored in the data warehouse, which records several data tables, the basic fields included in each data table, and the graph relationships between each data table. In addition, it may also record the binding relationships between at least some basic fields and multi-dimensional operation fields. That is to say, the constellation data model does not record the specific business data of the fields of each data table, but only records the information related to generating query statements. These information mainly include the table names of the data tables, the field information of the data tables, and the multi-dimensional operation field information bound to the basic fields. In addition, a graph relationship between each data table is established in the constellation data model, and the graph relationship includes at least a directed graph relationship.

[0036] Among them, the basic field is generally a field in the database that records the original business data. More specifically, the basic field is generally divided into a measure field and a dimension field. The measure field is generally a field in the fact table, and the dimension field is generally a field in the dimension table. Among them, the fact table and the dimension table are concepts in data modeling in the data warehouse field. Currently, the main data models involved in the data warehouse field are the star data model, the snowflake data model generated on the basis of the star data model (which can actually be considered a kind of star data model), and the constellation data model. In the star data model, it includes a data table located in the center and other data tables connected to the central data table. Among them, the central data table is the fact table, and the other data tables connected to it are the dimension tables.

[0037] In this step, by matching the keywords and phrases in the user's query statement, determine, from the constellation data model, a query field that matches the query intent as the target query field. In this example, the target query fields may include "brand", "sales amount", "customer", and "proportion", etc.

[0038] In one embodiment, the step of determining, in step S102, a query field that matches the query intent of the natural query statement from a preset constellation data model as the target query field includes:

[0039] Step S1021, determine, from a preset constellation data model, a query field associated with the natural query statement as a candidate query field;

[0040] Step S1022: Input the natural query statement, the candidate query fields, and a preset second task processing text into a second text processing model to obtain a number of target query fields. Among them, the second task processing text is used to prompt the second text processing model to understand the query intent of the natural query statement and determine the candidate query fields that conform to the query intent as the target query fields.

[0041] Among them, in step S1021, traverse the preset constellation data model to find the query fields that may be related to the natural query statement. These fields are called candidate query fields, and they form the initial set for the subsequent screening process.

[0042] In step S1022, input the natural query statement, the candidate query fields, and the preset second task processing text into the second text processing model. This model is specially trained to understand the intent of natural language queries and can select the fields that best match the query intent from the candidate fields. The second task processing text plays a key role. It provides additional information to the model on how to understand the query intent. The model will deeply analyze the natural query statement based on this information and screen out the fields that most closely match the query intent from the candidate query fields as the target query fields.

[0043] In summary, in this embodiment, by introducing the second text processing model, and combining the preliminary screening of candidate query fields and the auxiliary understanding of the second task processing text, the accuracy and efficiency of determining query fields are significantly improved.

[0044] In one embodiment, the second task processing text in step S1022 is further used to prompt the second text processing model to output a question statement for the entity segmentation with unclear semantics when the natural query statement has such a situation. The second task processing text is also used to prompt the second text processing model to re-understand the query intent of the natural query statement in combination with the answer information input by the user for the question statement.

[0045] In this embodiment, a prompt text is added to the second task processing text, which is used to prompt the second text processing model on how to handle the situation where semantic ambiguity occurs in entity word segmentation during the processing of natural query statements. Specifically, when there are entity word segmentations with semantic ambiguity in a natural query statement, this usually means that some words or phrases in the query statement are unclear, ambiguous, or difficult to understand for the model. These entity word segmentations may be professional terms, personal names, place names, etc., which have no clear meaning or reference in the current context. The second task processing text will prompt the second text processing model that in this case, a question statement for these entity word segmentations with semantic ambiguity should be output. The purpose of this question statement is to ask the user for more information so that the model can more accurately understand the query intent.

[0046] Furthermore, the response information input by the user for the question statement will be provided to the second text processing model as additional context information. The second task processing text will also prompt the model to re-understand the query intent of the natural query statement in combination with this response information. This means that the model will adjust or correct its previous understanding based on the user's response, so as to more accurately answer the user's query.

[0047] In summary, in this embodiment, the second task processing text is used to enhance the ability of the second text processing model to process complex or ambiguous queries. When there are words with semantic ambiguity in a natural query statement, through interaction with the user, the model can obtain more clues about the query intent and, in combination with the user's response information, achieve an accurate understanding of the query intent.

[0048] In one embodiment, the step of determining, in step S1021, the query fields associated with the natural query statement from the preset constellation data model as candidate query fields includes:

[0049] Step S10211: Perform entity word segmentation recognition on the natural query statement to obtain a number of entity word segmentations;

[0050] Step S10212: Obtain each query field of the constellation data model; calculate the semantic similarity between each query field and the natural query statement as well as the entity word segmentations respectively, and determine the query fields with semantic similarity greater than the preset similarity threshold as candidate query fields.

[0051] Among them, in step S10211, entity word segmentation recognition is performed on the natural query statement. This process mainly uses entity recognition technology to extract the entity word segments in the query statement. These word segments represent the core information in the query statement and are the basis for subsequent matching of query fields. In step S10212, each query field in the preset constellation data model is obtained. Then, using the semantic similarity calculation algorithm, the semantic similarity between each query field and the natural query statement is calculated, and the semantic similarity between each query field and the recognized entity word segments is calculated. The purpose of this step is to find the query field that is semantically closest to the natural query statement or the entity word segments. Finally, according to the preset similarity threshold, the query fields with semantic similarity greater than the threshold are determined as candidate query fields. These fields have a high semantic matching degree with the natural query statement or the entity word segments and are therefore regarded as the fields that may meet the user's query requirements. In summary, in this embodiment, by introducing entity word segmentation recognition and semantic similarity calculation, the query fields associated with the natural query statement are accurately selected as candidate query fields.

[0052] In one embodiment, the second task processing text in step S1022 is further used to prompt the second text processing model to understand the query intention of the natural query statement in combination with the input business knowledge;

[0053] After the step of performing entity word segmentation recognition on the natural query statement in step S10211 to obtain a number of entity word segments, the following steps are further included:

[0054] Step S10213, retrieving relevant business knowledge from the preset business knowledge base based on the natural query statement and the entity word segments;

[0055] The step of inputting the natural query statement, the candidate query fields, and the preset second task processing text into the second text processing model in step S1022 to obtain a number of target query fields includes:

[0056] Step S10221, inputting the natural query statement, the candidate query fields, the business knowledge, and the preset second task processing text into the second text processing model to obtain a number of target query fields.

[0057] In this embodiment, the second task processing text guides the second text processing model to deeply understand the query intention of the natural query statement by combining the input business knowledge. After entity word segmentation recognition is performed on the natural query statement in step S10211 to obtain several entity word segments, in step S10213, based on the natural query statement and the recognized entity word segments, relevant business knowledge is retrieved from the preset business knowledge base. That is, business knowledge related to the natural query statement or entity word segments is obtained so that the query intention can be understood more accurately in subsequent steps. Then, in step S10221, the natural query statement, candidate query fields, retrieved business knowledge, and the preset second task processing text are input into the second text processing model together, thereby obtaining several target query fields. The improvement in this step is that business knowledge is introduced as an input, enabling the second text processing model to more comprehensively understand the query intention of the natural query statement and generate more accurate target query fields accordingly. In summary, by introducing business knowledge, the accuracy and efficiency of processing the natural query statement and obtaining the target query fields are improved.

[0058] In one embodiment, the second task processing text in step S1022 is further used to prompt the second text processing model to understand the query intention of the natural query statement by combining the input synonym information;

[0059] After the step of performing entity word segmentation recognition on the natural query statement in step S10211 to obtain several entity word segments, the following steps are further included:

[0060] Step S10214, determining the synonyms corresponding to any one of the entity word segments from the preset synonym information library; obtaining synonym information according to the synonyms of each entity word segment;

[0061] The step of inputting the natural query statement, the candidate query fields, and the preset second task processing text into the second text processing model in step S1022 to obtain several target query fields includes:

[0062] Step S10222, inputting the natural query statement, the candidate query fields, the synonym information, and the preset second task processing text into the second text processing model to obtain several target query fields.

[0063] In this embodiment, considering that when a user expresses a certain standard vocabulary (i.e., a common, widely accepted vocabulary that can be accurately understood by the large model), due to factors such as different expression habits, cultural backgrounds, and professional fields, the user may use a vocabulary that has the same or similar meaning but different form as the standard vocabulary. These vocabularies may be special expressions in a certain field or have specific meanings in certain contexts. Therefore, they may not be common vocabularies and cannot be directly and accurately understood by the large model. To solve this problem, it is necessary to provide the corresponding standard vocabulary, that is, synonyms, for these habitual vocabularies so that the large model can accurately understand the user's natural query statement.

[0064] Among them, the second task processing text also guides the second text processing model to understand the query intention of the natural query statement by combining the input synonym information. In step S10214, for each entity token, check whether there is a corresponding synonym in the preset synonym information library. If there is a corresponding synonym for the entity token, integrate it into the synonym information. In step S10222, input the natural query statement, the candidate query field, the synonym information, and the preset second task processing text into the second text processing model. In this way, the model can more deeply understand the query intention of the natural query statement by combining the synonym information and generate a more accurate target query field accordingly.

[0065] Please refer to Figure 2 , in one embodiment, the step of determining the synonym corresponding to any one of the entity tokens from the preset synonym information library in step S10214 includes:

[0066] Step S102141, obtain the user identity information of the input natural query statement;

[0067] Step S102142, determine the synonym information library corresponding to the user identity information;

[0068] Step S102143, determine the synonym corresponding to any one of the entity tokens from the synonym information library.

[0069] In this embodiment, considering the differences in the habitual expression terms of different types of users, different synonym information libraries are respectively set for different types of users. Specifically, this embodiment maintains multiple synonym information libraries, and each library is optimized for different types of users or user groups. For example, one library may contain more professional terms and is suitable for professionals; another library may contain more common vocabularies and is suitable for ordinary users.

[0070] In step S102141, user identity information can be obtained through various means such as user login information, session context, device information, IP address, etc. This information may directly point to a specific user account, or it may be necessary to determine the user's identity through further analysis and matching. In step S102142, by matching the user identity information, it is possible to determine which synonym information library should be used. In step S102143, in the given synonym information library, an attempt is made to find corresponding synonyms for each entity token in the natural query statement. This process may involve various techniques such as string matching and semantic similarity calculation.

[0071] In summary, this embodiment can determine the most appropriate synonyms for the entity tokens in the natural query statement from the preset synonym information library according to the user's identity information, improving the accuracy of the model's understanding of the query intent of the natural query statement and generating more accurate target query fields accordingly.

[0072] For step S103, according to the target query field, determine at least one target data table in the constellation data model that only includes the target query field.

[0073] In this step, according to the determined target query field, further screen out the target data tables in the constellation data model that contain these fields. For example, these target data tables may include brand sales data tables, customer information data tables, etc., to ensure that the selected data tables can cover all the information required by the user's query.

[0074] In one embodiment, the target query field includes a dimension field and a measure field; the step of step S103 of determining at least one target data table in the constellation data model that only includes the target query field according to the target query field includes:

[0075] Step S1031, determine the data tables in the constellation data model that contain any one or more of the measure fields as the root data tables;

[0076] Step S1032, using the root data table as the root node, according to the graph relationship, determine the other data tables in the constellation data model that can reach the root data table. Respectively, using each root data table as the fact table and the other data tables that can be reached corresponding to the fact table as the dimension tables, to obtain several target star data models;

[0077] Step S1033, remove the dimension tables on the edges of the target star data model that do not include the dimension field, remove the base fields other than the measure fields in the fact table of any one of the target star data models, and remove the base fields other than the dimension field in the dimension table of any one of the target star data models;

[0078] Step S1034: Obtain a target data table according to the target star data model.

[0079] In this embodiment, the target query fields include measure fields and dimension fields. First, use the data tables in the constellation data model that contain any one or more measure fields as root data tables; use the fact tables centered on the root data tables and the other data tables that can reach the fact tables as dimension tables to determine several target star data models; remove the dimension tables irrelevant to the query and the basic fields (measure fields or dimension fields) irrelevant to the query in each target star data model, so that the finally obtained target star data model only includes the fact tables, dimension tables, and basic fields related to the query. Through the method of this embodiment of the present application, it is realized to trim the pre-modeled constellation data model according to the target query fields, and determine the target data table based on the trimmed minimum available star data model, avoiding querying other irrelevant data tables and irrelevant query fields, reducing the complexity of the first text processing model's understanding, and making the generation of query instructions more efficient and accurate.

[0080] For step S104, input the natural query statement, the target data table, and a preset first task processing text into a first text processing model to obtain a composite query instruction; wherein, the first task processing text is used to prompt the first text processing model to understand the query intention of the natural query statement, determine a composite query task that conforms to the query intention, and generate a composite query instruction based on the target data table and the composite query task; the composite query instruction includes several query instructions with a hierarchical dependency relationship.

[0081] In this step, input the natural query statement, the target data table, and a preset first task processing text into the first text processing model. The first text processing model understands and parses the natural query statement, and generates a composite query instruction that conforms to the user's query intention according to the prompt of the task processing text. The task processing text contains preset prompts or instructions for guiding the model to understand the hierarchical structure and logical relationship of the query. Among them, the composite query instruction is composed of multiple query instructions with a hierarchical dependency relationship. In this example, the composite query instruction may include two levels of queries: first, query the top three brands in terms of sales volume, and then query the top three customers among these brands and calculate the proportion of these customers in the sales volume of each brand.

[0082] In one embodiment, the first task processing text in step S104 includes task decomposition prompt text; the task decomposition prompt text is used to prompt the first text processing model to split the natural query statement into several sub-query statements with a hierarchical dependency relationship according to the semantic structure characteristics of the natural query statement and a preset semantic structure splitting rule; understand each sub-query task corresponding to the sub-query statement from bottom to top according to the hierarchical dependency relationship, and generate sub-query instructions according to the sub-query task; obtain a composite query instruction according to the sub-query instructions of each sub-query task.

[0083] Among them, the semantic structure splitting rule is used to analyze and process natural language text and split the natural language text into units with clear semantic components and structures. Specifically, the semantic structure splitting rule involves multiple aspects of natural language processing, including syntactic analysis, semantic role labeling, entity recognition, relation extraction, etc., which can be achieved by fine-tuning existing pre-trained models (such as BERT, GPT, etc.). That is to say, in this embodiment, the semantic structure splitting rule can be to train the pre-trained model with training text and labeled semantic units, so that it is built into the processing logic of the pre-trained model. This pre-trained model after training is the first text processing model.

[0084] In this embodiment, the task decomposition prompt text provides clear guidance to the first text processing model, enabling it to decompose the original natural query statement into multiple sub-query statements with a hierarchical dependency relationship according to the semantic structure characteristics of the natural query statement and the preset semantic structure splitting rule. This process is similar to dividing a big problem into multiple small problems, and each small problem (i.e., sub-query statement) is more specific and easier to handle. For example, for the query "Please query the per capita income of the top five cities in terms of population", it is split into "Query the top five cities in terms of population" and "Query the per capita income of the top five cities in terms of population".

[0085] After the splitting is completed, the model will understand each sub-query task corresponding to the sub-query statement one by one from bottom to top according to the hierarchical dependency relationship between these sub-query statements. This processing method from concrete to abstract and from local to whole helps the model to more accurately grasp the user's query intention and generate corresponding sub-query instructions. Subsequently, the model will further integrate these instructions according to the sub-query instructions generated by each sub-query task to form a complete composite query instruction. This composite query instruction not only contains all the information of the user's original query, but also improves the accuracy and efficiency of the query through task decomposition and the generation of sub-query instructions.

[0086] In summary, after introducing the task decomposition prompt text in this embodiment, through the understanding of task decomposition and hierarchical dependencies, the model can more accurately grasp the user's query requirements and generate query instructions that better fit the user's intention. This not only improves the accuracy of the query but also greatly shortens the query response time and enhances the user experience.

[0087] In one embodiment, the first task processing text in step S104 is further used to prompt the first text processing model to understand the query intention of the natural query statement in combination with the input query semantic enhancement information;

[0088] The step of inputting the natural query statement, the target data table, and the preset first task processing text into the first text processing model in step S104 to obtain a composite query instruction includes:

[0089] Step S1041, obtain the preset query semantic enhancement information; input the natural query statement, the target data table, the business knowledge information, and the preset first task processing text into the first text processing model to obtain a composite query instruction.

[0090] In this embodiment, the first task processing text guides the first text processing model to more deeply understand the query intention of the natural query statement in combination with the input query semantic enhancement information.

[0091] The query semantic enhancement information includes business knowledge related to the natural query statement, data table processing, and relevant guidance for generating query instructions. Specifically, it may include date and time conditions. For example, it is required that all queries must include date and time conditions. If the natural query statement does not explicitly specify the date and time, the date (year-month-day) equal to "2024-10-20" is default added as a filtering condition; if it has been specified, there is no need to add. It may also include data table processing guidance. For example, when using sql_json for query, only query data from one table and need to check whether all fields are in the table. It may also include query field selection. For example, when words such as "each" or "every" are included in the user's request, the subsequent fields need to be added to the SELECT part to ensure that the query results include all necessary fields. When the user's request contains descriptions such as TOPN or "the first N", the subsequent fields also need to be added to the SELECT part. It may also include set operations and result display requirements. For example, when the word "list" is included in the user's request, the set needs to be completely retained, and null values should be filled in for records without values. It may also include special indicator calculations. For example, for indicators related to the number of customers (such as the number of investment consulting customers), year-on-year, same-period values, month-on-month, previous-period values, etc., cannot be directly obtained from the table but need to be calculated using a specified plugin.

[0092] In summary, in this embodiment, query semantic enhancement information is additionally obtained in step S104 and used as part of the input, enabling the first text processing model to generate a more accurate composite query instruction that meets business requirements.

[0093] In one embodiment, the step of obtaining the preset query semantic enhancement information in step S1041 includes:

[0094] Step S10411, obtaining the user identity information of the input natural query statement;

[0095] Step S10412, obtaining the query semantic enhancement information corresponding to the user identity information from a preset query semantic enhancement information library.

[0096] This embodiment takes into account the differences in the identity backgrounds and query permissions of different users. Therefore, different query semantic enhancement information is set for different users respectively. This embodiment maintains a query semantic enhancement information library, which can adjust the query semantic enhancement information for different users or user groups. The adjustment can involve any aspect. For example, different settings can be made for different users in any aspect related to the business knowledge of the natural query statement, data table processing, or generation of query instructions. Thus, the first text processing model can more appropriately understand the user's natural query statement and generate a query instruction according to the required specifications.

[0097] Specifically, step S10411 first identifies and obtains the identity information of the user of the input natural query statement. Step S10412 uses the user's identity information to search for the corresponding query semantic enhancement information in the preset query semantic enhancement information library. This information library may be a database, configuration file, or data structure in memory, which stores the mapping relationship between different user identity information and the corresponding query semantic enhancement information. In one embodiment, the query semantic enhancement information may include, but is not limited to: First, business term mapping: mapping informal or specific business terms that the user may use to standard terms used internally in the system. Second, data table preference: the data tables or fields that the user often queries. Third, query pattern: the query structure or pattern that the user commonly uses, such as sorting method, filtering conditions, etc. Fourth, permission restriction: determining the data tables and fields that the user can access according to the user's permission level.

[0098] In summary, this embodiment can obtain relevant enhancement information from a preset query semantic enhancement information library according to the user's identity information. These information will subsequently be used in step S104 and input into the first text processing model together with the natural query statement, the target data table, and the preset first task processing text, etc., to generate a more accurate composite query instruction that meets the user's needs and can better adapt to the needs and query habits of different users.

[0099] In this embodiment, the first text processing model generates corresponding query instructions. Specifically, the query instructions can be to generate a structured query statement such as an SQL statement for querying data in a database; or to generate a set of codes such as a Python script, which is executed by a code executor to query data from the database and perform specific calculations or processing.

[0100] In one embodiment, the first task processing text in step S104 further includes a preset task prompt text, and the task prompt text includes usage scenario information and call interfaces of each preset function plugin; the task prompt text is used to prompt the first text processing model to understand the usage scenario corresponding to the query task, determine a matching function plugin according to the usage scenario information of each function plugin, obtain the call interface corresponding to the function plugin, and generate an execution code (i.e., a query instruction) according to the call interface.

[0101] In this embodiment, the task prompt text details the usage scenario information and corresponding call interfaces of each preset function plugin. These information are directly related to whether the first text processing model can accurately understand the usage scenario of the query task and accordingly select the correct function plugin to execute the query task. Specifically, the usage scenario information in the task prompt text describes key elements such as the type of query task applicable to each function plugin, the input data format, and the expected output result. These information provide a clear framework for the first text processing model, enabling it to screen out the most suitable one or more from numerous preset function plugins according to the specific requirements of the query task to execute the query task. At the same time, the task prompt text also contains the call interface information of each function plugin. The call interface is the bridge for the query instruction executor to interact with the function plugin, which defines how the query instruction executor calls the function plugin, passes parameters, and receives return results, etc. Based on the call interface information, the first text processing model can accurately generate the execution code, thereby enabling the query instruction executor to implement the call and execution of the function plugin. Similarly, in other embodiments, when the first text processing model needs to process other tasks, the task prompt text can be set to prompt the first text processing model to process the corresponding tasks and prompt relevant processing requirements.

[0102] In summary, the task prompt text of this embodiment provides a clear and accurate guiding framework for the first text processing model, enabling it to more quickly understand the specific requirements and usage scenarios of the query task and accordingly select the most suitable function plugin to generate the query instruction.

[0103] In one embodiment, the task prompt text further includes usage prompt information for the call interfaces of the respective function plugins; the usage prompt information is used to prompt the first text processing model to understand the target parameter information of the query task when generating execution code according to the call interface, and pass the target parameter information into the call interface; wherein, the target parameter information includes a base field, a dimension field, a qualification condition, and an analysis metric; the target parameter information is used to instruct the corresponding function plugin to query target data according to the base field and the dimension field, and perform the analysis of the analysis metric on the target data under the qualification condition.

[0104] Among them, the usage prompt information is a key component of the task prompt text, which details the specific usage methods and related details of each function plugin call interface. These information are related to whether the first text processing model can accurately pass the target parameter information in the query task into the call interface in the specified format and order, so as to generate effective execution code.

[0105] Specifically, the usage prompt information not only defines the structure of the target parameter information, but also clarifies the specific meaning and usage of these parameter information in the query task.

[0106] The target parameter information includes multiple key elements such as a base field, a dimension field, a qualification condition, and an analysis metric. Among them, the base field is the basic data element involved in the query task, which constitutes the data basis of the query task. For example, when the query task is "analyze whether the sales volume in the Guangzhou area in the past 3 months is abnormal", "sales volume" is the base field. The dimension field is used to refine and classify the base field to more accurately locate and analyze the target data. In the above example, "monthly" (in the past 3 months) is the dimension field. The qualification condition is used to screen and filter the query results to ensure that the query results meet the actual requirements. The qualification condition can include a time range, a data range, etc. In the above example, "region = Guangzhou" limits the data range to the Guangzhou area only. The analysis metric is the core purpose of the query task, which indicates what kind of analysis needs to be performed on the target data. In the above example, "data anomaly" is the analysis metric, which requires the function plugin to analyze the sales volume in the Guangzhou area in the past 3 months and give a predicted value of whether it is abnormal.

[0107] In summary, the usage prompt information not only clearly describes the meaning and usage of these parameter information, but also guides the first text processing model on how to correctly pass them into the call interface. In this way, the function plugin can accurately query the target data according to the passed target parameter information, and perform the analysis of the corresponding analysis metric on the target data under the qualification condition.

[0108] In summary, using the usage hint information provides a clear, detailed, and specific guiding framework for the first text processing model, enabling it to more accurately understand the target parameter information in the query task and correctly pass them into the call interface in the specified format and order. Additionally, since the usage hint information details each element of the target parameter information, their relationships, and their specific meanings and uses in the query task, the first text processing model can easily handle various complex and diverse query tasks. At the same time, since the usage hint information of the call interface is preset in the task hint text, it is easy to update and expand the functional plug-in without making large-scale modifications and adjustments to the first text processing model.

[0109] For step S105, based on the composite query instruction, query the query result of the natural query statement from the data warehouse corresponding to the constellation data model.

[0110] Among them, the composite query instruction is a set of instructions generated during the process of processing composite query tasks. It is the structured information formed after a series of processing and transformation of the natural query statement input by the user. The composite query instruction transforms the complex, multi-dimensional, and hierarchically dependent information requirements expressed by the user in natural language into a sequence of operation instructions that can be understood and executed by the machine, aiming to guide the system or model to accurately execute the corresponding data query and processing operations to achieve the user's composite query task. Specifically, the composite query instruction generally has the following characteristics: First, the composite query instruction is composed of multiple query instructions, and each query instruction corresponds to a subtask in the composite query task. These query instructions are not simply listed, but are arranged in an orderly manner according to the logical and dependency relationships between the subtasks in the composite query task. For example, for a composite query task involving "finding departments with an annual sales volume exceeding 5 million and calculating the average employee turnover rate of these departments in the past two years", its composite query instruction will contain two or more specific query instructions, one for finding departments that meet the annual sales volume condition, and the other for calculating the average employee turnover rate of the departments that meet the condition. Second, there is an execution order corresponding to the hierarchical dependency relationship in the composite query task among the various query instructions in the composite query instruction. This execution order ensures the orderly execution of each subtask to ensure the coherence and accuracy of the entire query task. In the above example, the query instruction for calculating the average employee turnover rate must be executed after the query instruction for finding departments with an annual sales volume exceeding 5 million, because it depends on the department data filtered by the previous instruction. This reflects the strictness of the execution order of the composite query instruction to avoid incorrect results caused by chaotic execution order. Third, the composite query instruction is presented in a structured form, which can be in text form or in the form of a data structure representation inside the computer system. For example, in some systems, it may be stored in the form of a nested list, an object array, or a tree structure. This structured representation enables the system or model to clearly identify the specific content of each query instruction, its execution order, and the dependency relationships between them, facilitating subsequent execution and processing.

[0111] In this step, according to the generated composite query instruction, a query operation is performed on the data warehouse corresponding to the constellation data model, and the query result is returned. The query result may be a table or report containing the required information, and the user can view and analyze these results through the interface.

[0112] In one embodiment, the step of querying the query result of the natural query statement from the data warehouse corresponding to the constellation data model based on the composite query instruction in step S105 includes:

[0113] Step S1051: Input the composite query instruction into a preset query instruction executor. Based on the hierarchical dependency relationship through the query instruction executor, determine the sequential execution order of each sub-query instruction in the composite query instruction, and sequentially execute each sub-query instruction according to the sequential execution order to query and obtain the query result of the natural query statement from the data warehouse corresponding to the constellation data model.

[0114] In this embodiment, the carefully generated composite query instruction is input into a preset query instruction executor. This composite query instruction already contains all the information of the user's query requirements, and there is a clear hierarchical dependency relationship between each sub-query instruction. After receiving the composite query instruction, the query instruction executor determines the sequential execution order of each sub-query according to the hierarchical dependency relationship in the instruction, ensuring that the query process can proceed in a logical order and avoiding data chaos or query errors. After determining the execution order, the query instruction executor will sequentially execute each sub-query instruction according to this order. During the execution process, relevant data is retrieved and processed from the data warehouse corresponding to the constellation data model to ensure that each sub-query can obtain the correct result. As each sub-query instruction is sequentially executed, the query instruction executor will gradually collect and integrate the results of these sub-queries. Finally, these results are summarized into a complete query result, which will accurately reflect the user's original query requirements. In summary, this embodiment significantly improves the execution efficiency and accuracy of the composite query instruction by introducing a preset query instruction executor and implementing the determination of the sequential execution order of sub-query instructions based on the hierarchical dependency relationship.

[0115] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and the present application also intends to include these changes and modifications.

Claims

1. A method for processing a composite query task, characterized in that: The following steps are involved: Acquire a natural query statement input by a user; the query task pointed to by the natural query statement is a compound query task; Determining a query field that matches the query intent of the natural query statement from a preset constellation data model as a target query field; the constellation data model includes a structured data table and a query field in the data table; According to the target query field, determining at least one target data table including only the target query field from the constellation data model; The natural query statement, the target data table and the preset first task processing text are input into the first text processing model to obtain a compound query instruction; wherein the first task processing text is used to prompt the first text processing model to understand the query intent of the natural query statement, determine a compound query task that meets the query intent, and generate a compound query instruction based on the target data table and the compound query task; the compound query instruction includes a plurality of query instructions with a hierarchical dependency relationship; Based on the compound query instruction, a query result of the natural query statement is obtained by querying from a data warehouse corresponding to the constellation data model.

2. The method for processing a complex query task according to claim 1, characterized in that: The first task processing text includes a task decomposition prompt text; the task decomposition prompt text is used to prompt the first text processing model to split the natural query statement into a plurality of sub-query statements with hierarchical dependencies according to the semantic structure characteristics of the natural query statement and a preset semantic structure splitting rule; understand the sub-query tasks corresponding to each of the sub-query statements from bottom to top according to the hierarchical dependencies, and generate sub-query instructions according to the sub-query tasks; and obtain a compound query instruction according to the sub-query instructions of each of the sub-query tasks.

3. The method for processing a complex query task according to claim 2, characterized in that: The step of querying the data warehouse corresponding to the constellation data model based on the compound query instruction to obtain the query result of the natural query statement includes: The compound query instruction is input into a preset query instruction executor, and the query instruction executor determines the execution order of each sub-query instruction in the compound query instruction based on the hierarchical dependency relationship, and executes each sub-query instruction in sequence according to the execution order, and obtains the query result of the natural query statement from the data warehouse corresponding to the constellation data model.

4. The method for processing a complex query task according to claim 1, characterized in that: The step of determining, from the preset constellation data model, a query field that matches the query intent of the natural query sentence as a target query field comprises: Determining, from a preset constellation data model, a query field associated with the natural query statement as a candidate query field; The natural query statement, the candidate query fields and the preset second task processing text are input into a second text processing model to obtain a plurality of target query fields; wherein the second task processing text is used to prompt the second text processing model to understand the query intent of the natural query statement, and to determine the candidate query fields that meet the query intent as target query fields.

5. The method for processing a complex query task according to claim 4, characterized in that: The step of determining, from the preset constellation data model, a query field associated with the natural query statement as a candidate query field comprises: Perform entity segmentation recognition on the natural query sentence to obtain a number of entity segmentations; Acquire each query field of the constellation data model; calculate semantic similarity between each query field and the natural query statement and the entity segmentation respectively, and determine the query field whose semantic similarity is greater than a preset similarity threshold as a candidate query field.

6. The method for processing a complex query task according to claim 5, characterized in that: The second task processing text is also used to prompt the second text processing model to understand the query intent of the natural query statement in combination with the input business knowledge; After the step of performing entity segmentation recognition on the natural query sentence to obtain a plurality of entity segmentations, the step further includes: Based on the natural query sentence and the entity segmentation, relevant business knowledge is retrieved from a preset business knowledge base; The step of inputting the natural query statement, the candidate query fields, and the preset second task processing text into the second text processing model to obtain a plurality of target query fields includes: The natural query statement, the candidate query fields, the business knowledge and the preset second task processing text are input into a second text processing model to obtain a plurality of target query fields.

7. The method for processing a complex query task according to claim 5, characterized in that: The second task processing text is also used to prompt the second text processing model to understand the query intent of the natural query sentence in combination with the input synonym information; After the step of performing entity segmentation recognition on the natural query sentence to obtain a plurality of entity segmentations, the step further includes: Determine a synonym corresponding to any of the entity participles from a preset synonym information database; obtain synonym information based on the synonyms of each of the entity participles; The step of inputting the natural query statement, the candidate query fields, and the preset second task processing text into the second text processing model to obtain a plurality of target query fields includes: The natural query sentence, the candidate query fields, the synonym information and the preset second task processing text are input into a second text processing model to obtain a plurality of target query fields.

8. The method for processing a complex query task according to claim 5, characterized in that: The second task processing text is also used to prompt the second text processing model to output a question statement for the semantically unclear entity segmentation when the natural query statement contains semantically unclear entity segmentation; the second task processing text is also used to prompt the second text processing model to re-understand the query intent of the natural query statement in combination with the answer information input by the user for the question statement.

9. The method for processing a complex query task according to claim 1, characterized in that: The first task processing text is also used to prompt the first text processing model to understand the query intent of the natural query statement in combination with the input query semantic enhancement information; The step of inputting the natural query statement, the target data table and the preset first task processing text into the first text processing model to obtain a compound query instruction includes: Acquire preset query semantic enhancement information; input the natural query statement, the target data table, the business knowledge information and the preset first task processing text into a first text processing model to obtain a compound query instruction.

10. The method for processing a complex query task according to claim 1, characterized in that: The target query field includes a dimension field and a metric field; the constellation data model records a number of data tables, basic fields included in each data table, and graph relationships of each data table; the graph relationship includes at least a directed graph relationship; The step of determining at least one target data table including only the target query field from the constellation data model according to the target query field comprises: Determine a data table in the constellation data model that contains any one or more of the metric fields as a root data table; Taking the root data table as the root node, determining other data tables in the constellation data model that can reach the root data table according to the graph relationship, taking each of the root data tables as a fact table and taking the other data tables that can be reached by the fact table as dimension tables, to obtain several target star data models; Remove the edge dimension tables that do not include the dimension fields in the target star data model, remove the basic fields other than the metric fields in any fact table of the target star data model, and remove the basic fields other than the dimension fields in any dimension table of the target star data model; A target data table is obtained according to the target star data model.

Citation Information

Patent Citations

  • Data processing method and device, equipment and computer readable storage medium

    CN114064812A

  • Data query method and device based on text processing model

    CN118708704A

  • Custom multi-dimensional analysis configuration method, system and equipment based on constellation model and medium

    CN119311687A

  • Question Answering Framework for Structured Query Languages

    US20140149446A1

Cited By

  • Data query method and system

    CN121255828A