Query method and device based on user information and storage medium

By identifying the semantic entity participle in the user query statement and using user-specific semantic processing text library and constellation data model to optimize the query statement, the problem of query intention ambiguity caused by differences in user background information is solved, and the accuracy and correlation of query results are improved.

CN120470110APending Publication Date: 2025-08-12GUANGZHOU SMART SOFTWARE CO LTD

Patent Information

Application Number
CN202510449388.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Existing natural language query methods ignore differences in user background information, resulting in ambiguity in the understanding of query intentions, affecting the accuracy and relevance of query results.

Method used

By obtaining the query statements and user information input by the user, identifying entity participles with unclear semantics, and using the user-specific semantic processing text library to obtain target processing text with clear semantics, optimizing the query statements, and combining the constellation data model and query model to generate query instructions, so as to improve the accuracy and relevance of query results.

Benefits of technology

It significantly improves the accuracy and relevance of query results, enhances user query experience, and provides support for the practical application of natural language processing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470110A_ABST
    Figure CN120470110A_ABST
Patent Text Reader

Abstract

The invention relates to a query method and device based on user information and a storage medium, and the method comprises the steps: obtaining a first natural query statement input by a user and the user information of the user, and recognizing entity segmented words in the first natural query statement; entity segmented words with unknown semantics are determined, and a target processing text capable of accurately expressing the meanings of the entity segmented words is obtained according to the user information. And optimizing the first natural query statement according to the target processing text to obtain a second natural query statement. And based on the second natural query statement, determining a target data table in a preset constellation data model, and inputting the second natural query statement and the target data table into the query large model to generate a query instruction. And finally, querying according to the query instruction to obtain a query result. According to the method, the original query statement is optimized by combining the semantic processing text library of the user, so that the semantics of the original query statement is clarified, the meaning expressed by the user can be clarified, and the understanding accuracy of the large query model on the query intention is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language query technology, and in particular to a query method, device and storage medium based on user information. Background Art

[0002] In the field of natural language processing, user-entered queries often contain entities related to their specific context, such as department names and personal historical query records. The meaning of these entities in natural language sentences can be completely different for different users. For example, when different users mention entities such as "this department" or "last query" in their queries, they actually correspond to different semantics. Existing query processing methods typically directly use user-entered queries as input, ignoring differences in user context. This can lead to ambiguity in the natural language model's understanding of query intent, resulting in inaccurate query instructions, affecting the accuracy and relevance of query results. Summary of the Invention

[0003] Based on this, the purpose of this application is to provide a query method based on user information, which can effectively combine user information and clarify semantically ambiguous entities, so that the query is successful and accurate.

[0004] The query method based on user information described in the embodiment of the present application includes the following steps: Obtaining a first natural query statement input by a user and user information of the user; Performing entity segmentation recognition on the first natural query sentence to obtain a plurality of first entity segmentations; Determining a semantically ambiguous first entity participle among the plurality of first entity participles as a second entity participle; obtaining a target processing text for the second entity participle based on the second entity participle and the user information; the target processing text being a text with clear semantics corresponding to the second entity participle obtained by searching a semantic processing text library corresponding to the user information; Obtaining a second natural query statement according to the first natural query statement and the target processed text; Based on the second natural query statement, determining a target data table in a preset constellation data model; Inputting the second natural query statement and the target data table into a query macro model to obtain a query instruction; wherein the query macro model is used to understand the query intent of the second natural query statement and determine a query task that meets the query intent; and generating a query instruction based on the target data table and the query task; Based on the query instruction, query the data warehouse corresponding to the constellation data model to obtain a query result of the second natural query statement; and determine the query result as the query result of the first natural query statement.

[0005] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed, the device where the computer-readable storage medium is located is controlled to implement a method as described in any one of the embodiments of the present application.

[0006] An embodiment of the present application further provides an electronic device, comprising a processor, a memory, and a computer-readable program stored in the memory, wherein the computer-readable program, when executed by the processor, implements the steps of the method described in any one of the embodiments of the present application.

[0007] The embodiment of the present application first obtains a first natural query statement entered by a user and the user's user information, and identifies entity segmentation words in the first natural query statement. Entity segmentation words with unclear semantics are identified, and based on the user information, a target processing text that accurately expresses the meaning of these entity segmentation words is retrieved from the user's semantic processing text library. The first natural query statement is then optimized based on the target processing text to obtain a second natural query statement. Based on the second natural query statement, a target data table is determined in a preset constellation data model, and the second natural query statement and the target data table are input into a query model to generate a query instruction. Finally, a query result is obtained from the data warehouse based on the query instruction. By combining the user's semantic processing text library, the original natural query statement is optimized to make its semantics clear, clarify the meaning expressed by the user, and improve the accuracy of the query model's understanding of the query intent, thereby generating a more accurate query instruction. This not only significantly improves the accuracy and relevance of the query results, but also enhances the user's query experience, providing strong support for the promotion of natural language processing technology in practical applications.

[0008] For better understanding and implementation, the present application is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 A flowchart of a user information-based query method according to an embodiment of the present application; Figure 2 Schematic diagram of the steps of determining the second entity segmentation and obtaining the target processing text in an embodiment of the present application; Figure 3 A schematic diagram of the steps for creating the third processed text data in an embodiment of the present application; Figure 4 This is a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0010] To make the objectives, technical solutions, and advantages of this application more clear, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements.

[0011] It should be understood that the embodiments described in the following examples do not represent all embodiments consistent with this application. Rather, they are merely examples of devices and methods consistent with certain aspects of this application, as detailed in the appended claims. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of this application without inventive effort are intended to fall within the scope of protection of this application.

[0012] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "the" and "the" used in this application are also intended to include plural forms, unless the context clearly indicates otherwise. In addition, in the description of this application, unless otherwise stated, "a plurality" refers to two or more. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone; the character " / " generally indicates that the objects associated before and after are in an "or" relationship.

[0013] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, this information should not be limited to these terms. Moreover, these terms are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor can they be understood to indicate or imply relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to the specific circumstances. Depending on the context, the words "if" / "if" used in this application can be interpreted as "at the time of" or "when" or "in response to determining".

[0014] The present application relates to the field of natural language query technology. In the field of natural language processing, the query statements input by users often contain entities related to the user's specific background, such as department names, personal historical query records, etc. The meanings of these entities in natural statements may be completely different for different users. For example, when entities such as "this department" and "the most recent query" are mentioned in the query statements of different users, they actually correspond to different semantics. Existing query processing methods usually directly take the query statements input by users as input, ignoring the differences in user background information, which may cause ambiguity in the natural language model when understanding the query intent, and then generate inaccurate query instructions, affecting the accuracy and relevance of the query results.

[0015] In response to the deficiencies in the prior art, an embodiment of the present application provides a query method based on user information, which can effectively combine user information and clarify semantically ambiguous entities, making the query successful and accurate.

[0016] Please refer to Figure 1 The query method based on user information in the embodiment of the present application includes the following steps: S101: Obtain a first natural query sentence input by a user and user information of the user; S102: Perform entity segmentation recognition on the first natural query sentence to obtain a plurality of first entity segmentations; S103: Determine a semantically ambiguous first entity segmentation among the plurality of first entity segmentations as a second entity segmentation; obtain a target processing text for the second entity segmentation based on the second entity segmentation and the user information; the target processing text is a text with clear semantics corresponding to the second entity segmentation obtained by searching a semantic processing text library corresponding to the user information; S104: Obtain a second natural query statement according to the first natural query statement and the target processed text; S105: Determine a target data table in a preset constellation data model based on the second natural query statement; S106: Inputting the second natural query statement and the target data table into a query macro model to obtain a query instruction; wherein the query macro model is used to understand the query intent of the second natural query statement and determine a query task that meets the query intent; and generating a query instruction based on the target data table and the query task; S107: Based on the query instruction, query the data warehouse corresponding to the constellation data model to obtain a query result of the second natural query statement; and determine the query result as the query result of the first natural query statement.

[0017] The embodiment of the present application first obtains a first natural query statement entered by a user and the user's user information, and identifies entity segmentation words in the first natural query statement. Entity segmentation words with unclear semantics are identified, and based on the user information, a target processing text that accurately expresses the meaning of these entity segmentation words is retrieved from the user's semantic processing text library. The first natural query statement is then optimized based on the target processing text to obtain a second natural query statement. Based on the second natural query statement, a target data table is determined in a preset constellation data model, and the second natural query statement and the target data table are input into a query model to generate a query instruction. Finally, a query result is obtained from the data warehouse based on the query instruction. By combining the user's semantic processing text library, the original natural query statement is optimized to make its semantics clear, clarify the meaning expressed by the user, and improve the accuracy of the query model's understanding of the query intent, thereby generating a more accurate query instruction. This not only significantly improves the accuracy and relevance of the query results, but also enhances the user's query experience, providing strong support for the promotion of natural language processing technology in practical applications.

[0018] The query method based on user information in the embodiment of the present application is executed by a computer. The following first describes the query model of the embodiment of the present application and the first natural language model involved in subsequent embodiments, and then describes each step in detail.

[0019] In an embodiment of the present application, the query large model or the first natural language model can be based on one or more pre-trained language models, such as BERT, GPT, etc. (these models have been trained on a large amount of text data and can capture the deep semantic and syntactic features of the language), and are fine-tuned to adapt to the specific tasks in this application, such as query intent understanding, query instruction generation, query field matching, etc. Of course, the query large model or the first natural language model can also be an existing large language model. As long as it has the processing capability of the specific tasks of this application, it can be applied in the technical solution of this application. This embodiment of the present application does not limit this.

[0020] In order to enable the query large model or the first natural language model to complete the task as required, it is usually necessary to provide a task processing text. The task processing text is used to require or prompt the task requirements of the model, which includes specific requirements and format instructions for the model output, and may also include task-related scenario information or context information for the model to learn and understand the task to be processed, etc. In an embodiment of the present application, different sub-texts can be set in the task processing text according to the needs of the query, which are used to require or prompt the model for the corresponding task requirements. The task processing text provided to the model in different task processing requirements is also different, such as the first task processing text, the second task processing text, and so on.

[0021] In step S101, a first natural query sentence input by a user and user information of the user are obtained; In this step, a natural query sentence submitted by a user through an interactive interface or other input method is received, such as "Please query the sales data of this department last month." At the same time, user information related to the user is also obtained. The user information can be information that can represent the user's identity, such as a user ID and user name.

[0022] In step S102, entity segmentation recognition is performed on the first natural query sentence to obtain a plurality of first entity segmentations; This step uses the entity segmentation recognition algorithm in natural language processing technology to perform word segmentation and entity recognition on the first natural query. For example, the query "Please query the sales data of this department last month" is segmented into "Please / query / this department / last month / sales / data". Among them, "this department", "last month", "sales data", etc. are identified as entity segmentations, namely the first entity segmentation.

[0023] In step S103, a first entity segmentation with unclear semantics among the plurality of first entity segmentations is determined to be a second entity segmentation; a target processing text for the second entity segmentation is obtained based on the second entity segmentation and the user information; the target processing text is a text with clear semantics corresponding to the second entity segmentation obtained by searching in a semantic processing text library corresponding to the user information; This step further analyzes these first entity segmentations to determine which ones have clear semantics and which ones are unclear or ambiguous. For example, "last month" is clear to all users, while "headquarters" may have different meanings depending on the user's department. Therefore, "headquarters" is determined to be a second entity segmentation, meaning an entity segmentation with unclear semantics.

[0024] The semantic processing text library is pre-built based on the user's preset configuration information. The preset configuration information is information related to the user's identity background, including but not limited to the user's department, direct supervisor, recently purchased products, commonly used professional terms, preference settings, and any other information related to the user himself. Since users from different backgrounds may have different identity information, professional terms, work habits, or information needs, their preset configuration information will also be different. Furthermore, based on this preset configuration information, corresponding target processing texts can be generated in advance for some common entity segmentation words that may be ambiguous in different contexts. The target processing text can express the clear meaning of the entity segmentation word in the user's context. For example, for users in the sales department, the target processing text generated for the "headquarters" mentioned in their sentence should be "sales department." The correspondence between these semantically unclear entity segmentation words and the corresponding target processing texts is stored in the semantic processing text library. When actually processing user queries, if a second entity segmentation word with unclear semantics is encountered, the clear semantics corresponding to the entity segmentation word can be directly searched in the semantic processing text library.

[0025] Please refer to Figure 2 In one embodiment, the step of determining the semantically ambiguous first entity participle among the plurality of first entity participles as the second entity participle in step S103 includes: Step S1031: obtaining the user's semantic processing text library based on the user information; wherein the semantic processing text library records a plurality of third entity segmentations and third processed texts corresponding to the third entity segmentations; the third processed texts are texts used to indicate clear semantics; In this embodiment, the semantic processing text library records a number of third entity segmentations and the third processing texts corresponding to each third entity segmentation. These third entity segmentations are common entity segmentations that may cause ambiguity in different contexts of different users. That is to say, third entity segmentations usually may have different meanings for different users, such as department names, project codes, specific product names, etc. The corresponding third processing texts are texts that can accurately reflect the user's meaning of these words in a specific context. For example, if the user's configuration information indicates that the user belongs to the "sales department", the third processing text of "headquarters" is "sales department".

[0026] Step S1032: If the first entity segmentation matches any third entity segmentation in the semantic processing text library, it is determined that the semantics of the first entity segmentation is ambiguous, and the semantically ambiguous first entity segmentation is determined as a second entity segmentation; After obtaining the user's semantically processed text library, the first entity segmentation identified in step S102 is matched with the third entity segmentation in the library. If a first entity segmentation completely matches a third entity segmentation in the library or satisfies certain matching rules (such as partial match, fuzzy match, etc.), the first entity segmentation is considered to be semantically ambiguous in the current user context and requires further semantic clarification processing, and is therefore determined as the second entity segmentation.

[0027] The step of obtaining a target processing text of the second entity segmentation according to the second entity segmentation and the user information in step S103 includes: Step S1033: Based on the semantic processing text library corresponding to the user information, a third processing text corresponding to the third entity segmentation matched by the second entity segmentation is obtained as the target processing text.

[0028] After determining the second entity segmentation, the user's semantic processing text library is used to directly search for the target processing text corresponding to the third entity segmentation that matches the second entity segmentation. This text will serve as the final semantic interpretation of the second entity segmentation, replacing or assisting in the interpretation of the corresponding first entity segmentation in the original first natural query, thereby generating a more specific second natural query.

[0029] In summary, this embodiment establishes a corresponding semantic processing text library for each user, and stores semantically unclear entity segmentations and their third processing texts in the user's semantic processing text library for the relevant background of the user, so as to improve the original natural query statement input by the user. This semantic clarification processing mechanism of this embodiment based on user-specific background information enables the query model to understand the user's query intention more accurately, and even when faced with query statements containing a large number of professional terms or specific contexts, it can generate query results that meet the user's expectations. In addition, since the semantic processing text library can be dynamically updated and maintained based on user feedback and new user background information, the query method of this embodiment also has good adaptability and scalability, and can continuously meet the user's changing query needs.

[0030] In step S104, a second natural query is obtained according to the first natural query and the target text of the second entity segmentation. In this step, the target processed text obtained in step S103 can replace or supplement the corresponding second entity segmentation in the original first natural query to obtain a new second natural query. For example, replacing "headquarters" with "sales department" will result in the second natural query being: "Please query the sales data of the sales department last month."

[0031] In one embodiment, the target processing text is a first processing text or a second processing text; the first processing text is an entity segmentation with clear semantics of the second entity segmentation; the second processing text is a text with a clear semantic interpretation of the second entity segmentation; If the target processing text of the second entity segmentation is the first processing text, the step of obtaining the second natural query statement according to the first natural query statement and the target processing text in step S104 includes: Step S1041: changing the second entity segmentation in the first natural query statement into the first processed text to obtain a second natural query statement; If the target processing text of the second entity segmentation is the second processing text, the step of obtaining the second natural query statement according to the first natural query statement and the target processing text in step S104 includes: Step S1042: Concatenate the first natural query statement and the second processed text to obtain a second natural query statement.

[0032] In this embodiment, the target processing text includes two possible types: a first processing text and a second processing text. The first processing text is an entity segmentation with clear semantics, which can directly replace the second entity segmentation in the original query statement. This text form is suitable for those situations where another clear word can be directly replaced, such as replacing "headquarters" with "sales department". The second processing text is an explanatory text, which provides a detailed explanation of the clear semantics of the second entity segmentation. This text form is suitable for those situations where additional information is needed to explain its meaning. For example, when the word "headquarters" may have different meanings in different departments or companies, the second processing text may provide an explanation such as "headquarters refers to the sales department".

[0033] If the target processing text for the second entity segmentation is the first processing text (step S1041): In this case, the second entity segmentation in the first natural query is directly replaced with the first processing text. For example, if the original query is "Please query the sales data of this department last month," and the target processing text for "headquarters" is "sales department," then "headquarters" is replaced with "sales department," resulting in the second natural query: "Please query the sales data of the sales department last month."

[0034] If the target processing text of the second entity segmentation is the second processing text (step S1042): In this case, the first natural query statement is spliced with the second processing text. This splicing method involves using the explanation text as part of the query statement or using the explanation text as an additional explanation of the query statement. For example, if the original query statement is "Please query the sales data of this department last month", and the target processing text of "this department" is "This department refers to the sales department", then the explanation text needs to be integrated into the query statement in some way, such as "Please query the sales data of this department last month, where this department refers to the sales department", to ensure that the query model can understand and process this query.

[0035] In summary, this embodiment further enhances the flexibility and accuracy of query processing by introducing two forms of target processing texts, namely, first processing text and second processing text, and adopting different strategies according to the types of these texts to generate the second natural query statement.

[0036] Please refer to Figure 3 In one embodiment, before the step of obtaining the first natural query sentence input by the user and the user information of the user in step S101, the following steps are further included: Step S1001: In response to a third processed text adding instruction triggered by a user, user information of the user and third processed text data are obtained; wherein the third processed text data includes a third entity segmentation with ambiguous semantics and a third processed text corresponding to the third entity segmentation; This embodiment describes how to build and update a user-specific semantic processing text library. In this step, a user action triggers a third processing text addition instruction, which is used to add or update target processing text data. When the user needs to specify explicit semantics for certain entity segmentations, they can trigger this instruction through the user interface. User information and the third processing text data entered by the user are then retrieved.

[0037] Step S1002: Add the third processed text data to the semantic processing text library corresponding to the user information.

[0038] After acquiring the user information and the third processed text data, these data are added to the semantically processed text library corresponding to the user. When adding data, a corresponding text library entry is located or created based on the user information, and the target processed text data is then stored within that entry. This allows for quick retrieval and application of these semantically explicit processing rules when processing future queries from the user.

[0039] In summary, this embodiment provides users with a flexible mechanism to add and update their semantically processed text repositories. This not only enhances the personalization and customization capabilities of the query system, but also enables the system to be continuously supplemented and optimized as user needs and context change. Ultimately, this design will help improve the accuracy and efficiency of query processing, thereby enhancing overall user satisfaction and experience.

[0040] In one embodiment, the third processing text adding instruction includes the first processing text adding instruction and / or the second processing text adding instruction; the third processing text data includes the first processing text data and / or the second processing text data; Before the step S1001 of obtaining the user information of the user and the third processed text data in response to the third processed text adding instruction triggered by the user, the method includes the following steps: Step S10011: In response to a first processing text creation instruction, receiving a third entity segmentation and a corresponding first processing text input by a user; obtaining first processing text data based on the third entity segmentation and the first processing text; and in response to a creation completion instruction, generating a first processing text addition instruction based on the user information of the user and the first processing text data. In this step, the user first triggers a specific instruction through the user interface, indicating that they wish to add a first processing text for a third entity segmentation word. In response to this instruction, the third entity segmentation word and the corresponding first processing text input by the user are received. The third entity segmentation word and the first processing text input by the user are combined into a first processing text data record. When the user confirms that the data they entered is correct, they will trigger a creation completion instruction. A first processing text addition instruction is then generated based on the user's user information and the first processing text data record. This instruction will be used to add the first processing text data to the user-specific semantic processing text library.

[0041] Step S10012, in response to the second processing text creation instruction, receive the third entity segmentation and the corresponding second processing text input by the user; obtain the second processing text data based on the third entity segmentation and the second processing text; in response to the creation completion instruction, generate the second processing text addition instruction based on the user information of the user and the second processing text data.

[0042] In this embodiment, the user triggers a specific instruction through the user interface, indicating that they wish to add a second processing text for a third entity segmentation word. In response to this instruction, the third entity segmentation word and the corresponding second processing text input by the user are received. The third entity segmentation word and the second processing text input by the user are combined into a second processing text data record. When the user confirms that the data they entered is correct, they will trigger a creation completion instruction. Subsequently, a second processing text addition instruction is generated based on the user's user information and the second processing text data record. This instruction will be used to add the second processing text data to the user-specific semantic processing text library.

[0043] In summary, this embodiment provides users with an intuitive and efficient mechanism to add and update their semantic processing text library. This design allows users to specify how they want the system to understand and process specific third-entity segmentations in a simple and clear manner. In addition, by subdividing the addition process into two parts: the first processed text and the second processed text, different types of semantic clarification processing requirements can be handled more flexibly. Ultimately, this design will help improve the accuracy and efficiency of query processing, while enhancing the overall user satisfaction and experience.

[0044] In step S105, a target data table is determined in a preset constellation data model based on the second natural query statement; wherein the constellation data model includes a structured data table and a query field in the data table; The constellation data model is an abstract representation used to describe the database structure. It is a data logic layer established in advance based on the various data tables stored in the data warehouse. It records several data tables, the basic fields contained in each data table, and the graph relationships between each data table. In addition, it may also record the binding relationships between at least some basic fields and multidimensional operation fields. In other words, the constellation data model does not record the specific business data of the fields in each data table, but only records the relevant information used to generate query statements. This information mainly includes the table name of the data table, the field information of the data table, and the multidimensional operation field information bound to the basic fields. In addition, the constellation data model establishes graph relationships between the various data tables, and these graph relationships include at least directed graph relationships.

[0045] Among them, the basic fields are generally the fields that record the original business data in the database. More specifically, the basic fields are generally divided into metric fields and dimension fields. The metric fields generally refer to the fields in the fact table, and the dimension fields generally refer to the fields in the dimension table. Among them, the fact table and the dimension table are concepts in data modeling in the field of data warehouse. Currently, the main data models involved in the field of data warehouse are star data model, snowflake data model generated based on the star data model (which can also be considered as a type of star data model) and constellation data model. In the star data model, it includes a data table located in the center and other data tables connected to the central data table. Among them, the central data table is the fact table, and the other data tables connected to it are dimension tables.

[0046] This step matches and determines the target data table in the constellation data model based on the second natural query statement. For example, the query "sales department's sales data for last month" may correspond to the "sales data table."

[0047] In one embodiment, the step of determining the target data table in the preset constellation data model based on the second natural query statement in step S105 includes: Step S1051: determining, from a preset constellation data model, a query field associated with the second natural query statement as a candidate query field; Step S1052: Input the second natural query and the candidate query fields into a first natural language model to obtain a plurality of target query fields; wherein the first natural language model is used to understand the query intent of the second natural query and determine the candidate query fields that meet the query intent as target query fields; Step S1053: determining a target data table including only the target query field from the constellation data model.

[0048] The second embodiment of the present invention combines the natural query statement and the candidate query fields into an input set, and inputs it into the first natural language model. The first natural language model understands the query intent of the second natural query statement, and filters out the target query fields that meet the query intent from the candidate query fields. For example, for the second natural query statement "Please query the sales data of the sales department last month", the model may find out that fields such as "sales department", "time" and "sales data" are target query fields. Furthermore, based on the obtained target query fields, the data table containing only these target query fields is filtered out from the constellation data model as the target data table. These target data tables are the basis for subsequent query instruction generation and data query. For example, if the target query fields are "sales department", "time" and "sales data", the sales data table of the sales department may be filtered out as the target data table.

[0049] In one embodiment, the step of determining, from a preset constellation data model, a query field associated with the natural query statement as a candidate query field in step S1051 includes: Step S10511, obtaining each query field of the constellation data model; Step S10512: determine as a candidate query field a query field whose semantic similarity with the second natural query statement or the entity segmentation of the second natural query statement is greater than a preset similarity threshold.

[0050] In this embodiment, each query field in the preset constellation data model is obtained. Then, using a semantic similarity calculation algorithm, each query field is respectively subjected to semantic similarity calculation with the second natural query statement and semantic similarity calculation with the entity segmentation of the second natural query statement. The purpose of this step is to find the query field that is semantically closest to the second natural query statement or the entity segmentation. Finally, based on a preset similarity threshold, query fields with a semantic similarity greater than the threshold are determined as candidate query fields. These fields have a high degree of semantic matching with the second natural query statement or the entity segmentation, and are therefore considered to be fields that may meet the user's query requirements. In summary, this embodiment accurately selects query fields associated with the second natural query statement as candidate query fields by introducing entity segmentation recognition and semantic similarity calculation.

[0051] In one embodiment, the target query field includes a dimension field and a metric field; the constellation data model records a plurality of data tables, basic fields included in each data table, and graph relationships between the data tables; the graph relationships include at least directed graph relationships; The step of determining the target data table including only the target query field from the constellation data model in step S1053 includes: Step S10531, determining a data table containing any one or more metric fields in the constellation data model as a first target data table; Step S10532: Using the first target data table as a root node, determine other data tables in the constellation data model that can reach the first target data table based on the graph relationship, and use each of the first target data tables as a fact table and the other data tables that can be reached from the fact table as dimension tables to obtain multiple target star data models. Step S10533: removing edge dimension tables that do not include the dimension fields in the target star data model, removing basic fields other than the metric fields in any fact table of the target star data model, and removing basic fields other than the dimension fields in any dimension table of the target star data model. Step S10534: Obtain a target data table according to the target star data model.

[0052] The target query fields of this embodiment include metric fields and dimension fields. First, a data table containing any one or more metric fields in the constellation data model is used as the first target data table; a fact table centered on the first target data table and other data tables that can reach the fact table are used as dimension tables to determine several target star data models; dimension tables irrelevant to the query in each target star data model are removed, as well as basic fields (metric fields or dimension fields) irrelevant to the query, so that the target star data model finally obtained only includes fact tables, dimension tables and basic fields related to the query. Through the method of the embodiment of the present application, the constellation data model model modeled in advance is trimmed according to the target query field, and the target data table is determined based on the minimum available star data model obtained by trimming, thereby avoiding querying other irrelevant data tables and irrelevant query fields, reducing the complexity of understanding the query instruction generation model, and making the generation of query instructions more efficient and accurate.

[0053] In step S106, the second natural query statement and the target data table are input into a query macro model to obtain a query instruction; wherein the query macro model is used to understand the query intent of the second natural query statement, determine a query task that meets the query intent, and generate a query instruction based on the target data table and the query task; The second natural query and the target data table are input into the query model. The query model understands the query intent in the natural language and generates a query instruction that meets the query intent based on the structure and fields of the target data table. The generated query instruction can be SQL statements or other database query languages. The specific configuration can be based on actual needs and is not detailed here.

[0054] For step S107, based on the query instruction, query the data warehouse corresponding to the constellation data model to obtain the query result of the second natural query statement; and determine the query result as the query result of the first natural query statement.

[0055] Based on the query instructions, the query operation is executed in the data warehouse corresponding to the constellation data model to obtain the query results. For example, the sales department's sales data for the previous month can be obtained from the "Sales Data Table". Finally, this query result is used as the query result for the original first natural query statement, and the query results can be returned to the user.

[0056] Please refer to Figure 4, an embodiment of the present application also provides an electronic device 301, comprising: a processor 302, a memory 303, and a computer program 304 stored in the memory 303 and executable on the processor 302, wherein the processor 302 implements the steps of the method described in any one of the embodiments of the present application when executing the computer program 304.

[0057] The processor 302 may include one or more processing cores. The processor 302 connects to various components within the electronic device 301 using various interfaces and circuits. It executes instructions, programs, code sets, or instruction sets stored in the memory 303 and accesses data from the memory 303 to perform various functions and process data within the electronic device 301. Optionally, the processor 302 may be implemented in the form of at least one of a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 302 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the touchscreen display; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 302 and may be implemented as a separate chip.

[0058] The memory 303 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 303 includes a non-transitory computer-readable storage medium. The memory 303 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 303 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch control instructions), instructions for implementing the aforementioned method embodiments, and the data storage area may store data involved in the aforementioned method embodiments. The memory 303 may also optionally be at least one storage device located remotely from the aforementioned processor 302.

[0059] The present application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the methods described in any of the embodiments of the present application. That is, those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented by instructing the relevant hardware through a program. The program, stored in a storage medium, includes instructions for causing a device (such as a microcontroller or chip) or a processor to perform all or part of the steps of the methods described in each embodiment of the present application. The computer program may include computer program code, which may be in source code form, object code form, an executable file, or some intermediate form. The aforementioned storage medium includes any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a removable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium. It should be noted that the content contained in computer-readable media can be appropriately increased or decreased according to the requirements of legislation and patent practices in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practices, computer-readable media does not include electrical carrier signals and telecommunications signals.

[0060] The above embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present application, and the present application is intended to encompass such modifications and variations.

Claims

1. A query method based on user information, characterized in that: The following steps are involved: Obtaining a first natural query statement input by a user and user information of the user; Performing entity segmentation recognition on the first natural query sentence to obtain a plurality of first entity segmentations; determining a semantically ambiguous first entity participle among the plurality of first entity participles as a second entity participle; obtaining a target processing text for the second entity participle based on the second entity participle and the user information; The target processing text is a text with clear semantics corresponding to the second entity segmentation obtained by querying in the semantic processing text library corresponding to the user information; Obtaining a second natural query statement according to the first natural query statement and the target processed text; Based on the second natural query statement, determining a target data table in a preset constellation data model; Inputting the second natural query statement and the target data table into a query macro model to obtain a query instruction; wherein the query macro model is used to understand the query intent of the second natural query statement and determine a query task that meets the query intent; and generating a query instruction based on the target data table and the query task; Based on the query instruction, query the data warehouse corresponding to the constellation data model to obtain a query result of the second natural query statement; and determine the query result as the query result of the first natural query statement.

2. The query method based on user information according to claim 1, characterized in that: The step of determining a semantically ambiguous first entity participle among the plurality of first entity participles as a second entity participle comprises: Acquiring the user's semantic processing text library according to the user information; wherein the semantic processing text library records a plurality of third entity segmentations and third processing texts corresponding to the third entity segmentations; the third processing text is a text used to indicate clear semantics; If the first entity segmentation matches any third entity segmentation in the semantic processing text library, it is determined that the semantics of the first entity segmentation is unclear, and the first entity segmentation with unclear semantics is determined as the second entity segmentation.

3. The query method based on user information according to claim 2, characterized in that: The step of obtaining a target processing text of the second entity segmentation according to the second entity segmentation and the user information includes: Based on the semantic processing text library corresponding to the user information, a third processing text corresponding to the third entity segmentation matched by the second entity segmentation is obtained as the target processing text.

4. The query method based on user information according to claim 2, characterized in that: Before the step of obtaining the first natural query sentence input by the user and the user information of the user, the method further includes the following steps: In response to a third processed text adding instruction triggered by a user, user information of the user and third processed text data are obtained; wherein the third processed text data includes a third entity segmentation with ambiguous semantics and a third processed text corresponding to the third entity segmentation; The third processed text data is added to the semantically processed text library corresponding to the user information.

5. The query method based on user information according to claim 1, characterized in that: The target processing text is the first processing text or the second processing text; the first processing text is an entity segmentation with clear semantics of the second entity segmentation; the second processing text is a text with clear semantic interpretation of the second entity segmentation; If the target processed text is a first processed text, the step of obtaining a second natural query statement according to the first natural query statement and the target processed text includes: Changing the second entity segmentation in the first natural query statement into the first processed text to obtain a second natural query statement; If the target processed text is a second processed text, the step of obtaining a second natural query statement according to the first natural query statement and the target processed text includes: The first natural query statement and the second processed text are concatenated to obtain a second natural query statement.

6. The query method based on user information according to claim 1, characterized in that: The step of determining a target data table in a preset constellation data model based on the second natural query statement includes: Determining, from a preset constellation data model, a query field associated with the second natural query statement as a candidate query field; Inputting the second natural query and the candidate query fields into a first natural language model to obtain a plurality of target query fields; wherein the first natural language model is used to understand the query intent of the second natural query and determine the candidate query fields that meet the query intent as target query fields; A target data table including only the target query field is determined from the constellation data model.

7. The query method based on user information according to claim 6, characterized in that: The step of determining, from a preset constellation data model, a query field associated with the natural query statement as a candidate query field comprises: Obtaining various query fields of the constellation data model; A query field having a semantic similarity with the second natural query sentence or the entity segmentation of the second natural query sentence greater than a preset similarity threshold is determined as a candidate query field.

8. The query method based on user information according to claim 6, characterized in that: The target query field includes a dimension field and a metric field; the constellation data model records a plurality of data tables, basic fields included in each data table, and graph relationships between each data table; the graph relationship includes at least a directed graph relationship; The step of determining a target data table including only the target query field from the constellation data model comprises: Determine a data table in the constellation data model that includes any one or more of the metric fields as a first target data table; Taking the first target data table as a root node, determining other data tables in the constellation data model that can reach the first target data table based on the graph relationship, and using each of the first target data tables as a fact table and the other data tables that can be reached by the fact table as dimension tables, to obtain multiple target star data models; Removing edge dimension tables that do not include the dimension fields in the target star data model, removing basic fields other than the metric fields in any fact table of the target star data model, and removing basic fields other than the dimension fields in any dimension table of the target star data model; A target data table is obtained according to the target star data model.

9. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed, the device where the computer-readable storage medium is located is controlled to implement the method according to any one of claims 1 to 8.

10. An electronic device, characterized in that: The method comprises a processor, a memory and a computer-readable program stored in the memory, wherein the computer-readable program implements the steps of the method according to any one of claims 1 to 8 when executed by the processor.

Citation Information

Patent Citations

  • Missing semantic supplementing method for multi-round question-answering system

    CN105589844A

  • Method and device for intention recognition of user questions

    CN110413746A

  • Semantic recognition method and equipment thereof

    CN112035506A

  • Data query method and storage medium

    CN117573698A

  • Method and related device for converting natural language into SQL (Structured Query Language) statement

    CN118331996A

Cited By

  • Query method of to-be-queried data, electronic equipment and storage medium

    CN120994879A