Query statement generation method and system and storage medium

By obtaining user problem description text, word segmentation and index library in the interactive interface to determine the index category, and using a large language model to generate query statements, the problem that query statements in the existing technology cannot meet user needs and are prone to errors is solved, and query statement generation with high accuracy and strong adaptability is achieved.

CN120162428AActive Publication Date: 2025-06-17TIANJIN TIANHE COMPUTER TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510638164.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-17
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

In the prior art, query statements cannot meet user needs, and the generated statements are prone to errors, especially in large-scale data scenarios and applications in specific fields.

Method used

By obtaining the problem description text entered by the user in the interactive interface, word segmentation is performed to determine the first alternative indicator category, and all indicators are obtained in the pre-built indicator library to determine the second alternative indicator category. Then, based on these categories and problem description text, a recommended query statement is generated using a pre-trained large language model.

Benefits of technology

It realizes the generation of query statements that meet user needs, avoids the model generating incorrect statements based on its own experience, ensures that the generated statements are within the scope of specific indicator categories, and improves the model's understanding in specific scenarios and the accuracy of query statements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162428A_ABST
    Figure CN120162428A_ABST
Patent Text Reader

Abstract

The invention provides a query statement generation method and system and a storage medium, and the method comprises the steps: obtaining a problem description text input by a user in an interaction interface, carrying out the word segmentation of the problem description text, determining a first alternative index category according to a word segmentation result, and obtaining all indexes from a pre-constructed index library, and determining a second alternative index category according to the obtained index and the problem description text, thereby generating at least one recommended query statement according to the first alternative index category, the second alternative index category, the problem description text and a pre-trained large language model, and realizing the generation of the query statement. Therefore, the generated query statement better meets the query requirement of the user, prediction errors of the model on the query requirement of the user due to the difference of description input by different users are avoided, the understanding of the model in a specific scene can be increased, the model is helped to find more comprehensive related indexes, and the query statement needed by the user is accurately generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of querying time-series data, and particularly to a method, a system, and a storage medium for generating query statements. Background Art

[0002] PromQL (Prometheus Query Language) is a dedicated language used by Prometheus to query and aggregate time-series data. Currently, PromQL query statements mainly rely on manual writing, which has the following problems: 1. High learning threshold, it is difficult for non-professionals to master complex syntax (such as aggregation functions, label matching, time range operations); 2. Inefficient and error-prone, manually writing multi-metric association queries is time-consuming and prone to missing logic; 3. Cross-team collaboration obstacles, business personnel cannot directly obtain monitoring data through natural language and rely on the technical team for secondary development.

[0003] Using a model to generate PromQL query statements can solve the above problems to a certain extent. However, in the related art, incorrect corpus used by the model may cause the generated statements to fail to meet user requirements; moreover, due to the limitation of the input data volume of the model, it may lead to difficulties in transferring all metrics to the model in large-scale data scenarios, thus causing errors in the statements generated by the model; in addition, the statement descriptions required by users in different fields vary greatly, and it is difficult for the model to understand user requirements in specific scenarios. Summary of the Invention

[0004] In view of the above defects or deficiencies in the prior art, this application aims to provide a method, a system, and a storage medium for generating query statements to solve the problems that query statements in the prior art cannot meet user requirements and query statements are incorrect.

[0005] An embodiment of this application provides a method for generating a query statement, and the method includes:

[0006] Obtain the problem description text input by the user in the interaction interface, perform word segmentation on the problem description text, and determine the first alternative metric category according to the word segmentation result;

[0007] Obtain all metrics in the pre-constructed metric library, and determine the second alternative metric category according to the obtained metrics and the problem description text, where the metric is the data name of the time series and the metric category is the classification of the data name;

[0008] Generate at least one recommended query statement based on the first alternative metric category, the second alternative metric category, the problem description text, and the pre-trained large language model.

[0009] Optionally, determining the first alternative metric category according to the word segmentation result includes:

[0010] Update the word segmentation counts of each word in the word segmentation table according to the current word in the word segmentation result;

[0011] Determine the first alternative index category according to the current word and the hit rates of the words under each index category in the word segmentation table.

[0012] Optionally, determining the first alternative index category according to the current word and the hit rates of the words under each index category in the word segmentation table includes:

[0013] Determine the predicted selection probabilities of each index category according to the current word and the hit rates of the words under each index category in the word segmentation table;

[0014] Determine the index categories with predicted selection probabilities greater than the preset probability threshold as the first alternative index category.

[0015] Optionally, determining the predicted selection probabilities of each index category according to the current word and the hit rates of the words under each index category in the word segmentation table includes:

[0016] For each index category, determine the sum of the hit rates of each current word under the index category in the word segmentation table as the numerator, and determine the sum of the hit rates of all words under the index category in the word segmentation table as the denominator;

[0017] Based on the numerator and the denominator, calculate the probability of the index category being selected under the current word to obtain the predicted selection probability of the index category.

[0018] Optionally, determining the second alternative index category according to the obtained index and the problem description text includes:

[0019] Construct category context information according to the obtained index;

[0020] Input the category context information and the problem description text into the large language model to obtain the second alternative index category.

[0021] Optionally, generating at least one recommended query statement based on the first alternative index category, the second alternative index category, the problem description text, and the pre-trained large language model includes:

[0022] Obtain a list of alternative indexes in the index library based on the first alternative index category and the second alternative index category;

[0023] Construct index context information according to the list of alternative indexes, and input the index context information and the problem description text into the large language model to obtain at least one recommended query statement.

[0024] Optionally, after generating at least one recommended query statement, it further includes:

[0025] Display recommended query statements on the interactive interface;

[0026] In response to detecting that the user triggers a confirmation operation for the recommended query statement, update the hit count and hit rate of each word in the word segmentation table based on the current word in the metric classification and word segmentation result corresponding to the recommended query statement.

[0027] Optionally, the construction of the metric library includes:

[0028] Obtain a time series database and extract each metric and its corresponding label name from the time series database;

[0029] Determine the label description information corresponding to each label name and determine the metric category corresponding to each metric;

[0030] Construct a metric library according to each metric and its corresponding metric category and label description information.

[0031] The embodiment of the present application also provides a query statement generation system, which includes an interaction module, a word segmenter, a context builder, a model usage module, and a feedbacker, where:

[0032] The interaction module is used to obtain the problem description text input by the user in the interactive interface, and to display recommended query statements in the interactive interface, and in response to detecting that the user triggers a confirmation operation for the recommended query statement, send the metric classification corresponding to the recommended query statement and the current word in the word segmentation result to the feedbacker;

[0033] The word segmenter is used to perform word segmentation on the problem description text and determine the first alternative metric category according to the word segmentation result;

[0034] The context builder is used to obtain all metrics in the pre-constructed metric library, construct category context information according to the obtained metrics, input the category context information and the problem description text into the large language model to obtain the second alternative metric category, and, based on the first alternative metric category and the second alternative metric category, obtain an alternative metric list in the metric library, construct metric context information according to the alternative metric list, and input the metric context information and the problem description text into the large language model to obtain at least one recommended query statement;

[0035] The feedbacker is used to update the hit count and hit rate of each word in the word segmentation table.

[0036] The embodiment of the present application also provides an electronic device, and the electronic device includes:

[0037] A processor and a memory;

[0038] The processor is configured to execute the steps of the query statement generation method provided in any embodiment of the present application by calling the program or instructions stored in the memory.

[0039] An embodiment of the present application further provides a computer-readable storage medium storing a program or instructions, and the program or instructions cause a computer to execute the steps of the query statement generation method provided in any embodiment of the present application.

[0040] In summary, the query statement generation method proposed in the present application obtains the problem description text input by the user in the interaction interface, performs word segmentation on the problem description text, determines the first alternative index category according to the word segmentation result, and then obtains all indexes in the pre-constructed index library. The second alternative index category is determined based on the obtained indexes and the problem description text. Thus, at least one recommended query statement is generated according to the first alternative index category, the second alternative index category, the problem description text, and the pre-trained large language model, realizing the generation of the query statement. This method enables the large language model to generate query statements within a specific index category range by combining the first alternative index category, the second alternative index category, and the problem description text, avoiding the model generating query statements based on its own experience, resulting in other indexes outside the index library in the query statements, and making the generated query statements more in line with the user's query needs. Moreover, this method determines the alternative index categories through word segmentation and the index library respectively for the large language model to generate statements, which can avoid the prediction error of the model for the user's query needs caused by the differences in the input descriptions of different users, increase the understanding ability of the model in specific scenarios, help the model discover more comprehensive relevant indexes, and enable the model to accurately generate the query statements required by the user. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 is a flowchart of a query statement generation method provided by an embodiment of the present application;

[0043] Figure 2 is a schematic diagram of a query statement generation system provided by an embodiment of the present application;

[0044] Figure 3 is a schematic diagram of the interaction process of the query statement generation system provided by an embodiment of the present application;

[0045] Figure 4 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0046] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of convenience of description, only the parts related to the invention are shown in the drawings.

[0047] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.

[0048] Embodiment 1

[0049] As mentioned in the background art, in view of the problems in the prior art, the present application proposes a method for generating a query statement. Figure 1 This is a flowchart of a method for generating a query statement provided by an embodiment of the present application. Refer to Figure 1 , the method for generating the query statement specifically includes:

[0050] S110. Obtain the problem description text input by the user in the interaction interface, perform word segmentation processing on the problem description text, and determine the first alternative index category according to the word segmentation result.

[0051] Among them, the interaction interface may specifically be a web user interface (Web UI), and the user can interact through a browser. Specifically, the interaction interface can be used for the user to input the problem description text, and for presenting the generated recommended query statement to the user, and for detecting the operations triggered by the user on the recommended query statement, etc.

[0052] In the embodiment of the present application, the problem description text can describe the time series data that the user needs to query in the time series database, reflecting the user's query requirements. Among them, the time series database can store time series (i.e., time series data), such as TSDB (Time Series Database); the time series can be a numerical sequence in which the index changes with time, and is uniquely determined by the index and the label. For example, CPU_usage_active{cluster = "HPC", device = "ln1"} represents the CPU usage rate of the ln1 device on the HPC cluster.

[0053] Exemplarily, the user can enter the problem description text in the text box on the interaction interface; or, the user can select the problem description text from the popular description texts displayed on the interaction interface, where the popular description texts can be obtained by counting the number of times and sorting the problem description texts entered by each user historically.

[0054] Specifically, after obtaining the problem description text input by the user on the interaction interface, the problem description text can be segmented. For example, a Chinese word segmenter (such as Jieba) can be used to extract the Chinese word segments in the problem description text to obtain each current word.

[0055] Among them, the word segmenter can have the ability to extract the object in the Chinese information. For example, if the problem description text is "CPU usage rate of node ln1 in the HPC cluster", in this problem description text, the CPU usage rate is the index. Therefore, the object in the problem description text can be extracted through algorithm configuration.

[0056] In order to avoid the influence of invalid words on the generation efficiency of subsequent query statements, the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm can also be used to screen out the key information in all current words and remove the invalid words, such as invalid words like "need", "here", "I", etc.

[0057] After obtaining the word segmentation result, further, the first alternative index category can be determined through the word segmentation result. Among them, the index is the data name of the time series, and the index category is the classification of the data name.

[0058] Exemplarily, the index can be understood as the name of a measurable phenomenon, such as CPU usage rate. The index category can be understood as the classification where the index is located. For example, the category of CPU usage rate is CPU, and the category of the total number of bytes of device memory is device.

[0059] In a specific implementation manner, determining the first alternative index category according to the word segmentation result includes the following steps:

[0060] Step 11: Update the word segmentation times of each word in the word segmentation table according to the current words in the word segmentation result;

[0061] Step 12: Determine the first alternative index category according to the current words and the hit rates of the words under each index category in the word segmentation table.

[0062] Among them, in step 11, the word segmentation times of each word in the word segmentation table can be updated according to each current word in the word segmentation result. The word segmentation table can be used to record the association relationship between words and index categories in the index library; the word segmentation table can include the hit times, word segmentation times, and hit rates of each word under each index category.

[0063] Among them, the hit count is the number of times the word is hit in the recommended query statements generated by the large language model, the word segmentation count is the number of times the word is obtained by segmenting the problem description text input by the user, and the hit rate is used to describe the probability that the word is hit in the recommended query statements generated by the large language model, which can reflect the frequency of the word being hit. For example, the hit rate is equal to the ratio of the hit count to the word segmentation count. Exemplarily, Table 1 shows a word segmentation table.

[0064] Table 1 A word segmentation table

[0065]

[0066] After updating the word segmentation table, further, in step 12, the current word can be used to query the index categories related to the current word in the word segmentation table, and the first alternative index category can be determined by combining the hit rates of the current word under each index category.

[0067] For the above step 12, in one example, determining the first alternative index category according to the current word and the hit rates of the words under each index category in the word segmentation table includes the following steps:

[0068] Step 121: Determine the predicted selection probabilities of each index category according to the current word and the hit rates of the words under each index category in the word segmentation table;

[0069] Step 122: Determine the index categories with predicted selection probabilities greater than the preset probability threshold as the first alternative index categories.

[0070] Among them, in step 121, the current word can be used to query the index categories related to the current word in the word segmentation table to obtain the hit rates of the current word under the index categories it is related to, and then the predicted selection probabilities of the related index categories can be calculated according to the hit rates of the current word under the index categories it is related to.

[0071] For the above step 121, optionally, determining the predicted selection probabilities of each index category according to the current word and the hit rates of the words under each index category in the word segmentation table includes:

[0072] For each index category, the sum of the hit rates of each current word under the index category in the word segmentation table is determined as the numerator, and the sum of the hit rates of each word under the index category in the word segmentation table is determined as the denominator; based on the numerator and the denominator, the probability of the index category being selected under the current word is calculated to obtain the predicted selection probability of the index category.

[0073] Among them, the index categories related to the current word can be searched from the word segmentation table first, and then the corresponding predicted selection probabilities can be determined for each index category respectively.

[0074] Specifically, for each indicator category, the sum of the hit rates of each current word under this indicator category can be used as the numerator, and the sum of the hit rates of all words under the indicator category can be used as the denominator to calculate the probability that the indicator category is selected under the current word, that is, the predicted selection probability, as shown in the following formula:

[0075] ;

[0076] In the formula, represents the probability that the indicator category is selected under the current word ~ , that is, the predicted selection probability, is the hit rate of the k-th current word under the indicator category, is the number of current words under the indicator category, is the hit rate of the m-th word under the indicator category, is the number of words under the indicator category.

[0077] Exemplarily, if the current words include 、 , then the predicted selection probability of the CPU indicator category under the current words is:

[0078] ;

[0079] Through the above optional implementation manners, the accurate determination of the predicted selection probabilities of each indicator category can be achieved, which is convenient for subsequent screening of alternative categories for the large language model to generate statements, making the statements generated by the model more accurately meet the user's needs.

[0080] After obtaining the predicted selection probabilities of each indicator category, the predicted selection probability can reflect the probability that each indicator category is selected under the user's current word. Further, in step 122, the indicator categories with predicted selection probabilities greater than the preset probability threshold can be selected as the first alternative indicator categories.

[0081] Through the above steps 121 - step 122, the alternative indicator categories can be determined based on the predicted selection probabilities of each indicator category, ensuring the accuracy of the alternative indicator categories and obtaining the indicator categories that meet the user's query requirements, so as to help the model generate more accurate query statements.

[0082] S120. Obtain all indicators from a pre-constructed indicator library, and determine the second alternative indicator categories according to the obtained indicators and the problem description text.

[0083] In the embodiments of the present application, in addition to determining the first alternative index category based on the word segmentation result of the problem description text, the second alternative index category can also be determined according to the index library and the problem description text, so as to help the model generate query statements within the scope of the index library and avoid the model generating query statements for indexes that do not exist in the time series database.

[0084] It should be noted that in the embodiments of the present application, the purpose of determining the first alternative index category based on the word segmentation result of the problem description text and the word segmentation table is as follows: The first alternative index category and the second alternative index category can be jointly input into the large language model to avoid the model directly using the second alternative index category to generate reference statements, provide an error correction mechanism for the model, solve the deviation of the second alternative index category caused by the differences in the input descriptions of different users, so as to help the large language model better understand the user's semantics, provide relevant indexes that the model has not understood or discovered, and enable the model to better understand the user's needs in a specific scenario.

[0085] Among them, the index library can include various index categories and the indexes under each index category. In one example, the construction of the index library includes the following steps:

[0086] Step 21, obtain the time series database and extract each index and the corresponding label name from the time series database;

[0087] Step 22, determine the label description information corresponding to each label name and determine the index category corresponding to each index;

[0088] Step 23, construct the index library according to each index and the corresponding index category and label description information.

[0089] Among them, the time series database can be TSDB, which is used to store each time series. In step 21, the time series database can be obtained, and from each time series in the time series database, the index and the corresponding label name are extracted.

[0090] Furthermore, in step 22, the label description information corresponding to each label name can be determined. The label description information can describe the explanation of all label names and the explanation of the index. And the index category corresponding to each index can also be determined, that is, the classification involved in the index.

[0091] Furthermore, in step 23, the index library can be constructed according to each index and the corresponding index category and label description information. The index library can store the indexes under each index category and the label description information corresponding to the indexes.

[0092] Exemplarily, the definition of the index library is as follows:

[0093] {

[0094] "<Category 1>":

[0095] {

[0096] "metric": "<Metric>",

[0097] "labels": ["<Label 1>", "<Label 2>"],

[0098] "desc": "<Label description information>"

[0099] },

[0100] {

[0101] "metric": "<Metric>",

[0102] "labels": ["<Label 1>", "<Label 2>"],

[0103] "desc": "<Label description information>"

[0104] }, ...

[0106] ,

[0107] "<Category 2>":

[0108] {

[0109] "metric": "<Metric>",

[0110] "labels": ["<Label 1>", "<Label 2>"],

[0111] "desc": "<Label description information>"

[0112] },

[0113] {

[0114] "metric": "<Metric>",

[0115] "labels": ["<Label 1>", "<Label 2>"],

[0116] "desc": "<Label description information>"

[0117] }, ...

[0119] , ...

[0121] }

[0122] In the embodiments of the present application, taking the CPU metric category and the device metric category as examples, some information in the metric library is shown:

[0123] {

[0124] "CPU":

[0125] {

[0126] "metric": "CPU_usage_idle",

[0127] "labels": ["cluster", "device", "CPU"],

[0128] "desc": cluster represents the cluster, device represents the device reflected by the metric, CPU represents the CPU ID on the device or the overall CPU, and this metric represents the idle usage rate of a certain CPU of a certain device in a certain cluster.

[0129] },

[0130] {

[0131] "metric": "CPU _usage_user",

[0132] "labels": ["cluster", "device", "CPU"],

[0133] "desc": cluster represents the cluster, device represents the device reflected by the metric, CPU represents the CPU ID on the device or the overall CPU, and this metric represents the usage rate of the user state of a certain CPU of a certain device in a certain cluster.

[0134] }

[0135] ,

[0136] "MEM":

[0137] {

[0138] "metric": "mem_total",

[0139] "labels": ["cluster", "device"],

[0140] "desc": cluster represents the cluster, device represents the device reflected by the metric, and this metric represents the total number of bytes of memory of a certain device in a certain cluster.

[0141] },

[0142] {

[0143] "metric": "mem_used_percent",

[0144] "labels": ["cluster", "device"],

[0145] "desc": The cluster represents the cluster, the device represents the device reflected by the metric, and this metric represents the memory usage rate of a certain device in a certain cluster.

[0146] }

[0148] }

[0149] Through the above example, partial information can be extracted from the time series database of large-scale data to construct an index library. Compared with the large-scale time series database, due to the high cardinality, there are tens of millions of time series (TimeSeries) in the time series database. After refining, the data in the constructed index library can be reduced to more than a thousand times, which is convenient for quickly obtaining index categories from the index library later and improving the efficiency of the overall process.

[0150] For example, CPU_usage_active{cluster="HPC", device="ln1"} 90% means that the current CPU usage rate of the ln1 device in the HPC cluster is 90%. However, in actual situations, due to the existence of multiple clusters and tens of thousands of devices in supercomputing centers, intelligent computing centers, data centers, etc., there will be <number of clusters * number of devices> unique time series for the CPU_usage_active metric, resulting in data expansion in the TSDB. In contrast, the index library removes the label values and refines the TSDB by using the label description method, which can greatly reduce the data volume for use as category context information for large language models.

[0151] In the embodiments of the present application, all metrics can be obtained from the index library, and then the second alternative index category can be determined according to the obtained metrics and the problem description text.

[0152] In a specific implementation manner, determining the second alternative index category according to the obtained metrics and the problem description text includes:

[0153] Constructing category context information according to the obtained metrics; inputting the category context information and the problem description text into a large language model to obtain the second alternative index category.

[0154] ​Among them, category context information can be constructed based on the metrics obtained from the metric library. Context information refers to the range of input information that the model refers to when generating answers, which is used to help the model understand the current context and generate relevant content. Category context information can help the large language model identify the metric categories contained in the user's question description text.

[0155] Furthermore, the category context information and the question description text can be input into the large language model, so that the large language model can identify the metric categories contained in the question description text based on the category context information, and obtain the second alternative metric categories.

[0156] Through the above implementation, the context information can be constructed using the metrics in the metric library and input into the large language model together with the question description text, which helps the large language model identify the metric categories contained in the question description text, improves the accuracy of the second alternative metric categories, and can help the model understand the user's needs in specific query scenarios, avoiding the model from outputting metric categories not covered by TSDB, thus avoiding the model from generating query statements for metrics that do not exist in TSDB.

[0157] S130. Generate at least one recommended query statement based on the first alternative metric categories, the second alternative metric categories, the question description text, and the pre-trained large language model.

[0158] Among them, the large language model can be a neural network model trained with a large amount of text data, which can understand and generate natural language text.

[0159] Specifically, after obtaining the first alternative metric categories and the second alternative metric categories, context information can be constructed based on the first alternative metric categories and the second alternative metric categories, and the context information and the question description text can be input into the large language model together to obtain at least one recommended query statement.

[0160] In a specific implementation, generating at least one recommended query statement based on the first alternative metric categories, the second alternative metric categories, the question description text, and the pre-trained large language model includes:

[0161] Based on the first alternative metric categories and the second alternative metric categories, obtain an alternative metric list in the metric library; construct metric context information according to the alternative metric list, and input the metric context information and the question description text into the large language model to obtain at least one recommended query statement.

[0162] Specifically, all metrics under the first alternative metric categories can be queried from the metric library, and all metrics under the second alternative metric categories can be queried from the metric library, and an alternative metric list can be constructed according to all the queried metrics.

[0163] Furthermore, the indicator context information can be constructed according to the alternative indicator list, where the indicator context information can help the large language model generate query statements that conform to the user's problem description.

[0164] Specifically, the indicator context information and the problem description text can be input into the large language model together, so that the large language model can understand the problem description text in combination with the indicator context information and generate a recommended query statement that meets the user's needs, such as a PromQL statement.

[0165] The above implementation obtains the alternative indicator list through the first alternative indicator category and the second alternative indicator category, thereby constructing the indicator context information, which can help the large language model identify the indicators that the user needs to query from the problem description text, and thus accurately generate the corresponding recommended query statement, making the statement generated by the large language model more in line with the user's needs.

[0166] In the embodiment of the present application, in order to achieve more accurate query statement recommendation, after generating the recommended query statement, the word segmentation table can also be updated based on the user's feedback to ensure more accurate next recommendation.

[0167] In some embodiments, after generating at least one recommended query statement, it further includes:

[0168] Display the recommended query statement on the interaction interface; in response to detecting that the user triggers a confirmation operation for the recommended query statement, update the hit times and hit rates of each word in the word segmentation table based on the indicator classification corresponding to the recommended query statement and the current word in the word segmentation result.

[0169] Among them, after obtaining the recommended query statement, further, the recommended query statement can be displayed in the interaction interface.

[0170] After displaying the recommended query statement through the interaction interface, the operation triggered by the user for the displayed recommended query statement can be detected on the interaction interface. Specifically, if it is detected that the user triggers a confirmation operation for the recommended query statement, it means that the recommendation of the query statement this time is successful, and the word segmentation table can be updated according to the indicator classification corresponding to the recommended query statement and the current word in the word segmentation result. For example, update the hit times and hit rates of the current word under the indicator classification corresponding to the recommended query statement in the word segmentation table.

[0171] Through the above implementation, the user behavior can be recorded based on the feedback mechanism, and the word segmentation table can be updated through the feedback mechanism, which can improve the understanding ability of the large language model for the user's description during the subsequent generation of the recommended query statement, so that the large language model can recommend the query statement more accurately.

[0172] The method for generating a query statement provided by an embodiment of the present application obtains the problem description text input by the user in the interaction interface, performs word segmentation on the problem description text, determines the first alternative index category according to the word segmentation result, and then obtains all indexes in the pre-constructed index library. The second alternative index category is determined according to the obtained index and the problem description text. Thus, at least one recommended query statement is generated according to the first alternative index category, the second alternative index category, the problem description text, and the pre-trained large language model, realizing the generation of the query statement. This method enables the large language model to combine the first alternative index category, the second alternative index category, and the problem description text to generate a query statement within a specific index category range, avoiding the model generating a query statement according to its own experience, resulting in other indexes outside the index library in the query statement, and making the generated query statement more in line with the user's query needs. Moreover, this method determines the alternative index category through word segmentation and the index library respectively for the large language model to generate statements, which can avoid the difference in the input descriptions of different users from causing the model to mispredict the user's query needs, can increase the model's understanding ability in a specific scenario, help the model discover more comprehensive relevant indexes, and enable the model to accurately generate the query statement required by the user.

[0173] Embodiment 2

[0174] An embodiment of the present application further provides a query statement generation system, which includes an interaction module, a word segmenter, a context builder, a large language model, and a feedbacker, where:

[0175] The interaction module is used to obtain the problem description text input by the user in the interaction interface, and is used to display the recommended query statement in the interaction interface, and in response to detecting that the user triggers a confirmation operation for the recommended query statement, send the index classification corresponding to the recommended query statement and the current word in the word segmentation result to the feedbacker;

[0176] The word segmenter is used to perform word segmentation on the problem description text and determine the first alternative index category according to the word segmentation result;

[0177] The context builder is used to obtain all indexes in the pre-constructed index library, construct category context information according to the obtained indexes, input the category context information and the problem description text into the large language model to obtain the second alternative index category, and is used to obtain an alternative index list in the index library based on the first alternative index category and the second alternative index category, construct index context information according to the alternative index list, and input the index context information and the problem description text into the large language model to obtain at least one recommended query statement;

[0178] The feedbacker is used to update the hit times and hit rates of each word in the word segmentation table.

[0179] Based on the above embodiments, optionally, the tokenizer is further configured to update the tokenization times of each term in the tokenization table according to the current term in the tokenization result; and determine the first alternative index category according to the current term and the hit rates of the terms under each index category in the tokenization table.

[0180] Based on the above embodiments, optionally, the tokenizer is further configured to determine the predicted selection probability of each index category according to the current term and the hit rates of the terms under each index category in the tokenization table; and determine the index categories with predicted selection probabilities greater than the preset probability threshold as the first alternative index categories.

[0181] Based on the above embodiments, optionally, for each index category, the tokenizer determines the sum of the hit rates of the current terms under the index category in the tokenization table as the numerator, and determines the sum of the hit rates of all terms under the index category in the tokenization table as the denominator; and calculates the probability of the index category being selected under the current term based on the numerator and the denominator to obtain the predicted selection probability of the index category.

[0182] Based on the above embodiments, optionally, the system further includes an index library, wherein the construction of the index library includes:

[0183] Obtain a time series database, and extract each index and the corresponding label name from the time series database;

[0184] Determine the label description information corresponding to each label name, and determine the index category corresponding to each index;

[0185] Construct an index library according to each index and the corresponding index category and label description information.

[0186] Figure 2 It is a schematic diagram of a query statement generation system provided by an embodiment of the present application. As Figure 2 shown, the user's question description can be input into the tokenizer and the context builder at the same time. The tokenizer can update the tokenization times, etc. in the tokenization table. The context builder can obtain indexes from the index library. The context builder can also construct category context information and index context information and input them into the Large-Language Models (LLMs). After the large language model generates a recommended query statement, it can be sent to the interaction module for display. The interaction module can also send the user feedback to the feedbacker, and the feedbacker can update the hit times and hit rates in the tokenization table.

[0187] Figure 3It is a schematic diagram of the interaction process of the query statement generation system provided by the embodiments of the present application. As shown in Figure 3, first, the interaction module can obtain the problem description and push the problem description to the context builder and the tokenizer. Among them, the token table can obtain the index list and push the index list to the context builder. Then, the context builder constructs the category context and pushes the category context and the index category to the large language model. The large language model determines the second alternative index category and returns the second alternative index category to the context builder.

[0188] Moreover, the tokenizer can tokenize the problem description, push the tokens to the token table for update, and the token table returns the index category to the tokenizer. The tokenizer determines the first alternative index category and pushes the first alternative index category to the context builder. Then, the context builder constructs the index context according to the alternative index categories (including the first alternative index category and the second alternative index category) and pushes the problem description and the index context to the large language model. The large language model generates a recommended query statement.

[0189] Furthermore, the large language model returns the recommended query statement and the corresponding index category to the interaction module. The interaction module displays the result and feeds back the user operation, and returns the user operation to the feedbacker. The feedbacker updates the token table according to the feedback and returns the hit situation to the token table for update.

[0190] The query statement generation system provided by the embodiments of the present application can be applied to the steps in the query statement generation method provided by the method embodiments of the present application, and the implementation steps and beneficial effects are not described herein again.

[0191] Embodiment 3

[0192] Figure 4 It is a schematic diagram of the structure of an electronic device provided by the embodiments of the present application. As Figure 4 shown, the electronic device 400 includes one or more processors 401 and a memory 402.

[0193] The processor 401 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 400 to perform desired functions.

[0194] The memory 402 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 401 may run the program instructions to implement the method for generating query statements according to any embodiment of the present application described above and / or other desired functions. Various contents such as initial external parameters, thresholds, etc. may also be stored in the computer-readable storage media.

[0195] In one example, the electronic device 400 may further include: an input device 403 and an output device 404, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). The input device 403 may include, for example, a keyboard, a mouse, etc. The output device 404 may output various information to the outside, including warning prompt information, braking force, etc. The output device 404 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0196] Of course, for simplicity, Figure 4 only some of the components related to the present application in the electronic device 400 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 400 may further include any other appropriate components.

[0197] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps of the method for generating query statements provided by any embodiment of the present application.

[0198] The computer program product may be written in any combination of one or more programming languages to write program code for performing the operations of the embodiments of the present application. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0199] In addition, an embodiment of the present application may also be a computer-readable storage medium storing computer program instructions, which, when run by a processor, cause the processor to execute the steps of the method for generating a query statement provided in any embodiment of the present application.

[0200] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0201] It should be noted that the terms used in the present application are only for describing specific embodiments and do not limit the scope of the present application. As shown in the specification and claims of the present application, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. The term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, or device including the element.

[0202] It should also be noted that the orientation or positional relationship indicated by terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation of the present application. Unless otherwise clearly specified and limited, terms such as "installed", "connected", "connected to" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0203] In this text, specific examples are used to illustrate the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. The above is only the preferred implementation manner of the present application. It should be noted that due to the limited nature of written expression and the objectively infinite specific structures, for those of ordinary skill in the art, without departing from the principles of the present application, several improvements, refinements or changes can be made, or the above technical features can be combined in an appropriate manner; these improvements, refinements, changes or combinations, or directly applying the inventive concept and technical solution to other occasions without improvement, shall all be regarded as the protection scope of the present application.

Claims

1. A method for generating a query statement, characterized in that: include: Obtaining a problem description text input by a user in an interactive interface, performing word segmentation processing on the problem description text, and determining a first candidate indicator category according to the word segmentation result; Obtain all indicators in a pre-built indicator library, and determine a second candidate indicator category according to the obtained indicators and the problem description text, wherein the indicator is the data name of the time series and the indicator category is the classification of the data name; At least one recommended query statement is generated based on the first candidate indicator category, the second candidate indicator category, the question description text, and a pre-trained large language model.

2. The method according to claim 1, characterized in that: The step of determining the first candidate indicator category according to the word segmentation result includes: According to the current word in the word segmentation result, update the number of word segmentations of each word in the word segmentation table; A first candidate indicator category is determined according to the current word and the hit rates of words under each indicator category in the word segmentation table.

3. The method according to claim 2, characterized in that Determining a first candidate indicator category according to the current word and the hit rates of words under each indicator category in the word segmentation table includes: Determine the predicted selection probability of each indicator category according to the current word and the hit rate of the words under each indicator category in the word segmentation table; The indicator category with a predicted selection probability greater than a preset probability threshold is determined as the first candidate indicator category.

4. The method according to claim 3, characterized in that Determining the predicted selection probability of each indicator category according to the current word and the hit rate of the words under each indicator category in the word segmentation table includes: For each indicator category, the sum of the hit rates of each current word under the indicator category in the word segmentation table is determined as the numerator, and the sum of the hit rates of each word under the indicator category in the word segmentation table is determined as the denominator; Based on the numerator and the denominator, the probability of the indicator category being selected under the current word is calculated to obtain the predicted selection probability of the indicator category.

5. The method according to claim 1, characterized in that Determining the second candidate indicator category according to the acquired indicator and the problem description text includes: Build category context information based on the acquired indicators; The category context information and the question description text are input into the large language model to obtain a second candidate indicator category.

6. The method according to claim 1, characterized in that Based on the first candidate indicator category, the second candidate indicator category, the question description text, and a pre-trained large language model, generating at least one recommended query statement includes: Based on the first candidate indicator category and the second candidate indicator category, obtaining a candidate indicator list in the indicator library; Indicator context information is constructed according to the candidate indicator list, and the indicator context information and the question description text are input into the large language model to obtain at least one recommended query statement.

7. The method according to claim 2, characterized in that After generating at least one recommended query statement, the method further includes: Displaying the recommended query statement on the interactive interface; In response to detecting that a user triggers a confirmation operation on the recommended query statement, based on the indicator classification corresponding to the recommended query statement and the current word in the word segmentation result, the hit count and hit rate of each word in the word segmentation table are updated.

8. The method according to claim 1, characterized in that The construction of the indicator library includes: Obtain a time series database, and extract each indicator and a corresponding label name from the time series database; Determine the label description information corresponding to each label name, and determine the indicator category corresponding to each indicator; Construct an indicator library based on each indicator and the corresponding indicator category and label description information.

9. A query statement generation system, characterized in that: The system includes an interaction module, a word segmenter, a context builder, a large language model and a feedback device, wherein: An interactive module, used to obtain a question description text input by a user in an interactive interface, and to display a recommended query statement in the interactive interface, and in response to detecting that a user triggers a confirmation operation on the recommended query statement, send the indicator classification corresponding to the recommended query statement and the current word in the word segmentation result to a feedback device; The word segmenter is used to perform word segmentation processing on the problem description text, and determine the first candidate indicator category according to the word segmentation result; The context builder is used to obtain all indicators in a pre-built indicator library, construct category context information according to the obtained indicators, input the category context information and the question description text into the large language model to obtain a second candidate indicator category, and to obtain a candidate indicator list in the indicator library based on the first candidate indicator category and the second candidate indicator category, construct indicator context information according to the candidate indicator list, and input the indicator context information and the question description text into the large language model to obtain at least one recommended query statement; The feedback device is used to update the hit times and hit rates of each word in the word segmentation table.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program or instruction, and the program or instruction enables a computer to execute the steps of the method for generating a query statement as claimed in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Text generation method and device based on artificial intelligence, computer equipment and medium

    CN112395385A

  • Data acquisition method and device, equipment and storage medium

    CN115203367A

  • Structured query statement generation method and device, equipment and medium

    CN117149812A

  • Log query statement generation method and device, equipment and storage medium

    CN118152341A

  • Data query method and device, model training method and device, equipment and medium

    CN119046406A