Query statement generation method and device, equipment, readable storage medium and product

By using the scenario adaptation model in the query statement generation method, identifying the application scenarios required by users and generating target query knowledge, the problems of large knowledge base search scope and low positioning accuracy in the existing technology are solved, and the reliability of the query statement is improved.

CN120011393APending Publication Date: 2025-05-16CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510104855.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, the query statement generation method has the problem of large knowledge base search scope and low knowledge positioning accuracy, resulting in low reliability of the generated query statements.

Method used

By obtaining the user's current demand data and application scenario distribution characteristics, inputting them into the pre-trained scenario adaptation model, extracting semantic features and fusing application scenario features, predicting application scenarios, thereby generating target query knowledge and generating query statements.

Benefits of technology

Through the identification and screening of the scene adaptation model, the scope of knowledge retrieval is reduced, the accuracy of knowledge retrieval is improved, the generated query statements are adapted to user application scenarios, and the situations that do not meet user needs are reduced, thereby improving the reliability of query statement generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011393A_ABST
    Figure CN120011393A_ABST
Patent Text Reader

Abstract

The invention discloses a query statement generation method and device, equipment, a readable storage medium and a product, and relates to the technical field of natural language process.The method comprises the steps that current demand data and application scene distribution characteristics of a user are obtained; the application scene distribution features and the current demand data are input into a pre-trained scene adaptation model, an application scene corresponding to the current demand data is obtained through output, the scene adaptation model is used for extracting semantic features in the current demand data to obtain semantic vectors, and the semantic vectors are used for obtaining the application scene corresponding to the current demand data. Predicting and outputting after fusing the semantic vector and the application scene distribution characteristics; and generating target query knowledge based on the application scene and the current demand data, and generating a query statement according to the target query knowledge. The problems that the query statement is difficult to meet the requirements of a user and the reliability of the generated query statement is low are solved, and the effect of improving the reliability of the query statement generation result is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing technology, and in particular to a query statement generation method, device, equipment, readable storage medium and product. Background Art

[0002] As business becomes increasingly complex, the information that users need to query in the database becomes more diverse and complex. When users use the database to query data, they need to master the database-related query language. In the case of complex business scenarios, the threshold for users to use the database will be greatly increased.

[0003] In order to lower the threshold for using databases, human natural language can be converted into query statements that machines can understand, such as SQL (Structured Query Language) query statements, and the answers can be returned from database queries based on the query statements. This can serve as an intelligent bridge between language and database, allowing users who are not familiar with databases to quickly query the data they want using only natural language.

[0004] The traditional query statement generation method stores domain knowledge in a vectorized storage manner to build a knowledge base, and then retrieves knowledge that matches user needs by querying the knowledge base to generate query statements based on the matched knowledge. However, the knowledge base often stores a large amount of knowledge and has a large search range. The knowledge matching process is usually a whole-base search, which easily retrieves a large amount of knowledge for subsequent query statement generation. The knowledge positioning accuracy is low, which makes it difficult for subsequent query statements to meet user needs and the generated query statements have low reliability.

[0005] Therefore, how to improve the reliability of query statement generation is a technical problem that needs to be solved urgently in this technical field. Summary of the invention

[0006] The main purpose of the present application is to provide a query statement generation method, device, equipment, readable storage medium and product, aiming to solve the technical problem of how to improve the reliability of query statement generation.

[0007] To achieve the above object, the present application provides a query statement generation method, the query statement generation method comprising the following steps:

[0008] Obtain users’ current demand data and application scenario distribution characteristics;

[0009] Input the application scenario distribution features and the current demand data into a pre-trained scenario adaptation model, and output the application scenario corresponding to the current demand data, wherein the scenario adaptation model is used to extract semantic features from the current demand data to obtain a semantic vector, and fuse the semantic vector with the application scenario distribution features to predict the output;

[0010] Target query knowledge is generated based on the application scenario and the current demand data, and a query statement is generated according to the target query knowledge.

[0011] In one embodiment, the step of generating a query statement based on the target query knowledge includes:

[0012] Performing prompt word engineering processing on the target query knowledge to obtain knowledge prompt words, wherein the knowledge prompt words include one or more of entity information prompt words, query template example prompt words and table metadata information prompt words;

[0013] Constructing a thought chain according to the knowledge prompt words, wherein the thought chain includes multiple levels of prompt words, and the prompt words at each level are preset grammatical standard prompt words or one of the knowledge prompt words;

[0014] The thought chain and the current demand data are input into a preset large language model, and a generated query sentence is output, wherein the thought chain is interactively input into the large language model as a prompt word.

[0015] In one embodiment, the scenario adaptation model includes a feature extraction layer, a feature fusion layer and a prediction layer, and the step of inputting the application scenario distribution features and the current demand data into the pre-trained scenario adaptation model and outputting the application scenario corresponding to the current demand data includes:

[0016] Extracting semantic features in the current demand data based on a convolutional neural network (CNN) through the feature extraction layer;

[0017] The feature fusion layer fuses the semantic vector and the application scenario distribution feature based on a multimodal Tucker Fusion algorithm to obtain a fusion feature;

[0018] The prediction layer predicts the application scenario corresponding to the current demand data according to the fusion feature, and outputs the application scenario.

[0019] In one embodiment, the step of generating target query knowledge based on the application scenario and the current demand data includes:

[0020] Initial query knowledge matching the current demand data is retrieved from a preset knowledge base, and target query knowledge applicable to the application scenario is screened out from the initial query knowledge.

[0021] In one embodiment, the preset knowledge base includes one or more of a preset entity knowledge base, a preset table structure knowledge base, and a preset template knowledge base, and the step of retrieving initial query knowledge matching the current demand data from the preset knowledge base includes:

[0022] If the preset knowledge base includes a preset entity knowledge base, searching the preset entity knowledge base for a target entity that matches the current demand data, obtaining entity information corresponding to the target entity in the preset entity knowledge base, and determining that the initial query knowledge includes the entity information corresponding to the target entity;

[0023] If the preset knowledge base includes a preset table structure knowledge base, searching the preset table structure knowledge base for a target table that matches the current demand data, obtaining table metadata information of the target table in the preset table structure knowledge base, and determining that the initial query knowledge includes the table metadata information of the target table;

[0024] If the preset knowledge base includes a preset template knowledge base, the target demand data matching the current demand data is searched in the preset template knowledge base, and a query template example corresponding to the target demand data in the preset template knowledge base is obtained, and it is determined that the initial query knowledge includes the query template example corresponding to the target demand data.

[0025] In one embodiment, the preset table structure knowledge base is annotated with business scenarios corresponding to each table, the initial query knowledge includes table metadata information, and the step of screening target query knowledge from the initial query knowledge based on the application scenario includes:

[0026] Acquire the business scenario corresponding to each of the target tables, and filter out the interference table from the target table, wherein the business scenario corresponding to the selected table is inconsistent with the application scenario;

[0027] Interference query knowledge is determined based on the interference table, and the interference query knowledge is deleted from the initial query knowledge to obtain target query knowledge, wherein the interference query knowledge at least includes table metadata information of the interference table.

[0028] In addition, to achieve the above-mentioned purpose, the present application also provides a query statement generating device, the query statement generating device comprising:

[0029] The acquisition module is used to obtain the user's current demand data and application scenario distribution characteristics;

[0030] A scene adaptation module, used to input the application scene distribution characteristics and the current demand data into a pre-trained scene adaptation model, and output the application scene corresponding to the current demand data, wherein the scene adaptation model is used to extract semantic features in the current demand data to obtain a semantic vector, and fuse the semantic vector with the application scene distribution characteristics to predict the output;

[0031] A generation module is used to generate target query knowledge based on the application scenario and the current demand data, and generate a query statement according to the target query knowledge.

[0032] In addition, to achieve the above-mentioned purpose, the present application also provides a query statement generating device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the query statement generating method as described above.

[0033] In addition, to achieve the above-mentioned purpose, the present application also provides a readable storage medium, which is a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and the computer program is executed by a processor to implement the steps of the query statement generation method as described above.

[0034] The present application also provides a computer program product, including a computer program, which implements the steps of the query statement generation method as described above when executed by a processor.

[0035] One or more technical solutions proposed in this application have at least the following technical effects:

[0036] Acquire the user's current demand data and application scenario distribution characteristics; input the application scenario distribution characteristics and the current demand data into a pre-trained scenario adaptation model, and output the application scenario corresponding to the current demand data, wherein the scenario adaptation model is used to extract semantic features in the current demand data to obtain a semantic vector, and fuse the semantic vector with the application scenario distribution characteristics to predict the output; generate target query knowledge based on the application scenario and the current demand data, and generate a query statement based on the target query knowledge. Thus, in the embodiment of the present application, after obtaining the current demand data of the user and the distribution characteristics of the application scenario, the application scenario corresponding to the current demand data is identified using the pre-trained scenario adaptation model, and the target query knowledge is generated based on the application scenario and the current demand data, so as to use the application scenario to reduce the retrieval scope of the knowledge, realize the scenario screening capability in the process of knowledge retrieval, screen out the target query knowledge used in this application scenario, and generate a query statement based on the target query knowledge, instead of directly using the query knowledge obtained by retrieving only based on the demand data to generate a query statement, realize the precise positioning of knowledge through the scenario adaptation method, improve the accuracy of knowledge retrieval, make the generated query statement adapt to the user's application scenario, reduce the situation where the generated query statement does not meet the user's needs, and thus achieve the goal of improving the reliability of query statement generation. In addition, by screening the query knowledge by identifying the application scenario corresponding to the current demand data, a query basis with high accuracy can also be generated in complex application scenarios, which improves the adaptability to complex business scenarios in the query statement generation process. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0038] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0039] Figure 1 This is a flowchart of the first embodiment of the query statement generation method of the present application;

[0040] Figure 2 A schematic diagram of the scene adaptation model training and reasoning process involved in an embodiment of the query statement generation method of the present application;

[0041] Figure 3 A schematic diagram of the construction and update process of a preset knowledge base involved in an embodiment of the query statement generation method of the present application;

[0042] Figure 4 This is a schematic diagram of the overall query process involved in an embodiment of the query statement generation method of the present application;

[0043] Figure 5 A schematic diagram of the prompt word interaction process involved in an embodiment of the query statement generation method of the present application;

[0044] Figure 6 A schematic diagram of the structure of a query statement generating device for the present application;

[0045] Figure 7 Schematic diagram of the device structure of the hardware operating environment involved in the query statement generation method device in the embodiment of the present application.

[0046] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0047] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present invention.

[0048] The technology of generating SQL queries from natural language text refers to understanding the semantics of user questions through models, and training a large number of natural language and SQL pairs so that the model can understand and generate SQL queries. Traditional SQL query statement generation methods mainly include two types: SQL grammar writing rules and large language models.

[0049] Generation method based on grammar rule parsing: According to the preset unified mapping relationship, the original SQL statement issued by the central cluster is converted to automatically generate sub-SQL statements matching M edge clusters. The original SQL statement is parsed according to the preset mapping rules to generate a first syntax tree; based on the preset mapping relationship between the main table metadata and the M groups of sub-table metadata, the main table metadata in the first syntax tree is replaced with the M groups of sub-table metadata to obtain M second syntax trees; the M second syntax trees are reversed and M sub-SQL statements are generated.

[0050] Generation method based on large language model: There are two main types of SQL generation methods based on large models, namely, methods based on large model fine-tuning and methods based on enhanced retrieval. The former constructs domain knowledge as instruction data and inputs it into an open source large model for instruction fine-tuning; the latter generates prompt words through external knowledge retrieval and generates SQL using the context learning ability of the large model. In practical applications, the method based on enhanced retrieval is more widely used. Specifically, when the user enters a question, the user question is parsed and the domain knowledge stored in the database is text-matched using the parsed result combined with the original question. The retrieved knowledge is used to generate prompt words, which are then input into the large model for SQL generation. More in-depth, for example, a method based on an open source large language model that combines enhanced retrieval with fine-tuning technology generates SQL statements. This method achieves the result of SQL generation through five steps. First, build a knowledge base to store domain knowledge in a vectorized way. Second, build a knowledge base retrieval engine that supports word vector matching, which is used for user question matching and vectorized knowledge retrieval. Third, build the Text2SQL Prompt (Text to SQL Prompt) model of the knowledge base to complete the prompt word project. Fourth, fine-tune the large model based on the P-tuning (pre-training) method to align the large language model with the language in the domain. Finally, interactively parse and generate SQL for the large model, input the demand text containing domain knowledge and user needs into the large model for reasoning, and complete the SQL parsing and generation work.

[0051] However, the traditional query statement generation method has at least the following problems:

[0052] The generation method based on grammatical rule parsing has at least the problems of difficult template construction and semantic understanding. The large model generation method based on fine-tuning has the problem of insufficient knowledge real-time and catastrophic forgetting of the model; the large model method based on enhanced retrieval has difficulty in generating SQL when it needs to accurately screen from a large number of candidate data tables.

[0053] The process of building templates consumes a lot of manpower, and for tasks with complex processes, the quality of the built templates is difficult to control; for complex SQL logic, such as SQL nesting, multi-table association queries and other scenarios, template rules are not easy to expand, and the rules are prone to conflict, making modeling difficult and difficult to generate accurately. The rule matching method does not semantically understand the user's query needs, and only maps based on the template matching method, which cannot fully utilize the expert knowledge in the field, and the generated SQL has poor accuracy.

[0054] Although the generation method of fine-tuning a large model can permanently store the knowledge learned in the training data, if knowledge beyond the training data needs to be supplemented, fine-tuning training must be performed again, and the knowledge update lacks real-time performance. The fine-tuned model may have the problem of catastrophic forgetting. To alleviate this problem, it is necessary to mix general domain data for data, and the data ratio needs to be adjusted. The training is time-consuming and requires a lot of computing power.

[0055] When using enhanced retrieval methods to generate SQL, the retrieval algorithm is under great pressure when a large amount of knowledge is stored in the knowledge base, and it is easy to recall knowledge that does not match user needs, causing logic problems in large model reasoning. Especially when the table structure is similar or a field exists in multiple tables, it is more difficult for the knowledge retrieval algorithm to filter out the data table that matches the user's needs.

[0056] Based on this, the main solution of this application is: to obtain the user's current demand data and application scenario distribution characteristics; to input the application scenario distribution characteristics and the current demand data into a pre-trained scenario adaptation model, and output the application scenario corresponding to the current demand data, wherein the scenario adaptation model is used to extract the semantic features in the current demand data to obtain a semantic vector, and to fuse the semantic vector with the application scenario distribution characteristics to predict the output; to generate target query knowledge based on the application scenario and the current demand data, and to generate a query statement based on the target query knowledge.

[0057] After obtaining the user's current demand data and application scenario distribution characteristics, the application scenario corresponding to the current demand data is identified using the pre-trained scenario adaptation model, and target query knowledge is generated based on the application scenario and the current demand data, so as to use the application scenario to reduce the retrieval scope of knowledge, realize the scenario screening capability in the process of knowledge retrieval, screen out the target query knowledge used in this application scenario, and generate query statements based on the target query knowledge, instead of directly using the query knowledge obtained based on the demand data retrieval to generate query statements, realize the precise positioning of knowledge through scenario adaptation, improve the accuracy of knowledge retrieval, make the generated query statements adapt to the user's application scenario, reduce the situation where the generated query statements do not meet the user's needs, and thus achieve the goal of improving the reliability of query statement generation. In addition, by identifying the application scenario corresponding to the current demand data, the query knowledge is screened, so that a query basis with higher accuracy can be generated in complex application scenarios, and the adaptability to complex business scenarios in the query statement generation process is improved.

[0058] It should be noted that the execution subject of each embodiment of the query statement generation method of the present application can be a computing service device with data processing, network communication and program running functions, such as a server, a tablet computer, a personal computer, a mobile phone, etc., or a query statement generation device that can achieve the above functions. The embodiments of the query statement generation method of the present application do not make specific restrictions on this.

[0059] Based on this, the present application proposes a query statement generation method of the first embodiment, please refer to Figure 1 As shown, the query statement generation method includes the following steps S10 to S30:

[0060] Step S10, obtaining the user's current demand data and application scenario distribution characteristics;

[0061] The current demand data may specifically include but is not limited to the demand text input by the user (also called the user) that is expected to be converted into a query statement. For example, assuming that the user enters "the number of users with different user statuses and the average ARPU value in the campus user details table on May 18, 2024" in text form in the chat input box in the front-end interactive page, then "the number of users with different user statuses and the average ARPU value in the campus user details table on May 18, 2024" can be determined as the current demand data.

[0062] The application scenario feature may specifically be a distribution feature of the application scenario to which the user's historical questions belong, which may be obtained by analyzing historical demand data, and the historical demand data may specifically be demand data input by the user before the current demand data.

[0063] Exemplarily, the application scenarios in the historical demand data can be grouped according to the scenario categories, and the number of each group is counted to generate the application scenario distribution characteristics corresponding to the number of historical demand households. The feature number is the number of scenario categories. If there is no demand for some scenarios, it is supplemented with a preset value, such as 0. Exemplarily, the application scenario distribution characteristics can be {"Number of demand proposals for scenario A": "2", "Number of demand proposals for scenario B": "0", "Number of demand proposals for scenario C"; "7", ...}.

[0064] Step S20, inputting the application scenario distribution characteristics and the current demand data into a pre-trained scenario adaptation model, and outputting the application scenario corresponding to the current demand data, wherein the scenario adaptation model is used to extract semantic features in the current demand data to obtain a semantic vector, and then fuse the semantic vector with the application scenario distribution characteristics to predict the output;

[0065] When setting the modeling target of the scene adaptation model, it is found that there is a certain correlation between the application scenario characteristics in the user's historical data and the application scenario of the current data. In layman's terms, the application scenario of the user's question usually has a certain stability. Based on this, this embodiment obtains the user's application scenario distribution characteristics and inputs them into the scene adaptation model.

[0066] The scene adaptation model is pre-trained to identify the application scenario corresponding to the current demand data, so as to achieve efficient and high-accuracy scene recognition.

[0067] The scene adaptation model can be a pre-built classification model, and in order to reduce the demand for computing resources, the scene model can be a small model, such as a support vector machine (SVM), a decision tree model, etc., which is not specifically limited in this embodiment. Among them, a small model refers to a model with a relatively small number of parameters and a relatively simple structure, which is suitable for processing small-scale data sets or for use in situations where computing resources are limited.

[0068] The training method of the scene adaptation model can be a model trained under a supervised learning framework. Supervised learning is a machine learning paradigm in which the model learns by using labeled training data to predict or classify new data. Specifically, the scene adaptation model can be trained using the demand data and application scenario distribution characteristics as inputs and the annotated application scenarios as labels to obtain a pre-trained scene adaptation model.

[0069] After the scene adaptation model is trained, it can be used to identify the application scenario corresponding to the current demand data. The application scenario corresponding to the current demand data is identified based on the application scenario distribution characteristics and the current demand data. There is a certain correlation between the application scenario currently being asked by the user and the application scenario that has been asked in the past. In this way, the application scenario can be identified by combining the application scenario distribution characteristics with the current demand data to improve the recognition accuracy of the application scenario.

[0070] The application scenario may specifically be used to indicate a business scenario to which the current demand data is applied.

[0071] Relevant personnel can divide all business scenarios in advance. For example, in a specific implementation, business scenarios include channel analysis, terminal sales, 5G sales, complaint analysis and other scenarios, so as to identify the business scenarios to which the current demand data is applied in all business scenarios and obtain the application scenarios corresponding to the current demand data. For example, assuming that the business scenario to which the current demand data "the number of users and average ARPU values ​​of different user states in the campus user details table on May 18, 2024" is applied is identified as 5G sales, then the application scenario corresponding to the current demand data is determined to be 5G sales.

[0072] Step S30, generating target query knowledge based on the application scenario and the current demand data, and generating a query statement according to the target query knowledge.

[0073] It should be noted that the target query knowledge may specifically be query knowledge used in the application scenario.

[0074] A query statement refers to a statement used to retrieve, search or query data from a database. In this embodiment, the query statement generated may specifically be an SQL query statement. In the following, each embodiment of the present application is described and illustrated by generating an SQL query statement.

[0075] This embodiment obtains the user's current demand data and application scenario distribution characteristics; inputs the application scenario distribution characteristics and the current demand data into a pre-trained scenario adaptation model, and outputs the application scenario corresponding to the current demand data, wherein the scenario adaptation model is used to extract semantic features in the current demand data to obtain a semantic vector, and fuses the semantic vector with the application scenario distribution characteristics to predict the output; generates target query knowledge based on the application scenario and the current demand data, and generates a query statement based on the target query knowledge. Thus, in this embodiment, after obtaining the current demand data and application scenario distribution characteristics of the user, the application scenario corresponding to the current demand data is identified using the pre-trained scenario adaptation model, and the target query knowledge is generated based on the application scenario and the current demand data, so as to use the application scenario to reduce the retrieval scope of knowledge, realize the scenario screening capability in the process of knowledge retrieval, screen out the target query knowledge used in this application scenario, and generate query statements based on the target query knowledge, instead of directly using the query knowledge obtained by retrieving only based on the demand data to generate query statements, realize the precise positioning of knowledge through scenario adaptation, improve the accuracy of knowledge retrieval, make the generated query statements adapt to the user's application scenario, reduce the situation where the generated query statements do not meet the user's needs, and thus achieve the goal of improving the reliability of query statement generation. In addition, by screening the query knowledge by identifying the application scenario corresponding to the current demand data, a query basis with high accuracy can also be generated in complex application scenarios, and the adaptability to complex business scenarios in the query statement generation process is improved.

[0076] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated later. On this basis, the scene adaptation model includes a feature extraction layer, a feature fusion layer and a prediction layer, and the step of inputting the application scene distribution features and the current demand data into the pre-trained scene adaptation model and outputting the application scene corresponding to the current demand data includes:

[0077] Step A10, extracting semantic features in the current demand data based on a convolutional neural network (CNN) through the feature extraction layer;

[0078] The scene adaptation model includes a feature extraction layer, a feature fusion layer and a prediction layer connected in sequence. The feature extraction layer is used to extract the semantic features of the input data (current demand data in this embodiment), and the specific extraction method of extracting the semantic features can be pre-set. For example, in this embodiment, the current demand data can be vectorized first, and converted into a feature vector by vectorization to encode the current demand data into a vector of fixed length, and then the semantic features of the feature vector are extracted using a CNN network (Convolutional Neural Network).

[0079] Among them, vectorization refers to the process of converting data into numerical vectors. Similarly, the specific processing method of vectorization processing can also be pre-set. For example, the vectorization processing of data can be completed by embedding (such as word embedding). Embedding processing usually refers to converting data from its original form into a low-dimensional, continuous vector representation that can capture the intrinsic characteristics and structure of the data.

[0080] Step A20, fusing the semantic vector with the application scenario distribution feature through the feature fusion layer based on a multimodal Tucker Fusion algorithm to obtain a fusion feature;

[0081] The feature fusion layer is used to fuse semantic vectors with application scenario distribution features. Similarly, a specific fusion method can be pre-set. For example, in this embodiment, the semantic vectors and application scenario distribution features are fused based on the Tucker Fusion algorithm (based on multimodal Tucker fusion). The algorithm uses the application scenario distribution features and the feature vectors of the current demand data for outer product to construct a high-order vector, and obtains the core feature vector and factor matrix through the Tucker decomposition operator, and calculates the fused low-order features.

[0082] Step A30, predicting the application scenario corresponding to the current demand data according to the fusion features through the prediction layer, and outputting the application scenario.

[0083] It should be noted that the prediction result of the prediction layer may specifically be the probability distribution of the total scene category length of the application scene, wherein the scene category corresponding to the position with the largest probability value is the application scene output by the model prediction.

[0084] For example, in order to help understand the specific implementation of application scenario recognition in this embodiment, a specific embodiment is now listed. In this specific embodiment, refer to Figure 2 As shown, the specific implementation of application scenario recognition includes:

[0085] (1) Historical information retrieval: retrieve the historical demand data corresponding to the questions submitted by the user in the past according to the user's unique identifier. Figure 2 ), as the input of feature engineering. For example, the retrieved user historical demand data can be [{"demand text":"query the average ARPU (after discount) of users on 20240712","demand time":"2023-10-23 14:03:25","demand proposer ID":"X1"}, ...].

[0086] (2) Feature engineering: construct feature data based on the retrieved user historical demand data. The feature data required for scene adaptation model training and prediction is constructed through two dimensions: the distribution characteristics of the corresponding scenes of historical demands and the semantic characteristics of the current demand text in the current demand data. Feature engineering is mainly divided into three steps:

[0087] Step 1: Group the application scenarios in the historical demand data according to the scenario categories, count the number of each group, and generate the scenario distribution characteristics corresponding to the historical demand data (i.e., the application scenario distribution characteristics). The feature number is the number of scenario categories. If there is no demand for some scenarios, fill in 0 to supplement. For example, the scenario distribution characteristics can be {"Number of demand proposals for scenario A": "2", "Number of demand proposals for scenario B": "0", "Number of demand proposals for scenario C": "7", ...}.

[0088] Step 2: Based on the text vector embedding model Text2Vec model, embed the user's current demand text (that is, the current demand data), encode the demand text into a vector of fixed length, and select a pre-trained embedding model of 100 dimensions in this specific embodiment. Exemplarily, the text vector after vector embedding can be [-0.046396967, 0.0005419473, -0.0477318, 0.012477098, 0.019023571, -0.032702204, 0.022857154, 0.0064721783, 0.037023034, -0.02674591, -0.04008618 7, 0.009776611, 0.051318944, -0.014182999, -0.043979462, -0.04074635, 0.00025827286, 0.04813718, 0.067842305, 0.01381239, 0.030908015, -0.037887994, 0.008981815, ...].

[0089] Step 3: Feature fusion, fuse the features generated in step 1 and step 2 to generate the input data required for scene adaptation model training. It should be noted that in order to better fuse structured and text features, the Tucker Fusion feature fusion algorithm is used. This algorithm uses the scene distribution features and the demand text vector features for outer product to construct a high-order vector, obtains the core feature vector and factor matrix through the Tucker decomposition operator, and calculates the fused low-order features.

[0090] (3) Scene adaptation model training. The CNN model performs well in extracting text features. Since the goal of the scene adaptation model is to discover the correlation between scene features and labels in the user's historical data, CNN can better extract the local semantic features of the feature vector. The steps of CNN model modeling are as follows:

[0091] Step 1: Get the training data set, split it into input data and label data, and obtain the historical scene distribution characteristics through historical information retrieval and feature engineering ( Figure 2 The scene distribution features and semantic features shown in .

[0092] Step 2: Construct a CNN network. The CNN network consists of two feature extraction layers and a fully connected layer. The output layer performs a Softmax operation on the output of the fully connected layer.

[0093] It should be noted that each feature extraction layer consists of a convolutional layer and a pooling layer.

[0094] Step 3: Define the loss function and use Cross Entropy as the loss function. Since the scene adaptation model is essentially a classification model, the number of samples of each scene category in the business data used is relatively balanced. The Cross Entropy loss function has a good effect on this task while matching the classification task.

[0095] Step 4: Use the feature data and scene labels to iteratively train the model until the loss value converges. After convergence, the model parameters will be fixed and stored persistently.

[0096] (4) Prediction and reasoning of the scenario adaptation model: obtain the user's new demand text to be converted, that is, the current demand data to be converted into SQL query statements, generate historical application scenario features and semantic features through historical information retrieval and feature engineering steps, and input them into the scenario adaptation model for prediction and reasoning. The prediction result is the probability distribution of the total number of scenario categories, where the scenario category corresponding to the position with the largest probability value is the application scenario predicted by the model output ( Figure 2 Output scene categories shown in ).

[0097] It should be noted that the above examples are only used to assist in understanding the present application and do not constitute a limitation on the application scenario identification process of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0098] Based on this and the first embodiment and / or the second embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the above-mentioned embodiments 1 and 2 can be referred to the above introduction, and no further description will be given later. On this basis, the step of generating target query knowledge based on the application scenario and the current demand data includes:

[0099] Step B10: retrieve initial query knowledge matching the current demand data from a preset knowledge base, and select target query knowledge applicable to the application scenario from the initial query knowledge.

[0100] The preset knowledge base can be a knowledge base built in advance for storing knowledge. The data source for building the knowledge base can be data collected from the business system operation log or the data obtained from the data warehouse task configuration information. After collecting the original data according to the annotation classification, the data is annotated and integrated to build the preset knowledge base. Among them, the knowledge can be grammatical rules, database structures, entities, etc. required to help understand and apply the query language, so as to create effective database query statements.

[0101] It should be noted that the business scenarios to which each piece of knowledge in the preset knowledge base is applicable can be marked in advance. In this way, the business scenarios to which each initial query knowledge is applicable can be quickly obtained, thereby efficiently filtering out target query knowledge with consistent scenario categories from the initial query knowledge based on the application scenarios corresponding to the current demand data.

[0102] Furthermore, after the query statement is generated, the query statement can be output to the user for the user to verify whether the query statement meets the requirements, and the user can submit usage feedback to respond to the usage feedback request, and update the preset knowledge base according to the usage feedback request. For example, when the usage feedback request indicates that the output query statement meets the requirements, the generated query statement is stored in the preset knowledge base together with the current requirement data input by the user, and new business requirement knowledge is accumulated for the preset knowledge base; when the usage feedback request indicates that the output query statement does not meet the requirements, the generated example is analyzed, and the query knowledge that cannot be correctly identified is extracted to add new or existing knowledge is corrected, and saved to the preset knowledge base, and new knowledge is accumulated for the preset knowledge base.

[0103] In a possible implementation, the preset knowledge base includes one or more of a preset entity knowledge base, a preset table structure knowledge base, and a preset template knowledge base, and the step of retrieving initial query knowledge matching the current demand data from the preset knowledge base includes:

[0104] Step C10: if the preset knowledge base includes a preset entity knowledge base, searching the preset entity knowledge base for a target entity that matches the current demand data, obtaining entity information corresponding to the target entity in the preset entity knowledge base, and determining that the initial query knowledge includes the entity information corresponding to the target entity;

[0105] The preset knowledge base includes, but is not limited to, one or more of a preset entity knowledge base, a preset table structure knowledge base, and a preset template knowledge base. Relevant personnel can set other knowledge bases based on actual needs, such as a preset sample knowledge base for storing training data, wherein the training data is data for training the scene adaptation model. This embodiment does not impose specific restrictions on this.

[0106] The preset entity knowledge base is used to store entities, and each entity is specifically represented by an entity information. Each entity information includes but is not limited to entity name, entity explanation label, entity type, association table, field professional terminology explanation text, field indicator calculation method and statistical source description text. Relevant personnel can also mark other data as entity information, and this embodiment does not make specific restrictions on this.

[0107] Entity types can be divided into professional vocabulary explanations, field description information, table metadata information to which the field belongs, indicator calculation methods, etc.; entity explanation annotations are field application scenarios and field description texts respectively; the associated table is the name of the table to which the field belongs (if a field is used by multiple tables at the same time, the associated table contains all the table names of the field). Exemplarily, an entity information stored in the preset entity knowledge base can be {"Entity Name": "Month-on-month Growth Rate", "Entity Explanation Annotation": "Month-on-month Growth Rate = (Current Period Number - Previous Period Number) / Previous Period Number x 100%", "Entity Type": "Indicator Calculation Method", "Associated Table": "Table A"}.

[0108] Furthermore, an index can be constructed by entity name. When searching, first extract the entities in the current demand data, and match each entity extracted from the current demand data with each index of the preset knowledge base. If the match is successful, the entity represented by the entity name is determined to be a target entity that matches the current demand data, and repeat this process to obtain all target entities that match the current demand data.

[0109] Furthermore, the index column of the preset entity knowledge base is vectorized and encoded. The vectorization represents the semantic vector representing the index data. During retrieval, the entities extracted from the current demand data are also vectorized, and the similarity between the two vectors is matched by vector similarity calculation to obtain the target entity. Specifically, when the similarity between the two vectors is greater than the preset similarity threshold, it can be determined that the two vectors match.

[0110] Furthermore, if the retrieved entity information only contains table name information, the requirement text data input by the user in the previous time series may be obtained and used as the new current requirement text data, and the process returns to step S20 until the retrieved entity information contains more than just table name information.

[0111] Step C20: if the preset knowledge base includes a preset table structure knowledge base, searching the preset table structure knowledge base for a target table that matches the current demand data, obtaining table metadata information of the target table in the preset table structure knowledge base, and determining that the initial query knowledge includes the table metadata information of the target table;

[0112] The preset table structure database is used to store tables, and each table is represented by table metadata information. The table metadata information includes but is not limited to the table name, table description information, full field name, and text composed of one or more data in the field description information. Relevant personnel can also mark other data as table metadata information, and this embodiment does not make specific restrictions on this. For example, a table metadata information can be {"Data table name": "DWS_USER_SH_ARPU_MM_YYYYMM", "Table structure annotation": "**Table name**: User income table\n\n**Field name**: COUPON\n**Field Chinese name**: and package coupon\n**Field description**: and package coupon fee\n**Field caliber**: derived from cur_fee in cqxzcbd b.rpt_st_cw_mkt_hbq_div_month_detail\n\n**Field name**: CO UNTY_ID\n**Field Chinese name**: Branch code\n**Field description**: Branch code\n**Field caliber**: Branch code\n..."}.

[0113] Similarly, an index of a preset table structure knowledge base may be constructed (eg, using the table name as the index), and the index may be vectorized and stored, which will not be repeated here.

[0114] Step C30, if the preset knowledge base includes a preset template knowledge base, search the preset template knowledge base for target demand data that matches the current demand data, obtain the query template example corresponding to the target demand data in the preset template knowledge base, and determine that the initial query knowledge includes the query template example corresponding to the target demand data.

[0115] The preset template knowledge base is used to store demand data and corresponding query template examples, which can collect historical user demand data and annotate SQL query statements written according to user demand. For example, a piece of data in the preset template knowledge base can be {"user demand text": "query the average ARPU (after discount) of users on July 12, 2024", "query template example": "SELECT AVG(zk_arpu) AS avg_zk_arpu FROM DW_USER_ARPU_MOU_DOU_MM_20240712"}.

[0116] Similarly, an index of a preset template knowledge base may also be constructed (eg, using demand data as an index), and the index may be vectorized and stored, which will not be repeated here.

[0117] It should be noted that when constructing the index of the preset template knowledge base, in a preferred embodiment, the demand text vector that is vectorized and converted by the Text2vec model can be operated to construct an HNSW (HierarchicalNavigableSmallWorld) index. The HNSW index is a graph-based approximate nearest neighbor search algorithm index. By constructing a hierarchical small world network, efficient similarity search can be achieved in high-dimensional space. The index can effectively improve the efficiency of vectorized retrieval.

[0118] It should be noted that the initial query knowledge includes but is not limited to one or more of entity information, query template examples, and table metadata information, and may include query specifications, which is not specifically limited in this embodiment.

[0119] In this embodiment, based on the user's current demand data, a preset entity knowledge base, a preset table structure knowledge base and / or a preset template knowledge base are retrieved to obtain initial query knowledge such as entity information, table metadata information and / or query template examples, thereby ensuring that the generated initial query knowledge meets user needs and providing an effective screening data set for subsequent screening to improve the accuracy and relevance of the generated query statements.

[0120] Furthermore, when updating the preset knowledge base, the corresponding data may be stored in the corresponding knowledge base to complete the update of the corresponding knowledge base. For example, when the data to be stored is the current demand data and the corresponding query statement, the current demand data and the query statement may be stored in the preset template knowledge base to update the preset template knowledge base; and when the data to be stored is information related to an entity, such information may be stored in the preset entity knowledge base to update the preset entity knowledge base.

[0121] In a possible implementation, the preset table structure knowledge base is annotated with business scenarios corresponding to each table, the initial query knowledge includes table metadata information, and the step of screening target query knowledge from the initial query knowledge based on the application scenario includes:

[0122] Step D10, obtaining the business scenarios corresponding to the target tables, and filtering out interference tables from the target tables, wherein the business scenarios corresponding to the selected tables are inconsistent with the application scenarios;

[0123] The preset table structure knowledge base is marked with business scenarios corresponding to each table, so that interference tables inconsistent with the scenario categories of application scenarios of current demand data can be screened out based on the business scenarios marked in each table.

[0124] Step D20, determining interference query knowledge based on the interference table, deleting the interference query knowledge from the initial query knowledge to obtain target query knowledge, wherein the interference query knowledge at least includes table metadata information of the interference table.

[0125] The interference query knowledge may specifically be query knowledge associated with the interference table, including but not limited to table metadata information of the interference table. For example, if the initial query knowledge includes entity information, and the entity information includes an association table, then the interference entity information whose association table is the interference table may also be filtered out from the entity information included in the initial query knowledge, and the interference entity information may also be determined as interference query knowledge, so as to filter the entity information, further reduce the knowledge scope, and improve the accuracy of knowledge positioning.

[0126] Furthermore, in this embodiment, based on the business scenarios of the table and the application scenarios of the current demand data, interfering query knowledge is deleted from the initial query knowledge, thereby narrowing the scope of knowledge retrieval and improving the accuracy of knowledge provided for generating query statements, so that the generated query statements are more stable and reliable.

[0127] For example, in order to help understand the construction and update of the preset knowledge base in this embodiment, a specific embodiment is now listed. In this specific embodiment, the preset knowledge base includes an entity library (that is, a preset entity knowledge base), a sample library (that is, a preset sample knowledge base), a template library (that is, a preset template knowledge base), and a table structure library (that is, a preset table structure knowledge base). Based on this, refer to Figure 3 As shown, the construction and update of the preset knowledge base includes:

[0128] 1. Data annotation. First, collect the corpus data, then annotate and store the collected corpus data to build four libraries. The way to obtain data is to collect business system operation logs or data warehouse task configuration information. After collecting the corpus data according to the annotation classification, annotate and integrate the data respectively. The annotation method is as follows:

[0129] (1) Sample Library Figure 3The scene adaptation model feature library shown in the figure is annotated, and the corresponding scene category is annotated based on the user's historical demand data: the scene adaptation model feature data storage modeling feature is collected by collecting the historical demand description text, demand time, demand proposer and other information of the user in the business system, and feature generation is performed through feature engineering, and then the scene category corresponding to the demand text is annotated. Specifically, based on the word embedding model Text2vec-large-chinese, the demand description text is converted into a sentence vector feature, and the number of demand proposals for each scene is counted according to the demand proposer group. It should be noted that in order to better explore the potential of vector matching, this specific embodiment compares the actual effects of the Text2vec-large-chinese (large Chinese text vectorization) model and the BGE (Business Entity and Grammar model, business entity and grammar model) model. Since the data deposited in the actual business system is mostly of medium length and the text description is in Chinese, when training the scene-adapted CNN model, the Text2vec model has a higher accuracy rate in the test set, so this specific embodiment uses the Text2vec model as the vectorization model. Exemplarily, the data after the template library annotation can be {"user demand text":"query the average ARPU of users on 20240712 (after discount)","number of times demand for scenario A is raised":"2","number of times demand for scenario B is raised":"0","number of times demand for scenario C is raised":"7",...,"scene category label":"scenario A"}.

[0130] (2) Entity library annotation, using the descriptive text of the field explanation class, calculation method class, statistical caliber class, and professional term class entities as annotation content: The entity library stores the entity name text annotated with detailed explanations and descriptions, and builds an index based on the entity name.

[0131] (3) Template library annotation, taking the SQL query statements corresponding to the historical user demand texts (i.e., query template examples) as annotation content: The template library stores user demand texts annotated with SQL, builds an index by collecting historical user demand data, and annotates SQL query statements written according to user needs.

[0132] (4) Table structure information annotation, with the corresponding field information and table name description of each business table as the annotation content: the table structure library stores table structure text annotated with field names and descriptions, and builds an index by integrating the business system data dictionary. The annotation data is a text consisting of each table name and description information, and all field names and description information.

[0133] 2. Vectorized storage. First, apply the BCembedding algorithm to vectorize and store the three types of data marked in step 1. For the template library data, the vectorized text content is an index column. Specifically, the template library index is the user demand text and its vector; the entity library index is the entity name and its vector; the table structure library index is the table name and its vector. The vector fields in the three libraries are stored together with others. Specifically, the entity name in the entity library is vectorized and stored together with the entity explanation, entity type and other fields. The user demand text in the template library is vectorized and stored with the corresponding SQL query field. The table name in the table structure library is vectorized and stored together with the table metadata information.

[0134] Four knowledge bases are constructed through steps 1 and 2.

[0135] 3. Run SQL to generate applications and SQL query statements ( Figure 3 The result SQL shown in the figure is generated, and the knowledge base is updated through iterative learning based on user feedback: the user can run the SQL generation application to output the generated SQL query statement for the SQL generation application. The user's demand data generated in the process of using the SQL query generation application and the corresponding SQL query statement generated by the model can be marked by the user's satisfaction. If the SQL query statement generated by the model meets the user's needs, that is, the business logic is verified, then this set of data ( Figure 3 The embodiment data shown in the example is automatically vectorized and stored, input into the template library, enrich the annotation data set, provide richer annotation data for model context learning, and can also input high data into the scene adaptation model feature library (i.e., sample library). When the business logic check fails, the embodiment data of the generated example is analyzed, the wrong entities are analyzed, the entity knowledge that the model fails to correctly identify is extracted, new or existing knowledge is corrected, and saved to the entity library to accumulate new knowledge for the entity library.

[0136] It should be noted that the above examples are only used to assist in understanding the present application and do not constitute a limitation on the construction and update process of the preset knowledge base of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0137] Based on the first embodiment, the second embodiment and / or the third embodiment of the present application, in the fourth embodiment of the present application, the same or similar contents as those in the first embodiment, the second embodiment and the third embodiment can be referred to the above introduction, and will not be repeated later. On this basis, the step of generating a query statement based on the target query knowledge includes:

[0138] Step E10, performing prompt word engineering processing on the target query knowledge to obtain knowledge prompt words, wherein the knowledge prompt words include one or more of entity information prompt words, query template example prompt words and table metadata information prompt words;

[0139] Prompt Engineering is a technique used in the field of Natural Language Processing (NLP) to create and optimize input prompts to better guide the model to generate the desired output.

[0140] The entity information prompt word may specifically be a prompt word corresponding to the entity information, the query template example prompt word may specifically be a prompt word corresponding to the query template example, and the table metadata information may specifically be a prompt word corresponding to the table metadata information.

[0141] Step E20, constructing a thought chain according to the knowledge prompt words, wherein the thought chain includes multiple levels of prompt words, and the prompt words at each level are preset grammar standard prompt words or one of the knowledge prompt words;

[0142] Chain of Thought (CoT) describes a series of reasoning steps in a model when generating text or solving problems. It provides a way to understand how the model reasoned from input to output, which increases the interpretability of the model's decision-making process. For example, assuming that the prompt words included in the thought chain include entity information prompt words, query template example prompt words, table metadata information prompt words, and preset grammar specification prompt words, then the thought chain can be [preset grammar specification prompt words -> entity information prompt words -> query template example prompt words -> information prompt words].

[0143] The preset grammar standard prompt word can be a prompt word for indicating grammar standard. For example, in a specific implementation, the preset grammar standard prompt word can be "You are a data analysis expert and proficient in MySQL. Given an input question, create a grammatically correct MySQL query to run. You must query only the columns required to answer the question. The table name suffix {YYYYMMDD} represents the date, the granularity is day, and the table is a daily table; the suffix {YYYYMM} granularity is month, and the table is a monthly table. If the table has a date suffix, it is necessary to query the corresponding daily table or monthly table in combination with the date in the question. When generating SQL, you need to use the real English field name field_name, you cannot use the Chinese description, and only use the field name you can see in the table below. Be careful not to query non-existent fields. In addition, please pay attention to which column is in which table. If the question involves "today", please pay attention to using the CURDATE() function to get the current date.

[0144] Step E30, inputting the thought chain and the current demand data into a preset large language model, and outputting a generated query sentence, wherein the thought chain is interactively input into the large language model as a prompt word.

[0145] The preset large language model can be specifically a large language model with query sentence generation function. When the large language model is predicting, the model implicitly calculates information similar to gradients through the self-attention mechanism and complex nonlinear transformations. This information helps the model adjust its internal representation during reasoning, thereby making more effective predictions on the input.

[0146] The target query knowledge can be specifically input into the large language model as a prompt word, interact with the large language model, and enhance the contextual learning ability of the large language model, thereby ensuring that the generated query statements meet user needs and improving the accuracy and relevance of the generated query statements.

[0147] Inputting the query template examples corresponding to the target demand data retrieved from the preset template knowledge base into the large language model is equivalent to allowing the large language model to understand the user's needs and learn the format characteristics of the query statement using the corresponding query template examples as templates.

[0148] Inputting the entity information retrieved from the preset entity knowledge base into the large language model is equivalent to allowing the large language model to learn and understand user needs, allowing the large model to have relevant knowledge and the ability to analyze which entities in user needs correspond to query conditions or query fields.

[0149] Inputting the table metadata information retrieved from the preset table structure knowledge base into the large language model is equivalent to allowing the large language model to parse the table metadata information corresponding to the field by learning user needs. It is the process of finding the fields required by the query keyword.

[0150] Furthermore, the large language model can also construct a global memory object, which stores the conversation time, sequence number, user code, user input question and model output result corresponding to each interaction with the user, and is used to cope with the multi-round conversation scenarios generated when the user interacts with the large language model.

[0151] For example, it is assumed that the user has completed a round of query statement generation tasks through the application, and the user submits a new round of requirements based on the content output by the previous round of application. At this time, the memory object stores the content of 2 rounds of conversations, the content is: [{"Conversation time":"2024-06-19 10:20:31","Conversation sequence number":"000001","User code":"user_101","User demand text":"Analyze the number of users and average ARPU value of different user status on May 18, 2024","Model output":"User demand There are multiple tables containing ARPU fields. Please add the data table you want to analyze to the conversation content. The information of multiple tables is as follows: DWS_SCHOOL_USER_DETAIL_20240518-Campus user details table, DWS_USER_ARPU_AFTER_DISCOUNT_20240518-User discount income table, DWB_USER_NEW_DETAIL_COUNTY_DT_20240518-City income table"}, {"Conversation time":"2024-06-19 10:22:11","Dialogue Sequence Number":"000002","User Code":"user_101","User Requirement Text":"Campus User List","Model Output":"SELECT\n userstatus_id,\nCOUNT(user_id)AS user_count,\n AVG(arpu)AS avg_arpu\n FROM DWS_SCHOOL_USER_DETAIL_20240518\n GROUP BY\n userstatus_id\nORDERBY\n userstatus_id;"}].

[0152] Furthermore, when the target query knowledge includes table metadata information of multiple tables, that is, when multiple tables are retrieved, the demand data can be rewritten based on the large language model to reconstruct the user demand text, re-identify the application scenario based on the rewritten demand data, and filter out the query knowledge suitable for the application scenario from the target query knowledge again, and generate a query statement based on the re-filtered query knowledge and the rewritten demand data. Specifically, the user's historical needs can be extracted, and the needs can be rewritten by combining the historical user demand text with the preset prompt words. In a specific embodiment, the preset prompt words can be "You are a demand analysis expert, proficient in demand text analysis and refinement. Given several demand texts, understand the needs, extract key information, and rewrite a demand. Please note that the rewritten demand must contain all the important information of the demand text." The prompt words enable the large language model to understand the user's historical needs, understand, extract and rewrite the demand content contained in the user's historical conversation content, and the rewritten demand text contains the user's historical demand scenario elements. Combined with the scenario information contained in the user's historical demand information, the large language model is helped to select a table suitable for the scenario demand from multiple candidate tables.

[0153] In this embodiment, query statements are generated based on the large language model, and the task of converting the current demand data input by the user into query statement generation is completed with the help of external knowledge retrieval and the context learning ability of the large model. For the large language model, knowledge injection can be performed without complex parameter fine-tuning, the consumption of computing resources is relatively lower, and the real-time nature of knowledge is stronger.

[0154] For example, in order to help understand the technical concept or technical principle of the query statement generation method after the present embodiment is combined with the first embodiment, the second embodiment and the third embodiment, a specific embodiment is now listed. In this specific embodiment, refer to Figure 4 As shown in the figure, the query statement generation process includes:

[0155] First, a preset knowledge base is constructed and a scene adaptation model is trained. The construction of the preset knowledge base includes: (1) collecting historical data and annotating the collected historical data. The annotation content includes sample library annotation and large language model knowledge corpus annotation. The annotated sample library provides feature data and scene classification labels for scene adaptation model training; the annotated large language model knowledge corpus data includes entity library, template library, and table structure library, which are necessary external knowledge for large language models to use for context learning and reasoning. Each library table corresponds to different field information and indexes, and the annotation data is prepared according to the corresponding table structure. The specific annotation method can refer to the above-mentioned specific implementation content and will not be described in detail here. (2) The index column of each library table is vectorized and encoded. The vectorization represents the semantic vector representing the index data, which provides a matching medium for the large language model to retrieve the knowledge required for context learning. The text similarity is measured by the text matching similarity algorithm. The higher the similarity score value, the more similar the user's needs are to the stored external knowledge.

[0156] Based on the built preset knowledge base and the trained scenario adaptation model, query statement generation includes scenario adaptation and SQL generation. The specific implementation process includes:

[0157] The first step is to input the business requirement text to be converted, that is, the current data.

[0158] The second step is historical information retrieval, which obtains the user's historical demand data and generates application scenario distribution characteristics.

[0159] The third step is scenario adaptation model prediction. The scenario adaptation model is used to make predictions, output the application scenario corresponding to the user's demand text, and filter out the target query knowledge based on the application scenario.

[0160] The fourth step is the LLM (Large Language Model) thinking chain. The large language model infers and performs prompt word engineering, at least using the target query knowledge as the prompt word.

[0161] In the fifth step, LLM generates SQL, and the large language model context learns and generates SQL.

[0162] The sixth step is to collect SQL user feedback and manually verify the SQL generated by the large language model.

[0163] The seventh step is to generate result verification. If the verification passes, the sample library and template library are updated; if the verification fails, the entity library is updated.

[0164] Further, see Figure 5 As shown, the fourth step specifically includes:

[0165] (1) Obtain the user demand text to be converted (that is, the current demand data), and use the user demand text as subsequent input.

[0166] (2) Construct a memory object to store detailed data for each round of dialogue: Construct a global memory object. The memory object stores detailed data such as the dialogue time, sequence number, user code, user input question, and model output result corresponding to each interaction with the user. It is used to cope with the multi-round dialogue scenarios generated when the user interacts with the large language model.

[0167] (3) Construct a large language model generation module. The large language model generation module implements the SQL query generation task by building a thought chain. The construction process includes the following steps:

[0168] a. Construct the Thinking Chain 1 module, identify and extract entities in the user input demand text, and generate an entity list: This module is used to parse the user input demand text, perform entity parsing on the current round of user input, use the sequence recognition model to extract the entity names in the user requirements, form an entity name list, and generate an array object containing user requirements and entity name list information, and input the entity list into the Thinking Chain 2 module.

[0169] Exemplarily, when performing entity resolution, the large language model can concatenate the preset first prompt word text with the user demand text. For example, the concatenated text can be "You are an entity extractor and need to perform entity extraction on user needs. Background: I want to query data in the database, and you need to extract keywords from the query statement. Finally, a JSON structure is returned. Here are some examples: <Example> Question: What is the average user rate in Beijing in February? Answer: {"entity list": ["February", "Beijing", "User", "Average Rate"]}. Please perform entity extraction and only return the JSON structure. No other comments are allowed. User needs are as follows: The number of users and average ARPU values ​​of different user status in the campus user details table on May 18, 2024.". After reasoning by the large language model, the entity list extracted from the user demand text is {"entity list": ["20240518", "campus user", "details", "different", "user status", "number of users", "average", "ARPU value"]}.

[0170] b. Construct the Thinking Chain 2 module, search the entity library, perform vectorized search on each entity in the entity list, and extract entity explanations and entity type information from the entity library: This module is a search module, which parses and searches the user demand text and each element in the extracted entity list, respectively. By searching the entity library, each entity in the entity list is vectorized searched, and entity explanations and entity type information are extracted from the entity library. The template library is searched, and the user demand text is vectorized searched. The closest demand text and corresponding SQL query (that is, query template example) in the template library are extracted, and the field explanation type entity retrieved from the entity library is input, and the table name information in the entity explanation corresponding to the field explanation type entity retrieved from the entity library is extracted. The specific steps for retrieval are as follows:

[0171] First, perform an entity library search operation. The entity retrieval process is to traverse each entity in the entity list, convert each entity into a semantic vector through a vectorized model, and use a similarity algorithm to perform text matching with the entity name vector index column in the entity library to retrieve entity explanations and entity category data to provide domain knowledge for the large model. In this application, the preset domain knowledge category, that is, the entity category, includes professional vocabulary explanations, field description information, table metadata information of field attribution, and indicator calculation methods. For different entity categories, there are also differences in subsequent operations. In particular, if the retrieved entity category only has table name information, it is necessary to extract the user demand text of the previous conversation number through the memory module and re-retrieve the entity information.

[0172] Secondly, the template library is searched based on the user's demand text, and the SQL query corresponding to the user's demand is searched in the template library, that is, the query template example. It should be noted that the vectorized search method is the same, only the text to be searched and the vector library index column are different, so it will not be repeated.

[0173] Finally, based on the entity name of the field interpretation category retrieved from the entity library, the table name is extracted by regular expression filtering, and the table name index in the table structure library is retrieved based on the extracted table name to obtain the table metadata information. Exemplarily, the entity whose entity category is field interpretation can be {"entity name":"ARPU","entity interpretation annotation":"Average revenue per user is usually referred to as ARPU, corresponding table name: DWS_SCH OOL_USER_DETAIL_{yyyymmdd}-campus user details table, DWS_USER_ARPU_AF TER_DISCOUNT_{yyyymmdd}-user discount income table, DWB_USER_NEW_DET AIL_COUNTY_DT_{yyyymmdd}-city income table","entity type":"field interpretation"}. The table name string after the "table name:" character is parsed through regular expressions, and all table names and Chinese names of the table names contained in the string are extracted through symbol cutting. The extracted list is ["DWS_SCH OOL_USER_DETAIL_{yyyymmdd}","DWS_USER_ARPU_AFTER_DISCOUN T_{yyyymmdd}","DWB_USER_NEW_DETAIL_COUNTY_DT_{yyyymmdd}"]. After the table name is extracted, the list is traversed to retrieve the corresponding table metadata information in the table structure library.

[0174] It should be noted that when more than one table name is retrieved through the table structure, multiple tables can be counted. In order to meet the user's statistical needs, it is necessary to confirm the specific table name that the user wants to perform statistical analysis on. At this time, the table name information is spliced ​​and input into the Thinking Chain 3 module for processing. When only one table name is retrieved from the table structure, the table name information, user demand text, entity list and other information are input into the Thinking Chain 4 module for processing.

[0175] c. Construct the thinking chain 3 module, read the historical needs of the memory object, and rewrite the needs based on the large language model: This module is used to reconstruct the user's demand text. This module calls the memory module of historical storage, extracts the user's historical needs, and rewrites the needs through the second prompt word combined with the historical user demand text. The preset second prompt word can be "You are a demand analysis expert, proficient in demand text analysis and refinement. Given several demand texts, understand the needs, extract key information, and rewrite a demand. Please note that the rewritten demand must contain all important information in the demand text." This prompt word enables the large language model to understand the user's historical needs, understand, extract and rewrite the demand content contained in the user's historical conversation content. The rewritten demand text contains the user's historical demand scenario elements, combined with the scenario information contained in the user's historical demand information to help the large language model select a table suitable for the scenario needs from multiple candidate tables.

[0176] d. Construct the thinking chain 4 module, and perform prompt word engineering based on the knowledge retrieved from the three knowledge bases: This module is used for prompt word engineering processing. The prompt words combined by prompt word engineering are used to generate SQL query statements, where the combination information includes the third to sixth prompt words, which are used for SQL syntax specification prompt words, entity information prompt words, query template example prompt words, and table metadata information prompt words respectively. The preset third prompt word can be "You are a data analysis expert and proficient in MySQL. Given an input question, create a syntactically correct MySQL query to run. You must query only the columns required to answer the question. The table name suffix {YYYYMMDD} represents the date, the granularity is day, and the table is a daily table; the suffix {YYYYMM} granularity is month, and the table is a monthly table. If the table has a date suffix, you need to query the corresponding daily table or monthly table in combination with the date in the question. When generating SQL, you need to use the real English field name field_name, you cannot use the Chinese description, and only use the field name you can see in the table below. Be careful not to query non-existent fields. Also, please note which column is in which table. If the question involves "today", please note the use of the CURDATE() function to get the current date. In the process of generating SQL statements, please do not use aliases. In addition to giving SQL answers, after giving the answer, please explain yourself concisely and clearly using the same language as the question. Assume that there is a database with the following tables and columns: Given the following database schema, convert the following natural language request into a valid SQL query. ", this prompt word is to standardize the grammatical specifications and output form of the SQL queries generated by the large language model, to avoid the SQL query syntax generated by the native knowledge of the large model not meeting the user's requirements. The preset fourth prompt word can be "The following is some ner (entity) information to help generate SQL.<ner_info> {Entity List Information}< / ner_info> ", in order to allow the large language model to better understand the semantic information of the user's demand text, the entity list information is spliced ​​with the fourth prompt word, wherein the entity list information includes the entity name and the entity interpretation information and entity category retrieved through the entity library. The preset fifth prompt word can be "The following are some examples of SQL generated using natural language: user question natural language: {user demand text sample}; generate SQL: {SQL query sample}", the user demand text and the corresponding SQL query sample are spliced ​​with the fifth prompt word. The preset sixth prompt word can be "The following is some information about data tables related to user questions that helps generate SQL:<table_info> {Table metadata information}< / table_info> ", concatenate the table metadata information with the fifth prompt word. It should be noted that the fourth to sixth prompt words are referenced in the labeled data format because most of the code data in the large language model training samples are input in a labeled form, which is helpful for the large model to understand.

[0177] e. Construct the thinking chain 5 module, the large language model performs context learning, and generates SQL that meets the needs of users: This module is used for the large language model to perform context learning and SQL generation reasoning. Generate query statements based on the in context learning capability of the large language model, which enables the large language model to predict labels for unseen inputs without additional parameter updates.

[0178] 2. The text to be converted is input into the large language model. The large language model performs chain reasoning based on the content included in the user's demand text, and performs entity extraction, logical judgment, knowledge retrieval, and prompt word engineering operations in sequence according to the chain reasoning logic preset by the method. Finally, the combined prompt words are input into the large language model for SQL query generation operation.

[0179] 3. Output interaction, output the SQL query statement generated by the large language model, the user performs SQL query logic verification, and adds new knowledge based on the verification results. After the large language model completes the SQL reasoning generation steps that meet the user's needs, it outputs the generated SQL query to the user, who then performs SQL query logic verification and adds new knowledge based on the verification results.

[0180] It should be noted that the above examples are only used to assist in understanding the present application and do not constitute a limitation on the query statement generation method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0181] In addition, the embodiment of the present application also provides a query statement generating device, referring to Figure 6 As shown, the query statement generating device includes:

[0182] The acquisition module 10 is used to obtain the user's current demand data and application scenario distribution characteristics;

[0183] A scene adaptation module 20 is used to input the application scene distribution characteristics and the current demand data into a pre-trained scene adaptation model, and output the application scene corresponding to the current demand data, wherein the scene adaptation model is used to extract the semantic features in the current demand data to obtain a semantic vector, and fuse the semantic vector with the application scene distribution characteristics to predict the output;

[0184] The generating module 30 is used to generate target query knowledge based on the application scenario and the current demand data, and to generate a query statement according to the target query knowledge.

[0185] In addition, an embodiment of the present application also proposes a query statement generating device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the query statement generating method as described above.

[0186] refer to Figure 7 , which shows a schematic diagram of the structure of a query statement generation device suitable for implementing the embodiment of the present application. The query statement generation device in the embodiment of the present application may also include but is not limited to mobile terminals such as servers, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The query statement generating device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0187] like Figure 7 As shown, the query statement generating device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the query statement generating device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the query statement generation device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a query statement generation device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or provided instead.

[0188] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0189] The query statement generation device provided in the embodiment of the present application adopts the query statement generation method in the above embodiment, which can solve the technical problem of how to improve the reliability of query statement generation. Compared with the prior art, the beneficial effects of the query statement generation device provided by the present application are the same as the beneficial effects of the query statement generation method provided in the above embodiment, and other technical features in the query statement generation device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0190] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0191] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0192] In addition, to achieve the above-mentioned purpose, an embodiment of the present application also provides a readable storage medium having computer-readable program instructions (ie, computer program) stored thereon, and the computer-readable program instructions are used to execute the query statement generation method in the above-mentioned embodiment.

[0193] The computer-readable storage medium provided in the embodiment of the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: Random Access Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.

[0194] The computer-readable storage medium may be included in the query statement generating device; or may exist independently without being assembled into the query statement generating device.

[0195] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the query statement generation device, the query statement generation device: obtains the user's current demand data and application scenario distribution characteristics; inputs the application scenario distribution characteristics and the current demand data into a pre-trained scenario adaptation model, and outputs the application scenario corresponding to the current demand data, wherein the scenario adaptation model is used to extract semantic features in the current demand data to obtain a semantic vector, and predicts the output after fusing the semantic vector with the application scenario distribution characteristics; generates target query knowledge based on the application scenario and the current demand data, and generates a query statement based on the target query knowledge.

[0196] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0197] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0198] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of a module does not, in some cases, constitute a limitation on the module itself.

[0199] The readable storage medium provided in the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned query statement generation method, and can solve the technical problem of how to improve the reliability of query statement generation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the present application are the same as the beneficial effects of the query statement generation method provided in the above-mentioned embodiment, and will not be repeated here.

[0200] In addition, an embodiment of the present application also proposes a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the query statement generation method described above are implemented.

[0201] The specific implementation of the computer program product of the present application is basically the same as the above-mentioned query statement generation and / or speech enhancement method embodiments, and will not be repeated here.

[0202] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.

[0203] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0204] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software sensor, which is stored in a storage medium (such as ROM / RAM, disk, CD) as described above, including a number of instructions for a terminal device (which can be a mobile phone, computer, server or network device, etc.) to execute the methods described in each embodiment of the present application.

[0205] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A query statement generation method, characterized in that: The query statement generation method comprises the following steps: Obtain users’ current demand data and application scenario distribution characteristics; Input the application scenario distribution features and the current demand data into a pre-trained scenario adaptation model, and output the application scenario corresponding to the current demand data, wherein the scenario adaptation model is used to extract semantic features from the current demand data to obtain a semantic vector, and fuse the semantic vector with the application scenario distribution features to predict the output; Target query knowledge is generated based on the application scenario and the current demand data, and a query statement is generated according to the target query knowledge.

2. The query statement generating method according to claim 1, characterized in that: The step of generating a query statement based on the target query knowledge includes: Performing prompt word engineering processing on the target query knowledge to obtain knowledge prompt words, wherein the knowledge prompt words include one or more of entity information prompt words, query template example prompt words and table metadata information prompt words; Constructing a thought chain according to the knowledge prompt words, wherein the thought chain includes multiple levels of prompt words, and the prompt words at each level are preset grammatical standard prompt words or one of the knowledge prompt words; The thought chain and the current demand data are input into a preset large language model, and a generated query sentence is output, wherein the thought chain is interactively input into the large language model as a prompt word.

3. The query statement generating method according to claim 1, characterized in that: The scenario adaptation model includes a feature extraction layer, a feature fusion layer and a prediction layer. The step of inputting the application scenario distribution features and the current demand data into the pre-trained scenario adaptation model and outputting the application scenario corresponding to the current demand data includes: Extracting semantic features in the current demand data based on a convolutional neural network (CNN) through the feature extraction layer; The feature fusion layer fuses the semantic vector and the application scenario distribution feature based on a multimodal Tucker Fusion algorithm to obtain a fusion feature; The prediction layer predicts the application scenario corresponding to the current demand data according to the fusion feature, and outputs the application scenario.

4. The query statement generating method according to any one of claims 1 to 3, characterized in that: The step of generating target query knowledge based on the application scenario and the current demand data includes: Initial query knowledge matching the current demand data is retrieved from a preset knowledge base, and target query knowledge applicable to the application scenario is screened out from the initial query knowledge.

5. The query statement generating method according to claim 4, characterized in that: The preset knowledge base includes one or more of a preset entity knowledge base, a preset table structure knowledge base, and a preset template knowledge base. The step of retrieving initial query knowledge matching the current demand data from the preset knowledge base includes: If the preset knowledge base includes a preset entity knowledge base, searching the preset entity knowledge base for a target entity that matches the current demand data, obtaining entity information corresponding to the target entity in the preset entity knowledge base, and determining that the initial query knowledge includes the entity information corresponding to the target entity; If the preset knowledge base includes a preset table structure knowledge base, searching the preset table structure knowledge base for a target table that matches the current demand data, obtaining table metadata information of the target table in the preset table structure knowledge base, and determining that the initial query knowledge includes the table metadata information of the target table; If the preset knowledge base includes a preset template knowledge base, the target demand data matching the current demand data is searched in the preset template knowledge base, and a query template example corresponding to the target demand data in the preset template knowledge base is obtained, and it is determined that the initial query knowledge includes the query template example corresponding to the target demand data.

6. The query statement generating method according to claim 5, characterized in that: The preset table structure knowledge base is annotated with business scenarios corresponding to each table, the initial query knowledge includes table metadata information, and the step of screening target query knowledge from the initial query knowledge based on the application scenario includes: Acquire the business scenarios corresponding to the target tables, and filter out interference tables from the target tables, wherein the business scenarios corresponding to the selected tables are inconsistent with the application scenarios; Interference query knowledge is determined based on the interference table, and the interference query knowledge is deleted from the initial query knowledge to obtain target query knowledge, wherein the interference query knowledge at least includes table metadata information of the interference table.

7. A query statement generating device, characterized in that: The query statement generating device comprises: The acquisition module is used to obtain the user's current demand data and application scenario distribution characteristics; A scene adaptation module, used to input the application scene distribution characteristics and the current demand data into a pre-trained scene adaptation model, and output the application scene corresponding to the current demand data, wherein the scene adaptation model is used to extract semantic features in the current demand data to obtain a semantic vector, and fuse the semantic vector with the application scene distribution characteristics to predict the output; A generation module is used to generate target query knowledge based on the application scenario and the current demand data, and generate a query statement according to the target query knowledge.

8. A query statement generating device, characterized in that: The device comprises: a memory, a processor, and a query statement generation program stored in the memory and executable on the processor, wherein the query statement generation program is configured to implement the steps of the query statement generation method according to any one of claims 1 to 6.

9. A readable storage medium, characterized in that: The readable storage medium comprises a computer-readable storage medium, on which a query statement generation program is stored. When the query statement generation program is executed by a processor, the steps of the query statement generation method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a query statement generation program, and when the query statement generation program is executed by a processor, the steps of the query statement generation method according to any one of claims 1 to 6 are implemented.