Query statement generation method and device and computer equipment

By converting user problems and database table information into embedding vectors and selecting target embedding vectors using the minimum angle regression algorithm to generate query statements, the problem of low accuracy in generating query statements in the prior art is solved, and higher accuracy and efficiency are achieved.

CN120371850APending Publication Date: 2025-07-25CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510444083.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, the accuracy of generating query statements is low, especially when there are many database tables. The length limit of input text of large models and the consumption of computing resources leads to poor recall and sorting, high rule dependence and low accuracy of entity extraction.

Method used

By converting user problems into the first sentence embedding vector, preset database table information into the second sentence embedding vector, the minimum angle regression algorithm is used to select the target number of embedding vectors from the most similar embedding vectors to generate a query statement.

Benefits of technology

The accuracy of query statement generation is improved, and the table recall and sorting is optimized through the quadratic filtering of matching semantics and minimum angle regression algorithms, and the redundant information is reduced, which improves the accuracy and efficiency of query statements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371850A_ABST
    Figure CN120371850A_ABST
Patent Text Reader

Abstract

The invention discloses a query statement generation method and device and computer equipment. The method comprises the following steps: receiving a user question, and converting the user question into a first sentence embedding vector; converting table information in a preset database into a second sentence embedding vector; sequentially performing similarity calculation on the first sentence embedding vector and second sentence embedding vectors in a preset database, selecting a plurality of second sentence embedding vectors with the highest similarity from the plurality of second sentence embedding vectors, and determining the selected second sentence embedding vectors as screened embedding vectors; selecting a target number of target embedding vectors from the plurality of screened embedding vectors by adopting a minimum angle regression algorithm; and generating a query statement corresponding to the user question according to the target embedded vector. The method provided by the invention at least solves the technical problem that the accuracy of the query statement generated in the related technology is relatively low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and more particularly, to a method, apparatus, and computer device for generating query statements. Background Art

[0002] Currently, in the related art, the main way to convert a user's question into a query statement is through a large model. For example, NL2sql (Natural Language to Structured Query Language, converting natural language into a query statement). However, when there are many tables in the database, it is impossible to provide the schema (table information) of all tables in the database to the large model because the input text length of the large model is limited. Although there are many techniques that allow the input text of the large model to be very long, this is accompanied by a doubling of computing resource consumption. Also, some studies have pointed out that an overly long input text information will affect the output ability of the large model because the signal-to-noise ratio of the input will decrease as the amount of invalid text input increases. Therefore, it is also necessary to streamline the input of the large model, that is, it is necessary to recall, sort, and filter the tables according to the query statement, and use the filtered table information as the input of the large model regarding the schema. At the same time, if the recalled tables are not complete enough, the large model will not be able to generate the correct SQL due to the lack of necessary table information. Therefore, in the case of having a large model, the effects of table recall and sorting will greatly affect the effect of nl2sql. One way is rule-based, and the other way is to first extract entities from the query statement and then use the extracted entities to perform table recall. However, the above methods highly depend on the setting of rules and the correctness of entity extraction. In the case where the rule setting is incomplete or the entity extraction accuracy is low, the accuracy of the generated query statement is also low. Summary of the Invention

[0003] Embodiments of this application provide a method, apparatus, and computer device for generating query statements to at least solve the technical problem of low accuracy of the generated query statements in the related art.

[0004] According to one aspect of the embodiments of this application, a method for generating a query statement is provided, including: receiving a user's question and converting the user's question into a first sentence embedding vector; converting the table information in a preset database into a second sentence embedding vector; calculating the similarity between the first sentence embedding vector and the second sentence embedding vectors in the preset database in sequence, selecting the multiple second sentence embedding vectors with the highest similarity from the multiple second sentence embedding vectors, and determining them as the filtered embedding vectors; using the least angle regression algorithm to select a target number of target embedding vectors from the multiple filtered embedding vectors; and generating a query statement corresponding to the user's question according to the target embedding vectors.

[0005] Optionally, the minimum angle regression algorithm is used to select a target number of target embedding vectors from the multiple filtered embedding vectors, including: determining the multiple filtered embedding vectors as feature vectors, and determining the first sentence embedding vector as the target vector; obtaining an initial residual vector, where the initial residual vector is equal to the target vector; updating the initial residual vector to obtain multiple residual vectors, where the multiple residual vectors include: the initial residual vector and the updated residual vector; sequentially selecting the feature vectors most similar to the multiple residual vectors from the multiple feature vectors, and determining them as the target embedding vectors.

[0006] Optionally, updating the initial residual vector to obtain multiple residual vectors includes: obtaining the initial residual vector; obtaining two feature vectors most similar to the initial residual vector among the multiple feature vectors, which are respectively: the first vector and the second vector, where the similarity between the first vector and the initial residual vector is higher than the similarity between the second vector and the initial residual vector; moving the initial residual vector in the direction of the first vector by a preset step length to obtain the updated residual vector, where the direction of the updated residual vector is the same as the direction of the angular bisector of the first vector and the second vector.

[0007] Optionally, the method further includes: obtaining a current residual vector, where the current residual vector is not the initial residual vector; obtaining a remaining feature vector set, where the remaining feature vector set includes: the feature vectors among the multiple feature vectors except for the selected target embedding vectors; selecting the feature vector with the highest similarity to the current residual vector from the remaining feature vector set as the third vector; moving the current residual vector in the direction of the current residual vector by a preset step length to obtain the updated residual vector, where the similarity between the updated residual vector and the selected target embedding vector is equal to the similarity between the updated residual vector and the third vector.

[0008] Optionally, converting the table information in the preset database into a second sentence embedding vector includes: obtaining the table name, table comment, field name, and field comment in the table information; concatenating the table name, the table comment, the field name, and the field comment into a sentence using a preset character; converting the sentence into the second sentence embedding vector using a preset sentence embedding model.

[0009] Optionally, generating a query statement corresponding to the user question according to the target embedding vector includes: obtaining a preset prompt word, where the preset prompt word is determined according to the target embedding vector and the user question; analyzing the preset prompt word using a preset large model to obtain the query statement corresponding to the user question.

[0010] Optionally, obtain a preset prompt word, including: mapping the target embedding vector to target table information, and constructing a table information prompt word according to the target table information, where the table information prompt word at least includes: a selected database table and keyword fields included in the selected database table; constructing a user question prompt word according to the user question; constructing the preset prompt word from the table information prompt word and the user question prompt word.

[0011] According to another aspect of the embodiments of the present application, there is also provided a query statement generation device, including: a receiving module, configured to receive a user question and convert the user question into a first sentence embedding vector; a conversion module, configured to convert table information in a preset database into a second sentence embedding vector; a first screening module, configured to calculate the similarity between the first sentence embedding vector and the second sentence embedding vectors in the preset database in sequence, select multiple second sentence embedding vectors with the highest similarity from the multiple second sentence embedding vectors, and determine them as the screened embedding vectors; a second screening module, configured to select a target number of target embedding vectors from the multiple screened embedding vectors by using the least angle regression algorithm; a generation module, configured to generate a query statement corresponding to the user question according to the target embedding vector.

[0012] According to still another aspect of the embodiments of the present application, there is also provided a computer device, including: a memory and a processor, where the memory is used to store program instructions; the processor is connected to the memory and is configured to execute the above query statement generation method.

[0013] According to still another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium, where the non-volatile storage medium includes a stored computer program, and the device where the non-volatile storage medium is located executes the above query statement generation method by running the computer program.

[0014] According to still another aspect of the embodiments of the present application, there is also provided a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the above query statement generation method is implemented.

[0015] In the embodiments of the present application, a user question is received, and the user question is converted into a first sentence embedding vector; the table information in a preset database is converted into a second sentence embedding vector; the first sentence embedding vector is sequentially calculated for similarity with the second sentence embedding vectors in the preset database, and multiple second sentence embedding vectors with the highest similarity are selected from the multiple second sentence embedding vectors and determined as the screened embedding vectors; a least angle regression algorithm is used to select a target number of target embedding vectors from the multiple screened embedding vectors; a query statement corresponding to the user question is generated according to the target embedding vectors, thereby achieving the purpose of secondary screening through the least angle regression algorithm after screening by matching semantics, thus realizing the technical effect of improving the accuracy of generating query statements, and further solving the technical problem of low accuracy of the generated query statements in the related art. Brief Description of the Drawings

[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0017] Figure 1 is a hardware structure block diagram of a computer terminal for implementing a method for generating a query statement according to an embodiment of the present application;

[0018] Figure 2 is a flowchart of a method for generating a query statement according to an embodiment of the present application;

[0019] Figure 3 is a flowchart of another method for generating a query statement according to an embodiment of the present application;

[0020] Figure 4 is a structural diagram of a device for generating a query statement according to an embodiment of the present application. Detailed Description of the Embodiments

[0021] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0022] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0023] The information collected in the embodiments of this application is information and data authorized by the user or fully authorized by all parties. And the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure and application, etc., all comply with the relevant laws, regulations and standards of the relevant region, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or reject the results of automated decision-making; if the user chooses to reject, it enters the expert decision-making process.

[0024] To solve the problems existing in the related art, the embodiments of this application provide a method for generating a query statement, and this method can run in Figure 1 the computer terminal shown below, and the following explains this computer terminal.

[0025] The method embodiments for generating query statements provided by the embodiments of this application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The following shows a hardware structure block diagram of a computer terminal for implementing the method for generating a query statement. As Figure 1 shown, the computer terminal 10 may include one or more processors (the processors may include, but are not limited to, processing devices such as a microprocessor MCU or a field programmable gate array FPGA, shown as 102a, 102b,..., 102n in the figure), a memory 104 for storing data, and a transmission module 106 for communication functions connected by wired and / or wireless networks. In addition, it may further include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, and a BUS bus. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown in, or have the same asFigure 1 The different configurations shown.

[0026] It should be noted that one or more of the above processors and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or incorporated in whole or in part into any one of the other elements in the computer terminal 10. As involved in the embodiments of the present application, the data processing circuit is a processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0027] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for generating query statements in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above method for generating query statements. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely set relative to the processor, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.

[0028] The transmission module 106 is used to receive or send data via a network. Specific examples of the above network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission module 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission module 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0029] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10.

[0030] It should be noted here that in some alternative embodiments, the above Figure 1 shown computer terminal can include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1This is only an example of a specific concrete instance and is intended to illustrate the types of components that may exist in the above computer terminal.

[0031] Under the above operating environment, an embodiment of a method for generating a query statement is provided in an embodiment of the present application. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And, although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0032] Figure 2 is a flowchart of a method for generating a query statement according to an embodiment of the present application, as Figure 2 shown, the method includes the following steps:

[0033] Step S202, receive a user question and convert the user question into a first sentence embedding vector;

[0034] In step S202, the way to convert the user question into a first sentence embedding vector can be to use a preset sentence embedding model. The user question can be querying data, creating a table, transforming data, etc.

[0035] Step S204, convert the table information in the preset database into a second sentence embedding vector;

[0036] In step S204, the table information, for example: table schema. That is, in the field of databases, Schema refers to the definition of the organization and structure of data in a database. It includes the definitions of all tables in the database, the names and types of fields, and the relationships between data (such as primary key and foreign key constraints).

[0037] Step S206, calculate the similarity between the first sentence embedding vector and the second sentence embedding vectors in the preset database in sequence, and select the second sentence embedding vectors with the highest similarity from the multiple second sentence embedding vectors, and determine them as the filtered embedding vectors;

[0038] Step S208, use the least angle regression algorithm to select a target number of target embedding vectors from the multiple filtered embedding vectors;

[0039] Step S210, generate a query statement corresponding to the user question according to the target embedding vector.

[0040] Through the above steps S202 to S210, receive the user's question, convert the user's question into a first sentence embedding vector; convert the table information in the preset database into a second sentence embedding vector; calculate the similarity between the first sentence embedding vector and the second sentence embedding vectors in the preset database in sequence, select multiple second sentence embedding vectors with the highest similarity from the multiple second sentence embedding vectors, and determine them as the screened embedding vectors; use the least angle regression algorithm to select a target number of target embedding vectors from the multiple screened embedding vectors; generate a query statement corresponding to the user's question according to the target embedding vector, thereby achieving the purpose of secondary screening through the least angle regression algorithm after screening by matching semantics, thus realizing the technical effect of improving the accuracy of generating query statements, and further solving the technical problem of low accuracy of the query statements generated in the related art. The following is a detailed description.

[0041] It should be noted that when there are a large number of tables in the database, there will be duplicate information in the content of some tables. For example: when the user's question is "Which enterprises have not been financed? Give the enterprise name and legal representative." According to text matching and vector matching, the following tables will be recalled according to the matching score: Enterprise (Chinese name, legal representative), Enterprise Finance (enterprise name, legal representative), Enterprise Basic Information (enterprise name, legal representative), Enterprise Information Summary Table (name, legal representative), Enterprise Business Information Table (name, legal representative), Enterprise Summary Statistics Table (name, legal representative), Enterprise Information Details Table (name, legal representative), Enterprise Financing, Enterprise Investment. It can be seen that since the weights of name and legal representative are relatively high, if the recalled tables have these two fields, they will always rank at the forefront of the sorted recalled tables. From the above example, it can be seen that there are already tables containing the two fields required by the demand. Although the subsequent tables containing these two fields have a relatively high relevance to the user's question, they do not provide more other information, resulting in duplicate and redundant recalled tables, and further causing too much duplicate information to be input into the large model.

[0042] To solve the problems existing in the related art, in the implementation of this application, the least angle regression algorithm is used to select a target number of target embedding vectors from multiple screened embedding vectors. The specific steps include: determining multiple screened embedding vectors as feature vectors, and determining the first sentence embedding vector as the target vector; obtaining an initial residual vector, where the initial residual vector is equal to the target vector; updating the initial residual vector to obtain multiple residual vectors, where the multiple residual vectors include: the initial residual vector and the updated residual vector; sequentially selecting the feature vectors most similar to the multiple residual vectors from multiple feature vectors and determining them as the target embedding vectors. Through the above method, the feature vectors most relevant to the target vector Y are gradually selected, and the correlation between features is considered, avoiding the selection of redundant features, ensuring that the finally selected feature vectors can maximize the explanation of the target vector Y with the least redundant information, thereby improving the efficiency and quality of feature selection.

[0043] Among them, the process of updating the initial residual vector to obtain multiple residual vectors is as follows: obtaining the initial residual vector; obtaining two feature vectors most similar to the initial residual vector among multiple feature vectors, which are respectively: the first vector and the second vector, where the similarity between the first vector and the initial residual vector is higher than the similarity between the second vector and the initial residual vector; moving the initial residual vector in the direction of the first vector by a preset step length to obtain the updated residual vector, where the direction of the updated residual vector is the same as the angular bisector direction of the first vector and the second vector.

[0044] In the case where the current residual vector is not the initial residual vector, obtaining the current residual vector, where the current residual vector is not the initial residual vector; obtaining a remaining feature vector set, where the remaining feature vector set includes: the feature vectors among multiple feature vectors except the selected target embedding vectors; selecting the feature vector with the highest similarity to the current residual vector from the remaining feature vector set and determining it as the third vector; moving the current residual vector in the direction of the current residual vector by a preset step length to obtain the updated residual vector, where the similarity between the updated residual vector and the selected target embedding vector is equal to the similarity between the updated residual vector and the third vector.

[0045] Specifically, in step 1, when the current residual vector is the initial residual vector, select a direction of a feature vector xi that has the highest correlation (e.g., the smallest cosine angle) with the initial residual vector y(0) from multiple feature vectors. Move a preset step size θi in the direction of xi such that the residual vector at this time, y(0) - θixi (the updated residual vector relative to the initial residual vector), has the same correlation with the feature vector xi (the first vector) and another feature vector xj with the highest correlation (the second vector) (or, in other words, such that y(0) - θixi is exactly on the angular bisector of xi and xj).

[0046] In step 2, move a preset step size θi in the direction of the above angular bisector (which is also the current residual vector direction) as the new feature vector search direction, such that the correlation between the residual vector y(0) - ∑θixi (the updated residual vector relative to the current residual vector) and each previously selected feature vector is equal to the correlation with the feature vector with the highest correlation (the third vector) in the remaining feature set (such that y(0) - ∑θixi is on the spatial angular bisector of the above feature vectors).

[0047] Repeat step 2 until the target number of target embedding vectors is selected.

[0048] After each iteration, make the correlation between the residual vector y(0) - ∑θixi and all the selected feature vectors equal to the correlation with the feature vector with the highest correlation in the remaining feature set. The algorithm adjusts θi to ensure that the "correlation" of the residual vector with all the feature vectors in the current model is comparable, thus fairly handling all features and avoiding giving overly high weights to some features too early due to their particularly high correlation. In each iteration, select the feature vector that is "most relevant" to the current residual vector, and then update the residual vector and weights along the angular bisector direction with an appropriate step size θi until the stopping condition is met, such as the magnitude of the residual vector being lower than a certain threshold or all features having been considered. This method not only helps with feature selection but also effectively handles multicollinearity, avoids unnecessary fluctuations in feature weights, and thus obtains a stable and interpretable regression model.

[0049] Specifically in the embodiment of this solution, through the above method, the most representative and informative tables can be selected in a systematic manner from multiple potentially relevant tables, while taking into account the possible redundant information between the tables, thereby optimizing the table recall and sorting in the query statement generation process and improving the accuracy and efficiency of the query.

[0050] In some embodiments of the present application, the specific steps of converting the table information in the preset database into the second sentence embedding vector are as follows: Obtain the table name, table comment, field name, and field comment in the table information; Concatenate the table name, the table comment, the field name, and the field comment into a sentence using a preset character; Use a preset sentence embedding model to convert the sentence into the second sentence embedding vector.

[0051] For example: A preset database contains a table named Employee Information, and its related information is as follows:

[0052] Table name: Employee Information;

[0053] Table comment: Stores the basic information of company employees, including employees' personal profiles and work-related data;

[0054] Field names: Employee ID, Name, Department, Hire Date, Salary;

[0055] Field comments: Employee ID (a number that uniquely identifies an employee), Name (the full name of the employee), Department (the company department to which the employee belongs), Hire Date (the specific date when the employee joined the company), Salary;

[0056] Concatenate the key data into a sentence using a preset character:

[0057] Select a preset character, such as a comma, to concatenate the above table name, table comment, field name, and field comment into a coherent sentence.

[0058] In an alternative manner, the concatenated sentence may be: "Employee Information, Employee ID, Name, Department, Hire Date, Salary".

[0059] In some embodiments of the present application, generating the query statement corresponding to the user question according to the target embedding vector includes: Obtaining a preset prompt word, where the preset prompt word is determined according to the target embedding vector and the user question; Using a preset large model to analyze the preset prompt word to obtain the query statement corresponding to the user question.

[0060] Among them, the specific steps of obtaining the preset prompt word are as follows: Map the target embedding vector to target table information, and construct a table information prompt word according to the target table information. The table information prompt word includes at least: the selected database table and the key fields included in the selected database table; Construct a user question prompt word according to the user question; Construct the preset prompt word from the table information prompt word and the user question prompt word.

[0061] An alternative preset prompt word is as follows:

[0062] Q: "Which companies have never been financed? Give the names and legal representatives."

[0063] CREATE TABLE (Create Table) Enterprises (

[0064] Entry ID text,

[0065] Chinese name double,

[0066] Establishment time text,

[0067] Legal representative time,

[0068] Province of affiliation text,

[0069] Registered capital text,

[0070] PRIMARY KEY (Primary Key) (Entry ID) );

[0072] CREATE TABLE Enterprise Financing (

[0073] Enterprise ID double,

[0074] Financing round double,

[0075] Total financing amount text,

[0076] Year double,

[0077] FOREIGN KEY (Enterprise ID) REFERENCES Enterprises (Entry ID) );

[0079] CREATE TABLE Investment Company (

[0080] Enterprise ID time,

[0081] Investment company double,

[0082] Financing round text,

[0083] Financing amount text,

[0084] Investment company shareholding ratio double,

[0085] FOREIGN KEY (Enterprise ID) REFERENCES Enterprises (Entry ID));

[0086] The generated query statement is as follows:

[0087] 'select T1. Chinese name, T1. Legal representative from enterprise as T1 where T1. Entry id not in (select enterprise id from enterprise financing as T2)' means: Select the 'Chinese name' and 'Legal representative' of all enterprises from the 'enterprise' table, but only include those enterprises whose 'Entry id' is not in the 'enterprise id' column of the 'enterprise financing' table.

[0088] Taking the above user question as an example, the update process of the residual vector r is as follows: Select the most relevant features: In the first step, the algorithm will find the feature vector X1 that has the highest correlation with r (the current residual vector). For example, X1 may be the vector representation of the 'enterprise' table because the fields 'name' and 'legal person' in 'enterprise' are most relevant to the query.

[0089] Update the residual vector: Next, the algorithm will move a step size θ1 in the direction of X1 to reduce the angle between r and X1 until the angle between r and another feature vector X2 is equal. This usually means that r already contains all the information that X1 can provide, and further moving in the direction of X1 will not significantly reduce the residual. Introduce the second feature: At this time, X2 is added to the regression model, and r will move along the angle bisector direction of X1 and X2, and update the step sizes θ1 and θ2 again until the angle between r and the third feature X3 is equal. This ensures that while introducing X2, the influence of X1 still exists, and r is optimally updated in both directions.

[0090] Repeat the process: The algorithm continues to repeat the above steps, introducing a new feature vector each time and updating the residual vector r until the termination conditions are met. These termination conditions may include that all feature vectors have been considered, the residual vector r is small enough, or the model has introduced enough features (for example: 5 table embedding vectors).

[0091] Filter the feature vectors: Finally, the algorithm will filter out those feature vectors whose weights are not 0 during the regression process, that is, retain the table vectors that are most explanatory for Y (the original query vector). In our example, these may be the tables that are most relevant to financing and enterprise information.

[0092] Figure 3 Another method for generating a query statement is shown, as Figure 3 shown:

[0093] s1. Embedding vector of the table schema: Concatenate the table name, table comment, field name, and field comment into a sentence with ','. Then vectorize the sentence through a sentence embedding model to obtain sentence embeddings, and one table corresponds to one sentence vector to get the embedding vector of the table information;

[0094] s2. Write the table information and sentence vector information into a database that supports vector matching;

[0095] s3. Also vectorize the user's query through the same sentence embedding model;

[0096] s4. Through vector matching technology, find the top 30 with the highest similarity between the query sentence vector and the database;

[0097] s5. Use these 30 vectors as X and the query sentence vector as Y. Perform LARS least angle regression and select 5 vectors with non-zero weights; - Innovation;

[0098] s6. Use the tables corresponding to the selected vectors as the tables for recall ranking. Construct a prompt based on the table information and the user's question;

[0099] s7. Use the prompt as input to the large model, obtain the result returned by the large model, and extract the generated query statement;

[0100] s8. Clean the results and extract the query statements.

[0101] Figure 4 It is a query statement generation device according to an embodiment of the present application. The device includes:

[0102] A receiving module 40 for receiving a user question and converting the user question into a first sentence embedding vector;

[0103] A conversion module 42 for converting the table information in the preset database into a second sentence embedding vector;

[0104] A first screening module 44 for calculating the similarity between the first sentence embedding vector and the second sentence embedding vectors in the preset database in sequence, and selecting multiple second sentence embedding vectors with the highest similarity from the multiple second sentence embedding vectors to be determined as the screened embedding vectors;

[0105] A second screening module 46 for selecting a target number of target embedding vectors from the multiple screened embedding vectors by using the least angle regression algorithm;

[0106] A generation module 48 for generating a query statement corresponding to the user question according to the target embedding vector.

[0107] Through the generation device of the above query statement, the user question is received, and the user question is converted into a first sentence embedding vector; the table information in the preset database is converted into a second sentence embedding vector; the similarity between the first sentence embedding vector and the second sentence embedding vectors in the preset database is calculated in sequence, and multiple second sentence embedding vectors with the highest similarity are selected from the multiple second sentence embedding vectors and determined as the filtered embedding vectors; the least angle regression algorithm is used to select a target number of target embedding vectors from the multiple filtered embedding vectors; a query statement corresponding to the user question is generated according to the target embedding vectors, thereby achieving the purpose of secondary screening through the least angle regression algorithm after screening by matching semantics, thus realizing the technical effect of improving the accuracy of generating query statements, and further solving the technical problem of low accuracy of the generated query statements in the related art.

[0108] The second screening module 46 includes: a selection sub-module for using the least angle regression algorithm to select a target number of target embedding vectors from the multiple filtered embedding vectors, including: determining the multiple filtered embedding vectors as feature vectors, and determining the first sentence embedding vector as the target vector; obtaining an initial residual vector, where the initial residual vector is equal to the target vector; updating the initial residual vector to obtain multiple residual vectors, where the multiple residual vectors include: the initial residual vector and the updated residual vector; sequentially selecting the feature vectors most similar to the multiple residual vectors from the multiple feature vectors and determining them as the target embedding vectors.

[0109] The selection sub-module includes: a first update unit and a second update unit, where the first update unit is used to update the initial residual vector to obtain multiple residual vectors, including: obtaining the initial residual vector; obtaining two feature vectors most similar to the initial residual vector among the multiple feature vectors, which are respectively: a first vector and a second vector, where the similarity between the first vector and the initial residual vector is higher than the similarity between the second vector and the initial residual vector; moving the initial residual vector in the direction of the first vector by a preset step length to obtain the updated residual vector, where the direction of the updated residual vector is the same as the direction of the angular bisector of the first vector and the second vector.

[0110] A second update unit is configured to obtain a current residual vector, where the current residual vector is not the initial residual vector; obtain a set of remaining feature vectors, where the set of remaining feature vectors includes: feature vectors among the multiple feature vectors except the selected target embedding vector; select a feature vector with the highest similarity to the current residual vector from the set of remaining feature vectors as the third vector; move the current residual vector in the direction of the current residual vector by a preset step length to obtain the updated residual vector, where the similarity between the updated residual vector and the selected target embedding vector is equal to the similarity between the updated residual vector and the third vector.

[0111] The transformation module 42 includes: a transformation sub-module for transforming the table information in a preset database into a second sentence embedding vector, including: obtaining the table name, table annotation, field name, and field annotation in the table information; concatenating the table name, the table annotation, the field name, and the field annotation into a sentence using a preset character; and transforming the sentence into the second sentence embedding vector using a preset sentence embedding model.

[0112] The generation module 48 includes: a generation sub-module for generating a query statement corresponding to the user question according to the target embedding vector, including: obtaining a preset prompt word, where the preset prompt word is determined according to the target embedding vector and the user question; and analyzing the preset prompt word using a preset large model to obtain the query statement corresponding to the user question.

[0113] The generation sub-module includes: an obtaining unit for obtaining a preset prompt word, including: mapping the target embedding vector to target table information, and constructing a table information prompt word according to the target table information, where the table information prompt word at least includes: the selected database table and the key fields included in the selected database table; constructing a user question prompt word according to the user question; and constructing the preset prompt word from the table information prompt word and the user question prompt word.

[0114] It should be noted that Figure 4 the query statement generation device shown is used to execute Figure 2 the query statement generation method shown. Therefore, the relevant explanations in the above query statement generation method also apply to this query statement generation device, and will not be elaborated here.

[0115] An embodiment of the present application further provides a computer device, including: a memory and a processor, where the memory is used to store program instructions; the processor is connected to the memory and is used to execute the above query statement generation method.

[0116] Through the method for generating a query statement executed by the above computer device, a user question is received, and the user question is converted into a first sentence embedding vector; the table information in the preset database is converted into a second sentence embedding vector; the first sentence embedding vector and the second sentence embedding vectors in the preset database are sequentially subjected to a similarity calculation, and multiple second sentence embedding vectors with the highest similarity are selected from the multiple second sentence embedding vectors and determined as the filtered embedding vectors; the least angle regression algorithm is used to select a target number of target embedding vectors from the multiple filtered embedding vectors; and a query statement corresponding to the user question is generated according to the target embedding vectors, thereby achieving the purpose of secondary screening through the least angle regression algorithm after screening by matching semantics, thus realizing the technical effect of improving the accuracy of generating the query statement, and further solving the technical problem of low accuracy of the generated query statement in the related art.

[0117] An embodiment of the present application further provides a non-volatile storage medium, which includes a stored computer program. Wherein, the device where the non-volatile storage medium is located executes the above method for generating a query statement by running the computer program.

[0118] Through the method for generating a query statement stored in the above non-volatile storage medium, a user question is received, and the user question is converted into a first sentence embedding vector; the table information in the preset database is converted into a second sentence embedding vector; the first sentence embedding vector and the second sentence embedding vectors in the preset database are sequentially subjected to a similarity calculation, and multiple second sentence embedding vectors with the highest similarity are selected from the multiple second sentence embedding vectors and determined as the filtered embedding vectors; the least angle regression algorithm is used to select a target number of target embedding vectors from the multiple filtered embedding vectors; and a query statement corresponding to the user question is generated according to the target embedding vectors, thereby achieving the purpose of secondary screening through the least angle regression algorithm after screening by matching semantics, thus realizing the technical effect of improving the accuracy of generating the query statement, and further solving the technical problem of low accuracy of the generated query statement in the related art.

[0119] An embodiment of the present application further provides a computer program product, which includes computer instructions. When the computer instructions are executed by a processor, the steps of the method for generating a query statement in the present application are implemented.

[0120] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0121] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0122] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0123] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0124] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0125] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks or optical discs that can store program codes.

[0126] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for generating a query statement, characterized in that, including: receiving a user question and converting the user question into a first sentence embedding vector; converting the table information in a preset database into a second sentence embedding vector; calculating the similarity between the first sentence embedding vector and the second sentence embedding vectors in the preset database in sequence, selecting multiple second sentence embedding vectors with the highest similarity from the multiple second sentence embedding vectors, and determining them as the filtered embedding vectors; selecting a target number of target embedding vectors from the multiple filtered embedding vectors by using the least angle regression algorithm; generating a query statement corresponding to the user question according to the target embedding vectors.

2. The method according to claim 1, wherein selecting a target number of target embedding vectors from the multiple filtered embedding vectors by using the least angle regression algorithm, including: determining the multiple filtered embedding vectors as feature vectors, and determining the first sentence embedding vector as a target vector; obtaining an initial residual vector, where the initial residual vector is equal to the target vector; updating the initial residual vector to obtain multiple residual vectors, where the multiple residual vectors include: the initial residual vector and the updated residual vector; sequentially selecting the feature vectors most similar to the multiple residual vectors from the multiple feature vectors, and determining them as the target embedding vectors.

3. The method according to claim 2, wherein updating the initial residual vector to obtain multiple residual vectors, including: obtaining the initial residual vector; obtaining two feature vectors most similar to the initial residual vector among the multiple feature vectors, which are respectively: a first vector and a second vector, where the similarity between the first vector and the initial residual vector is higher than the similarity between the second vector and the initial residual vector; moving the initial residual vector in the direction of the first vector by a preset step length to obtain the updated residual vector, where the direction of the updated residual vector is the same as the direction of the angular bisector of the first vector and the second vector.

4. The method according to claim 2, wherein The method further includes: obtaining a current residual vector, where the current residual vector is not the initial residual vector; obtaining a remaining feature vector set, where the remaining feature vector set includes: the feature vectors in the multiple feature vectors except the selected target embedding vectors; selecting a feature vector with the highest similarity to the current residual vector from the remaining feature vector set and determining it as a third vector; moving the current residual vector in the direction of the current residual vector by a preset step length to obtain the updated residual vector, where the similarity between the updated residual vector and the selected target embedding vectors is equal to the similarity between the updated residual vector and the third vector.

5. The method according to claim 1, wherein converting the table information in a preset database into a second sentence embedding vector, including: obtaining the table name, table comment, field name, and field comment in the table information; concatenating the table name, the table comment, the field name, and the field comment into a sentence by using a preset character; converting the sentence into the second sentence embedding vector by using a preset sentence embedding model.

6. The method according to claim 1, characterized in that, generating a query statement corresponding to the user question according to the target embedding vectors, including: Obtain a preset prompt word, where the preset prompt word is determined according to the target embedding vector and the user question; Analyze the preset prompt word using a preset large model to obtain a query statement corresponding to the user question.

7. The method according to claim 6, wherein Obtaining a preset prompt word includes: Map the target embedding vector to target table information, and construct a table information prompt word according to the target table information. The table information prompt word at least includes: a selected database table and keyword fields included in the selected database table; Construct a user question prompt word according to the user question; Construct the preset prompt word from the table information prompt word and the user question prompt word.

8. A generating device for query statements, characterized in that It includes: A receiving module, configured to receive a user question and convert the user question into a first sentence embedding vector; A conversion module, configured to convert table information in a preset database into a second sentence embedding vector; A first screening module, configured to calculate the similarity between the first sentence embedding vector and the second sentence embedding vectors in the preset database in sequence, select multiple second sentence embedding vectors with the highest similarity from the multiple second sentence embedding vectors, and determine them as the screened embedding vectors; A second screening module, configured to select a target number of target embedding vectors from the multiple screened embedding vectors using the least angle regression algorithm; A generation module, configured to generate a query statement corresponding to the user question according to the target embedding vector.

9. A computer device, characterized in that, It includes: A memory and a processor, where the memory is used to store program instructions; The processor, connected to the memory, is configured to execute the method for generating a query statement according to any one of claims 1 to 7.

10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, the method for generating a query statement according to any one of claims 1 to 7 is implemented.