Method and related device for generating database statements based on large language models

By generating multiple variant statements based on a large language model and obtaining the target data structure from the database, the problem of small coverage of schema in the prior art is solved, and the accuracy and confidence of database statements are improved.

CN119201979BActive Publication Date: 2025-06-17SICHUAN WUJI SMART TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411534562.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-06-17
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

In the prior art, when searching schemas from the database through natural statements entered by users, there is a problem that the schema coverage is small, resulting in poor accuracy of generated database statements.

Method used

Using a method based on a large language model, a more accurate database statement is generated by obtaining the pending statements input by the user, and a multiple corresponding variant statements are obtained from the database based on these variant statements, thereby generating more accurate database statements.

Benefits of technology

By expanding the coverage of the retrieved data structures, the accuracy of the generated database statements is improved, ensuring that the target database statements can be successfully executed and have the highest confidence level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119201979B_ABST
    Figure CN119201979B_ABST
Patent Text Reader

Abstract

An embodiment of this application proposes a method and related device for generating database statements based on a large language model, which relates to the technical field of data processing. Obtain the statement to be processed input by the user, and generate multiple variant statements corresponding to the statement to be processed through the large language model; both the statement to be processed and the variant statements are natural languages, and the variant statements have different syntactic structures and the same semantic logic as the statement to be processed; obtain the target data structure from the database through the large language model according to the multiple variant statements, and generate database statements corresponding to the respective variant statements according to each variant statement and the target data structure; determine the target database statement from the multiple database statements; the target database statement is the database statement that can be successfully executed and has the highest confidence level, so the coverage range of the retrieved data structure can be expanded, and the accuracy of the finally generated target database statement can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing. Specifically, it relates to a method for generating database statements based on a large language model and related devices. Background Art

[0002] With the development of natural language processing technology, computers can better understand and process natural language. Therefore, to improve the user experience, users are allowed to directly input natural statements when they want to perform relevant operations on a database, and the computer generates corresponding database statements according to the natural statements input by the user to achieve operations on the database, such as data query.

[0003] Currently, a computer can directly obtain relevant data structures (schemas) from a database through natural statements input by a user, and then generate corresponding database statements based on the retrieved schemas. However, this method also has certain limitations.

[0004] Specifically, the method of directly retrieving schemas from a database through natural statements input by a user may have the problem of a relatively small coverage range of schemas, resulting in poor accuracy of the generated database statements. Summary of the Invention

[0005] In view of this, the purpose of this application is to provide a method for generating database statements based on a large language model and related devices to expand the coverage range of retrieved data structures, thereby improving the accuracy of the generated database statements.

[0006] To achieve the above purpose, the technical solutions adopted in the embodiments of this application are as follows:

[0007] In a first aspect, the present invention provides a method for generating database statements based on a large language model, the method including:

[0008] Obtain a to-be-processed statement input by a user, and generate multiple variant statements corresponding to the to-be-processed statement through a large language model; both the to-be-processed statement and the variant statements are natural languages, and the variant statements have different syntactic structures but the same semantic logic as the to-be-processed statement;

[0009] Obtain a target data structure from the database through the large language model according to the multiple variant statements, and generate database statements corresponding to the variant statements according to each variant statement and the target data structure;

[0010] Determine a target database statement from the multiple database statements; the target database statement is the database statement that can be successfully executed and has the highest confidence level.

[0011] In an alternative embodiment, obtaining, by the large language model, a target data structure from a database according to a plurality of the variant statements, and generating database statements corresponding to the respective variant statements according to each of the variant statements and the target data structure, includes:

[0012] Inputting each of the variant statements into the large language model for processing to obtain a first semantic structure corresponding to each of the variant statements;

[0013] Obtaining, from the database, a target table and target fields corresponding to each of the variant statements according to the first semantic structure corresponding to each of the variant statements, taking the union of the target table and the target fields to obtain the target data structure;

[0014] Inputting the target data structure and each of the variant statements into the large language model for processing to obtain database statements corresponding to the respective variant statements.

[0015] In an alternative embodiment, the method further includes:

[0016] Inputting the statement to be processed into the large language model for processing to obtain a second semantic structure corresponding to the statement to be processed, and determining whether a target first semantic structure semantically identical to the second semantic structure meets a first preset condition;

[0017] If the target first semantic structure does not meet the first preset condition, adding a negative label to other semantic structures except the target first semantic structure, and re-inputting the other first semantic structures with the added negative label and the variant statements corresponding to the other first semantic structures into the large language model for processing to obtain new first semantic structures corresponding to the variant statements, and determining a target first semantic structure from the new first semantic structures until the target first semantic structure meets the first preset condition.

[0018] In an alternative embodiment, the obtaining, from the database, a target table and target fields corresponding to each of the variant statements according to the first semantic structure corresponding to each of the variant statements, taking the union of the target table and the target fields to obtain the target data structure, includes:

[0019] If the target first semantic structure meets the first preset condition, obtaining, from the database, a target table and target fields corresponding to the variant statements corresponding to the target first semantic structure according to each of the target first semantic structures;

[0020] Taking the union of the target table and the target fields to obtain the target data structure;

[0021] Inputting the target data structure and each of the variant statements into the large language model for processing to obtain database statements corresponding to each of the variant statements, includes:

[0022] Inputting the target data structure and variant statements corresponding to each of the target first semantic structures into the large language model for processing to obtain database statements corresponding to the variant statements corresponding to each of the target first semantic structures.

[0023] In an alternative embodiment, determining the target database statement from the multiple database statements includes:

[0024] Executing each of the database statements respectively, and determining multiple first candidate database statements from the multiple database statements according to the execution results corresponding to each of the database statements;

[0025] Generating regeneration statements corresponding to each of the first candidate database statements, and determining second candidate database statements from the multiple first candidate database statements according to the regeneration statements corresponding to each of the first candidate database statements;

[0026] Wherein, the regeneration statement is the natural language corresponding to the first candidate database statement;

[0027] If there are multiple second candidate database statements, determining the target database statement from the multiple second candidate database statements according to the statement to be processed, the regeneration statements corresponding to each of the second candidate database statements, and the execution results;

[0028] If there is one second candidate database statement, determining the second candidate database statement as the target database statement.

[0029] In an alternative embodiment, generating regeneration statements corresponding to each of the first candidate database statements, and determining second candidate database statements from the multiple first candidate database statements according to the regeneration statements corresponding to each of the first candidate database statements, includes:

[0030] Inputting each of the first candidate database statements into the large language model for processing respectively to obtain regeneration statements corresponding to each of the first candidate database statements;

[0031] Determining the first candidate database statements corresponding to the regeneration statements that are semantically and logically consistent with the statement to be processed as the second candidate database statements.

[0032] In an alternative embodiment, before generating regeneration statements corresponding to each of the first candidate database statements, the method further includes:

[0033] Determine whether the first candidate database statement meets the second preset condition;

[0034] If the first candidate database statement does not meet the second preset condition, determine a first variant statement according to other database statements in the database statement except the first candidate database statement, regenerate a new database statement corresponding to the first variant statement, and determine a first candidate database statement from the new database statement until the first candidate database statement meets the second preset condition;

[0035] The generating of the regenerated statements corresponding to each of the first candidate database statements includes:

[0036] If the first candidate database statement meets the second preset condition, generate the regenerated statements corresponding to each of the first candidate database statements.

[0037] In an alternative embodiment, the method further includes:

[0038] If there is no regenerated statement with the same semantic logic as the statement to be processed, determine the variant statements corresponding to each of the first candidate database statements as second variant statements, and use each of the first candidate database statements as negative examples to regenerate new database statements corresponding to each of the second variant statements, and obtain a first candidate database statement from the new database statements, and generate the regenerated statements corresponding to each of the first candidate database statements until at least one regenerated statement with the same semantic logic as the statement to be processed is obtained.

[0039] In an alternative embodiment, the determining of the target database statement from multiple second candidate database statements according to the statement to be processed, the regenerated statements corresponding to each of the second candidate database statements, and the execution results includes:

[0040] Calculate the similarity scores corresponding to each of the second candidate database statements respectively according to the statement to be processed and the regenerated statements corresponding to each of the second candidate database statements; the similarity score represents the similarity between the regenerated statement of the second candidate database statement and the statement to be processed;

[0041] Vote and score each of the second candidate database statements according to the execution results of the second candidate database statements to obtain the vote scores corresponding to each of the second candidate database statements; the vote score represents the credibility of the second candidate database statement;

[0042] For each of the second candidate database statements, calculate the confidence score corresponding to the second candidate database statement according to the similarity score and the vote score;

[0043] Determine the second candidate database statement with the highest confidence score as the target database statement.

[0044] In a second aspect, the present invention provides a database statement generation device based on a large language model, the device comprising:

[0045] A generation module, configured to obtain a to-be-processed statement input by a user, and generate a plurality of variant statements corresponding to the to-be-processed statement through a large language model; both the to-be-processed statement and the variant statements are natural languages, and the variant statements have different syntactic structures from the to-be-processed statement, but are semantically consistent;

[0046] The generation module is further configured to obtain a target data structure from a database through the large language model according to the plurality of variant statements, and generate database statements corresponding to the respective variant statements according to each of the variant statements and the target data structure;

[0047] A determination module, configured to determine a target database statement from the plurality of database statements; the target database statement is a database statement that can be successfully executed and has the highest confidence.

[0048] In a third aspect, the present invention provides an electronic device, comprising a processor and a memory, the memory storing a computer program executable by the processor, and the processor being capable of executing the computer program to implement the method according to any one of the foregoing embodiments.

[0049] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program, when executed by a processor, implements the method according to any one of the foregoing embodiments.

[0050] After obtaining the to-be-processed statement input by the user, the database statement generation method and related device provided in the embodiments of the present application can generate a plurality of variant statements corresponding to the to-be-processed statement through a large language model, so that a target data structure can be obtained from the database according to the plurality of variant statements. Since the variant statements are logically consistent with the to-be-processed statement input by the user but have different syntactic structures, the target data structure can be obtained from the database through the plurality of variant statements, thereby expanding the coverage of the retrieved data structures. Based on this, the corresponding search space can be expanded when generating database statements based on the target data structure, thereby improving the accuracy of the finally generated target database statement.

[0051] The features and advantages will be more obvious and understandable. The following specific embodiments are given and described in detail in conjunction with the accompanying drawings as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0053] Figure 1 Shows a block diagram of an electronic device provided by an embodiment of the present application;

[0054] Figure 2 Shows a schematic flowchart of a method for generating database statements based on a large language model provided by an embodiment of the present application;

[0055] Figure 3 Shows an example diagram of selecting a target database statement;

[0056] Figure 4 Shows another schematic flowchart of a method for generating database statements based on a large language model provided by an embodiment of the present application;

[0057] Figure 5 Shows a functional module diagram of a device for generating database statements based on a large language model provided by an embodiment of the present application.

[0058] Icons: 100 - Memory; 110 - Processor; 120 - Communication Module; 200 - Generation Module; 210 - Determination Module. Detailed Embodiments

[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Usually, the components of the embodiments of the present application described and shown in the accompanying drawings here can be arranged and designed in various different configurations.

[0060] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0061] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0062] With the development of natural language processing technology, computers can better understand and process natural language. Therefore, in order to improve the user experience, users are allowed to directly input natural sentences when they want to perform relevant operations on the database. The computer generates relevant database statements based on the natural sentences input by the user to achieve operations on the database, such as data query.

[0063] Currently, the computer can directly obtain relevant data structures (schemas) from the database through the natural sentences input by the user, and then generate corresponding database statements based on the retrieved schemas. However, this method has the following problems:

[0064] 1. It has certain retrieval limitations. Specifically, since the natural sentences input by the user may be ambiguous or have unclear meanings, and there may be many similar tables or fields in the database, this will lead to a problem of a small coverage range when directly retrieving schemas from the database through the natural sentences input by the user. Therefore, the search space will be small when generating database statements, resulting in poor accuracy of the generated database statements.

[0065] 2. In the prior art, when validating the generated database statements, it often relies on explicit errors of the database statements, such as execution errors, etc., but does not pay attention to implicit errors of the database statements, such as logical errors although the execution is passed, etc. This will also lead to poor accuracy of the generated database statements.

[0066] 3. In the prior art, when an error occurs during the generation of database statements, it is often unable to correct itself, resulting in a high error rate in the generation of database statements.

[0067] 4. In the prior art, the error rate is often high when retrieving schemas, and since this part of the work is generally borne by large language models, it is easy to cause the hallucination problem of large language models, resulting in poor recall rate of the generated database statements.

[0068] It should be noted that the hallucination problem of the model refers to the situation where the model generates or predicts information or results that do not conform to reality when processing data.

[0069] Based on this, the embodiments of the present application provide a method and related device for generating database statements based on a large language model to solve the above problems.

[0070] Specifically, Figure 1 The following is a block diagram of the electronic device provided by the embodiments of the present application. Please refer to Figure 1 , this electronic device includes a memory 100, a processor 110, and a communication module 120. Each element of the memory 100, the processor 110, and the communication module 120 is electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines.

[0071] Among them, the memory 100 is used to store computer programs or data that can be executed by the processor. The memory 100 can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.

[0072] The processor 110 is used to read / write the data or computer programs stored in the memory and execute the computer program to implement the method for generating database statements based on a large language model provided by the embodiments of the present application.

[0073] The communication module is used to establish a communication connection between the electronic device and other communication terminals through the network and is used to send and receive data through the network.

[0074] Optionally, the electronic device can be a device such as a PC or a server, or a mobile terminal device such as a mobile phone or a tablet.

[0075] It should be understood that Figure 1 The structure shown is only a schematic diagram of the electronic device, and the electronic device may also include more or fewer components than those shown in Figure 1 , or have a configuration different from that shown in Figure 1 . Figure 1Each component shown in the figure may be implemented by hardware, software, or a combination thereof.

[0076] Next, taking the electronic device in the above Figure 1 as the execution subject, the method for generating database statements based on a large language model provided by the embodiments of the present application will be introduced exemplarily in combination with the process schematic diagram. Specifically, Figure 2 is a process schematic diagram of the method for generating database statements based on a large language model provided by the embodiments of the present application. Please refer to Figure 2 The method includes:

[0077] Step S20: Obtain the statement to be processed input by the user, and generate multiple variant statements corresponding to the statement to be processed through a large language model.

[0078] Among them, both the statement to be processed and the variant statements are natural languages, and the variant statements have different syntactic structures from the statement to be processed but the same semantic logic.

[0079] Optionally, the statement to be processed may be a query statement of the user for some data. It can be understood that the statement to be processed can represent the query intention of the user.

[0080] In this embodiment, the syntactic structure may include the specific words used, the sentence structure, and the sentence complexity. That is to say, a variant statement refers to a sentence that is at least different from the statement to be processed in terms of word usage, sentence complexity, and sentence structure, but has the same meaning logic.

[0081] In one example, if the statement to be processed is "Find the model with the highest sales volume between July 15, 2024 and July 20, 2024", the variant statements may include "Which car sold the best between July 15, 2024 and July 20, 2024", "During the period from July 15, 2024 to July 20, 2024, which model's sales data is the most prominent", "If we focus on the sales period from July 15, 2024 to July 20, 2024, which car is the most popular in the market", "If we examine the car sales between July 15, 2024 and July 20, 2024, which car's sales volume is far ahead", and so on.

[0082] In this embodiment, the statement to be processed can be processed by a pre-trained large language model to generate multiple corresponding variant statements. Optionally, relevant prompts can be generated to guide the large language model to perform the expected processing.

[0083] It should be noted that Prompt refers to the input text used to guide or stimulate the model to generate a specific type of output. A Prompt can be a question, an instruction, a partial sentence, or any form of text, and its purpose is to stimulate the model to produce the expected response or behavior.

[0084] In a possible implementation, the electronic device can process the statement to be processed through a pre-set variant statement generation template, thereby generating a prompt (Prompt) corresponding to the statement to be processed, and inputting this Prompt into the large language model to output multiple variant statements.

[0085] Step S21, obtain the target data structure from the database according to multiple variant statements through the large language model, and generate database statements corresponding to each variant statement according to each variant statement and the target data structure.

[0086] Optionally, the target data structure can be used to guide the generation of corresponding database statements.

[0087] Optionally, the type of the database statement can be set by the user according to the actual application situation. In a possible implementation, the database statement can be SQL.

[0088] Optionally, since the multiple generated database statements may be inconsistent, the electronic device also needs to select a suitable target database statement from the multiple database statements and return it.

[0089] Step S22, determine the target database statement from the multiple database statements.

[0090] Among them, the target database statement is the database statement that can be successfully executed and has the highest confidence.

[0091] In the database statement generation method based on the large language model provided by the embodiments of the present application, after the electronic device obtains the statement to be processed input by the user, it can generate multiple variant statements corresponding to the statement to be processed, and thus can obtain the target data structure from the database according to the multiple variant statements. Since the variant statements are logically consistent with the statement to be processed input by the user but have different syntactic structures, the target data structure can be obtained from the database through the multiple variant statements, thereby expanding the coverage of the retrieved data structures. Based on this, the corresponding search space can be expanded when generating database statements based on the target data structure, thereby improving the accuracy of the finally generated target database statement.

[0092] Next, a possible implementation manner is provided for how the electronic device obtains the target data structure from the database according to multiple variant statements through the large language model, and generates database statements corresponding to each variant statement according to each variant statement and the target data structure.

[0093] In this embodiment, the electronic device can input each variant statement into a large language model for processing to obtain the first semantic structure corresponding to each variant statement. Then, according to the first semantic structure corresponding to each variant statement, the target table and target field corresponding to each variant statement are obtained from the database, and the union of the target table and target field is taken to obtain the target data structure.

[0094] Optionally, the first semantic structure can represent the semantic logic of the variant statement.

[0095] In a possible implementation manner, the first semantic structure can be in a semi-structured form, such as a dictionary structure of {entity: […], relationship: […]}, where the entity can be a noun in the variant statement, and the relationship can be the relationship between nouns.

[0096] For example, if the variant statement is "Find the students in Class 1 of School A", then the entities can include School A, Class 1, and students, and the relationship is the relationship among these three.

[0097] Optionally, the electronic device can process each variant statement according to a pre-stored semantic structure generation template to obtain the Prompt corresponding to each variant statement, and then input the Prompt corresponding to each variant statement into the large language model respectively to generate the first semantic structure corresponding to each variant statement.

[0098] Optionally, the electronic device can input the first semantic structure corresponding to each variant statement into the large language model for processing respectively, and the large language model obtains the target table and target field corresponding to each variant statement from the database according to each first semantic structure.

[0099] It can be understood that before inputting the first semantic structure into the large language model, the electronic device also needs to process the first semantic structure through a pre-stored query template to obtain the Prompt corresponding to each first semantic structure.

[0100] It can be understood that since the target data structure includes the target tables and target fields corresponding to all variant statements, the coverage range of the data structure can be expanded, and thus the search space when generating database statements can be expanded.

[0101] Optionally, after obtaining the target data structure, the electronic device can also input the target data structure and each variant statement into the large language model for processing to obtain the database statement corresponding to each variant statement.

[0102] Optionally, the electronic device can process each variant statement according to a template generated based on a pre-stored database statement to obtain the Prompt corresponding to each variant statement, and then input the target data structure and the Prompt corresponding to each variant statement into the large language model for processing, so as to obtain the database statement corresponding to each variant statement.

[0103] In a possible implementation manner, the electronic device can input a variant statement and the target data structure into the large language model each time, so as to generate the database statement corresponding to the variant statement.

[0104] Optionally, in order to effectively eliminate the hallucination problem of the large language model, the electronic device can also preprocess the first semantic structure before obtaining the target table and target field corresponding to each variant statement from the database according to the first semantic structure corresponding to each variant statement.

[0105] Specifically, the electronic device can input the statement to be processed into the large language model for processing, obtain the second semantic structure corresponding to the statement to be processed, and determine whether the target first semantic structure with the same semantics as the second semantic structure meets the first preset condition.

[0106] In a possible implementation manner, the first preset condition can be that the proportion of the number of target first semantic structures in the total number of all first semantic structures reaches the first preset value.

[0107] Optionally, the first preset value can be set according to the actual application situation, for example, 80%, and the present application does not make too many limitations on this.

[0108] In another possible implementation manner, the first preset condition can also be that the number of target first semantic structures reaches the first preset number.

[0109] Optionally, the first preset number can be set by the user according to the actual application situation, for example, set to 10, etc., and the present application does not make too many limitations on this.

[0110] Optionally, the electronic device can process the statement to be processed according to a pre-stored semantic structure generation template, so as to obtain the Prompt corresponding to the statement to be processed, and then input the Prompt corresponding to the statement to be processed into the large language model to generate the second semantic structure corresponding to the statement to be processed.

[0111] In this embodiment, the second semantic structure is in a semi-structured form, such as a dictionary structure of {entity: […], relationship: […]}, where the entity can be a noun in the variant statement, and the relationship can be the relationship between nouns.

[0112] Optionally, the electronic device may input each of the first semantic structures and the second semantic structures into a large language model for processing respectively, so as to determine whether the first semantic structure is a target first semantic structure with the same semantics as the second semantic structure.

[0113] It can be understood that if the semantics of the first semantic structure is the same as that of the second semantic structure, it means that the variant statement corresponding to the first semantic structure is logically consistent with the statement to be processed, and the variant statement is successfully parsed. If the semantics of the first semantic structure is different from that of the second semantic structure, it means that the variant statement corresponding to the first semantic structure is not logically consistent with the statement to be processed, and there may be a situation where the variant statement fails to be parsed.

[0114] It can be understood that if the number of successfully parsed variant statements is small and does not meet the first preset condition, the variant statements that fail to be parsed need to be parsed again until the number of successfully parsed variant statements reaches a certain amount, so as to ensure the coverage and accuracy of the retrieved target data structure at the same time.

[0115] Specifically, if the target first semantic structure does not meet the first preset condition, the electronic device may add a negative label to other semantic structures except the target first semantic structure, and re-input the other first semantic structures with the negative label and the variant statements corresponding to the other first semantic structures into the large language model for processing, obtain new first semantic structures corresponding to the variant statements, and determine the target first semantic structure from the new first semantic structures until the target first semantic structure meets the first preset condition.

[0116] Optionally, other semantic structures except the target first semantic structure refer to the first semantic structures corresponding to the variant statements that fail to be parsed.

[0117] Optionally, the negative label may indicate that the first semantic structure is a wrong negative example, so as to prompt the large language model to avoid this result and output a new result.

[0118] Optionally, after obtaining the new first semantic structures, the electronic device may determine the target first semantic structure again among these new first semantic structures, and combine the previously generated target first semantic structures to determine whether all the target first semantic structures meet the first preset condition.

[0119] For example, if there are 10 variant statements, 2 successfully parsed variant statements, and 8 variant statements that fail to be parsed, the electronic device may use the first semantic structures corresponding to these 8 variant statements that fail to be parsed as negative examples, re-input these 8 variant statements that fail to be parsed into the large language model for processing, generate 8 new first semantic structures, and thus determine the target first semantic structure from these 8 new first semantic structures.

[0120] In this example, if the proportion of the number of target first semantic structures in all first semantic structures reaches 80%, and among these 8 new first semantic structures, 6 target first semantic structures are re-determined, then combining the 2 previously successfully parsed variant statements, it can be determined that 8 out of these 10 variant statements are successfully parsed. Therefore, it can be determined that the target first semantic structure meets the first preset condition.

[0121] In this example, if the number of target first semantic structures re-determined among these 8 new first semantic structures does not reach 6, for example, only 4 new first semantic structures are target first semantic structures, then combining the 2 previously successfully parsed variant statements, it can be determined that only 6 out of these 10 variant statements are successfully parsed, and the target first semantic structure does not meet the first preset condition. Therefore, for these 4 variant statements that are still not successfully parsed, it is necessary to continue to input the corresponding first semantic structures as negative examples into the large language model for processing until the number of finally determined target first semantic structures reaches 8.

[0122] Optionally, considering that there may be a situation where the target first semantic structure still does not meet the first preset condition after repeating the above steps many times, a possible processing method for this situation is given below.

[0123] In a possible implementation manner, if the number of times of repeating the above steps reaches the preset number of times and the target first semantic structure still does not meet the first preset condition, an error can be directly reported.

[0124] In another possible implementation manner, if the number of times of repeating the above steps reaches the preset number of times and the target first semantic structure still does not meet the first preset condition, it can also be considered that there may be a problem with the incorrect parsing of the second semantic structure. Therefore, the second semantic structure can be used as a negative example, a negative label can be added, and the second semantic structure and the statement to be processed can be input into the large language model for re-processing, so as to obtain a new second semantic structure, and re-determine whether the target first semantic structure with the same semantics as the new second semantic structure meets the first preset condition.

[0125] In this embodiment, the electronic device verifies the first semantic structure, that is, determines whether the first semantic structure is the target first semantic structure, and determines whether the number of target first semantic structures reaches a certain number. If the verification fails, a correction loop can be constructed to achieve self-correction by re-inputting other first semantic structures outside the target first semantic structure as negative examples into the large language model to guide the large language model to re-generate new first semantic structures, so as to reduce the error rate of database statement generation.

[0126] In this embodiment, if the target first semantic structure meets the first preset condition, the electronic device can obtain the target table and target field corresponding to the variant statement of the target first semantic structure from the database according to each target first semantic structure, so as to take the union of the target table and target field to obtain the target data structure. Then, the target data structure and the variant statements corresponding to each target first semantic structure are input into the large language model for processing to obtain the database statements corresponding to the variant statements of each target first semantic structure.

[0127] Optionally, if the target first semantic structure meets the first preset condition, the electronic device can directly delete the variant statements with parsing failures, and obtain the corresponding target table and target field according to the target first semantic structure corresponding to the successfully parsed variant statements.

[0128] Please continue to refer to the above example. If the number of target first semantic structures reaches 8, the electronic device can delete two variant statements with parsing failures, and obtain the target tables and target fields corresponding to the 8 successfully parsed variant statements according to these 8 target first semantic structures. Thus, by taking the union of the target tables and target fields corresponding to these 8 variant statements, the target data structure corresponding to the successfully parsed variant statements as a whole is obtained.

[0129] In this example, the electronic device can also input the target data structure and these 8 variant statements into the large language model for processing to obtain 8 database statements.

[0130] In this embodiment, by using the semantic structure corresponding to the variant statement with parsing failure as a negative example and re-inputting it into the large language model to re-extract the new semantic structure corresponding to the variant statement, the hallucination problem of the large language model can be eliminated to the greatest extent. And in this way, the burden of the large language model in retrieving the target table and target field can be effectively reduced, the recall rate of the target table and target field can be improved, and then the recall rate of the generated database statements can be improved.

[0131] Optionally, since multiple database statements are generated at this time, and there may be certain differences between the database statements, it is also necessary to select the most appropriate one from the multiple database statements as the result to return.

[0132] Next, a possible implementation manner is provided for how to determine the target database statement from multiple database statements.

[0133] Specifically, the electronic device can first execute each database statement respectively, and determine multiple first candidate database statements from the multiple database statements according to the execution results corresponding to each database statement.

[0134] Optionally, the first candidate database statement is a database statement that has been successfully executed.

[0135] In this embodiment, the electronic device may further generate a regeneration statement corresponding to each first candidate database statement, and determine a second candidate database statement from multiple first candidate database statements according to the regeneration statements corresponding to the first candidate database statements.

[0136] Among them, the regeneration statement is the natural language corresponding to the first candidate database statement.

[0137] Optionally, in order to verify whether there are hidden problems such as logical errors in the successfully executed database statements, the electronic device may generate the corresponding regeneration statements for each successfully executed first candidate database statement, that is, generate natural statements according to the database statements.

[0138] Optionally, since the regeneration statement can be used to determine whether there are hidden problems such as logical errors in the successfully executed first candidate database statement, the second candidate database statement refers to a database statement that has neither an obvious execution error problem nor a hidden logical error problem.

[0139] In this embodiment, if there are multiple second candidate database statements, the electronic device may determine a target database statement from multiple second candidate database statements according to the statement to be processed, the regeneration statements corresponding to the second candidate database statements, and the execution result; if there is only one second candidate database statement, the second candidate database statement is determined as the target database statement.

[0140] Optionally, if the number of second candidate database statements is multiple, the electronic device may calculate a measurement score, such as a confidence score, for each second candidate database statement according to the statement to be processed, the regeneration statements corresponding to the second candidate database statements, and the execution result, so as to select a target database statement according to the confidence score. It can be understood that in the process of screening the target database statement, the electronic device can not only verify whether the database statement has been successfully executed to avoid obvious error problems, but also verify whether the database statement has logical errors to avoid hidden error problems, so the accuracy of the finally screened target database statement can be improved.

[0141] Optionally, considering that if the number of first candidate database statements is too small, it may be impossible to screen out suitable second candidate database statements, so before generating the regeneration statements corresponding to each first candidate database statement, the electronic device may also determine whether the first candidate database statement meets the second preset condition.

[0142] In a possible implementation, the second preset condition may be that the proportion of the number of the first candidate database statements to the number of all database statements reaches a second preset value.

[0143] Optionally, the second preset value may be set according to the actual application scenario. For example, it is 80%, and the present application does not make excessive limitations on this.

[0144] In another possible implementation, the first preset condition may also be that the number of target first semantic structures reaches a first preset number.

[0145] Optionally, the first preset number may be set by the user according to the actual application scenario. For example, it is set to 10, etc., and the present application does not make excessive limitations on this.

[0146] Optionally, if the first candidate database statement does not meet the second preset condition, then a first variant statement is determined according to the other database statements except the first candidate database statement in the database statements, a new database statement corresponding to the first variant statement is regenerated, and a first database statement is determined from the new database statement until the first candidate database statement meets the second preset condition.

[0147] Optionally, all the other database statements except the first candidate database statement are database statements with execution failures.

[0148] In this embodiment, if the first candidate database statement does not meet the second preset condition, the electronic device may determine a variant statement with execution failure, that is, the first variant statement, according to the database statements with execution failures, and regenerate a new database statement corresponding to the first variant statement, so as to screen out the first candidate database statement from the new database statement until the newly screened first candidate database statement and the previously screened first candidate database statement jointly meet the second preset condition.

[0149] In an example, if there are 10 database statements, the second preset condition is that the proportion of the number of the first candidate database statements to the number of all database statements reaches 80%, and there are only 2 first candidate database statements screened out from these 10 database statements, and there are 8 database statements with execution failures, then the electronic device may determine the variant statements corresponding to these 8 database statements as the first variant statements, regenerate new database statements for these 8 first variant statements, and re-determine the first candidate database statement from these 8 new database statements.

[0150] In this example, if the number of the first candidate database statements determined from these 8 new database statements reaches 6, then combining the 2 first candidate database statements determined previously, it can be determined that the proportion of the number of the first candidate database statements that are successfully executed among these 10 database statements reaches 80%. Therefore, it can be determined that the first candidate database statement meets the second preset condition.

[0151] In this example, if the number of the first candidate database statements determined from these 8 new database statements does not reach 6, for example, there are only 4, then combining the 2 first candidate database statements determined previously, it can be determined that there are only 6 first candidate database statements that are successfully executed among these 10 database statements, and its proportion does not reach 80%, still not meeting the second preset condition. Therefore, new database statements can be generated again for the variant statements corresponding to the 4 database statements that are not successfully executed, and the first candidate database statements are determined again from these 4 new database statements until the proportion of the finally determined first candidate database statements reaches 80%.

[0152] Optionally, considering that there may be a situation where the first candidate database statement cannot meet the first preset condition even after repeating the above steps many times, then at this time, in order to prevent entering an infinite loop, the electronic device can report an error.

[0153] In this embodiment, the electronic device can verify the database statement, that is, determine whether the database statement is a first candidate database statement and determine whether the number of the first candidate database statements reaches a certain quantity. If the verification fails, the database statements other than the first candidate database statements can be re - input into the large - language model as negative examples to guide the large - language model to regenerate new database statements, thereby constructing a correction loop to achieve self - correction and reduce the error rate of database statement generation.

[0154] Optionally, for the query scenario, considering that the specific situations of execution failure may include the following two: 1. Query field error, that is, there is no relevant field or relevant table; 2. Return is empty, that is, no relevant data is queried. Therefore, different methods can be used to generate new database statements for these two situations respectively.

[0155] In this embodiment, if the query field error occurs, it may be because the target data structure retrieved from the database is incorrect. Therefore, for the first variant statement corresponding to the database statement with a query field error, the electronic device can add a negative label to the database statement, and re - input the error reason (field error), error field, this database statement, and this first variant statement into the large - language model for processing, so as to re - obtain the target table and target field corresponding to this first variant statement, and further re - obtain the target data structure to regenerate the database statement.

[0156] In this embodiment, if the returned result is empty, it may be because there is a problem with the generated database statement itself. Therefore, for the first variant statement corresponding to the database statement with an empty return, the electronic device can add a negative label to the database statement, and re-enter the error reason (the return is empty), the database statement, and the first variant statement into the large language model for processing, so as to obtain the target data structure again to regenerate the database statement.

[0157] It can be understood that by regenerating the corresponding database statements in different ways for different execution failure reasons, the execution passing rate of the regenerated database statements can be maximally improved.

[0158] In this embodiment, if the first candidate database statement meets the second preset condition, the electronic device can generate a regeneration statement corresponding to each first candidate database statement.

[0159] Next, a possible implementation method is provided for how to generate the regeneration statements corresponding to each first candidate database statement, and how to determine the second candidate database statement from multiple first candidate database statements according to the regeneration statements corresponding to each first candidate database statement.

[0160] Specifically, the electronic device can respectively input each first candidate database statement into the large language model for processing to obtain the regeneration statement corresponding to each first candidate database statement, and then determine the first candidate database statement corresponding to the regeneration statement with the same semantic logic as the statement to be processed as the second candidate database statement.

[0161] Optionally, the electronic device can process each first candidate database statement according to a pre-stored regeneration template to obtain the Prompt corresponding to each first candidate database statement, and then respectively input the Prompt corresponding to each first candidate database statement into the large language model for processing to obtain the regeneration statement corresponding to each first candidate database statement.

[0162] Optionally, the electronic device can also determine the regeneration statement with the same semantic logic as the statement to be processed through the large language model, so as to determine the first candidate database statement corresponding to the regeneration statement as the second candidate database statement.

[0163] In a possible implementation manner, the large language model can represent the parsing result as a dichotomous pseudo label, such as 0 and 1, where 0 represents that the regeneration statement has the same semantic logic as the statement to be processed, and 1 represents that the regeneration statement has a different semantic logic from the statement to be processed. Then, the electronic device can determine whether the first candidate database statement corresponding to the regeneration statement is the second candidate database statement according to the output pseudo label.

[0164] Understandably, if there is no second candidate database statement, it means that all the successfully executed database statements have logical errors, that is, these database statements are generated incorrectly. Therefore, the electronic device can regenerate database statements for the variant statements corresponding to all the first candidate database statements in the absence of a second candidate database statement to re-screen the second candidate database statement.

[0165] In a possible implementation, if there is no regenerated statement with the same semantic logic as the statement to be processed, the electronic device can determine the variant statements corresponding to each first candidate database statement as the second variant statements, and use each first candidate database statement as a negative example to regenerate new database statements corresponding to each second variant statement, and obtain the first candidate database statement from the new database statements to generate the regenerated statement corresponding to each first candidate database statement until at least one regenerated statement with the same semantic logic as the statement to be processed is obtained.

[0166] Optionally, the electronic device can use each first candidate database statement as a negative example, add a negative label to each first candidate database statement, and re-enter each first candidate database statement and its corresponding second variant statement into the large language model for processing to regenerate the database statements corresponding to each second variant statement.

[0167] Optionally, the negative label can indicate that there is a logical error in the first candidate database statement.

[0168] Optionally, to ensure the executability of the regenerated database statement, the electronic device can determine whether the regenerated database statement is a first candidate database statement, and then generate the regenerated statement corresponding to the first candidate database statement to determine whether the semantic logic of the regenerated statement is consistent with the statement to be processed until at least one regenerated statement with the same semantic logic as the statement to be processed is obtained.

[0169] In this embodiment, the electronic device can verify the first candidate database statement, that is, determine whether there is a second candidate database statement in the first candidate database statement. If the verification fails, the electronic device can re-enter the first candidate database statement as a negative example into the large language model to guide the large language model to regenerate new database statements, thereby constructing a correction loop to achieve self-correction and reduce the error rate of database statement generation.

[0170] In this embodiment, if only one second candidate database statement is obtained, the second candidate database statement can be directly determined as the target database statement. If multiple second candidate database statements are obtained, the electronic device still needs to screen out the target database statement from the multiple second candidate database statements.

[0171] Next, a possible implementation manner is provided for how to determine a target database statement from multiple second candidate database statements according to the statement to be processed, the regenerated statements corresponding to each second candidate database statement, and the execution results.

[0172] Specifically, the electronic device can calculate the similarity scores corresponding to each second candidate database statement respectively according to the statement to be processed and the regenerated statements corresponding to each second candidate database statement, and vote and score each second candidate database statement according to the execution results of the second candidate database statements to obtain the vote scores corresponding to each second candidate database statement.

[0173] Among them, the similarity score represents the similarity between the regenerated statement of the second candidate database statement and the statement to be processed, and the vote score represents the credibility of the second candidate database statement.

[0174] Optionally, the electronic device can calculate the similarity score through vector similarity technology and calculate the vote score through the result voting and scoring method.

[0175] In this embodiment, the vote scoring method accurately counts the occurrence frequencies of each second candidate database statement, automatically identifies and assigns scores corresponding to each execution result according to the occurrence frequencies, so as to calculate the vote score through the scores corresponding to each execution result.

[0176] Specifically, the more the execution result appears, the higher the score corresponding to the execution result.

[0177] In this embodiment, through result voting and scoring, it can be determined whether the execution result of the second candidate database statement is consistent with the execution results of other second candidate database statements.

[0178] It can be understood that the higher the similarity score, the higher the similarity between the semantic logic of the second candidate database statement and the semantic logic of the statement to be processed input by the user; the higher the vote score, the more times the execution result corresponding to the second candidate database statement appears, and the more credible the second candidate database statement is.

[0179] In this embodiment, the electronic device can calculate the confidence score corresponding to the candidate database statement respectively for each second candidate database statement according to the similarity score and the vote score, so as to determine the second candidate database statement with the highest confidence score as the target database statement.

[0180] Optionally, corresponding weight values can be set for the similarity score and the voting score in advance. Then, the electronic device can calculate the confidence score of each second candidate database statement according to its similarity score, the weight value corresponding to the similarity score, the voting score, and the weight value corresponding to the voting score.

[0181] In a possible implementation manner, the confidence score can be calculated according to the following formula:

[0182] X = aM + bN

[0183] Where X represents the confidence score, a represents the weight value corresponding to the similarity score, b represents the weight value corresponding to the voting score, M represents the similarity score, and N represents the voting score.

[0184] In an example, Figure 3 For an example diagram of selecting the target database statement, please refer to Figure 3 , for the to-be-processed statement "Find the model with the highest sales volume between July 15th and July 20th" input by the user, four variant statements can be generated. Among them, the database statement corresponding to the 3rd variant statement fails to execute, and the semantic logic of the regenerated statement corresponding to the database statement corresponding to the 4th variant statement is different from that of the to-be-processed statement. Therefore, it can be determined that the second candidate database statements are the database statements generated corresponding to the 1st variant statement and the 2nd variant statement.

[0185] In this example, the electronic device can calculate the confidence scores of the database statements generated corresponding to the 1st variant statement and the 2nd variant statement respectively, so as to finally determine that the target database statement is the database statement generated corresponding to the 2nd variant statement.

[0186] It can be understood that in this embodiment, the electronic device can perform a convergence process on the expanded search space by calculating the confidence, so as to further improve the recall rate of the database statement.

[0187] Next, taking the to-be-processed statement input by the user as a query problem and the database language as SQL as an example, combined with Figure 4 an overall exemplary introduction to the database statement generation method based on the large language model provided in the embodiments of the present application will be given. Specifically, Figure 4 For another flowchart of the database statement generation method based on the large language model provided in the embodiments of the present application, please refer to Figure 4 .

[0188] First, the user can input a user question, and through the large language model (LLM), multiple variant questions are output.

[0189] Optionally, the electronic device can also extract the semantic structure, i.e., entities and relationships, from the user question and the generated multiple variant questions through a large language model, so as to judge whether the extracted entities and relationships are logically consistent with the original user question according to the first semantic structure corresponding to each variant question and the second semantic structure corresponding to the user question.

[0190] If the number of consistent target first semantic structures does not reach a certain number, the first semantic structure corresponding to the variant question that is not logically consistent with the original user question is used as a negative example, and the first semantic structure and the variant question are re-input into the large language model to extract entities and relationships until the number of consistent target first semantic structures reaches a certain number.

[0191] If the number of consistent target first semantic structures reaches a certain number, the target first semantic structure is input into the large language model, and the large language model retrieves the target table and target field corresponding to the variant statement corresponding to each target first semantic structure from the database, and takes the union of the multiple target tables and target fields to obtain the target data structure.

[0192] The union (target data structure) and the variant questions corresponding to each target first semantic structure are input into the large language model for processing, so as to generate a corresponding SQL statement for each such variant question and execute each SQL statement.

[0193] If the number of SQL statements that fail to execute reaches a certain number, the SQL statements can be regenerated respectively according to the reasons for passing. That is, if the execution fails due to a field error, it can be used as a negative example and input into the large language model to reselect the target table and target field; if the execution fails due to an empty return, it can be used as a negative example to re-guide the generation of the corresponding SQL statement until the number of SQL statements that execute successfully reaches a certain number.

[0194] If the number of SQL statements that execute successfully reaches a certain number, the large language model can generate a regenerated question for the SQL statements that execute successfully, and the large language model can judge whether the semantic logic of the user question and each regenerated question is consistent.

[0195] If all the regenerated questions are not semantically logically consistent with the user question, they can be used as negative examples to re-guide the generation of the corresponding SQL statement.

[0196] If there is at least one regenerated question that is semantically logically consistent with the user question (mainly for the case where there are multiple regenerated questions that are semantically logically consistent with the user question), the results of the SQL statements corresponding to each regenerated question can be voted on to obtain the voting scores of each SQL statement.

[0197] In addition, it is also necessary to calculate the embedding similarity between each regenerated question and the user question to determine the similarity score of each SQL statement.

[0198] Finally, by performing a weighted calculation on the voting score and the similarity score, the final target SQL statement is determined, and the data queried according to the target SQL statement is output.

[0199] To execute the corresponding steps in the above embodiments and various possible ways, an implementation manner of a database statement generation device based on a large language model is given below. Optionally, the database statement generation device based on a large language model can adopt the Figure 1 device structure of the electronic device shown above. Further, please refer to Figure 5 , Figure 5 which is a functional module diagram of a database statement generation device based on a large language model provided by an embodiment of the present application. It should be noted that the basic principle and the technical effects produced by the database statement generation device based on a large language model provided in this embodiment are the same as those in the above embodiments. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above embodiments. The database statement generation device based on a large language model includes: a generation module 200 and a determination module 210.

[0200] The generation module 200 is used to obtain the statement to be processed input by the user and generate multiple variant statements corresponding to the statement to be processed through a large language model; both the statement to be processed and the variant statements are natural languages, and the variant statements have different syntactic structures from the statement to be processed but the same semantic logic.

[0201] It can be understood that the generation module 200 can also be used to execute the above step S20.

[0202] The generation module 200 is also used to obtain a target data structure from the database through a large language model according to multiple variant statements, and generate database statements corresponding to each variant statement according to each variant statement and the target data structure.

[0203] It can be understood that the generation module 200 can also be used to execute the above step S21.

[0204] The determination module 210 is used to determine a target database statement from multiple database statements; the target database statement is a database statement that can be successfully executed and has the highest confidence level.

[0205] It can be understood that the determination module 210 can also be used to execute the above step S22.

[0206] Optionally, the generation module 200 is further configured to input each variant statement into a large language model for processing to obtain a first semantic structure corresponding to each variant statement; obtain a target table and target fields corresponding to each variant statement from a database according to the first semantic structure corresponding to each variant statement, take the union of the target table and the target fields to obtain a target data structure; input the target data structure and each variant statement into the large language model for processing to obtain a database statement corresponding to each variant statement.

[0207] Optionally, the generation module 200 is further configured to input a statement to be processed into a large language model for processing to obtain a second semantic structure corresponding to the statement to be processed, and determine whether a target first semantic structure semantically identical to the second semantic structure meets a first preset condition; if the target first semantic structure does not meet the first preset condition, add a negative label to other semantic structures except the target first semantic structure, and re-input the other first semantic structures with the negative label added and the variant statements corresponding to the other first semantic structures into the large language model for processing to obtain a new first semantic structure corresponding to the variant statement, and determine the target first semantic structure from the new first semantic structure until the target first semantic structure meets the first preset condition.

[0208] Optionally, if the target first semantic structure meets the first preset condition, the generation module 200 is further configured to obtain a target table and target fields corresponding to the variant statement corresponding to the target first semantic structure from the database according to each target first semantic structure; take the union of the target table and the target fields to obtain a target data structure; input the target data structure and the variant statements corresponding to each target first semantic structure into the large language model for processing to obtain a database statement corresponding to the variant statement corresponding to each target first semantic structure.

[0209] Optionally, the determination module 210 is further configured to execute each database statement respectively, and determine a plurality of first candidate database statements from the plurality of database statements according to the execution results corresponding to each database statement; generate a regenerated statement corresponding to each first candidate database statement, and determine a second candidate database statement from the plurality of first candidate database statements according to the regenerated statement corresponding to each first candidate database statement; wherein, the regenerated statement is the natural language corresponding to the first candidate database statement; if there are multiple second candidate database statements, determine a target database statement from the multiple second candidate database statements according to the statement to be processed, the regenerated statements corresponding to each second candidate database statement, and the execution results; if there is one second candidate database statement, determine the second candidate database statement as the target database statement.

[0210] Optionally, the determination module 210 is further used to input each first candidate database statement into the large language model for processing, to obtain a regenerated statement corresponding to each first candidate database statement; and to determine the first candidate database statement corresponding to the regenerated statement that is semantically logically consistent with the statement to be processed as the second candidate database statement.

[0211] Optionally, the determination module 210 is further used to determine whether the first candidate database statement satisfies a second preset condition; if the first candidate database statement does not satisfy the second preset condition, determine the first variant statement according to other database statements in the database statement except the first candidate database statement, regenerate a new database statement corresponding to the first variant statement, and determine the first candidate database statement from the new database statement until the first candidate database statement satisfies the second preset condition; if the first candidate database statement satisfies the second preset condition, generate the regenerated statements corresponding to each first candidate database statement.

[0212] Optionally, the determination module 210 is further used to determine the variant statement corresponding to each first candidate database statement as a second variant statement if there is no regenerated statement that is semantically logically consistent with the statement to be processed, and to use each first candidate database statement as a negative example to regenerate a new database statement corresponding to each second variant statement, and to obtain the first candidate database statement from the new database statement, and to generate a regenerated statement corresponding to each first candidate database statement, until at least one regenerated statement that is semantically logically consistent with the statement to be processed is obtained.

[0213] Optionally, the determination module 210 is further used to calculate the similarity score corresponding to each second candidate database statement according to the statement to be processed and the regenerated statement corresponding to each second candidate database statement; the similarity score represents the similarity between the regenerated statement of the second candidate database statement and the statement to be processed; according to the execution result of the second candidate database statement, each second candidate database statement is voted and scored to obtain the voting score of each second candidate database statement; the voting score represents the credibility of the second candidate database statement; for each second candidate database statement, according to the similarity score and the voting score, the confidence score corresponding to the second candidate database statement is calculated; and the second candidate database statement with the highest confidence score is determined as the target database statement.

[0214] Optionally, the above modules can be stored in the form of software or firmware. Figure 1 The memory shown in the figure or solidified in the operating system (OS) of the electronic device, and can be Figure 1 Meanwhile, the data and program codes required for executing the above modules can be stored in the memory.

[0215] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the method for generating database statements based on a large language model provided by the embodiment of the present application.

[0216] In several embodiments provided by the present application, it should be understood that the disclosed apparatus and method can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of apparatuses, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0217] In addition, in each embodiment of the present application, the functional modules can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.

[0218] If the above functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0219] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for generating database sentences based on a large language model, characterized in that: The method comprises: Obtaining a sentence to be processed input by a user, and generating multiple variant sentences corresponding to the sentence to be processed through a large language model; the sentence to be processed and the variant sentences are both natural languages, and the variant sentences have different grammatical structures from the sentence to be processed but consistent semantic logic; Obtaining a target data structure from a database according to the plurality of variant sentences using the large language model, and generating a database sentence corresponding to each variant sentence according to each variant sentence and the target data structure; The step of obtaining a target data structure from a database according to the plurality of variant sentences by using the large language model, and generating a database statement corresponding to each variant sentence according to each variant sentence and the target data structure, comprises: Inputting each of the variant sentences into a large language model for processing to obtain a first semantic structure corresponding to each of the variant sentences; According to the first semantic structure corresponding to each of the variant statements, the target table and the target field corresponding to each of the variant statements are acquired from the database, and a union of the target table and the target field is taken to obtain the target data structure; Inputting the target data structure and each of the variant sentences into the large language model for processing to obtain a database sentence corresponding to each of the variant sentences; A target database statement is determined from the multiple database statements; the target database statement is a database statement that can be successfully executed and has the highest confidence.

2. The method according to claim 1, characterized in that The method further comprises: Input the sentence to be processed into the large language model for processing, obtain a second semantic structure corresponding to the sentence to be processed, and determine whether a target first semantic structure with the same semantics as the second semantic structure meets a first preset condition; the first preset is reduced to a ratio of the number of the target first semantic structures to the number of all the first semantic structures reaching a first preset value, or the number of the target first semantic structures reaches a first preset number; If the target first semantic structure does not meet the first preset condition, negative labels are added to other semantic structures except the target first semantic structure, and the other first semantic structures with the negative labels added and the variant sentences corresponding to the other first semantic structures are re-input into the large language model for processing to obtain a new first semantic structure corresponding to the variant sentence, and the target first semantic structure is determined from the new first semantic structure until the target first semantic structure meets the first preset condition.

3. The method according to claim 2, characterized in that The step of acquiring a target table and a target field corresponding to each of the variant statements from the database according to the first semantic structure corresponding to each of the variant statements, taking a union of the target table and the target field, and obtaining the target data structure includes: If the target first semantic structure satisfies the first preset condition, acquiring a target table and a target field corresponding to a variant statement corresponding to the target first semantic structure from the database according to each target first semantic structure; Taking a union of the target table and the target field to obtain the target data structure; The step of inputting the target data structure and each of the variant sentences into the large language model for processing to obtain a database sentence corresponding to each of the variant sentences includes: The target data structure and the variant sentences corresponding to each of the target first semantic structures are input into the large language model for processing to obtain database sentences corresponding to each of the variant sentences corresponding to the target first semantic structures.

4. The method according to claim 1, characterized in that: The step of determining a target database statement from the plurality of database statements comprises: Executing each of the database statements respectively, and determining a plurality of first candidate database statements from the plurality of database statements according to the execution results corresponding to each of the database statements; generating a regenerated statement corresponding to each of the first candidate database statements, and determining a second candidate database statement from a plurality of the first candidate database statements according to the regenerated statement corresponding to each of the first candidate database statements; Wherein, the regenerated sentence is the natural language corresponding to the first candidate database sentence; If there are multiple second candidate database statements, determine the target database statement from the multiple second candidate database statements according to the to-be-processed statement, the regenerated statements corresponding to each of the second candidate database statements, and the execution results; If there is one second candidate database statement, the second candidate database statement is determined as the target database statement.

5. The method according to claim 4, characterized in that The generating of the regenerated statements corresponding to the first candidate database statements, and determining the second candidate database statement from the plurality of the first candidate database statements according to the regenerated statements corresponding to the first candidate database statements, comprises: Inputting each of the first candidate database sentences into the large language model for processing, respectively, to obtain a regenerated sentence corresponding to each of the first candidate database sentences; A first candidate database statement corresponding to the regenerated statement that is semantically logically consistent with the statement to be processed is determined as the second candidate database statement.

6. The method according to claim 4, characterized in that Before generating the regenerated statements corresponding to the first candidate database statements, the method further includes: Determine whether the first candidate database statements meet a second preset condition; the second preset condition is that the ratio of the number of the first candidate database statements to the number of all the database statements reaches a second preset value; If the first candidate database statement does not satisfy the second preset condition, determining a first variant statement according to other database statements in the database statements except the first candidate database statement, regenerating a new database statement corresponding to the first variant statement, and determining a first candidate database statement from the new database statement, until the first candidate database statement satisfies the second preset condition; The generating of the regenerated statements corresponding to the first candidate database statements includes: If the first candidate database statements meet the second preset condition, then a regenerated statement corresponding to each of the first candidate database statements is generated.

7. The method according to claim 4, characterized in that The method further comprises: If there is no regenerated statement that is semantically logically consistent with the statement to be processed, the variant statement corresponding to each of the first candidate database statements is determined as a second variant statement, and each of the first candidate database statements is used as a negative example to regenerate a new database statement corresponding to each of the second variant statements, and the first candidate database statement is obtained from the new database statement to generate a regenerated statement corresponding to each of the first candidate database statements until at least one regenerated statement that is semantically logically consistent with the statement to be processed is obtained.

8. The method according to claim 4, characterized in that The step of determining a target database statement from a plurality of second candidate database statements according to the statement to be processed, the regenerated statements corresponding to each of the second candidate database statements, and the execution results includes: Calculating similarity scores corresponding to the statements of the second candidate database respectively according to the statements to be processed and the regenerated statements corresponding to the statements of the second candidate database; the similarity scores represent the similarity between the regenerated statements of the second candidate database statements and the statements to be processed; According to the execution result of the second candidate database statement, voting and scoring each of the second candidate database statements to obtain a voting score of each of the second candidate database statements; the voting score represents the credibility of the second candidate database statement; For each of the second candidate database statements, respectively, calculate a confidence score corresponding to the second candidate database statement according to the similarity score and the voting score; The second candidate database statement with the highest confidence score is determined as the target database statement.

9. A database statement generation device based on a large language model, characterized in that: The device comprises: A generation module is used to obtain a sentence to be processed input by a user, and generate multiple variant sentences corresponding to the sentence to be processed through a large language model; the sentence to be processed and the variant sentences are both natural languages, and the variant sentences have different grammatical structures from the sentence to be processed, but have consistent semantic logic; The generating module is further configured to obtain a target data structure from a database according to the plurality of variant sentences through the large language model, and generate a database sentence corresponding to each variant sentence according to each variant sentence and the target data structure; The generation module is further configured to input each variant statement into the large language model for processing to obtain a first semantic structure corresponding to each variant statement; obtain a target table and a target field corresponding to each variant statement from the database according to the first semantic structure corresponding to each variant statement, and obtain a target data structure by taking a union of the target table and the target field; input the target data structure and each variant statement into the large language model for processing to obtain a database statement corresponding to each variant statement; The determination module is used to determine a target database statement from the multiple database statements; the target database statement is a database statement that can be successfully executed and has the highest confidence.

10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program executable by the processor, and the processor can execute the computer program to implement the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Similar statement generation method and device based on pre-training language model

    CN113807074A

  • SQL (Structured Query Language) statement generation method and device based on large language model, equipment and medium

    CN117743371A