Statement generation method and device for structured query language, equipment and medium

By pre-constructing syntax alignment corpus and hybrid expert network model, the syntax normative problem of cross-database systems is solved, the accuracy and efficiency of SQL statement generation is improved, and the needs of complex query are adapted to.

CN120296034APending Publication Date: 2025-07-11TRANSWARP TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510467209.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to ensure syntax normability when generating structured query language statements across different database systems, resulting in query execution failure and training large models is expensive and ineffective.

Method used

Through the pre-constructed syntax alignment corpus and mixed expert network model, the standard alignment learning efficiency and accuracy of the statement generation model are improved, and the input data is dynamically allocated to the expert network using rule routing and gated networks to generate SQL statements that comply with specific syntax specifications.

Benefits of technology

It improves the syntax compliance and accuracy of statement generation, reduces the complexity and cost of model training, and enhances the adaptability to complex queries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296034A_ABST
    Figure CN120296034A_ABST
Patent Text Reader

Abstract

The invention discloses a statement generation method and device of a structured query language, equipment and a storage medium. The statement generation method comprises the steps of obtaining user query information; outputting a target statement generation result corresponding to the user query information based on the user query information and a statement generation model of a structured query language; the statement generation model of the structured query language is generated by training based on different historical user query information and corresponding statement generation results extracted from a pre-constructed grammar alignment corpus according to the historical user query information. By means of the method, the efficiency and precision of standard alignment learning of the statement generation model are improved through the pre-constructed grammar alignment corpus, the grammar following ability during statement generation can be effectively improved through the trained statement generation model of the structured query language, and the grammar accuracy of statements is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of artificial intelligence, and in particular, to a method, device, equipment and storage medium for generating statements of Structured Query Language. Background Art

[0002] Natural Language to Structured Query Language (NL2SQL) is a technology that converts natural language text into Structured Query Language (SQL), aiming to enable non-technical users to query databases using natural language, reducing the threshold for non-professionals to enter the database field and promoting the popularization of data decision-making. However, despite the great potential of the statement generation technology of Structured Query Language, it still faces many challenges in practical applications. For example, different syntax rules and characteristics of different database systems increase the complexity of the statement generation method of Structured Query Language, making it difficult to ensure that the generated SQL statements comply with the specifications, resulting in query execution failures.

[0003] The existing technical solutions to solve the above problems can be classified into two categories: one is to support the generation of multiple SQL specifications by relying on rules, templates, knowledge, etc. For example, the rule engine method: guiding the generation of SQL statements through predefined rules, this method can ensure that the generated SQL statements comply with specific syntax specifications, but the cost of formulating and maintaining the rules is relatively high, and it is difficult to adapt to complex query requirements. Another example is the template matching method: using preset templates to generate SQL statements, this method is more effective when dealing with common query patterns, but it may not be flexible enough for non-standardized or complex queries. Therefore, this type of method relies on a large amount of manual work, with complex configuration and difficulty in being comprehensive. The other category is to support multiple specifications based on model learning. For example, the model fine-tuning method: the model learns SQL syntax specifications from a large amount of SQL statement data, this method can better adapt to different SQL syntax query requirements, but it may require a large amount of training data, and it is difficult to train models that support multiple coding specifications. Another example is the syntax enhancement method: combining domain expert knowledge and large model technology, constructing a system that can understand complex query requirements and generate compliant SQL statements through the RAG technology, this method performs well when dealing with complex queries, but the construction and maintenance costs are relatively high. Therefore, due to the large commonality and small differences of different SQL specifications, the model learning effect of this type of method is not good.

[0004] Therefore, how to enable the trained large model to accurately generate SQL statements with specified syntax specifications when the commonality of different SQL specifications is large and the differences are small has become an urgent task. Summary of the Invention

[0005] An embodiment of the present invention provides a method, apparatus, device, and storage medium for generating Structured Query Language (SQL) statements. By using a pre-constructed syntax alignment corpus, the efficiency and accuracy of the expert network for canonical alignment learning are improved. Through the trained SQL statement generation model, the ability to follow syntax during statement generation can be effectively enhanced, and the syntax accuracy of the statements can be improved.

[0006] In a first aspect, an embodiment of the present invention provides a method for generating Structured Query Language (SQL) statements, including:

[0007] Obtaining user query information;

[0008] Based on the user query information and the SQL statement generation model, outputting a target statement generation result corresponding to the user query information; the SQL statement generation model is trained based on different historical user query information and the corresponding statement generation results extracted from a pre-constructed syntax alignment corpus according to each historical user query information.

[0009] In a second aspect, an embodiment of the present invention further provides a device for generating Structured Query Language (SQL) statements, including:

[0010] An information acquisition module, configured to obtain user query information;

[0011] A statement generation result module, configured to output a target statement generation result corresponding to the user query information based on the user query information and the SQL statement generation model; the SQL statement generation model is trained based on different historical user query information and the corresponding statement generation results extracted from a pre-constructed syntax alignment corpus according to each historical user query information.

[0012] In a third aspect, an embodiment of the present disclosure further provides an electronic device, including:

[0013] One or more processors;

[0014] A storage device, configured to store one or more programs,

[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating Structured Query Language (SQL) statements provided by the embodiments of the present disclosure.

[0016] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium containing computer-executable instructions, where the computer-executable instructions are used to execute the method for generating Structured Query Language (SQL) statements provided by the embodiments of the present disclosure when executed by a computer processor.

[0017] Fifth aspect, an embodiment of the present disclosure provides a computer program product, which includes a computer program that, when executed by a processor, implements the statement generation method of the structured query language provided in the embodiment of the first aspect above.

[0018] The present invention discloses a statement generation method, device, equipment and storage medium for structured query language, which improves the efficiency and accuracy of the statement generation model for canonical alignment learning through a pre-constructed syntax alignment corpus. Through the trained statement generation model for structured query language, the ability to follow syntax during statement generation can be effectively improved, and the syntax accuracy of statements can be increased. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent. Throughout the accompanying drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the original components and elements are not necessarily drawn to scale.

[0020] Figure 1 It is a flowchart of a statement generation method for a structured query language provided by an embodiment of the present disclosure;

[0021] Figure 2 It is a flowchart of a statement generation method for a structured query language provided by the second embodiment of the present disclosure;

[0022] Figure 3 It is a flowchart of a construction process of a syntax alignment corpus provided by the present embodiment of the present disclosure;

[0023] Figure 4 It is a flowchart of generating a second user query information-structured query language statement pair for other syntax specifications except the seed specification provided by the present embodiment of the present disclosure;

[0024] Figure 5 It is a flowchart of a training method for a statement generation model of a structured query language provided by the present embodiment of the present disclosure;

[0025] Figure 6 It is a schematic structural diagram of a basic model of a statement generation model for a structured query language provided by the present embodiment of the present disclosure;

[0026] Figure 7 It is a flowchart of a training method for a basic model of a statement generation model for a structured query language provided by the present embodiment of the present disclosure;

[0027] Figure 8 It is a schematic structural diagram of a statement generation device for a structured query language provided by an embodiment of the present disclosure;

[0028] Figure 9 The structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. Specific implementation manners

[0029] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not used to limit the protection scope of the present disclosure.

[0030] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0031] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0032] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0033] It can be understood that the data involved in the technical solution of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related regulations.

[0034] Embodiment 1

[0035] Figure 1 The flowchart of the statement generation of a structured query language provided by an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to the situation of providing a solution to the problem that SQL statements that cannot accurately generate specified grammar specifications. This method can be executed by a statement generation device for a structured query language, and this device can be implemented in the form of software and / or hardware.

[0036] As Figure 1 shown, a method for generating a statement of a structured query language provided by an embodiment of the present disclosure may specifically include the following steps:

[0037] S101. Obtain user query information.

[0038] In this embodiment, the user query information may be natural language text directly input by the user or natural language text generated by converting the voice input by the user. The natural language text includes a functional description implemented by a generated SQL statement. Among them, the user query information may include information indicating that the generated SQL statement meets a specific SQL specification according to the requirements, or may not include information on a specific SQL specification.

[0039] Specifically, the method for obtaining user query information may include: obtaining the information obtained by parsing the user's voice or text input through a voice assistant as the user query information; or obtaining the query information input by the user in software or application programs for generating SQL statements.

[0040] S102. Output a target statement generation result corresponding to the user query information based on the user query information and a statement generation model of the structured query language; wherein, the statement generation model of the structured query language is trained and generated based on different historical user query information and the corresponding statement generation results extracted from a pre-constructed syntax alignment corpus according to each historical user query information.

[0041] In this embodiment, the statement generation model of the structured query language may be a generative artificial intelligence model for outputting a target statement generation result corresponding to the input user query information. The statement generation model of the structured query language may include: an input layer, a rule router, a mixture of experts network, and an output layer. Among them, the input layer is used for the model to input data, and the output layer is used for outputting the result of the model. The rule router is connected to the mixture of experts network. The mixture of experts network includes a gating network and at least two expert networks. The expert network can be SQL domain experts with various different structures. The rule router may be a data allocation mechanism based on preset logical conditions for allocating the input data of the model to each expert network in the mixture of experts network. The output end of the gating network is connected to each expert network. The gating network will determine which expert networks the data screened by the rule router is allocated to for processing. The specific allocation strategy of the gating network dynamically learns the mapping relationship between the input features and the expert weights. The mixture of experts network may also be replaced by a modular neural network. Multiple expert networks or multiple modular neural networks are integrated in the statement generation model of the structured query language. By dynamically selecting some networks to process the input through the rule router and the gating network, a large-scale parameter pool is retained to ensure the model's ability, while reducing the actual computational amount, achieving a balance between computational efficiency and model ability.

[0042] Continuing from the above, the training dataset for the statement generation model to output user query information consists of the user query information and the statement generation results extracted from the pre-built syntax alignment corpus corresponding to the user query information. The syntax alignment corpus can be pre-built, and it includes: historical user query information, and SQL statements with different syntax specifications corresponding to the historical user query information. The historical user query information can be the user's historical query information, that is, the natural language text entered by the user at a historical moment for generating SQL statements. The target statement generation result can be the SQL statement that the user aims to generate, and this SQL statement can be put into the SQL engine for execution and implement the function corresponding to what the user query information describes or requires.

[0043] Specifically, input the user query information into the statement generation model of the structured query language. The statement generation model of the structured query language outputs the target statement generation result corresponding to the user query information. Further, the user query information can also be input into a general natural language processing model or a natural language processing model in the SQL field. Through the above two models, the user query information is processed to output a query text that is more easily recognized by the model. Then, input the query text into the statement generation model of the structured query language to output the target statement generation result corresponding to the user query information.

[0044] The technical solution of the embodiment of the present invention can directly predict and generate the target statement generation result corresponding to the user query information through the trained statement generation model of the structured query language, effectively improving the syntax compliance ability and syntax accuracy during statement generation. At the same time, the pre-built syntax alignment corpus improves the efficiency and accuracy of the statement generation model for specification alignment learning.

[0045] Embodiment Two

[0046] Figure 2 FIG. is a flowchart of a method for generating a structured query language statement provided in Embodiment Two of the present invention. It is further optimized and extended based on the above embodiments and can be combined with various optional technical solutions in the above embodiments. As Figure 2 shown, a method for generating a structured query language statement provided in Embodiment Two of the present invention specifically includes the following steps:

[0047] S201. Obtain user query information.

[0048] S202. Obtain the pre-built structured query language generation prompt instruction. Among them, the preset role of the structured query language generation prompt instruction is a structured query language expert who knows the syntax alignment corpus; the structured query language generation prompt instruction is an instruction for generating prompt words of the structured query language.

[0049] S203. Use the user query information as the operation parameter for generating the prompt instruction of the structured query language to generate the prompt word of the structured query language.

[0050] In this embodiment, the structured query language generation prompt instruction can be an instruction information for generating the prompt word of the structured query language, which is used to generate the prompt word of the structured query language. The prompt word can guide the generative artificial intelligence model to perform the cognitive task of structured query language generation through natural language or code. The preset role of the structured query language generation prompt instruction can be an instruction parameter that endows the structured query language expert identity of the known grammar alignment corpus through semantic constraints and knowledge boundaries, and is used to control the output style and professional depth. The operation parameter can be an industry specialization parameter for structured query language generation, and the operation parameter includes the user's required function and the specified grammar specification for realizing the required function. The structured query language generation prompt can enable the statement generation model of the structured query language to better identify the input data.

[0051] Specifically, obtain the structured query language generation prompt instruction pre-built for generating the prompt word of the structured query language, fill the user query information into the template of the structured query language generation prompt instruction as the operation parameter, and use the filled template as the prompt word of the structured query language. Preferably, the user query information can also be parsed in natural language first, the parsing result can be converted into machine-readable operation parameters, and then the machine-readable operation parameters can be filled into the template of the structured query language generation prompt instruction.

[0052] Exemplarily, the template format of the pre-built Structured Query Language (SQL) generation prompt instruction is "You are now an SQL expert. Given a syntax-aligned corpus {input}, please write for me..." When the user's query information is "Generate an SQL statement for {instruction} with syntax specification 1", perform natural language parsing. The parsing result is "An SQL statement for {instruction} that conforms to SQL syntax specification 1". Fill the parsing result into the template of the SQL generation prompt instruction, and the output is "You are now an SQL expert. Given a syntax-aligned corpus {input}, please write for me an SQL statement for {instruction} that conforms to SQL syntax specification 1". When there are no specification requirements in the user's query information, the final output prompt can be: "You are now an SQL expert. Given a syntax-aligned corpus {input}, please write for me an SQL statement for {instruction}". When there are specific specification requirements in the user's query information, such as the specific specification requirement is SQL syntax specification 1, the input data at this time is beneficial for the SQL statement generation model to select the correct expert and improve the accuracy of generation. The form can be: "You are now an SQL expert. Given a syntax-aligned corpus {input}, please write for me an SQL statement for {instruction} that conforms to SQL syntax specification 1".

[0053] S204. Input the SQL prompt into the SQL statement generation model to output the target statement generation result.

[0054] Specifically, input the SQL prompt into the SQL statement generation model. The SQL statement generation model outputs the target statement generation result according to the SQL prompt.

[0055] Using this method, generate instructions through pre-built SQL prompts to guide the generative statement generation model to correctly understand the user's intention and improve the accuracy of the model output.

[0056] Based on the above embodiments, an SQL statement generation method further includes: the construction process of the syntax-aligned corpus. Figure 3 This is a flowchart of the construction process of the syntax-aligned corpus provided by the embodiments of the present invention, as Figure 3 shown, and specifically includes the following steps:

[0057] S301. Obtain at least two SQL syntax specifications.

[0058] In this embodiment, the syntax specification may be the basic framework and specification scope of the Structured Query Language. The Structured Query Language may include: American National Standards Institute Structured Query Language, Transactional Structured Query Language, Procedural SQL Language, Multi-Model Extended SQL, and Graph Query Language, etc.

[0059] Exemplarily, the syntax specification of the Procedural SQL Language in the Structured Query Language includes block structures and control flow statements, etc. The syntax specification of the Transactional SQL in the Structured Query Language includes transaction basic syntax, concurrency control and isolation levels, and distributed transactions.

[0060] Specifically, obtain a configuration file or document configured with the syntax specification of the Structured Query Language, read the configuration file or document, and obtain at least two syntax specifications of the Structured Query Language. The way to obtain the configuration file or document can be: for example, download the syntax specification of the Structured Query Language from the official website and configure it into a configuration file or document; or obtain the syntax specification through a developer community such as an interactive learning platform and configure it into a configuration file or document; or implement it through a standardized SQL parser, which can reversely learn the syntax rules, obtain the syntax specification, and configure it into a configuration file or document.

[0061] S302. Use the syntax specification of any one Structured Query Language as the seed specification.

[0062] In this embodiment, the seed specification may be the specification used as a benchmark.

[0063] This step is used to select the syntax specification of one Structured Query Language as the seed specification. Among them, any one can be arbitrarily selected as the seed specification. Preferably, the one with the most comprehensive commonality with other multiple specifications can be selected as the seed specification.

[0064] S303. Extract the syntax features of each syntax specification respectively, and determine the differences between the seed specification and the other syntax specifications except the seed specification according to the syntax features.

[0065] In this embodiment, the syntax feature may be the formal rule shown by the syntax specification in the combination and aggregation relationships, including attributes such as keywords, syntactic structures, and semantic constraints. The difference may be the differences between the seed specification and other syntax specifications. The differences may be specifically manifested in types such as keywords, function methods, and syntax structure characteristics. The syntax features of the syntax specifications of different Structured Query Languages are not exactly the same. Of course, there are also many common principles, such as the transaction control syntax being highly consistent, etc.

[0066] Specifically, the method for extracting the syntactic features of each syntax specification may include at least one of the following: for example, extracting the syntactic features of each syntax specification through a static analysis tool chain; or extracting the syntactic features of each syntax specification through a dynamic analysis method or quantum syntax analysis. And the differences between the seed specification and other syntax specifications except the seed specification are determined by comparing the seed specification with other syntax specifications one by one.

[0067] S304. Generate a statement in the first structured query language of the seed specification corresponding to the historical user query information according to the historical user query information and the syntactic features of the seed specification.

[0068] In this embodiment, the statement in the first structured query language may be an SQL statement corresponding to the historical user query information under the syntax of the seed specification.

[0069] Specifically, taking the generated SQL statement corresponding to the historical user query information as the statement in the first structured query language according to the syntactic features of the seed specification can be generated manually or automatically. The way to automatically generate the statement in the first structured query language may be: first perform entity recognition on the historical user query information through the SQL generation method, and then complete rule engine conversion and operator selection according to the historical user query information to generate an SQL statement. It can also be that the latest neuro-symbolic engine automatically generates an SQL statement corresponding to the historical user query information under the syntax of the seed specification, where the neuro-symbolic engine is a model based on the self-attention mechanism injected with logical programming language rules.

[0070] Based on this embodiment, it is also possible to perform verification on the generated statement in the first structured query language based on an automated method, such as putting the generated statement in the first structured query language into an SQL engine for execution and comparing whether the result is correct with the expected result.

[0071] S305. Generate a statement in the second structured query language of the other syntax specifications except the seed specification according to the differences between the seed specification and the other syntax specifications.

[0072] In this embodiment, the statement in the second structured query language may be a structured query language statement corresponding to the historical user query information under the syntax of the other syntax specifications.

[0073] Specifically, according to the differences between the seed specification and the other syntax specifications, some statements in the statement in the first structured query language are replaced, and the statements common to the seed specification and the other syntax specifications are retained, and finally a statement in the second structured query language of each syntax specification except the seed specification is generated.

[0074] Exemplarily, in this step, the method of generating by using a pre-trained large model can be adopted. The differences between the seed specification and other syntax specifications are used as enhancement information to guide the generation of other syntax specification statements.

[0075] Further, based on the above-mentioned invention embodiments, the method for determining the differences between the seed specification and other syntax specifications except the seed specification according to each syntax feature includes:

[0076] Compare the syntax features in the seed specification with the corresponding syntax features in other syntax specifications; if the syntax features in the seed specification are inconsistent with the corresponding syntax features in other syntax specifications, then use the syntax feature in the seed specification as a differential syntax feature, and determine the difference points between the differential syntax feature and the corresponding syntax features of other syntax specifications; the types of difference points include at least one of the following: keywords; function methods; syntax structure characteristics.

[0077] In this embodiment, the differential syntax feature can be the syntax feature that the seed specification is different from other syntax specifications, and can be determined by comparing the syntax features in the seed specification with the corresponding syntax features in other syntax specifications one by one. During the comparison process, the commonalities and differences between the syntax features in the seed specification and other syntax specifications can be determined. The commonalities can be the places where each syntax specification is consistent in use, including: all databases follow the basic query structure of the American National Standards Institute SQL standard and the transaction control syntax is highly consistent. The difference points can be the different points in multiple aspects relative to the seed specification and the corresponding mapping relationships. Such as keywords, function methods, and syntax structure characteristics, etc.

[0078] S306. Use the historical user query information, the statements of the first structured query language, and the statements of the second structured query language as the content in the syntax alignment corpus.

[0079] Specifically, make the statements of the first structured query language corresponding to the historical user query information be in one-to-one or one-to-many correspondence with the statements of the second structured query language, and store the corresponding historical user query information, the statements of the first structured query language, and the statements of the second structured query language into the syntax alignment corpus. The syntax alignment corpus can also include metadata descriptions of data tables or libraries, as well as enhancement information such as thought chains and background knowledge. The metadata descriptions and enhancement information can be used as parameter information to generate a prompt word for generating training samples when generating training set samples.

[0080] Further, based on the above-mentioned invention embodiments, to generate the statement pairs of the second user query information - structured query language for the remaining syntax specifications except the seed specification according to the syntax features, as Figure 4 shown, the following steps are further included:

[0081] S3051. Use the statement content corresponding to the differential syntax feature in the statement of the first Structured Query Language as the first statement.

[0082] S3052. Call the syntax features of other syntax specifications according to the differential points to generate corresponding replacement information, and replace the differential points with the replacement information within the first statement, and use the first statement after replacement as the second statement.

[0083] S3053. Replace the first statement in the statement of the first Structured Query Language with the second statement to obtain the statement of the second Structured Query Language.

[0084] In this embodiment, the first statement may be part of the statement content in the statement of the first Structured Query Language, and this part of the statement content is the statement content corresponding to the differential syntax feature. The second statement may be an SQL statement generated according to the syntax features of other syntax specifications corresponding to the differential syntax feature. The replacement information may be the SQL statement content for replacing the statement content corresponding to the differential syntax feature.

[0085] Specifically, first extract the statement content corresponding to the differential syntax feature in the statement of the first Structured Query Language as the first statement, then obtain the syntax features of other syntax specifications corresponding to the differential syntax feature in the statement of the first Structured Query Language according to the above differential points, generate corresponding replacement information according to the syntax features of other syntax specifications, replace the differential points with the replacement information within the first statement, use the first statement after replacement as the second statement, and finally replace the first statement in the statement of the first Structured Query Language with the second statement to obtain the statement of the second Structured Query Language.

[0086] Exemplarily, when the differential syntax feature is a keyword or a function, obtain the syntax features of other syntax specifications to generate the syntax for this part of the SQL statement. Among them, the syntax feature may also be a keyword or a function. When the differential syntax feature is a structure, obtain the structure of other syntax specifications to generate this part of the SQL statement.

[0087] Based on the above embodiment, a method for generating a statement of a Structured Query Language further includes: a training process of a statement generation model of a Structured Query Language, as Figure 5 shown, the training process of the statement generation model of a Structured Query Language includes:

[0088] S401. Divide the training dataset into a training set and a test set according to a preset dataset splitting ratio, and initialize and construct an initial model of a structured query language statement generation model. Among them, the initial model of the structured query language statement generation model includes a rule router and a mixture of experts network; the rule router is connected to the mixture of experts network; the mixture of experts network includes a gating network and at least two expert networks; the output end of the gating network is connected to each expert network; the rule router is used to allocate the input data of the initial model of the structured query language statement generation model to each expert network in the mixture of experts network according to predefined rule conditions.

[0089] In this embodiment, the preset dataset splitting ratio can be a preset splitting ratio, which is set according to the actual situation. The data in the preset dataset can be obtained from a syntactic alignment corpus, and the obtaining method can be manual or generated by a model. Exemplarily, the data in the training set of the structured query language statement generation model can be generated by a trained artificial intelligence model. The input of the model can use the following natural language description, for example: Please act as a database query assistant. Your task is to generate an SQL statement that conforms to the syntax specification according to the given historical user query information and enhancement class information. The initial model of the structured query language statement generation model can be a model that is built but not trained. The structure of the initial model of the structured query language statement generation model successively includes an input layer, a rule router, a mixture of experts network, and an output layer. Among them, the input layer is used for the model to input data, and the output layer is used to output the results of the model. The rule router is connected to the mixture of experts network. The mixture of experts network includes a gating network and at least two expert networks. The rule router can be a data allocation mechanism based on preset logical conditions, which is used to allocate the input data of the model to each expert network in the mixture of experts network. The output end of the gating network is connected to each expert network. The gating network will determine which expert networks the data filtered by the rule router is allocated to for processing. The specific allocation strategy of the gating network dynamically learns the mapping relationship between the input features and the expert weights.

[0090] Specifically, divide the training dataset into a training set and a test set according to a preset dataset splitting ratio, and initialize and construct an initial model of a structured query language statement generation model.

[0091] Exemplarily, Figure 6 is a schematic structural diagram of a basic model of a structured query language statement generation model provided in this embodiment of the present disclosure, as Figure 6As shown, the basic model of the statement generation model of Structured Query Language includes rule routing, the gating network in the Mixture of Experts network, and at least two expert networks. The rule routing and the gating network are connected to at least two expert networks. The rule routing and the gating network are jointly responsible for distributing the input data to each expert network. The priority of the rule routing is higher than that of the gating network. Therefore, first, the rule routing distributes the input data to each expert network in the Mixture of Experts network according to the predefined rule conditions. When the predefined rule conditions are not met, the gating network decides which expert networks to process different data. The specific distribution strategy of the gating network determines which expert networks to process different samples by dynamically learning the mapping relationship between the input features and the expert weights. Among them, the predefined rule conditions are that the input data meeting specific features correspond to the specific expert networks to which they are input.

[0092] S402. Train the initial model of the statement generation model of Structured Query Language according to the training set to obtain the basic model of the statement generation model of Structured Query Language.

[0093] In this embodiment, the basic model of the statement generation model of Structured Query Language can be a basic model that iterates multiple rounds on the training set to gradually optimize the model parameters.

[0094] Specifically, through sparse gradient update, parameter sharing and regularization, or other optimization algorithms, select a suitable optimization function to adjust the model parameters of the Mixture of Experts network. Among them, the model parameters include: the routing network parameters of the gating network and the network model parameters of each expert network. Iterate multiple rounds on the training set to gradually optimize the model parameters to obtain the basic model of the statement generation model of Structured Query Language.

[0095] S403. Perform a performance test on the basic model of the statement generation model of Structured Query Language using the test set; if the test result meets the preset model test passing conditions, use the basic model of the statement generation model of Structured Query Language as the final statement generation model of Structured Query Language, otherwise retrain the initial model of the statement generation model of Structured Query Language according to the training set.

[0096] Specifically, use the test set to evaluate the performance of the basic model of the statement generation model of Structured Query Language, such as task accuracy, generalization ability, and computational efficiency, etc. The preset test passing conditions include but are not limited to prediction error threshold, convergence speed, and resource consumption, etc. If the test result meets the passing conditions, use the basic model of the statement generation model of Structured Query Language as the statement generation model of Structured Query Language; otherwise, retrain the initial model until the conditions are met.

[0097] On the basis of the above embodiment, the initial model of the structured query language statement generation model is trained according to the training set to obtain the basic model of the structured query language statement generation model, such as Figure 7 As shown, the training steps include:

[0098] S4031. Input the training set into the initial model of the statement generation model of the structured query language.

[0099] S4032. Determine whether the data in the input training set triggers the rule of rule routing. If the rule of rule routing is triggered, skip the gating network and directly distribute the data in the training set to the corresponding expert network according to the predefined rule conditions. If the rule of rule routing is not triggered, the data in the training set is distributed to the corresponding expert network by the gating network.

[0100] In this embodiment, the rule routing and the gated network have a higher priority in allocating data to each expert network. When there are standard requirements in the user query instruction, the rule routing is prioritized for allocation. When there are no standard requirements in the user query instruction, it is allocated to the expert network through the gated network.

[0101] In this method, when the input data can be controlled by rule routing or gated network and input into different expert networks, in actual implementation, either rule routing or gated network can be selected, or both rule routing and gated network can be used in combination. Among them, the specific implementation method of rule routing is to first construct a rule mapping table, and the mapping relationship in the rule mapping table is respectively the specification requirements extracted from the input and the number of expert networks selected corresponding to the specification requirements. The difference between the gated network and rule routing is that there is no need to pre-define the rule mapping table, and the expert network is selected by parsing the set field content through one-hot encoding mapping and top K screening technology. When there is no specification requirement, general processing is performed, and when there is a specification requirement, the expert network is selected by the top K screening technology.

[0102] Specifically, the user query instruction is input into the statement generation model of the structured query language. Whether the rule routing rule is triggered is determined based on whether there are specification requirements in the user query instruction. When there are specification requirements, the rule routing rule is triggered, and the gated network routing is skipped to allocate the input data to the corresponding specification expert network. When there are no specification requirements, the gated network is used to allocate the input data to the expert network. Among them, the gated network can be used to allocate to a set expert network, or to a random expert network. The allocation result is completely constrained by the input data and model parameters. In theory, the same input will generate the same expert allocation result under fixed parameters.

[0103] S4033. Use the weighted function of the expert loss function, the syntax specification discrimination function, and the perplexity loss function according to the preset weight values as the loss function, and optimize the model parameters with the goal of minimizing the loss function to obtain the basic model of the structured query language statement generation model.

[0104] Among them, the expert loss function is a function that quantifies the prediction error by calculating the cross-entropy between the predicted probability distribution output by each expert network and the true label; the syntax specification discrimination function is a contrast loss function whose optimization goal includes minimizing the distance between sample pairs with the same syntax specification in the training set and maximizing the distance between sample pairs with different syntax specifications in the training set; the perplexity loss function is a function that calculates the average branching factor of the probability distribution output by the structured query language generation model.

[0105] In this embodiment, when using rule routing to assign experts, that is, assigning experts through predefined rules, there is no parameter dependence, so training is not required. When assigning experts through a gating network, the rule dependence can be configured with parameters (such as thresholds, weights), and adjustment is required through training. The loss function of this model is the weighted function of the expert loss function, the syntax specification discrimination function, and the perplexity loss function according to the preset weight values as the loss function. The specific setting of the weight values is determined according to the actual situation and is not specifically limited in this embodiment.

[0106] Exemplarily, for the expert loss function: for each expert network, the cross-entropy loss is used to calculate the output loss of the expert, which can effectively measure the difference between the probability distribution predicted by the model and the true label. The form of the expert loss function L1 is as follows:

[0107]

[0108] where y i represents the true value of the i-th label, and p i represents the probability of the i-th label predicted by the model, and N represents the number of samples participating in the calculation of the loss.

[0109] The syntax specification discrimination function: used to ensure that the feature representation learned by the model can distinguish data under different syntax specifications, and the contrast loss function L2 is used to minimize the distance of samples with the same syntax specification and maximize the distance of samples with different syntax specifications at the same time.

[0110]

[0111] Among them, d represents the Euclidean distance of the sample features, y represents the label of whether the samples match, margin represents the set threshold, and N is the number of samples participating in the calculation of the loss.

[0112] Perplexity loss function L3: Used to measure the performance of the model under specific grammar specifications. The uncertainty of the model is measured by calculating the average branching factor of the probability distribution output by the model.

[0113]

[0114] Among them, p i represents the probability of the i-th label predicted by the model, and N represents the number of samples participating in the calculation of the loss.

[0115] The total loss function L4 is the weighted sum of the above three loss functions, specifically:

[0116] L4 = αL1 + βL2 + γL3

[0117] Among them, α, β, and γ represent the first weight parameters, used to balance the contributions of different losses.

[0118] Based on the above embodiments, for the data containing different grammar specifications in the same batch of samples, the loss function can be further improved as follows: The weighted sum of the defined multi-label cross-entropy loss and Hamming loss function is used as the loss function.

[0119] Exemplarily, define the multi-label cross-entropy loss L5, calculate the binary classification loss of each label, and add them up to obtain the overall loss.

[0120]

[0121] Among them, y i represents the true value of the i-th label, p i represents the probability of the i-th label predicted by the model, and σ represents the Sigmoid function.

[0122] The Hamming loss function measures the proportion of the number of mispredicted labels in the total number of labels among all samples. For data containing different grammar specifications, the Hamming loss function L6 can intuitively represent the error rate of the model on each label.

[0123]

[0124] Among them, m represents the number of samples, q represents the number of categories, and I represents the indicator function. When the predicted value of the j-th category is different from the true value y j of the j-th category, take 1, otherwise take 0.

[0125] The total loss function L7 is the weighted sum of the above three loss functions, specifically:

[0126] L7 = θL5 + μL6

[0127] Among them, θ and μ represent the second weight parameters, which are used to balance the contributions of different losses. And the training accuracy of different grammar specifications is optimized by adjusting the weights. For example, if the prediction error of a certain grammar specification has a greater impact on the final result, the weight of the corresponding loss function can be increased.

[0128] Embodiment III

[0129] Figure 8 This embodiment of the present invention also provides a schematic structural diagram of a statement generation device for a structured query language, as Figure 8 shown, the device includes: an information acquisition module 501 and a statement generation result module 502.

[0130] The information acquisition module 501 is used to acquire user query information;

[0131] The statement generation result module 502 is used to output the target statement generation result corresponding to the user query information based on the user query information and the statement generation model of the structured query language; the statement generation model of the structured query language is trained and generated based on different historical user query information and the corresponding statement generation results extracted from the pre-constructed grammar alignment corpus according to each historical user query information.

[0132] The technical solution provided by this embodiment of the present disclosure improves the efficiency and accuracy of the statement generation model for canonical alignment learning through the pre-constructed grammar alignment corpus. Through the trained statement generation model of the structured query language, the grammar compliance ability and grammar accuracy during statement generation are effectively improved.

[0133] Further, the device further includes: a construction module for the grammar alignment corpus;

[0134] The construction module for the grammar alignment corpus is used for:

[0135] acquiring at least two grammar specifications of the structured query language;

[0136] taking any one of the grammar specifications of the structured query language as the seed specification;

[0137] extracting the grammar features of each grammar specification respectively, and determining the differences between the seed specification and the other grammar specifications except the seed specification according to the grammar features;

[0138] generating statements of the first structured query language corresponding to the historical user query information according to the historical user query information and the grammar features of the seed specification;

[0139] generating statements of the second structured query language of the other grammar specifications except the seed specification according to the differences between the seed specification and the other grammar specifications;

[0140] Use the historical user query information, the statement of the first structured query language, and the statement of the second structured query language as the content in the syntax alignment corpus.

[0141] Furthermore, the construction module of the syntax alignment corpus is also used for:

[0142] Compare the syntax features in the seed specification with the corresponding syntax features in the other syntax specifications;

[0143] If the syntax features in the seed specification are inconsistent with the corresponding syntax features in the other syntax specifications, then use the syntax features in the seed specification as the differential syntax features, and determine the difference points between the differential syntax features and the corresponding syntax features in the other syntax specifications; the types of the difference points include at least one of the following: keywords; function methods; syntax structure characteristics.

[0144] Furthermore, the construction module of the syntax alignment corpus is also used for:

[0145] Retain the first statement content corresponding to the first identifier in the first user query information-structured query language statement;

[0146] Use the statement content corresponding to the differential syntax feature in the first structured query language statement as the first statement;

[0147] Call the syntax feature of the other syntax specification according to the difference point to generate the corresponding replacement information, and replace the difference point with the replacement information in the first statement, and use the first statement after replacement as the second statement;

[0148] Replace the first statement in the first structured query language statement with the second statement to obtain the second structured query language statement.

[0149] Furthermore, the statement generation result module 502 can also be used for:

[0150] Obtain a pre-built structured query language generation prompt instruction; wherein, the preset role of the structured query language generation prompt instruction is a structured query language expert who knows the syntax alignment corpus; the structured query language generation prompt instruction is an instruction for generating a prompt word of the structured query language;

[0151] Use the user query information as the operation parameter of the structured query language generation prompt instruction to generate a prompt word of the structured query language;

[0152] Input the prompt words of the structured query language into the statement generation model of the structured query language to output the target statement generation result.

[0153] Furthermore, the device further includes: a training module for the statement generation model of the structured query language;

[0154] The training module for the statement generation model of the structured query language can be used for:

[0155] Divide the training dataset into a training set and a test set according to a preset dataset division ratio, and construct an initial model of the statement generation model of the structured query language; wherein, the initial model of the statement generation model of the structured query language includes a rule router and a mixture-of-experts network; the rule router is connected to the mixture-of-experts network; the mixture-of-experts network includes a gating network and at least two expert networks; the output end of the gating network is connected to each of the expert networks; the rule router is used to allocate the input data of the initial model of the statement generation model of the structured query language to each of the expert networks in the mixture-of-experts network according to predefined rule conditions;

[0156] Train the initial model of the statement generation model of the structured query language according to the training set to obtain a basic model of the statement generation model of the structured query language;

[0157] Perform performance testing on the basic model of the statement generation model of the structured query language using the test set; if the test result meets the preset model test passing condition, then use the basic model of the statement generation model of the structured query language as the final statement generation model of the structured query language, otherwise retrain the initial model of the statement generation model of the structured query language according to the training set.

[0158] Furthermore, the training module for the statement generation model of the structured query language can be used for:

[0159] Input the training set into the initial model of the statement generation model of the structured query language;

[0160] Judge whether the data in the input training set triggers the rules of the rule router. If the rules of the rule router are triggered, skip the gating network and directly allocate the data in the training set to the corresponding expert network according to the predefined rule conditions. If the rules of the rule router are not triggered, allocate it to the corresponding expert network by the gating network;

[0161] The function obtained by weighting the expert loss function, the syntax specification discrimination function, and the perplexity loss function according to preset weight values is used as the loss function, and the model parameters are optimized with the goal of minimizing the loss function to obtain the basic model of the statement generation model for the structured query language; wherein, the expert loss function is a function that quantifies the prediction error by calculating the cross-entropy between the predicted probability distribution output by each expert network and the true label; the syntax specification discrimination function is a contrast loss function whose optimization goal includes minimizing the distance between sample pairs with the same syntax specification in the training set and maximizing the distance between sample pairs with different syntax specifications in the training set; the perplexity loss function is a function that calculates the average branching factor of the probability distribution output by the structured query language generation model.

[0162] The above device can execute the methods provided in all the foregoing embodiments of the present invention, and has corresponding functional modules and beneficial effects for executing the above methods. For technical details not described in detail in this embodiment, reference may be made to the methods provided in all the foregoing embodiments of the present invention.

[0163] Embodiment 4

[0164] Figure 9 FIG. 10 shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0165] As Figure 9 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program executable by at least one processor, and the processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0166] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0167] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for generating statements of structured query language.

[0168] In some embodiments, the method for generating statements of structured query language can be implemented as a computer program, which is tangibly included in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method for generating statements of structured query language described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the method for generating statements of structured query language in any other suitable manner (e.g., by means of firmware).

[0169] The various embodiments of the systems and technologies described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0170] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0171] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0172] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0173] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0174] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0175] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0176] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for generating statements of a structured query language, characterized in that, Including: Obtain user query information; Generate a target statement generation result corresponding to the user query information based on the user query information and a statement generation model of a structured query language; The statement generation model of the structured query language is trained and generated based on different historical user query information and statement generation results corresponding thereto extracted from a pre-constructed syntax alignment corpus according to each of the historical user query information.

2. The method according to claim 1, wherein The construction process of the syntax alignment corpus includes: Obtain the syntax specifications of at least two structured query languages; Use the syntax specification of any one structured query language as a seed specification; Extract the syntax features of each of the syntax specifications respectively, and determine the differences between the seed specification and other syntax specifications except the seed specification according to each of the syntax features; Generate a statement of the first structured query language of the seed specification corresponding to the historical user query information according to the historical user query information and the syntax features of the seed specification; Generate a statement of the second structured query language of the other syntax specifications except the seed specification according to the differences between the seed specification and the other syntax specifications; Use the historical user query information, the statement of the first structured query language, and the statement of the second structured query language as the content in the syntax alignment corpus.

3. The method according to claim 2, wherein The determination of the differences between the seed specification and other syntax specifications except the seed specification according to each of the syntax features includes: Compare the syntax features in the seed specification with the corresponding syntax features in the other syntax specifications; If the syntax features in the seed specification are inconsistent with the corresponding syntax features in the other syntax specifications, use the syntax feature in the seed specification as a differential syntax feature, and determine the difference points between the differential syntax feature and the corresponding syntax features of the other syntax specifications; the types of the difference points include at least one of the following: keywords; function methods; syntax structure characteristics.

4. The method according to claim 3, wherein The generation of a statement of the second structured query language of the other syntax specifications except the seed specification according to the differences between the seed specification and the other syntax specifications includes: Use the statement content corresponding to the differential syntax feature in the statement of the first structured query language as a first statement; Call the syntax feature of the other syntax specification according to the difference point to generate corresponding replacement information, and replace the difference point with the replacement information in the first statement, and use the first statement after replacement as a second statement; Replace the first statement in the statement of the first structured query language with the second statement to obtain the statement of the second structured query language.

5. The method according to claim 1, characterized in that, The output of the target statement generation result corresponding to the user query information based on the user query information and the statement generation model of the structured query language includes: Obtain a pre-constructed structured query language generation prompt instruction; wherein, the preset role of the structured query language generation prompt instruction is a structured query language expert who knows the syntax alignment corpus; the structured query language generation prompt instruction is an instruction for generating a prompt word of a structured query language. Use the user query information as the operation parameter for generating the prompt instruction of the structured query language to generate the prompt word of the structured query language. Input the prompt word of the structured query language into the statement generation model of the structured query language to output the target statement generation result.

6. The method according to claim 1, wherein The training process of the statement generation model of the structured query language includes: Divide the training data set into a training set and a test set according to a preset data set division ratio, and construct an initial model of the statement generation model of the structured query language; wherein, the initial model of the statement generation model of the structured query language includes a rule router and a mixture of experts network; the rule router is connected to the mixture of experts network; the mixture of experts network includes a gating network and at least two expert networks; the output end of the gating network is connected to each of the expert networks; the rule router is used to allocate the input data of the initial model of the statement generation model of the structured query language to each of the expert networks in the mixture of experts network according to predefined rule conditions. Train the initial model of the statement generation model of the structured query language according to the training set to obtain a basic model of the statement generation model of the structured query language. Use the test set to perform a performance test on the basic model of the statement generation model of the structured query language; if the test result meets the preset model test pass condition, use the basic model of the statement generation model of the structured query language as the final statement generation model of the structured query language, otherwise retrain the initial model of the statement generation model of the structured query language according to the training set.

7. The method according to claim 6, wherein The training of the initial model of the statement generation model of the structured query language according to the training set to obtain a basic model of the statement generation model of the structured query language includes: Input the training set into the initial model of the statement generation model of the structured query language. Judge whether the data in the input training set triggers the rules of the rule router. If the rules of the rule router are triggered, skip the gating network and directly allocate the data in the training set to the corresponding expert network according to the predefined rule conditions. If the rules of the rule router are not triggered, allocate it to the corresponding expert network by the gating network. Take the weighted function of the expert loss function, the grammar specification discrimination function, and the perplexity loss function according to the preset weight values as the loss function, and optimize the model parameters with the goal of minimizing the loss function to obtain the basic model of the statement generation model for the structured query language; wherein, the expert loss function is a function that quantifies the prediction error by calculating the cross-entropy between the predicted probability distribution output by each expert network and the true label; the grammar specification discrimination function is a contrast loss function whose optimization goal includes minimizing the distance between sample pairs with the same grammar specification in the training set and maximizing the distance between sample pairs with different grammar specifications in the training set; the perplexity loss function is a function that calculates the average branching factor of the probability distribution output by the structured query language generation model.

8. A statement generation device for a structured query language, characterized in that, Including: An information acquisition module for acquiring user query information; A statement generation result module for outputting the target statement generation result corresponding to the user query information based on the user query information and the statement generation model of the structured query language; The statement generation model of the structured query language is trained and generated based on different historical user query information and the corresponding statement generation results extracted from a pre-constructed grammar alignment corpus according to each historical user query information.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the statement generation method of the structured query language according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to implement the statement generation method of the structured query language according to any one of claims 1-7 when executed.