Statement conversion method and device based on fine tuning model, equipment and medium

By preprocessing the original query statements and reconstructing the slot relationship graph based on a fine-tuning model method, the accuracy and flexibility issues of converting natural language into SQL statements in traditional methods are solved, and more efficient SQL statement generation is achieved.

CN120723802APending Publication Date: 2025-09-30CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510881130.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Traditional methods of converting natural language into SQL statements rely on manually designed grammar and logic, which lack flexibility. This results in an inability to accurately identify query intent when processing diverse natural language expressions. This is especially true when faced with complex query natural statements, long sentences, or grammatically incorrect natural statements, which easily lead to inaccurate or incomplete SQL statements being generated.

Method used

Through a fine-tuning model-based method, the context information of the original query statement is obtained for preprocessing, and a multi-task model is used to perform slot identification and operation type prediction. The slot relationship graph is reconstructed, and the target SQL template is matched and assembled to generate SQL statements that meet the query intent and semantic information.

Benefits of technology

It improves the accuracy of converting natural language into SQL statements, reduces the difficulty of writing query statements, and improves the convenience of query statement generation and the completeness of semantic analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723802A_ABST
    Figure CN120723802A_ABST
Patent Text Reader

Abstract

According to the statement conversion method based on the fine tuning model, semantic information and query intention of the query text are analyzed from two aspects of slot positions and operation types through the multi-task model, so that the accuracy of the query statement is improved, and semantic completion is performed on the query statement by reconstructing the slot position relation graph, so that the accuracy of the query statement is improved. Therefore, a relation topology basis is provided for subsequent template matching, and the semantic analysis integrity of the complex query statement is improved. And finally, according to the slot position relation graph, an SQL template conforming to the query intention and semantic information corresponding to the original query statement is determined, and the target SQL statement is obtained by splicing and assembling the target SQL template based on the slot position relation graph, so that the writing difficulty of the query statement is reduced. When the method is applied to a data query function of a system in the financial business field or the medical field, the input convenience of query statements can be improved, and the statement conversion accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of natural language and text classification, and in particular to a sentence conversion method, apparatus, computer device and computer-readable storage medium based on a fine-tuning model. Background Art

[0002] With the rapid development of informatization, data-driven decision-making has become a strategic priority for many businesses and organizations. Faced with massive amounts of data, quickly and accurately obtaining the required information not only improves business efficiency but also provides strong support for decision-making. However, traditional data query methods often require users to master Structured Query Language (SQL), which poses a significant barrier for non-technical users. This means that users facing complex databases must spend considerable time learning SQL syntax and database structure to effectively retrieve the required data. To address this issue, natural language queries can be converted to SQL queries (the text2sql task). Current text2sql query rules rely on manually designed syntax and logic, lacking flexibility. This results in the system being unable to accurately identify query intent when processing diverse natural language expressions. Furthermore, complex, long, or grammatically incorrect query statements can easily lead to inaccurate or incomplete SQL statements. Therefore, improving the accuracy of natural language conversion to SQL has become a pressing issue. Summary of the Invention

[0003] The present application provides a statement conversion method, apparatus, computer equipment, and storage medium based on a fine-tuning model to improve the accuracy of converting natural language into SQL statements.

[0004] In a first aspect, the present application provides a sentence conversion method based on a fine-tuning model, the method comprising:

[0005] Based on the original query statement corresponding to the target user, obtaining context information corresponding to the original query statement, and preprocessing the original query statement based on the context information to obtain a standardized query text;

[0006] Based on the multi-task model, slot identification and operation type prediction are performed on the standardized query text to obtain a slot label sequence and an operation type probability distribution;

[0007] Reconstructing a slot graph relationship based on the slot label sequence and the operation type probability distribution to obtain a slot relationship graph;

[0008] Based on the slot relationship graph, determining a target SQL template that matches the original query statement in a template knowledge base;

[0009] The slot relationship graph is spliced ​​and assembled with the target SQL template to generate a target SQL statement that conforms to the query intent and semantic information corresponding to the original query statement.

[0010] In a second aspect, the present application further provides a sentence conversion device based on a fine-tuning model, the device comprising:

[0011] A text preprocessing module is used to obtain context information corresponding to the original query statement based on the original query statement corresponding to the target user, and preprocess the original query statement based on the context information to obtain a standardized query text;

[0012] A statement slot identification module is used to perform slot identification and operation type prediction on the standardized query text based on a multi-task model to obtain a slot label sequence and an operation type probability distribution;

[0013] A relationship graph reconstruction module, configured to reconstruct a slot graph relationship based on the slot label sequence and the operation type probability distribution to obtain a slot relationship graph;

[0014] An SQL template matching module is configured to determine a target SQL template matching the original query statement in a template knowledge base based on the slot relationship graph;

[0015] The SQL statement assembly module is used to splice and assemble the slot relationship diagram with the target SQL template to generate a target SQL statement that conforms to the query intent and semantic information corresponding to the original query statement.

[0016] In a third aspect, the present application also provides a computer device comprising a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the above-mentioned fine-tuning model-based statement conversion method when executing the computer program.

[0017] In a fourth aspect, the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the above-mentioned statement conversion method based on the fine-tuning model.

[0018] The present application discloses a statement conversion method, device, computer equipment and storage medium based on a fine-tuning model. Based on the original query statement corresponding to the target user, the context information corresponding to the original query statement is obtained, and based on the context information, the original query statement is preprocessed to obtain a standardized query text; based on the multi-task model, the standardized query text is slot identified and the operation type is predicted to obtain a slot label sequence and an operation type probability distribution; based on the slot label sequence and the operation type probability distribution, the slot graph relationship is reconstructed to obtain a slot relationship graph; based on the slot relationship graph, a target SQL template matching the original query statement is determined in a template knowledge base; the slot relationship graph is spliced ​​and assembled with the target SQL template to generate a target SQL statement that conforms to the query intent and semantic information corresponding to the original query statement. In the above manner, the present application preprocesses the original query statement through context to eliminate expression problems such as ambiguity and unclear reference in the original query statement, thereby improving the accuracy of the converted statement. Then, a multi-task model is used to identify slots and predict operation types for standardized query text. This model analyzes the semantic information and query intent of the query text from both the slot and operation type perspectives, thereby improving the accuracy of the query statement. A slot relationship graph is reconstructed based on the slot label sequence and the probability distribution of the operation type. This graph is then used to semantically complete the query statement, providing a relational topology basis for subsequent template matching and improving the semantic parsing completeness of complex query statements. Finally, based on the slot relationship graph, an SQL template that matches the query intent and semantic information corresponding to the original query statement is determined. This template is then assembled with the target SQL template based on the slot relationship graph to generate the target SQL statement. This not only improves the query conversion accuracy but also reduces the difficulty of query writing and improves the convenience of query generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 This is a flowchart of a first embodiment of the sentence conversion method based on the fine-tuning model of the present application;

[0021] Figure 2 This is a flowchart of a second embodiment of the sentence conversion method based on the fine-tuning model of the present application;

[0022] Figure 3It is a schematic diagram of the sentence conversion device based on the fine-tuning model of the present application;

[0023] Figure 4 This is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0024] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0025] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0026] It should be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0027] It should be further understood that the term “and / or” used in this specification and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0028] Embodiments of the present application provide a statement conversion method, apparatus, computer device, and storage medium based on a fine-tuning model. The statement conversion method based on the fine-tuning model can be applied to a server, where a statement conversion program deployed in the server converts natural language query statements into SQL query statements. The server can be a standalone server or a server cluster.

[0029] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.

[0030] See also Figure 1 , Figure 1 This is a schematic flow chart of a sentence conversion method based on a fine-tuning model provided in an embodiment of the present application.

[0031] like Figure 1As shown, the sentence conversion method based on the fine-tuning model specifically includes steps S101 to S105.

[0032] Step S101: Based on an original query statement corresponding to a target user, context information corresponding to the original query statement is obtained, and based on the context information, the original query statement is preprocessed to obtain a standardized query text;

[0033] In this embodiment, upon receiving the natural language query text corresponding to the original query statement input by the user, the original natural language query text and the colloquial expression text in the corresponding context cache are converted into a standard structure (e.g., "the past three years" is converted to "last_3_years"). Based on the context information, the pronouns or aliases in the original natural language query text are standardized (e.g., "the company" is converted to "Ping An Pu Hui"; "Company A" is converted to "A"). This yields the preprocessed standard query text cleaned_query (e.g., "List the net profit growth rates of Company A last_3_years, grouped by year"). The context includes historical conversation information and / or user-related information corresponding to the target user.

[0034] The preprocessing of the original query statement based on the context information to obtain a standardized query text includes:

[0035] Identifying referents of pronouns in the original query based on a coreference resolution model and the context information, and replacing the pronouns based on the referents;

[0036] Based on regular rules and a temporal parser, the time fuzzy expression in the original query statement is standardized and converted to obtain the standardized query text.

[0037] Specifically, a coreference resolution model (such as SpanBERT-based mention-linking) is used to identify the referent of a pronoun (e.g., associating "its" with "Company A" in the context cache). Regular rules and a temporal parser (such as SUTime) are used to transform ambiguous expressions (e.g., "nearly three years" → [start_time: NOW-3YEARS, end_time: NOW]).

[0038] It is understandable that the user can input the original natural language query text through various methods such as text input or voice input.

[0039] Step S102: Based on the multi-task model, slot identification and operation type prediction are performed on the standardized query text to obtain a slot label sequence and an operation type probability distribution;

[0040] In this embodiment, a multi-task model is used to identify slots and predict operation types for standard query text, generating a slot label sequence and an operation type probability distribution. Slots include: query indicators (e.g., "nii"); query condition types (ctypes) (e.g., "cx_flag"); query condition values ​​(cvalues) (e.g., "default"); sort types (otypes) (e.g., "default"); default condition types (dtypes) (e.g., "dvalues"), and default condition values ​​(dvalues). The query indicator (indicator) is the core query object (e.g., "nii," "roe"); the condition type (ctype) is the filter condition type (e.g., "time range," "institution type"); the condition value (cvalue) is the specific filter value (e.g., "2023," "Ping An Property & Casualty Insurance"); the sort type (otype) is the result sorting method (e.g., "by amount in descending order"); and the default condition (dtype / dvalue) is the implicit condition for the business (e.g., "only display valid data").

[0041] Specifically, through the slot sequence standard layer of the multi-task model, CRF decoding is used to output the BIEO label sequence corresponding to the slot in the standard query text to obtain the slot label sequence; through the operation classification layer of the multi-task model, Softmax is used to output the probability distribution of the operation type.

[0042] Step S103: reconstructing a slot graph relationship based on the slot label sequence and the operation type probability distribution to obtain a slot relationship graph;

[0043] In this embodiment, the slot graph relationship is reconstructed based on the graph attention network, slot label sequence and operation type probability distribution.

[0044] Specifically, each slot in the slot label sequence is used as a node to initialize the slot graph structure, wherein node feature = word vector + type code.

[0045] The multi-head attention in the graph attention network is used to calculate the dependency weights between nodes (such as the correlation strength between "net profit growth rate" and "the past three years") and complete the implicit relationship edges of the slot graph structure (such as automatically adding the OWNED_BY relationship of "company → indicator");

[0046] Based on the business metadata, necessary edges are added to the slot graph structure (for example, the time range field must be connected to the indicator node) to inject constraint rules into the slot graph structure and obtain a weighted slot relationship graph.

[0047] Among them, the slot relationship graph with weights can also be constructed through the following steps:

[0048] Perform dependency syntactic analysis on the query syntactic structure of the standard query text to identify core verbs (such as "list") and their relationships with other components (such as the subject-predicate relationship between "list" and "Company A", the compound relationship between "growth rate" and "net profit", etc.), and obtain a dependency graph dependency_graph containing the relationship types between each word in the standard query text (such as the relationship between "list" and "A" is "nsubj").

[0049] Based on the dependency graph, the operation intent of the original natural language query text is roughly extracted. That is, the operation type corresponding to the original natural language query text is identified according to the core verb (such as "list" corresponds to the SELECT operation), and the auxiliary operations are marked (such as "group" corresponds to the GROUP BY operation). The initial operation intent operation_hints including the main operation type (such as SELECT) and its confidence, as well as the auxiliary operation list (such as GROUP BY) are obtained.

[0050] Perform slot identification on the dependency graph dependency_graph, that is, identify the entity types in the dependency graph, obtain the information of each node (such as operation node (SELECT), company node, metric node, grouping node, etc.), and map it to a standard value (such as "net profit growth rate" is mapped to "net_prof it_growth"), and obtain the entity mapping table (used to record the original form, standard ID, data type and other information of each entity).

[0051] Deep semantic fusion is performed on the composite indicators of the dependency graph (for example, "net profit growth rate" is converted into the calculation formula "(net_profit_2023 / net_profit_2022-1)*100") to obtain the association between each node. Based on the association between each node and the corresponding standard value of each node, the tree structure of the dependency graph for ing is normalized, that is, the semantic tree corresponding to the dependency graph is constructed, specifically with QUERY_ROOT as the root node, connecting the operation node (SELECT) and its child nodes (such as company nodes, measurement nodes, grouping nodes, etc.), and expressing the relationship between nodes (such as the association between measurement nodes and time filter nodes). This results in a semantic tree structure that includes a node list (node ​​type, standard value, child node and other information).

[0052] The data type (data_type) and reference ID (ref_id) are added to the slot information in the semantic tree to obtain a structured slot list, including parameter type, value, confidence, data type, reference ID, etc.

[0053] A relationship graph is constructed based on the semantic tree, and relationship types (such as "AGGREGATE_FOR") and constraint levels (such as "REQUIRED") are added to the edges to obtain a relationship graph containing nodes and edges, where the edges have attributes such as relationship type, confidence, and constraint level.

[0054] Calculate the confidence for each slot in the relationship graph, detect data type conflicts (type_conflicts) and missing required relationships (missing_constraints), complete the confidence evaluation, and obtain cross-step verification marks (such as data type conflicts and missing constraints).

[0055] Based on the graph relationship reconstruction results and the original output of the model, the uncertainty of the slot relationship graph is quantified to obtain the uncertainty score, that is, the confidence score of the key indicators in the slot relationship graph (including slot filling confidence and main operation type confidence) is calculated.

[0056] Specifically, during multi-task model inference, the Dropout layer was kept activated (p = 0.1) to simulate random neuron failures. Then, 20 independent forward propagations were performed on the same query. The fluctuations of key metrics during the forward propagation process (such as slot fluctuations, relationship weight entropy, and operation type entropy) were analyzed. Finally, based on the confidence mapping rules, slot-level confidence (used for contextual parameter binding, such as locating the character position of "net profit growth rate" in the original query to ensure accurate template parameter mapping) and relationship edge confidence (used for graph similarity calculation) were derived.

[0057] In a specific embodiment, FGSM adversarial disturbance (ε=0.01) may be injected during the forward propagation of the multi-task model to improve noise stability.

[0058] Therefore, ambiguity in query text is eliminated through preprocessing, reference ambiguity is resolved through context binding, and computational benchmarks are unified through time normalization. The graph attention network (GAT) dynamically reconstructs the slot relationship graph based on the features of adjacent slot nodes, repairing the structural deficiencies in independent slot identification. Uncertainty scores are then used as a weighting factor for template matching (for example, slots with a confidence score < 0.8 trigger manual review). This improves the completeness of semantic parsing for complex queries, reduces the rate of missed slot relationships, and enhances template matching accuracy.

[0059] Step S104: Based on the slot relationship graph, determine a target SQL template that matches the original query statement in a template knowledge base;

[0060] In this embodiment, candidate templates are screened based on the main operation type (such as SELECT) and the complexity of the relationship graph in the structured slot list.

[0061] Specifically, through the analysis results of the operation intention (including the main operation type and the secondary operation type), the templates corresponding to the main operation type are filtered from the template library, and the above-mentioned filtered templates are further filtered based on the secondary operation type, that is, the templates containing at least one secondary operation type are retained to obtain a list of candidate template IDs.

[0062] Based on the graph neural network GNN, the relationship graph is neurally encoded to encode the relationship graph into a vector, and the graph vectors of each candidate template are pre-stored.

[0063] Among them, graph neural network encoding includes:

[0064] Node features = slot type embedding (128 dimensions) + confidence (normalized)

[0065] Edge features = relation type embedding (64 dimensions) + edge confidence

[0066] Using node features, edge features, and GAT (graph attention network), the graph-level representation vector is calculated as the graph vector of the relationship graph.

[0067] Based on the graph vector of the relationship graph and the graph vector of each candidate template, the multidimensional similarity between the relationship graph and each candidate template is calculated, wherein the multidimensional similarity includes structural similarity, semantic coverage and constraint satisfaction rate, and the total similarity is obtained by weighting the above similarities.

[0068] Specifically, semantic coverage = |matching slot type| / |query slot type|, and the complexity alignment is obtained based on the ratio of the number of edges in the query to the number of required edges in the template.

[0069] In a specific embodiment, a semantic compatibility check may also be performed, that is, for each slot type, it is checked whether there is a corresponding parameter placeholder in the template (such as COMPANY→${company}).

[0070] Thus, a matching score of each candidate template is obtained, and the candidate templates are sorted based on the matching score.

[0071] The candidate template with the highest matching score is determined as the best matching template and the target SQL template. Based on the best matching template and constraint rules, adaptive template patching is implemented, that is, missing components (such as the GROUP BY clause) are patched. Low-confidence slots use soft binding (retaining placeholders) and a multi-template voting mechanism. When high-risk slots exist, multi-template voting is enabled to generate the final result. Specifically, this includes:

[0072] If the confidence level is low, the multi-template voting mechanism is activated;

[0073] If the best template matching score is lower than the threshold, template repair is initiated, including node completion (if the template is missing a slot parameter (such as TIME_RANGE), calling the business rule library to automatically add a WHERE clause) and edge repair (if the key relationship edge is missing (such as the OWNED_BY edge is not mapped), inserting the JOIN condition expression).

[0074] Step S105: Assemble the slot relationship diagram and the target SQL template to generate a target SQL statement that meets the query intent and semantic information corresponding to the original query statement.

[0075] In this embodiment, the slot template assembly is used to match and assemble the structured slots parsed from the user query with the predefined SQL template to generate a parameterized SQL statement.

[0076] Specifically, the identified slot values ​​(such as company name, time range, indicator name) are filled into the corresponding parameter positions in the template to realize the splicing and assembly of the slot relationship diagram and the target SQL template.

[0077] This embodiment provides a statement conversion method, device, computer equipment and storage medium based on a fine-tuning model. Based on the original query statement corresponding to the target user, the context information corresponding to the original query statement is obtained, and based on the context information, the original query statement is preprocessed to obtain a standardized query text; based on the multi-task model, the standardized query text is slot identified and the operation type is predicted to obtain a slot label sequence and an operation type probability distribution; based on the slot label sequence and the operation type probability distribution, the slot graph relationship is reconstructed to obtain a slot relationship graph; based on the slot relationship graph, a target SQL template matching the original query statement is determined in the template knowledge base; the slot relationship graph and the target SQL template are spliced ​​and assembled to generate a target SQL statement that conforms to the query intent and semantic information corresponding to the original query statement. In the above manner, the present application preprocesses the original query statement through context to eliminate expression problems such as ambiguity and unclear reference in the original query statement, thereby improving the accuracy of the converted statement. Then, a multi-task model is used to identify slots and predict operation types for standardized query text. This model analyzes the semantic information and query intent of the query text from both the slot and operation type perspectives, thereby improving the accuracy of the query statement. A slot relationship graph is reconstructed based on the slot label sequence and the probability distribution of the operation type. This graph is then used to semantically complete the query statement, providing a relational topology basis for subsequent template matching and improving the semantic parsing completeness of complex query statements. Finally, based on the slot relationship graph, an SQL template that matches the query intent and semantic information corresponding to the original query statement is determined. This template is then assembled with the target SQL template based on the slot relationship graph to generate the target SQL statement. This not only improves the query conversion accuracy but also reduces the difficulty of query writing and improves the convenience of query generation.

[0078] Reference Figure 2 , Figure 2 2 is a flow chart of the second embodiment of the sentence conversion method based on the fine-tuning model of the present invention.

[0079] like Figure 2 As shown, before step S102, the following steps are also included:

[0080] Step S110: obtaining a labeled query sample by labeling the historical query statements with slots and operation types, wherein the historical query statements include the historical original user query text and its corresponding SQL query statement and business metadata;

[0081] Step S120: replacing synonyms in the annotated query sample with a financial term library, and performing adversarial enhancement on the annotated query sample to obtain an enhanced training set;

[0082] Step S130 , performing LoRA fine-tuning on the initial multi-task model based on the enhanced training set to obtain the multi-task model.

[0083] The step of performing LoRA fine-tuning on the initial multi-task model based on the enhanced training set to obtain the multi-task model includes:

[0084] Based on the enhanced training set and multi-task learning framework, the initial multi-task model shares the underlying Transformer representation and is connected to the slot labeling and operation classification through the upper layer;

[0085] A trainable low-rank matrix is ​​injected into each layer of the initial multi-task template, the model weights are frozen, and the low-rank matrix is ​​trained and optimized to obtain the multi-task model.

[0086] In this embodiment, training data is prepared for fine-tuning by performing multi-dimensional annotation (including slot annotation (BIOES tags) and operation type annotation (SELECT / WHERE, etc.)) on sample data (including original user query text, corresponding SQL query statements, and business metadata); specifically, the query text and corresponding slot information can be organized into training sample pairs (q, s), that is, a mapping from questions to slots.

[0087] Synonyms in sample data are replaced by financial terminology library to perform adversarial enhancement on sample data and obtain enhanced training set.

[0088] The business indicators in the enhanced training set are mapped to the database fields to obtain a business rule library, which is used for constraint checking in subsequent template assembly.

[0089] Based on the enhanced training set and multi-task learning framework, the multi-task model was fine-tuned using LoRA. This involves sharing the underlying Transformer representation of the multi-task model and connecting it to CRF (slot labeling) and MLP (operation classification) through the upper layers. Trainable low-rank matrices were injected into each layer of the model, freezing the model weights and optimizing only the matrices. The model was trained in stages, starting from easy to difficult, gradually increasing the query complexity of the model to obtain a fine-tuned multi-task model.

[0090] Among them, by injecting a low-rank adapter matrix into the base model (such as ChatGLM-6B), that is, ΔW = A × B T (A∈R d×r ,B∈R r×k ) to fine-tune the multi-task model with LoRA.

[0091] The multi-task model is trained using a multi-task training strategy, where the main task is CRF layer sequence labeling (slot identification) and the auxiliary task is MLP head operation classification (SELECT / WHERE, etc.).

[0092] When training the model in stages, from easy to difficult, the training samples are graded according to complexity in advance, such as:

[0093] Phase 1: Simple Query (Single Entity + Single Index)

[0094] Phase 2: Compound Query (Multiple Entities + Time Filtering)

[0095] Phase 3: Nested query (subquery + multi-table join).

[0096] Furthermore, the slot graph relationship reconstruction is performed based on the slot label sequence and the operation type probability distribution to obtain a slot relationship graph, including:

[0097] Initializing a slot map structure using each slot in the slot label sequence as a node;

[0098] Based on the multi-head attention in the graph attention network, the dependency weights between each node are calculated to obtain the initial relationship graph with weights;

[0099] The implicit relationship edges in the initial relationship graph are completed, and based on the preset business metadata and preset constraint rules, necessary edges are added to the initial relationship graph to obtain a weighted slot relationship graph.

[0100] In this embodiment, each slot in the slot label sequence is used as a node to initialize the slot graph structure, wherein node feature = word vector + type code.

[0101] The multi-head attention in the graph attention network is used to calculate the dependency weights between nodes (such as the correlation strength between "net profit growth rate" and "the past three years") and complete the implicit relationship edges of the slot graph structure (such as automatically adding the OWNED_BY relationship of "company → indicator");

[0102] Based on the business metadata, necessary edges are added to the slot graph structure (for example, the time range field must be connected to the indicator node) to inject constraint rules into the slot graph structure and obtain a weighted slot relationship graph.

[0103] Furthermore, determining a target SQL template matching the original query statement in a template knowledge base based on the slot relationship graph includes:

[0104] Based on the main operation type and the secondary operation type in the slot relationship diagram, screening a list of candidate templates in a preset template library;

[0105] Encoding the slot relationship graph based on a graph neural network to obtain a graph vector of the slot relationship graph;

[0106] Calculating the similarity between the slot relationship graph and each candidate template based on the graph vector of the slot relationship graph and the graph vector of each candidate template in the candidate template list;

[0107] The candidate template corresponding to the highest similarity is obtained as the target SQL template that matches the original query statement.

[0108] In this embodiment, candidate templates are screened based on the main operation type (such as SELECT) and the complexity of the relationship graph in the structured slot list.

[0109] Specifically, through the analysis results of the operation intention (including the main operation type and the secondary operation type), the templates corresponding to the main operation type are filtered from the template library, and the above-mentioned filtered templates are further filtered based on the secondary operation type, that is, the templates containing at least one secondary operation type are retained to obtain a list of candidate template IDs.

[0110] Based on the graph neural network GNN, the relationship graph is neurally encoded to encode the relationship graph into a vector, and the graph vectors of each candidate template are pre-stored.

[0111] Among them, graph neural network encoding includes:

[0112] Node features = slot type embedding (128 dimensions) + confidence (normalized)

[0113] Edge features = relation type embedding (64 dimensions) + edge confidence

[0114] Using node features, edge features, and GAT (graph attention network), the graph-level representation vector is calculated as the graph vector of the relationship graph.

[0115] Based on the graph vector of the relationship graph and the graph vector of each candidate template, the multidimensional similarity between the relationship graph and each candidate template is calculated, wherein the multidimensional similarity includes structural similarity, semantic coverage and constraint satisfaction rate, and the total similarity is obtained by weighting the above similarities.

[0116] Specifically, semantic coverage = |matching slot type| / |query slot type|, and the complexity alignment is obtained based on the ratio of the number of edges in the query to the number of required edges in the template.

[0117] In a specific embodiment, a semantic compatibility check may also be performed, that is, for each slot type, it is checked whether there is a corresponding parameter placeholder in the template (such as COMPANY→${company}).

[0118] Thus, a matching score of each candidate template is obtained, and the candidate templates are sorted based on the matching score.

[0119] The candidate template with the highest matching score is determined as the best matching template.

[0120] Furthermore, the target SQL template includes a parameterized SQL skeleton component and a parameter placeholder. The slot relationship diagram is assembled with the target SQL template to generate a target SQL statement that conforms to the query intent and semantic information corresponding to the original query statement, including:

[0121] Obtaining a preset slot template parameter mapping table and each slot in the slot relationship diagram;

[0122] Based on each slot type, dynamic parameter injection is performed on the parameterized SQL skeleton component, and the slot value corresponding to each slot is bound to the parameter placeholder to realize the splicing assembly of the slot relationship diagram and the target SQL template;

[0123] Perform multi-dimensional conflict detection on the assembled SQL statements, wherein the multi-dimensional conflict detection includes syntax conflict detection, business rule conflict detection, and relationship integrity verification;

[0124] When the assembled SQL statement passes the multi-dimensional conflict detection, the assembled SQL statement is used as the target SQL statement.

[0125] In this example, each slot in the slot relationship graph, such as [Type: COMPANY, Value: A], and a slot template parameter mapping table (i.e., slot-parameter mapping rules) are obtained. For example, slot type COMPANY → template parameter ${company}, slot type METRIC → template parameter ${metric}, and slot type TIME_RANGE → template parameter ${time_filter}. The slot type is then matched to the template parameter based on the mapping rules. For undefined slot types, an unknown type warning message is generated to facilitate manual processing.

[0126] Among them, dynamic parameter injection includes: basic value injection, composite value conversion and low confidence processing.

[0127] Furthermore, before performing template matching, a template library including SQL templates is pre-built, specifically including the following steps:

[0128] 1. Business scenario classification: Classify SQL templates according to business needs (such as financial analysis and customer statistics).

[0129] 2. Template structure construction: Each template contains a fixed SQL skeleton (such as the SELECT statement structure) and dynamic parameter placeholders.

[0130] The SQL skeleton components include: the SELECT module, which contains the core indicator slots; the WHERE module, which contains the time filter slots; and the GROUP BY module, which is used to enable when dimension slots exist.

[0131] 3. Dynamic parameter injection point: Mark the variable part in the template as a placeholder (such as ${company}).

[0132] Among them, dynamic parameter injection includes:

[0133] 1) Single value binding (confidence > 0.9), such as slot type: COMPANY → replace ${company} with 'A', and the slot value is directly written to SQL.

[0134] 2) Composite value binding, such as time range conversion: last_3_years→report_date BETWEEN'2021-01-01'AND'2023-12-31'; multi-company processing: [A,B,C]→company IN('A','B','C').

[0135] 3) For low-confidence slot type parameters (0.6≤confidence<0.9), keep the parameterization: ${metric} (do not replace the specific value) and add the comment tag: --[[LOW_CONFIDENCE_SLOT:METRIC]].

[0136] 4. Constraint rule binding: Add semantic constraints to the template (such as "METRIC type must be associated with TIME_FILTER").

[0137] The specific constraints include:

[0138] 1) Execution constraint rules:

[0139] Detection <opt:groupby>Placeholder. If a dimension slot exists, it is replaced by GROUP BY${dimension}; if there is no dimension slot, the entire GROUP BY clause is removed.

[0140] 2) Inject security rules:

[0141] Add data permissions, such as AND department IN('FINANCE');

[0142] Numeric null protection, such as COALESCE(${metric},0).

[0143] 3) Resolving conflicts

[0144] For multiple filtering conditions, the parameters corresponding to the highest confidence are retained;

[0145] To solve the conflict of indicator formula, the standard formula defined by business metadata is adopted.

[0146] This completes the generation of SQL statements bound to the rules.

[0147] Perform structural integrity checks on the generated target SQL statements, including:

[0148] Required element detection: If the template requires a WHERE clause, check whether the filter condition slot exists. If a required element is missing, the template is marked as "incomplete structure".

[0149] Nested level validation checks whether the GROUP BY clause and aggregate functions coexist. Checks the consistency of associated fields in multi-layer nested queries.

[0150] The target SQL statements that pass the verification are marked as "assembly-ready", and the target SQL statements that fail the verification are adaptively repaired.

[0151] After obtaining the parameterized SQL and database schema, an abstract syntax tree (AST) is first constructed using the SQLGlot parser. The type system verifies field value type compatibility (preventing strings from participating in numerical calculations), metadata checks ensure the existence of associated fields, and the permissions module injects data access constraints. When syntax defects are detected, the association path is automatically completed for missing JOINs, ambiguous fields are contextually resolved, and parameterized escapes are implemented to prevent SQL injection attacks. The resulting SQL includes a validation report and executable SQL, specifically annotating repair operations (such as "ADDED_GROUP_BY") and parameterized versions to ensure the syntactic correctness and security of the generated SQL.

[0152] Based on SQL execution logs and user satisfaction feedback, we first cluster and analyze error patterns (categorized by syntax, logic, permissions, etc.). We then use active learning algorithms to filter high-value samples for incremental fine-tuning. This includes online updates to LoRA adapter weights and dynamic expansion of the template library to cover new scenarios.

[0153] Therefore, through the above method, component update items (such as adapter weights and template versions), performance gains and new sample sizes are recorded in detail, forming a closed-loop optimization mechanism to further improve the recognition and prediction accuracy of the model.

[0154] See also Figure 3 , Figure 3 The embodiment of the present application provides a schematic block diagram of a sentence conversion apparatus based on a fine-tuning model, which is used to perform the aforementioned sentence conversion based on the fine-tuning model. The sentence conversion apparatus based on the fine-tuning model can be configured on a server.

[0155] like Figure 3 As shown, the sentence conversion device 300 based on the fine-tuning model includes:

[0156] The text preprocessing module 310 is configured to obtain context information corresponding to the original query statement based on the original query statement corresponding to the target user, and preprocess the original query statement based on the context information to obtain a standardized query text;

[0157] A statement slot identification module 320 is configured to perform slot identification and operation type prediction on the standardized query text based on a multi-task model to obtain a slot label sequence and an operation type probability distribution;

[0158] A relationship graph reconstruction module 330 is configured to reconstruct a slot graph relationship based on the slot label sequence and the operation type probability distribution to obtain a slot relationship graph;

[0159] An SQL template matching module 340 is configured to determine a target SQL template matching the original query statement in a template knowledge base based on the slot relationship graph;

[0160] The SQL statement assembly module 350 is used to splice and assemble the slot relationship graph with the target SQL template to generate a target SQL statement that conforms to the query intent and semantic information corresponding to the original query statement.

[0161] It should be noted that those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0162] The above-mentioned device can be realized in the form of a computer program. The computer program can be used in Figure 4 Runs on the computer equipment shown.

[0163] See also Figure 4 , Figure 4 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device may be a server.

[0164] See Figure 4 The computer device includes a processor, a memory, and a network interface connected through a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.

[0165] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, which, when executed, can cause a processor to perform any statement conversion based on a fine-tuning model.

[0166] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.

[0167] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can perform any statement conversion based on the fine-tuning model.

[0168] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0169] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0170] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:

[0171] Based on the original query statement corresponding to the target user, obtaining context information corresponding to the original query statement, and preprocessing the original query statement based on the context information to obtain a standardized query text;

[0172] Based on the multi-task model, slot identification and operation type prediction are performed on the standardized query text to obtain a slot label sequence and an operation type probability distribution;

[0173] Reconstructing a slot graph relationship based on the slot label sequence and the operation type probability distribution to obtain a slot relationship graph;

[0174] Based on the slot relationship graph, determining a target SQL template that matches the original query statement in a template knowledge base;

[0175] The slot relationship graph is spliced ​​and assembled with the target SQL template to generate a target SQL statement that conforms to the query intent and semantic information corresponding to the original query statement.

[0176] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:

[0177] Identifying referents of pronouns in the original query based on a coreference resolution model and the context information, and replacing the pronouns based on the referents;

[0178] Based on regular rules and a temporal parser, the time fuzzy expression in the original query statement is standardized and converted to obtain the standardized query text.

[0179] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:

[0180] By annotating historical query statements with slots and operation types, annotated query samples are obtained. The historical query statements include the original user query text and its corresponding SQL query statement and business metadata.

[0181] Replacing synonyms in the labeled query samples with a financial term library, and performing adversarial enhancement on the labeled query samples to obtain an enhanced training set;

[0182] The initial multi-task model is fine-tuned by LoRA based on the enhanced training set to obtain the multi-task model.

[0183] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:

[0184] Based on the enhanced training set and multi-task learning framework, the initial multi-task model shares the underlying Transformer representation and is connected to the slot labeling and operation classification through the upper layer;

[0185] A trainable low-rank matrix is ​​injected into each layer of the initial multi-task template, the model weights are frozen, and the low-rank matrix is ​​trained and optimized to obtain the multi-task model.

[0186] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:

[0187] Initializing a slot map structure using each slot in the slot label sequence as a node;

[0188] Based on the multi-head attention in the graph attention network, the dependency weights between each node are calculated to obtain the initial relationship graph with weights;

[0189] The implicit relationship edges in the initial relationship graph are completed, and based on the preset business metadata and preset constraint rules, necessary edges are added to the initial relationship graph to obtain a weighted slot relationship graph.

[0190] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:

[0191] Based on the main operation type and the secondary operation type in the slot relationship diagram, screening a list of candidate templates in a preset template library;

[0192] Encoding the slot relationship graph based on a graph neural network to obtain a graph vector of the slot relationship graph;

[0193] Calculating the similarity between the slot relationship graph and each candidate template based on the graph vector of the slot relationship graph and the graph vector of each candidate template in the candidate template list;

[0194] The candidate template corresponding to the highest similarity is obtained as the target SQL template that matches the original query statement.

[0195] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:

[0196] Obtaining a preset slot template parameter mapping table and each slot in the slot relationship diagram;

[0197] Based on each slot type, dynamic parameter injection is performed on the parameterized SQL skeleton component, and the slot value corresponding to each slot is bound to the parameter placeholder to realize the splicing assembly of the slot relationship diagram and the target SQL template;

[0198] Perform multi-dimensional conflict detection on the assembled SQL statements, wherein the multi-dimensional conflict detection includes syntax conflict detection, business rule conflict detection, and relationship integrity verification;

[0199] When the assembled SQL statement passes the multi-dimensional conflict detection, the assembled SQL statement is used as the target SQL statement.

[0200] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement any one of the fine-tuning model-based statement conversions provided in the embodiments of the present application.

[0201] The computer-readable storage medium may be an internal storage unit of the computer device described in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc., equipped on the computer device.

[0202] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.< / opt:groupby>

Claims

1. A sentence conversion method based on a fine-tuning model, characterized in that: include: Based on the original query statement corresponding to the target user, obtaining context information corresponding to the original query statement, and preprocessing the original query statement based on the context information to obtain a standardized query text; Based on the multi-task model, slot identification and operation type prediction are performed on the standardized query text to obtain a slot label sequence and an operation type probability distribution; Reconstructing a slot graph relationship based on the slot label sequence and the operation type probability distribution to obtain a slot relationship graph; Based on the slot relationship graph, determining a target SQL template that matches the original query statement in a template knowledge base; The slot relationship graph is spliced ​​and assembled with the target SQL template to generate a target SQL statement that conforms to the query intent and semantic information corresponding to the original query statement.

2. The sentence conversion method according to claim 1, characterized in that: The preprocessing of the original query statement based on the context information to obtain a standardized query text includes: Identifying referents of pronouns in the original query based on a coreference resolution model and the context information, and replacing the pronouns based on the referents; Based on regular rules and a temporal parser, the time fuzzy expression in the original query statement is standardized and converted to obtain the standardized query text.

3. The sentence conversion method according to claim 1, wherein: Before performing slot identification and operation type prediction on the standardized query text based on the multi-task model to obtain a slot label sequence and an operation type probability distribution, the method further includes: By annotating historical query statements with slots and operation types, annotated query samples are obtained. The historical query statements include the original user query text and its corresponding SQL query statement and business metadata. Replacing synonyms in the labeled query samples with a financial term library, and performing adversarial enhancement on the labeled query samples to obtain an enhanced training set; The initial multi-task model is fine-tuned by LoRA based on the enhanced training set to obtain the multi-task model.

4. The sentence conversion method according to claim 3, characterized in that: The performing LoRA fine-tuning on the initial multi-task model based on the enhanced training set to obtain the multi-task model includes: Based on the enhanced training set and multi-task learning framework, the initial multi-task model shares the underlying Transformer representation and is connected to the slot labeling and operation classification through the upper layer; A trainable low-rank matrix is ​​injected into each layer of the initial multi-task template, the model weights are frozen, and the low-rank matrix is ​​trained and optimized to obtain the multi-task model.

5. The sentence conversion method according to claim 1, wherein: The reconstructing a slot graph relationship based on the slot label sequence and the operation type probability distribution to obtain a slot relationship graph includes: Initializing a slot map structure using each slot in the slot label sequence as a node; Based on the multi-head attention in the graph attention network, the dependency weights between each node are calculated to obtain the initial relationship graph with weights; The implicit relationship edges in the initial relationship graph are completed, and based on the preset business metadata and preset constraint rules, necessary edges are added to the initial relationship graph to obtain a weighted slot relationship graph.

6. The sentence conversion method according to claim 1, characterized in that: Determining a target SQL template that matches the original query statement in a template knowledge base based on the slot relationship graph includes: Based on the main operation type and the secondary operation type in the slot relationship diagram, screening a list of candidate templates in a preset template library; Encoding the slot relationship graph based on a graph neural network to obtain a graph vector of the slot relationship graph; Calculating the similarity between the slot relationship graph and each candidate template based on the graph vector of the slot relationship graph and the graph vector of each candidate template in the candidate template list; The candidate template corresponding to the highest similarity is obtained as the target SQL template that matches the original query statement.

7. The sentence conversion method according to any one of claims 1 to 6, characterized in that: The target SQL template includes a parameterized SQL skeleton component and a parameter placeholder. The slot relationship diagram is assembled with the target SQL template to generate a target SQL statement that conforms to the query intent and semantic information corresponding to the original query statement, including: Obtaining a preset slot template parameter mapping table and each slot in the slot relationship diagram; Based on each slot type, dynamic parameter injection is performed on the parameterized SQL skeleton component, and the slot value corresponding to each slot is bound to the parameter placeholder to realize the splicing assembly of the slot relationship diagram and the target SQL template; Perform multi-dimensional conflict detection on the assembled SQL statements, wherein the multi-dimensional conflict detection includes syntax conflict detection, business rule conflict detection, and relationship integrity verification; When the assembled SQL statement passes the multi-dimensional conflict detection, the assembled SQL statement is used as the target SQL statement.

8. A sentence conversion device based on a fine-tuning model, characterized in that: include: A text preprocessing module is used to obtain context information corresponding to the original query statement based on the original query statement corresponding to the target user, and preprocess the original query statement based on the context information to obtain a standardized query text; A statement slot identification module is used to perform slot identification and operation type prediction on the standardized query text based on a multi-task model to obtain a slot label sequence and an operation type probability distribution; A relationship graph reconstruction module, configured to reconstruct a slot graph relationship based on the slot label sequence and the operation type probability distribution to obtain a slot relationship graph; An SQL template matching module is configured to determine a target SQL template matching the original query statement in a template knowledge base based on the slot relationship graph; The SQL statement assembly module is used to splice and assemble the slot relationship diagram with the target SQL template to generate a target SQL statement that conforms to the query intent and semantic information corresponding to the original query statement.

9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the sentence conversion method based on the fine-tuning model according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to implement the sentence conversion method based on a fine-tuning model according to any one of claims 1 to 7.

Citation Information

Cited By

  • NL2SQL generation method based on large language model

    CN120910089A

  • NL2SQL generation method based on large language model

    CN120910089B

  • Question and answer query method, system and equipment for data table and medium

    CN121560925A