Cross-dialect query automatic conversion method and device based on large language model, equipment and storage medium
By applying a large language model combined with syntax analysis and feature embedding methods in cross-database dialect query conversion, the problems of inaccurate analysis, inaccurate matching and unreliable conversion of cross-dialect query conversion in the existing technology are solved, and high-quality and reliable cross-dialect query conversion are achieved.
Patent Information
- Application Number
- CN202510125596.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-26
AI Technical Summary
The prior art has problems such as inaccurate resolution, inaccurate matching, unreliable conversion and lack of effective verification and optimization mechanisms in query conversion across database dialects, which makes it difficult to guarantee the correctness and reliability of the conversion results.
The automatic cross-dial dialect query conversion method based on the large language model is adopted, combining grammatical analysis, feature embedding, semantic matching and large language model, which can accurately identify and convert fragments in the query statement that are incompatible with the target dialect, and realize high-quality cross-dialect query conversion.
It realizes a cross-dial dialect query conversion scheme with high resolution and recognition accuracy, strong matching accuracy, reliable conversion results and strong adaptability, reducing maintenance costs.
Smart Images

Figure CN120045584A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information retrieval technology, and particularly to a cross-dialect query automatic conversion method, apparatus, device, and storage medium based on a large language model. Background Art
[0002] With the development of database technology, different database management systems have formed their own unique SQL dialects. These dialects have differences in grammar rules and function implementations. For example, the AGE function in PostgreSQL may not exist in other databases, resulting in incompatible query statements when migrating or integrating between different database systems. Especially in distributed database systems and hybrid cloud environments, the demand for cross-dialect query conversion is increasing.
[0003] Existing dialect conversion tools have serious defects. For example, open-source tools such as SQLGlot have problems such as missing rules (unable to convert the AGE function in PostgreSQL), conversion errors (wrongly converting the JSON_EXTRACT_PATH function in PostgreSQL to the non-existent JSON_EXTRACT function in Oracle), and not supporting specifying the database version (converting "LIMIT 1" to "FETCH FIRST 1 ROW ONLY" which is only supported in Oracle 19c but not in Oracle 11g).
[0004] Traditional methods mainly include: rule-based methods, template-based methods, and machine learning-based methods. Rule-based methods require a large number of conversion rules to be written manually, with high maintenance costs and difficulty in dealing with complex conversion scenarios brought about by dialect differences. Template-based methods perform matching and replacement through predefined templates, but the template coverage is limited, and it is difficult to handle conversion tasks that require context understanding. Although machine learning-based methods have a certain ability of automatic learning, due to the lack of in-depth understanding of SQL syntax structures and semantics, the accuracy and reliability of the conversion results are difficult to guarantee. These methods are particularly prone to errors when dealing with complex syntax structures such as nested queries and window functions.
[0005] In addition, the existing methods generally have the following problems: First, the syntax parsing of query statements is not precise enough, making it difficult to accurately identify the fragments that need to be converted; second, the characterization of the functional features of syntax elements is not sufficient enough, affecting the matching accuracy of syntax elements during conversion; third, there is a lack of effective verification and optimization mechanisms, making it difficult to ensure the correctness of the conversion results; fourth, the utilization of unstructured information such as the functional specifications and usage documents of database dialects is insufficient, and domain knowledge cannot be fully utilized to guide the conversion process. Traditionally, solving these problems requires engineers to write a large amount of code, which is both expensive and lacks flexibility, especially when the database version is updated.
[0006] Therefore, a technical solution that can accurately understand the syntax structure and semantic features of query statements and can effectively perform cross-dialect conversion is needed. Summary of the Invention
[0007] The present disclosure provides a cross-dialect query automatic conversion solution based on a large language model. This solution combines syntax parsing, feature embedding, semantic matching, and a large language model, and can accurately identify and convert fragments in a query statement that are incompatible with the target dialect, achieving high-quality cross-dialect query conversion.
[0008] According to an embodiment of the present disclosure, a cross-dialect query automatic conversion method based on a large language model is proposed, including:
[0009] Extract syntax elements offline from the syntax files and database documents of the source dialect and the target dialect, generate a syntax tree of the syntax elements, and annotate the corresponding function specifications;
[0010] Parse the original query based on the source dialect to obtain the source dialect syntax tree of the original query;
[0011] Match the subtrees of the source dialect syntax tree of the original query with the syntax tree of the source dialect syntax elements to divide the original query into multiple source dialect function fragments;
[0012] Parse the original query based on the target dialect, identify the fragments that are incompatible with the target dialect, and correspondingly determine which source dialect function fragment the incompatible fragment belongs to, and use this source dialect function fragment as the fragment to be converted;
[0013] Use a cross-dialect syntax embedding model to convert the fragment to be converted and the corresponding function specification into a first embedding vector. The cross-dialect syntax embedding model is used to convert the input into an embedding vector in a vector space measured by functional equivalence;
[0014] Perform similarity matching between the first embedding vector and the pre-stored embedding vectors of the target dialect syntax elements to obtain the matched target dialect syntax elements;
[0015] Input the fragment to be converted and the corresponding function specification, the matched target dialect syntax elements and the corresponding function specification into the large language model to generate a converted fragment in the form of the target dialect;
[0016] Replace the fragment to be converted in the original query with the converted fragment to obtain the converted query.
[0017] In some embodiments, extracting syntax elements offline from the syntax files and database documents of the source dialect and the target dialect, generating a syntax tree of the syntax elements, and annotating the corresponding function specifications includes:
[0018] Extract grammar elements offline from the grammar files of the source dialect and the target dialect, and use BNF to generate corresponding syntax trees for each grammar element, where the leaf nodes represent terminal symbols and the internal nodes represent non-terminal symbols;
[0019] Perform a depth-first traversal of the syntax tree of each grammar element, match each node in the syntax tree with the grammar element nodes in the database document, and label the function specifications obtained from the database document through the matching on the syntax tree.
[0020] In some embodiments, before inputting the fragment to be converted into the cross-dialect grammar embedding model, the method further includes performing the following simplification processing on the fragment to be converted:
[0021] Replace the custom functions of the database system in the fragment to be converted with equivalent standard SQL functions;
[0022] Replace the components irrelevant to the conversion in the fragment to be converted with non-terminal symbols;
[0023] Replace the query clauses irrelevant to the conversion in the fragment to be converted with non-terminal symbols.
[0024] In some embodiments, using the cross-dialect grammar embedding model to convert the fragment to be converted and the corresponding function specifications into a first embedding vector, including:
[0025] Generate a structural embedding vector of the fragment to be converted through the syntax structure encoder in the cross-dialect grammar embedding model;
[0026] Generate a multi-dimensional feature embedding vector of the function specification through the syntax specification encoder in the cross-dialect grammar embedding model;
[0027] Generate the first embedding vector by comprehensively combining the structural embedding vector and the multi-dimensional feature embedding vector through a multi-head cross-attention mechanism.
[0028] In some embodiments, generating a multi-dimensional feature embedding vector of the function specification through the syntax specification encoder in the cross-dialect grammar embedding model, including:
[0029] Use multiple encoders to capture features in the function specification from different dimensions and mix expert knowledge to generate multiple feature embedding vectors of different dimensions;
[0030] Dynamically calculate the feature weights of each dimension according to the input function specification through a gating mechanism;
[0031] Perform weighted aggregation on each feature embedding vector based on the feature weights to obtain the multi-dimensional feature embedding vector.
[0032] In some embodiments, the method further includes training the cross-dialect grammar embedding model using contrastive learning, including:
[0033] Training the cross-dialect grammar embedding model using the following contrastive loss function to increase the scaled cosine similarity between the output embedding vector and the embedding vector of the positive sample, and decrease the scaled cosine similarity between the output embedding vector and the embedding vector of the hard negative sample:
[0034]
[0035] Wherein, is the loss value of the i-th sample s i ; represents the set of positive samples paired with the sample s i ; represents the set of positive samples in the positive sample set; represents the set of hard negative samples paired with the sample s i ; zip represents the scaled cosine similarity between the embedding vector of the sample s i and the embedding vector of the matching positive sample s p ; z in represents the scaled cosine similarity between the embedding vector of the sample s i and the embedding vector of the matching hard negative sample s n .
[0036] In some embodiments, the method further includes:
[0037] Constructing the set of positive samples of the sample s i in the following manner:
[0038] Performing synonym replacement or word order adjustment on the functional specification corresponding to the sample s i to obtain a specification variant as a positive sample;
[0039] Using the grammatical elements in other database systems that have the same keywords and functions as the sample s i as positive samples;
[0040] Using the equivalent grammatical elements obtained by converting the sample s i to other database systems using a database dialect conversion tool as positive samples;
[0041] Constructing the set of hard negative samples of the sample s i in the following manner:
[0042] Clustering the grammatical elements that are positive samples of each other, and selecting those that belong to different clusters from the sample s i and have a similarity higher than that of the sample s i with the embedding vector of the sample si The syntactic elements with similarity to the embedding vectors of all positive samples are used as hard negative samples.
[0043] In some embodiments, the method further includes:
[0044] Verify the transformed query;
[0045] If the verification fails, repeat the steps of generating the transformed segment using the large language model, updating the transformed query, and verifying;
[0046] If the verification still fails after the number of repeated executions reaches a preset maximum retry number, expand the range of the segment to be transformed to obtain an extended segment to be transformed, input the syntax tree and corresponding function specification of the extended segment to be transformed into the cross-dialect syntax embedding model to obtain a new first embedding vector, and perform similarity matching and subsequent steps again based on the new first embedding vector.
[0047] In some embodiments, verifying the transformed query includes:
[0048] Perform syntax verification on the transformed query using a syntax parser that supports the target dialect;
[0049] Perform semantic verification on the transformed query using the large language model.
[0050] In some embodiments, expanding the range of the segment to be transformed to obtain an extended segment to be transformed includes:
[0051] Select additional segments from adjacent source dialect functional segments of the source dialect functional segment to which the original segment to be transformed belongs;
[0052] Merge the original segment to be transformed with the selected additional segments to obtain an extended segment to be transformed.
[0053] According to an embodiment of the present disclosure, a cross-dialect query automatic conversion device based on a large language model is proposed, including:
[0054] An offline information extraction unit for offline extracting syntactic elements from the syntax files and database documents of the source dialect and the target dialect, generating a syntax tree of the syntactic elements, and annotating the corresponding function specifications;
[0055] A first parsing unit for parsing the original query based on the source dialect to obtain the source dialect syntax tree of the original query;
[0056] A function segmentation unit for matching the subtrees of the source dialect syntax tree of the original query with the syntax tree of the source dialect syntactic elements to divide the original query into multiple source dialect functional segments;
[0057] An incompatible recognition unit for parsing the original query based on the target dialect, identifying fragments incompatible with the target dialect, and correspondingly determining which source dialect functional fragment the incompatible fragment belongs to, and taking the source dialect functional fragment as the fragment to be converted;
[0058] A cross-dialect grammar embedding model for converting the fragment to be converted and the corresponding functional specification into a first embedding vector, and the cross-dialect grammar embedding model is used to convert the input into an embedding vector in a vector space measured by functional equivalence;
[0059] A target dialect matching unit for performing similarity matching between the first embedding vector and the embedding vectors of the pre-stored target dialect grammar elements to obtain the matching target dialect grammar elements;
[0060] A large language model conversion unit for inputting the fragment to be converted and the corresponding functional specification, the matching target dialect grammar elements and the corresponding functional specification into a large language model to generate a converted fragment in the form of the target dialect;
[0061] A query replacement unit for replacing the fragment to be converted in the original query with the converted fragment to obtain a converted query.
[0062] In some embodiments, the device further includes:
[0063] A verification unit for verifying the converted query, and if the verification fails, repeating the execution of the large language model conversion unit, the query replacement unit and the verification unit;
[0064] A fragment expansion unit, if the verification unit still indicates verification failure after the number of repeated executions reaches a preset maximum retry number, the fragment expansion unit is used to expand the range of the fragment to be converted to obtain an expanded fragment to be converted, and the expanded fragment to be converted is input into the cross-dialect grammar embedding model, and the cross-dialect grammar embedding model, the target dialect matching unit, the large language model conversion unit and the query replacement unit are executed again.
[0065] According to an embodiment of the present disclosure, an electronic device is provided, the device includes a memory and a processor, the memory is used to store computer instructions that can be run on the processor, and the processor is used to implement the method described in any one of the above when executing the computer instructions.
[0066] According to an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and the program implements the method described in any one of the above when executed by a processor.
[0067] The technical solution proposed by the present disclosure has at least the following beneficial effects:
[0068] (1) High parsing and recognition accuracy
[0069] The present disclosure divides a query into functional fragments through syntax tree matching, and based on the target dialect, parses and identifies incompatible parts to accurately determine the range to be converted;
[0070] (2) Strong matching accuracy
[0071] The present disclosure enhances the representation of syntax elements using functional specifications, performs matching in the functional similarity vector space, and combines structural and semantic features to improve the matching accuracy;
[0072] (3) High conversion reliability
[0073] Some embodiments of the present disclosure introduce double verification of syntax and semantics, support multiple retries and range extension optimization, and use large language models to ensure the rationality of conversion;
[0074] (4) Strong adaptability
[0075] The present disclosure makes full use of the domain knowledge in the database documents, improves the generalization ability of the model through contrastive learning, and can handle the differences in different database versions;
[0076] (5) Low maintenance cost
[0077] Applying the cross-dialect query automatic conversion scheme proposed by the present disclosure, there is no need to manually write and maintain conversion rules, and it can automatically learn new syntax features and can easily adapt to database version updates.
[0078] Other features and advantages of the technical solution proposed by the present disclosure will be described in detail below. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the specification, and are used together with the specification to explain the principles of the specification.
[0080] Figure 1 Shows a flowchart of an automatic cross-dialect query conversion method based on a large language model according to an embodiment of the present disclosure.
[0081] Figure 2 Shows a schematic diagram of the architecture of a cross-dialect syntax embedding model according to an exemplary embodiment of the present disclosure.
[0082] Figure 3 Shows an overall architecture diagram of a database dialect conversion process based on a large language model according to an exemplary embodiment of the present disclosure.
[0083] Figure 4Shows a flowchart of the operation of database dialect conversion based on a large language model according to an exemplary embodiment of the present disclosure.
[0084] Figure 5 Is a schematic structural diagram of an electronic device shown in at least one embodiment of the present disclosure. Detailed implementation manners
[0085] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0086] Embodiments of the present disclosure can be applied to a computer system / server, which can operate together with many other general or special computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with a computer system / server include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, small computer systems, large computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.
[0087] The computer system / server can be described in the general context of computer system-executable instructions (such as program modules) executed by the computer system. Generally, program modules can include routines, programs, object programs, components, logics, data structures, and so on, which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment, where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0088] The cross-dialect query automatic conversion scheme based on a large language model proposed by the present disclosure innovatively combines grammar parsing, embedding matching, and a large language model to construct a unified cross-dialect conversion framework. The query function is segmented through grammar tree matching, the grammar elements are mapped to a function similarity vector space for matching using a cross-dialect grammar embedding model, and then the large language model is used for accurate conversion based on function specifications. At the same time, a verification and optimization mechanism is introduced to ensure the correctness of the conversion result, effectively solving the problems of existing methods in terms of parsing accuracy, matching accuracy, and conversion reliability.
[0089] Figure 1A flowchart of a method for automatic conversion of cross-dialect queries based on a large language model according to an embodiment of the present disclosure is shown. As shown in the figure, the method includes steps 1 to 8.
[0090] Step 1: Extract grammatical elements from the grammar files and database documents of the source dialect and the target dialect offline, generate a syntax tree of the grammatical elements, and annotate the corresponding functional specifications.
[0091] A grammar file is a file that describes the grammatical rules of a database query language. Syntax elements usually refer to the basic components that make up a query statement, including terminal symbols (such as keywords, operators, etc.) and non-terminal symbols (such as expressions, clauses, etc.). Functional specifications refer to the description of the functions and specifications of grammatical elements, which may include information such as functional descriptions, usage restrictions, parameters and default values, and usage examples.
[0092] In some embodiments, grammatical elements may first be extracted offline from grammar files of the source dialect and the target dialect, and a corresponding grammar tree may be generated for each grammatical element using the BNF (Backus-Naur Form) definition, in which leaf nodes represent terminal symbols and internal nodes represent non-terminal symbols; then the grammar tree of each grammatical element is traversed depth-first, each node in the grammar tree is matched with the grammar element nodes in the database document, and the functional specifications obtained from the database document through matching are annotated onto the grammar tree.
[0093] According to this embodiment, the syntax tree structure is accurately constructed using BNF definition, and the functional specification information matching each node is obtained through depth-first traversal, thereby effectively associating the domain knowledge in the database document with the syntax structure.
[0094] This step provides basic data for subsequent query parsing and conversion, including complete syntax tree structure information and rich functional specification information.
[0095] Step 2: parse the original query based on the source dialect to obtain the source dialect grammar tree of the original query.
[0096] The existing SQL grammar parsing technology can be used to generate the source dialect grammar tree of the original query.
[0097] For example, if the source dialect is PostgreSQL, the open source PostgreSQL syntax parser pg_query can be used to generate the corresponding syntax tree. pg_query is based on the PostgreSQL native parser and can accurately identify PostgreSQL-specific syntax structures and functions (such as DATE_PART). The generated syntax tree contains the root node and the subtrees under it, which completely retains the syntax structure information of the query statement.
[0098] For example, if the source dialect is MySQL, the syntax parser provided by MySQL Workbench can be adopted. For example, if the source dialect is Oracle, a general SQL parser such as JSqlParser can be adopted and the corresponding dialect mode can be configured.
[0099] These parsers can convert the query statement into a canonical syntax tree form, providing a basis for subsequent function analysis and conversion.
[0100] Step 3: Match the subtrees of the source dialect syntax tree of the original query with the syntax trees of the source dialect syntax elements to divide the original query into multiple source dialect function segments.
[0101] The syntax tree of the original query can be traversed from bottom to top. For each visited subtree, the subtree can be matched with the syntax tree of the source dialect syntax elements. The matching process for each subtree can include: obtaining the nodes of the subtree and their hierarchical relationship information, comparing this information with the nodes and their hierarchical relationships in the syntax trees of each syntax element. If there is a complete match, the subtree is marked as a function segment. For the successfully matched function segments, the position information in the original query syntax tree can be recorded, and the function specifications of the corresponding syntax elements can be marked on the function segments.
[0102] During the traversal process, if a subtree does not match any syntax element, the traversal can continue to the subtree corresponding to its parent node. In this way, each function segment in the query statement can be completely identified and its corresponding function specification can be marked.
[0103] Conventional string matching methods (such as regular expressions) are difficult to accurately match syntax elements in complex queries, especially when it is necessary to match across multiple clauses containing the same keywords. This embodiment proposes a matching based on the syntax tree. Through the matching based on the syntax tree, the boundaries of each function segment in the query can be identified and the corresponding function specifications can be obtained. This decomposition method based on the syntax structure enables subsequent targeted processing of the function segments that need to be converted, improving the accuracy of the conversion. At the same time, the function specification information marked on the function segments also provides a basis for subsequent processing.
[0104] Step 4: Parse the original query based on the target dialect, identify the segments that are incompatible with the target dialect, and correspondingly determine which source dialect function segment the incompatible segment belongs to, and use this source dialect function segment as the segment to be converted.
[0105] A grammar parser that supports the target dialect can be used to attempt to parse the original query. When the parser fails to recognize certain grammar structures or reports a grammar error, these parts are considered incompatible with the target dialect. For example, when submitting a query with the AGE function of PostgreSQL to the grammar parser of Oracle, an error that the AGE function cannot be recognized will be reported. Any grammar element that causes an incompatibility warning during the generation of the syntax tree can be considered incompatible.
[0106] For the identified incompatible fragments, the corresponding source dialect functional fragments can be determined. Based on the position information of the incompatible fragment in the original query, the source dialect functional fragment containing this position can be found and used as the fragment to be converted, ensuring that the complete grammar element containing the incompatible fragment can be obtained for subsequent high-quality conversion.
[0107] Step 5, use a cross-dialect grammar embedding model to convert the fragment to be converted and the corresponding function specification into a first embedding vector. The cross-dialect grammar embedding model is used to convert the input into an embedding vector in a vector space measured by functional equivalence.
[0108] In some embodiments, before using the cross-dialect grammar embedding model to process the fragment to be converted, various simplification processes can be performed on the fragment to be converted first.
[0109] Replace the custom functions of the database system in the fragment to be converted with equivalent standard SQL functions. Database systems usually implement various custom functions to enhance usability, but these functions are often not supported by other database systems. A preset mapping relationship can be used to map the custom function to an equivalent expression using a general function. For example, the case-insensitive ILIKE function of PostgreSQL is normalized to a standard SQL expression that applies LIKE and performs a lowercase conversion through the preset mapping relationship "str ILIKE pattern::=LOWER(str)LIKE LOWER(pattern)".
[0110] Replace the components irrelevant to the conversion in the fragment to be converted with non-terminal symbols. Some functions or expressions nested in a fragment are usually not used when converting the fragment and can be abstracted as non-terminal symbols. For example, for the fragment "CAST(CONCAT('DR.′',”,name)AS TEXT)", the internal CONCAT function does not affect the conversion of the CAST operation, so it can be abstracted as CAST(column_expr AS TEXT), where column_expr represents a non-terminal symbol.
[0111] Replace the query clauses unrelated to the transformation in the fragment to be transformed with non-terminal symbols. Some query clauses may be irrelevant to the main functions or behaviors of a specific grammar and can be abstracted as non-terminal symbols. For example, in the MySQL query "SELECT * FROM child WHERE age >= 10 LIMIT 2 OFFSET 10", the SELECT clause and the WHERE condition are irrelevant to the transformation of the LIMIT clause. Therefore, these parts can be abstracted, and only the LIMIT part is retained, i.e., select_stmt LIMIT 2 OFFSET 10.
[0112] Through the above simplification process, the grammatical features that need to be focused on can be highlighted, improving the efficiency and accuracy of subsequent matching. At the same time, the basic structure of the fragment to be transformed is retained, without affecting the realization of the overall function.
[0113] The syntax tree and the corresponding function specification of the simplified fragment to be transformed can be input into the cross-dialect grammar embedding model, which converts them into embedding vectors in a space measured by functional equivalence. In some embodiments, this transformation may include:
[0114] Generating a structural embedding vector of the fragment to be transformed through the syntax structure encoder in the cross-dialect grammar embedding model;
[0115] Generating a multi-dimensional feature embedding vector of the function specification through the syntax specification encoder in the cross-dialect grammar embedding model;
[0116] Generating the first embedding vector by comprehensively integrating the structural embedding vector and the multi-dimensional feature embedding vector through a multi-head cross-attention mechanism.
[0117] Specifically, the syntax structure encoder module is used to capture the structural features of the fragment to be transformed. It can use a code encoder model (such as StarEncoder) that contains SQL statements in the training data to extract the structural representation of the query. The input fragment to be transformed can first be tokenized into a sequence of tokens t = [t 1 , t 2 ,..., t n , and then input into the encoder to obtain the structural embedding vector
[0118] The syntax specification encoder can adopt multiple natural language-based encoders to capture the multi-dimensional features of the specification, and then calculate the embedding vector of the specification by weighting in a way of mixed expert knowledge. For example, let d denote the function specification associated with the fragment to be transformed S. M encoders that capture different aspects of the specification can be used to calculate and generate the embedding vectors respectively After that, a gating mechanism can be used to aggregate these embedding vectors to obtain the final canonical embedding vector where the gating weights are determined by the gating function Here, g m (d) is the output of the gating network of the m-th encoder, which allows the model to dynamically assign weights according to the input functional specifications. For example, when elaborating on the CONCAT function in MySQL, a higher weight can be assigned to the encoder that specifically understands and encodes the MySQL specifications
[0119] The structural-specification aggregator can use the multi-head cross-attention mechanism to synthesize the syntactic structure embedding vector and the canonical embedding vector to ensure that the cross-dialect syntactic embedding model focuses on the most relevant canonical parts of the syntactic structure and generates the final first embedding vector E(q):
[0120]
[0121] According to the cross-dialect syntactic embedding model of this embodiment, it can fully consider the information at both the syntactic structure and functional specification semantics levels simultaneously, effectively improving the expression ability of the embedding vector for functional equivalence
[0122] The training of the cross-dialect syntactic embedding model will be introduced later
[0123] Step 6: Perform similarity matching between the first embedding vector and the pre-stored embedding vectors of the target dialect syntactic elements to obtain the matching target dialect syntactic elements
[0124] The pre-stored embedding vectors of the target dialect syntactic elements can be obtained through the cross-dialect syntactic embedding model. The syntax trees and corresponding functional specifications of the syntactic elements of the target dialect can be input into the cross-dialect syntactic embedding model in advance, and the cross-dialect syntactic embedding model converts them into embedding vectors and stores these embedding vectors for subsequent matching. The embedding vectors of these syntactic elements and the first embedding vector generated by the fragment to be converted are in the same vector space, so the similarity can be directly calculated
[0125] In the target dialect, a similarity calculation method can be used to retrieve the syntactic elements that are functionally equivalent to the fragment to be converted. For example, cosine similarity, vector dot product, or Euclidean distance can be used to measure the similarity between the embedding vectors
[0126] The target dialect syntactic element with the highest similarity can be selected as the matching result. When there are multiple candidates with similar similarities, they can be screened in combination with the syntax rules
[0127] Through vector similarity matching, the functional equivalence features learned by the cross-dialect grammar embedding model can be effectively utilized, avoiding the work of manually writing matching rules. In addition, since the model considers the features at both the syntactic structure and functional specification levels when generating embedding vectors, the accuracy of matching can be significantly improved.
[0128] Step 7: Input the to-be-converted fragment and its corresponding functional specification, the matched target dialect grammar element and its corresponding functional specification into the large language model to generate a converted fragment in the target dialect form.
[0129] The to-be-converted fragment and its corresponding functional specification, the matched target dialect grammar element and its corresponding functional specification, the source dialect name, the target dialect name, the prompt, etc. can be input into the large language model together through a prompt template. The large language model synthesizes this information to generate the converted fragment. This method makes full use of the language understanding and generation capabilities of the large language model and can flexibly handle complex conversion scenarios based on the context.
[0130] The large language model performs the conversion according to the information in the prompt template. Through the names of the source dialect and the target dialect, the large language model can understand the context of the conversion. The to-be-converted fragment and its functional specification enable the model to understand the content and functional requirements to be converted. The grammar elements of the target dialect and their functional specifications provide the target form and constraints for the conversion. The prompt can clearly require the model to maintain functional equivalence and follow the grammar norms of the target dialect.
[0131] This embodiment proposes to perform conversion based on the large language model, which has significant advantages. The model can obtain rich SQL knowledge through pre-training, fully understand the syntactic characteristics of different dialects, can perform conversion in combination with functional specifications to ensure functional equivalence, can generate results based on target grammar elements to ensure grammar correctness, and can also handle complex conversion scenarios, such as cases where the syntactic structure needs to be adjusted.
[0132] In particular, when there are no exactly corresponding grammar elements in the target dialect, the model can understand the requirements based on the functional specification and adopt the method of combining multiple grammar elements to achieve equivalent functions; for cases that require context adaptation, the model can make appropriate adjustments according to semantic understanding; when it comes to data type conversion, the model can select the appropriate type based on the understanding of the characteristics of different dialects. These are all the advantages of using the large language model for cross-dialect conversion.
[0133] Step 8: Replace the to-be-converted fragment in the original query with the converted fragment to obtain a converted query.
[0134] Based on the position information of the segment to be converted in the original query, the range to be replaced can be located. Then, the converted segment generated by the large language model is replaced at that position, keeping the other parts of the query unchanged, so as to obtain the complete converted query.
[0135] In some embodiments, the method further includes:
[0136] Verify the converted query;
[0137] If the verification fails, then repeat the steps of using the large language model to generate the converted segment, updating the converted query, and verifying;
[0138] If the verification still fails after the number of repeated executions reaches the preset maximum number of retries, then expand the range of the segment to be converted to obtain an extended segment to be converted, input the syntax tree and corresponding function specifications of the extended segment to be converted into the cross-dialect grammar embedding model to obtain a new first embedding vector, and repeat the similarity matching and subsequent steps based on the new first embedding vector.
[0139] In some embodiments, verifying the converted query includes: performing syntax verification on the converted query using a syntax parser that supports the target dialect, and performing semantic verification on the converted query using a large language model.
[0140] In some embodiments, expanding the range of the segment to be converted to obtain an extended segment to be converted includes:
[0141] Select new segments from adjacent source dialect function segments of the source dialect function segment to which the original segment to be converted belongs;
[0142] Merge the original segment to be converted with the selected new segments to obtain an extended segment to be converted.
[0143] The following is an exemplary illustration.
[0144] Verifying the converted query can include two aspects: syntax verification and semantic verification. For syntax verification, a syntax parser that supports the target dialect can be used to check whether a query that conforms to the syntax specifications of the target dialect is successfully generated. For example, use the BNF definition of the target dialect to generate a syntax tree and detect whether there are still syntax warnings. If a warning is generated, it can be considered that the syntax verification fails; if no warning is generated and the syntax tree is correctly generated, it can be considered that the syntax verification passes. An appropriate existing SQL parser can be selected for syntax verification. For example, ANTL, JSqlParser, or the native parser of a specific database (such as pg_query for PostgreSQL).
[0145] Semantic verification can use large language models to analyze the functional equivalence of the queries before and after transformation. The large language model can reason and verify the functional equivalence of the two queries based on the two fragments before and after transformation and their corresponding functional specifications. If the large language model reasons that the functions are equivalent, the semantic verification passes; otherwise, the semantic verification fails. Syntax verification can be performed first, and then semantic verification can be performed on the transformed query that passes the syntax verification. If both semantic verification and syntax verification pass, the transformation can be considered successful. A large language model that supports SQL understanding, such as the GPT series of models, can be used for semantic verification.
[0146] When the verification fails, an iterative optimization strategy can be adopted. It is possible to return to step 7 above, use the large language model again to generate a new transformed fragment in the target dialect form, and replace the fragment to be transformed in the original query with the new transformed fragment to obtain a new transformed query, and then perform semantic verification and syntax verification on the new transformed query. If, after the preset maximum number of retries, a transformed query that can pass the verification has still not been obtained, consider determining a new fragment to be transformed, also known as the extended fragment to be transformed, through fragment expansion, and return to step 5 to obtain a new first embedding vector based on the extended fragment to be transformed, and repeat processes such as similarity matching, large language model transformation, and syntax verification and semantic verification.
[0147] The fragment expansion proposed in this embodiment is a progressive strategy from local to global. Starting from the source dialect functional fragment to which the original fragment to be transformed belongs, select its adjacent source dialect functional fragments for merging to expand the transformation scope. The selection of adjacent fragments is based on the structural relationship of the syntax tree, and preference is given to functional fragments that have a direct syntax dependency on the current fragment.
[0148] For example, for the SQL query:
[0149]
[0150] If the AGE function fails to be transformed multiple times, the scope can be extended to the entire SELECT clause, and the AGE function and the adjacent EXTRACT function can be processed together to expand the transformation scope.
[0151] The extended fragment is used as the new unit to be transformed. As described above, return to step 5 to obtain a new first embedding vector based on the extended fragment to be transformed, and repeat processes such as similarity matching, large language model transformation, and syntax verification and semantic verification. This process can be iterated multiple times until a suitable transformation result is found or the maximum expansion scope is reached.
[0152] The verification and optimization mechanism proposed in the above embodiments ensures the correctness of the conversion result through double verification, improves the conversion success rate through the retry mechanism, adopts a progressive expansion strategy to enable handling of complex dependencies, and avoids unnecessary large-scale conversions at the same time.
[0153] Figure 2 FIG. shows a schematic structural diagram of a cross-dialect grammar embedding model according to an exemplary embodiment of the present disclosure. As Figure 2 shown, the cross-dialect grammar embedding model adopts a design architecture of a double encoder plus an aggregator. The grammar structure encoder at the lower left processes the input grammar tree through a code encoder and a feedforward neural network, such as grammar tree tokens like "CONCAT", etc. The grammar specification encoder at the lower right includes M parallel natural language encoders for processing grammar specifications, such as specification tokens like "Concatenates", "Text", "Arguments", etc. The encoding results are processed through N feedforward neural networks and optimized through residual connections and layer normalization.
[0154] At the top layer, the structure-specification aggregator receives the outputs of the two encoders, calculates the attention scores through dot product operations, and uses the softmax function to obtain the attention weights, and finally generates the SQL embedding vector. The contrastive learning module on the right side of the aggregator optimizes the representation ability of the vector space by strengthening the connection with positive samples (green) and pulling away from negative samples (red).
[0155] The hierarchical architecture design proposed in this embodiment enables the embedding model to capture both grammar structure features and functional semantic information simultaneously, providing a good vector representation for subsequent similarity matching.
[0156] The cross-dialect grammar embedding model can be trained using contrastive learning.
[0157] For grammar element samples collected from the official documents (which can be abbreviated as database documents) and grammar files of each database system where represents the keyword or structure of the grammar element, represents the functional specification, and positive samples and hard negative samples can be generated for subsequent training.
[0158] In some embodiments, the positive sample set of sample s i can be constructed in the following way:
[0159] Perform synonym replacement or word order adjustment on the functional specification corresponding to sample s i to obtain a specification variant as the positive sample, which has a different expression from sample s i but the same semantics;
[0160] Take the syntax elements with the same keywords and functions in other database systems as positive samples as those in sample s i For example, the same built-in functions in MySQL and Oracle may have different specification styles to help the model identify equivalent functions in different database dialects;
[0161] Use a database dialect conversion tool to convert sample s i The equivalent syntax elements obtained by converting to other database systems are used as positive samples. For example, using a tool such as SQLGlot to convert the query of s i to other database dialects to generate a new query Each converted query is combined with its corresponding specification to form a new binary tuple of syntax elements s k i For s i is a positive sample of s
[0162] In some embodiments, the difficult negative sample set of sample s i can be constructed in the following way: cluster the syntax elements that are positive samples of each other, and select the syntax elements that belong to different clusters from sample s i and have a higher similarity between the embedding vectors of sample s i than the similarity between the embedding vectors of sample s i and all positive samples as difficult negative samples.
[0163] Specifically, after generating positive samples, they can be clustered according to the functional equivalence of the syntax elements. The syntax elements that are positive samples of each other are grouped into the same cluster, and the syntax elements that belong to different clusters from syntax element s i are used as candidate negative samples; calculate the similarities between the embedding vectors of syntax element s i and the embedding vectors of each candidate negative sample and the embedding vector of the syntax element respectively, and take the maximum similarity between the embedding vector of syntax element s i and the embedding vectors of its positive samples as the measurement criterion. If the similarity between syntax element s i and a certain candidate negative sample is greater than this measurement criterion, then this candidate negative sample is used as the difficult negative sample of syntax element s i
[0164] If only negative samples are selected based on belonging to different clusters, it may introduce syntax elements with low similarity, irrelevance, or large differences in the embedding vectors of syntax element s i as negative samples, resulting in significant noise. According to this embodiment, those negative samples with high similarity in the vector space but non-equivalent actual functions can be selected, which helps to improve the discrimination ability of the model.
[0165] In some embodiments, the cross-dialect grammar embedding model is trained using the following contrastive loss function to increase the scaled cosine similarity between the embedding vectors of the positive samples and decrease the scaled cosine similarity between the embedding vectors of the hard negative samples:
[0166]
[0167] Where, is the loss value of the i-th sample s i , represents the set of positive samples paired with the sample s i , represents the set of positive samples in the positive sample set , represents the set of hard negative samples paired with the sample s i , zip represents the scaled cosine similarity between the embedding vector of the sample s i and the embedding vector of the matching positive sample s p , z in represents the scaled cosine similarity between the embedding vector of the sample s i and the embedding vector of the matching hard negative sample s n .
[0168] By minimizing the above loss function, the cross-dialect grammar embedding model can be trained to increase the similarity between the sample and its positive sample, while decreasing the similarity with the hard negative sample, so as to learn a better representation of functional equivalence.
[0169] Figure 3 FIG. shows an overall architecture diagram of a database dialect conversion process based on a large language model according to an exemplary embodiment of the present disclosure. The upper part is a schematic diagram of function-based query preprocessing, and the lower part is a schematic diagram of local-to-global conversion. This figure shows the overall architecture of the database dialect conversion method based on a large language model.
[0170] In the query preprocessing stage, the system first parses the input SQL query (Q) to obtain a syntax tree, which contains node structures such as select_smt, func_expr, where_clause, etc. Subsequently, it enters the query partitioning link, and based on the syntax tree annotated with function specifications, the query is functionally segmented to obtain multiple function segments based on the source dialect. In the query simplification stage, the following two types of rules are adopted: Rule 1 "custom function standardization", such as mapping the ILIKE operation (q 3 ) to a standard form; Rule 2 "abstract conversion-irrelevant components", such as converting the CONCAT structure to a non-terminal symbol name_expr. Finally, a query syntax tree that is segmented and simplified according to functions is obtained.
[0171] In the local-to-global conversion phase, first, the segments to be converted are identified through a verification-driven process for selecting operations to be converted, such as the AGE function in q shown in the figure. When the local conversion fails, the system expands the scope to associated segments, such as expanding from q 1 to q 1 to q 2 . The currently determined segment q i to be converted enters the grammar enhancement conversion link, where grammar matching is performed based on a cross-dialect grammar embedding model, and the converted segment q i (including function specifications) and the matched target grammar (including function specifications) are converted using a large language model. The conversion result is verified through a hybrid query to ensure correctness. The architecture of the cross-dialect embedding model is shown on the far right, including a grammar structure encoder, a grammar specification encoder, and a structure-specification aggregator, and the model performance is optimized through contrastive learning. The entire process demonstrates a progressive conversion strategy from local to global, ensuring the accuracy and reliability of the conversion.
[0172] Figure 4 The flowchart of the database dialect conversion based on a large language model according to an exemplary embodiment of the present disclosure is shown. As shown in the figure, first, the input SQL query statement is parsed to generate a syntax tree. Then, it enters the query preprocessing stage, including two steps of query segmentation and query simplification, where the query is divided into multiple functional segments and standardized.
[0173] In the conversion processing stage, first, it is detected through a hybrid query verification whether there are segments to be converted. If the verification fails, the segment to be converted is located. If the verification fails and exceeds the attempt limit, an expansion operation is triggered to expand the conversion scope and reprocess.
[0174] For the determined segment to be converted, the system generates a prompt (Prompt) based on its syntax tree and function specifications and inputs it into the large language model. The large language model generates the converted result, replaces it in the original query, and obtains the complete converted query. Finally, the correctness of the conversion result is ensured through a hybrid query verification. If the verification passes, the final SQL query is output; otherwise, the conversion processing process is repeated.
[0175] According to an embodiment of the present disclosure, there is also provided a cross-dialect query automatic conversion device based on a large language model, including:
[0176] An offline information extraction unit for offline extracting grammar elements from the grammar files and database documents of the source dialect and the target dialect, generating a syntax tree of the grammar elements, and annotating the corresponding function specifications;
[0177] A first parsing unit for parsing the original query based on the source dialect to obtain the source dialect syntax tree of the original query;
[0178] A function segmentation unit, which is used to match the sub - tree of the source dialect syntax tree of the original query with the syntax tree of the source dialect grammar elements, so as to divide the original query into multiple source dialect function segments;
[0179] An incompatibility recognition unit, which is used to parse the original query based on the target dialect, recognize the segments that are incompatible with the target dialect, and correspondingly determine which source dialect function segment the incompatible segment belongs to, and use the source dialect function segment as the segment to be converted;
[0180] A cross - dialect grammar embedding model, which is used to convert the segment to be converted and the corresponding function specification into a first embedding vector, and the cross - dialect grammar embedding model is used to convert the input into an embedding vector in a vector space measured by functional equivalence;
[0181] A target dialect matching unit, which is used to perform a similarity match between the first embedding vector and the embedding vectors of the pre - stored target dialect grammar elements to obtain the matching target dialect grammar elements;
[0182] A large - language model conversion unit, which is used to input the segment to be converted and the corresponding function specification, the matching target dialect grammar elements and the corresponding function specification into a large - language model to generate a converted segment in the form of the target dialect;
[0183] A query replacement unit, which is used to replace the segment to be converted in the original query with the converted segment to obtain a converted query.
[0184] In some embodiments, the device further includes:
[0185] A verification unit, which is used to verify the converted query;
[0186] Wherein, if the verification fails, the large - language model conversion unit, the query replacement unit and the verification unit are repeatedly executed;
[0187] If the verification still fails after the number of repeated executions reaches the preset maximum number of retries, the range of the segment to be converted is expanded by a segment expansion unit to obtain an expanded segment to be converted, and the expanded segment to be converted is input into the cross - dialect grammar embedding model, and the cross - dialect grammar embedding model, the target dialect matching unit, the large - language model conversion unit and the query replacement unit are executed again.
[0188] For other details and features of this embodiment, please refer to the relevant description above.
[0189] Figure 5An electronic device provided by at least one embodiment of the present disclosure, the device includes a memory and a processor, the memory is used to store computer instructions that can run on the processor, and the processor is used to implement the cross-dialect query automatic conversion method based on the large language model described in any embodiment or implementation manner of the present disclosure when executing the computer instructions.
[0190] At least one embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the cross-dialect query automatic conversion method based on the large language model described in any embodiment or implementation manner of the present disclosure.
[0191] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0192] Each embodiment in this specification is described in a progressive manner. The same or similar parts between each embodiment can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the embodiment of the data processing device, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.
[0193] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0194] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or one or more of them in combination. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier to be executed by, or to control the operation of, data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode and transmit information to the appropriate receiver apparatus for execution by the data processing apparatus. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0195] The processes and logical flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform the functions by operating on input data and generating output. The processes and logical flows can also be performed by, or the apparatus can be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0196] Suitable computers for executing computer programs include, by way of example, general and / or special purpose microprocessors, or any other type of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory and / or a random access memory. Basic components of a computer include a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, etc., or the computer will be operatively coupled to such mass storage devices to receive data therefrom or to transfer data thereto, or both. However, a computer need not have such devices. In addition, a computer may be embedded in another device, such as a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name just a few.
[0197] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0198] Although this specification contains many specific implementation details, these should not be construed as limiting the scope of any invention or the scope of what is claimed, but rather as mainly describing the features of specific embodiments of a particular invention. Certain features that are described in multiple embodiments in this specification may also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although features may operate in certain combinations as described above and even be claimed initially in such a combination, one or more features from the claimed combination may in some cases be removed from the combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.
[0199] Similarly, although operations are depicted in the figures in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or sequentially, or that all illustrated operations be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0200] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims may be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures are not necessarily in the particular order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0201] The above description is only the preferred embodiment of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope protected by one or more embodiments of this specification.
Claims
1. A method for automatic conversion of cross-dialect queries based on a large language model, characterized in that: include: Extract grammatical elements from the grammar files and database documents of the source and target dialects offline, generate syntax trees of the grammatical elements, and annotate the corresponding functional specifications; Parsing the original query based on the source dialect to obtain the source dialect grammar tree of the original query; matching a subtree of a source dialect syntax tree of an original query with a syntax tree of a source dialect syntax element to divide the original query into a plurality of source dialect functional fragments; Parsing the original query based on the target dialect, identifying segments that are incompatible with the target dialect, and correspondingly determining to which source dialect functional segment the incompatible segment belongs, and using the source dialect functional segment as the segment to be converted; Converting the to-be-converted segment and the corresponding functional specification into a first embedding vector using a cross-dialect grammar embedding model, wherein the cross-dialect grammar embedding model is used to convert the input into an embedding vector in a vector space measured by functional equivalence; Performing similarity matching between the first embedding vector and the pre-stored embedding vector of the target dialect grammatical element to obtain a matched target dialect grammatical element; Inputting the segment to be converted and the corresponding functional specification, the matched target dialect grammatical element and the corresponding functional specification into the large language model to generate a converted segment in the target dialect form; The to-be-converted segment in the original query is replaced with the converted segment to obtain a converted query.
2. The method according to claim 1, characterized in that Extract grammatical elements from the grammar files and database documents of the source and target dialects offline, generate syntax trees for the grammatical elements, and annotate the corresponding functional specifications, including: Extract grammatical elements from the grammar files of the source dialect and the target dialect offline, and use BNF definition to generate a corresponding grammar tree for each grammatical element, where leaf nodes represent terminal symbols and internal nodes represent non-terminal symbols; The syntax tree of each syntax element is traversed depth-first, each node in the syntax tree is matched with the syntax element node in the database document, and the functional specification obtained from the database document through matching is annotated on the syntax tree.
3. The method according to claim 1, characterized in that Before inputting the to-be-converted segment into the cross-dialect grammar embedding model, the method further includes performing the following simplification processing on the to-be-converted segment: Replace the database system custom functions in the fragment to be converted with equivalent standard SQL functions; Replace the components irrelevant to the conversion in the fragment to be converted with non-terminal symbols; Replace query clauses in the fragment to be converted that are not relevant to the conversion with non-terminal symbols.
4. The method according to claim 1, characterized in that The to-be-converted segment and the corresponding functional specification are converted into a first embedding vector using a cross-dialect syntax embedding model, including: Generate a structural embedding vector of the segment to be converted through the grammatical structure encoder in the cross-dialect grammar embedding model; Generate a multi-dimensional feature embedding vector of functional norms through the grammatical norm encoder in the cross-dialect grammar embedding model; The structure embedding vector and the multi-dimensional feature embedding vector are integrated through a multi-head cross attention mechanism to generate the first embedding vector.
5. The method according to claim 4, characterized in that The grammatical specification encoder in the cross-dialect grammar embedding model generates a multi-dimensional feature embedding vector of functional specifications, including: Use multiple encoders to capture features in functional specifications from different dimensions and mix expert knowledge to generate multiple feature embedding vectors of different dimensions; Dynamically calculate the feature weights of each dimension based on the input functional specifications through a gating mechanism; The feature embedding vectors are weightedly aggregated based on the feature weights to obtain the multi-dimensional feature embedding vector.
6. The method according to claim 1, characterized in that The method also includes training the cross-dialect grammar embedding model using contrastive learning, including: The cross-dialect grammatical embedding model is trained using the following contrastive loss function to increase the scaled cosine similarity between the output embedding vector and the embedding vector of the positive sample, and to reduce the scaled cosine similarity between the output embedding vector and the embedding vector of the hard negative sample: in, is the i-th sample s i The loss value, Represents the sample s i The set of paired positive samples, Represents the positive sample set The number of positive samples, Represents the sample s i The set of paired hard negative examples, z ip Represents sample s i The embedding vector of p The scaled cosine similarity between the embedding vectors of in Represents sample s i The embedding vector of and the matching hard negative sample s n The scaled cosine similarity between the embedding vectors of .
7. The method according to claim 6, characterized in that The method further comprises: The samples are constructed by i The positive sample set: For sample s i The corresponding functional specifications are replaced with synonyms or word order is adjusted to obtain the standard variants as positive samples; The other database systems have i Syntactic elements with the same keywords and functions are used as positive samples; Use the database dialect conversion tool to convert the sample s i The equivalent grammatical elements obtained by conversion to other database systems are taken as positive samples; The samples are constructed by i The set of hard negative samples: Cluster the grammatical elements that are mutually positive samples and select i Belong to different clusters and have the same i The similarity between the embedding vectors of is higher than that of sample s i The syntactic element with the similarity with the embedding vectors of all positive samples is regarded as a hard negative sample.
8. The method according to claim 1, characterized in that The method further comprises: Verify the converted query; If the verification fails, the steps of generating the transformed segment using the large language model, updating the transformed query, and verifying are repeated; If the verification still fails after the number of repeated executions reaches the preset maximum number of retries, the range of the segment to be converted is expanded to obtain an extended segment to be converted, and the syntax tree and corresponding functional specifications of the extended segment to be converted are input into the cross-dialect grammar embedding model to obtain a new first embedding vector, and similarity matching and subsequent steps are performed again based on the new first embedding vector.
9. The method according to claim 8, characterized in that Verify the converted query, including: Use a parser that supports the target dialect to validate the syntax of the converted query; Use a large language model to semantically verify the converted queries.
10. The method according to claim 8, characterized in that The range of the segments to be converted is expanded to obtain the expanded segments to be converted, including: Selecting a new segment from adjacent source dialect functional segments of the source dialect functional segment to which the original segment to be converted belongs; The original segment to be converted is merged with the selected newly added segment to obtain an extended segment to be converted.
11. A cross-dialect query automatic conversion device based on a large language model, characterized in that: include: An offline information extraction unit, used for offline extracting grammatical elements from grammar files and database documents of the source dialect and the target dialect, generating a grammar tree of the grammatical elements, and marking corresponding functional specifications; A first parsing unit, configured to parse an original query based on a source dialect to obtain a source dialect grammar tree of the original query; a functional segmentation unit, for matching a subtree of a source dialect grammar tree of an original query with a grammar tree of a source dialect grammar element, so as to divide the original query into a plurality of source dialect functional segments; An incompatible identification unit, used for parsing the original query based on the target dialect, identifying the segment incompatible with the target dialect, and correspondingly determining to which source dialect functional segment the incompatible segment belongs, and using the source dialect functional segment as the segment to be converted; A cross-dialect syntax embedding model, used to convert the to-be-converted segment and the corresponding functional specification into a first embedding vector, wherein the cross-dialect syntax embedding model is used to convert the input into an embedding vector in a vector space measured by functional equivalence; a target dialect matching unit, configured to perform similarity matching between the first embedding vector and a pre-stored embedding vector of a target dialect grammatical element to obtain a matched target dialect grammatical element; A large language model conversion unit, configured to input the segment to be converted and the corresponding functional specification, the matched target dialect grammatical element and the corresponding functional specification into the large language model, and generate a converted segment in the target dialect form; The query replacement unit is used to replace the to-be-converted segment in the original query with the converted segment to obtain the converted query.
12. The device according to claim 11, characterized in that The device also includes: A verification unit, configured to verify the converted query, and if the verification fails, repeatedly executing the large language model conversion unit, the query replacement unit and the verification unit; A fragment expansion unit, if the verification unit still indicates verification failure after the number of repeated executions reaches a preset maximum number of retries, the fragment expansion unit is used to expand the range of the fragment to be converted to obtain an extended fragment to be converted, the extended fragment to be converted is input into the cross-dialect grammar embedding model, and the cross-dialect grammar embedding model, the target dialect matching unit, the large language model conversion unit and the query replacement unit are executed again.
13. An electronic device, characterized in that: The device comprises a memory and a processor, wherein the memory is used to store computer instructions executable on the processor, and the processor is used to implement the method according to any one of claims 1 to 10 when executing the computer instructions.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Video generation method and device based on virtual object and electronic equipment
CN117292022A
Method for converting natural language to SQL (Structured Query Language) based on large language model
CN118643050A
Method for constructing data visual query based on DSL and SQL abstract syntax tree
CN119149581A
System and Method for Transpilation of Machine Interpretable Languages
US20230129994A1
Cited By
Contract detail extraction and classification method and system
CN121435939A
Structured query statement generation method and device and related product
CN121658508A