Conversational text-to-SQL pre-training method and apparatus based on fine-grained link information

By using a pre-training method of fine-grained link information in conversational Text-to-SQL tasks, the link relationship between statements and patterns is obtained, and the neural network model is pre-trained, which solves the problems of low accuracy, cross-domain failure and dialogue context forgetting in the existing technology, and achieves more efficient Text-to-SQL semantic analysis.

WO2025118247A1PCT designated stage expired Publication Date: 2025-06-12SHENZHEN INST OF ADVANCED TECH

Patent Information

Application Number
PCT/CN2023/137151
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

In the prior art, the dialogue Text-to-SQL semantic analysis task has problems such as low accuracy, cross-domain failure and dialogue context forgetting.

Method used

The dialogue Text-to-SQL pre-training method based on fine-grained link information is adopted. By obtaining the text statements and database patterns of the conversation, the statement link relationship between the statements and the pattern link relationship between the statements and the database, these loss functions are used to pre-train the neural network model to obtain the pre-trained Text-to-SQL model.

Benefits of technology

It effectively improves the performance of the conversational Text-to-SQL model, improves accuracy, enhances cross-domain adaptability and dialogue context memory, and solves the problems of low accuracy, cross-domain failure and dialogue context forgetting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023137151_12062025_PF_FP_ABST
    Figure CN2023137151_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the technical field of computers. Disclosed are a conversational Text-to-SQL pre-training method and apparatus based on fine-grained link information. The method comprises: obtaining text statements of a conversation and a database schema, wherein the text statements comprise the current statement and a historical statement of the conversation; obtaining a statement link relationship between the statements in the conversation on the basis of the text statements of the conversation and the database schema, and obtaining a first loss function; obtaining a schema link relationship between the conversation and a database on the basis of the current statement of the conversation and the database schema, and obtaining a second loss function; and pre-training a neural network model on the basis of the first loss function and the second loss function to obtain a pre-trained Text-to-SQL model. The present application solves the problems in the prior art of low accuracy, cross-domain failure, and conversation context forgetting.
Need to check novelty before this filing date? Find Prior Art

Description

Conversational Text-to-SQL pre-training method and device based on fine-grained link information Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a conversational Text-to-SQL pre-training method and device based on fine-grained link information. Background Art

[0002] The main function of the Text-to-SQL semantic parsing task is to convert the descriptive text related to a piece of tabular data into a corresponding SQL query. To better meet the needs of real-life applications, this technology is generally integrated into conversational robots, that is, in a question-and-answer format. At the same time, the conversational Text-to-SQL semantic parsing task has also been expanded. That is, in a multi-round dialogue scenario, the text describing people's needs for the current tabular data is converted into a corresponding SQL query. The conversational Text-to-SQL semantic parsing task enables people to continuously interact with a piece of tabular data. People do not need to remember or fully describe historical information. They can also continue to ask more in-depth questions to obtain the necessary knowledge and information.

[0003] This shows that conversational Text-to-SQL semantic parsing tasks are more in line with actual needs and more convenient for people's daily lives. One real-life application of conversational Text-to-SQL semantic parsing tasks is a customer service robot that uses tabular data as a knowledge base.

[0004] However, there are few related inventions in the existing technology, and most of them are still in the research stage. The existing related inventions have problems such as low accuracy, cross-domain failure, and forgetting of dialogue context.

[0005] Therefore, there is an urgent need for a conversational Text-to-SQL pre-training method based on fine-grained link information that has high accuracy, cross-domain adaptability, and strong conversation context memory.

[0006] Summary of the Invention

[0007] The embodiments of the present application provide a conversational Text-to-SQL pre-training method, device, electronic device, and storage medium based on fine-grained link information to address the problems of low accuracy, cross-domain failure, and forgetting of conversation context in related technologies.

[0008] In order to solve the above technical problems, the technical solutions adopted in this application are:

[0009] According to one aspect of the present application, a conversational Text-to-SQL pre-training method based on fine-grained link information includes obtaining text sentences and database schemas of a conversation; the text sentences include current sentences and historical sentences of the conversation; obtaining statement link relationships between sentences in the conversation based on the text sentences and database schemas of the conversation, and obtaining a first loss function; obtaining a schema link relationship between the conversation and the database based on the current sentence and database schema of the conversation, and obtaining a second loss function; and pre-training a neural network model based on the first and second loss functions to obtain a pre-trained Text-to-SQL model.

[0010] In an exemplary embodiment, a sentence link relationship between sentences in the conversation is obtained based on the text sentences and database schema of the conversation, and obtaining a first loss function is achieved by the following steps: obtaining a word-level syntactic relationship between the current sentence and the text sentence in the conversation through a coreference resolution tool to obtain a sentence supervision label; obtaining a sentence heuristic representation of the conversation and the database based on the text sentences and database schema of the conversation; and obtaining a first loss function based on the sentence supervision label and the heuristic representation.

[0011] In an exemplary embodiment, obtaining a sentence heuristic representation of the conversation and the database based on the text sentences and database schema of the conversation is achieved by the following steps: obtaining token representations of all sentences in the conversation based on the text sentences and database schema of the conversation; mapping the token representations to sentence representations to obtain sentence representations of all sentences in the conversation; and performing matrix multiplication on the sentence representations to obtain a sentence heuristic representation.

[0012] In an exemplary embodiment, a pattern link relationship between the conversation and the database is obtained based on the current statement of the conversation and the database pattern, and a second loss function is obtained by the following steps: calculating the SQL structure similarity between the current statement and the historical statements in the database pattern; obtaining a pattern supervision label based on the SQL structure similarity and a set similarity threshold; obtaining a pattern heuristic representation of the conversation and the database based on the text statement of the conversation and the database pattern; and obtaining a second loss function based on the pattern supervision label and the pattern heuristic representation.

[0013] In an exemplary embodiment, obtaining a pattern heuristic representation of the conversation and the database based on the text sentences and database schema of the conversation is achieved by the following steps: obtaining token representations of all sentences in the conversation based on the text sentences and database schema of the conversation; mapping the token representations to sentence representations and pattern representations respectively to obtain sentence representations and pattern representations of all sentences in the conversation; and obtaining a pattern heuristic representation based on the sentence representations and pattern representations.

[0014] In an exemplary embodiment, the neural network model is a trained machine learning model with the ability to learn the context of natural language text; the method also includes the following steps: randomly selecting tokens to replace with mask tokens during the pre-training process, and obtaining a third loss function based on the mask tokens and the supervisory label, so that the third loss function participates in the pre-training; the supervisory label is the true label of the mask token.

[0015] In an exemplary embodiment, pre-training a neural network model according to the first loss function and the second loss function to obtain a pre-trained Text-to-SQL model includes the following steps: summing the first loss function, the second loss function, and the third loss function according to the homoscedastic uncertainty of the first loss function, the second loss function, and the third loss function to obtain a total loss function; optimizing the parameters of the neural network model according to the total loss function until the parameters meet the set conditions, thereby obtaining a pre-trained Text-to-SQL model.

[0016] According to one aspect of the present application, a conversational Text-to-SQL pre-training device based on fine-grained link information includes a data acquisition module for acquiring text statements and database schemas of a conversation; the text statements include current statements and historical statements of the conversation; a statement linking module for acquiring statement link relationships between statements in the conversation based on the text statements and database schemas of the conversation, and obtaining a first loss function; a pattern linking module for acquiring a pattern link relationship between the conversation and the database based on the current statement and database schema of the conversation, and obtaining a second loss function; and a model pre-training module for pre-training a neural network model based on the first loss function and the second loss function to obtain a pre-trained Text-to-SQL model.

[0017] According to one aspect of the present application, an electronic device includes at least one processor and at least one memory, wherein the memory stores program instructions or codes; the program instructions or codes are loaded and executed by the processor, so that the electronic device implements the conversational Text-to-SQL pre-training method based on fine-grained link information as described above.

[0018] According to one aspect of the present application, a storage medium stores program instructions or codes thereon, which are loaded and executed by a processor to implement the conversational Text-to-SQL pre-training method based on fine-grained link information as described above.

[0019] According to one aspect of the present application, a computer program product includes program instructions or codes, which are stored in a storage medium. A processor of an electronic device reads the program instructions or codes from the storage medium, loads and executes the program instructions or codes, so that the electronic device implements the conversational Text-to-SQL pre-training method based on fine-grained link information as described above.

[0020] The beneficial effects of the technical solution provided by this application are:

[0021] In the above technical solution, this application solves the problems of low accuracy, cross-domain failure, and forgetting of conversation context in related technologies.

[0022] Specifically, this application first obtains the text sentences and database schema of the conversation, where the text sentences include the current sentences and historical sentences of the conversation, and then obtains the sentence link relationship between the sentences in the conversation based on the text sentences and database schema to obtain a first loss function, learns the complex syntactic relationship between natural languages ​​in the conversational Text-to-SQL scenario, and then obtains the schema link relationship between the conversation and the database based on the current sentence and database schema to obtain a second loss function, captures the fine-grained link relationship between the sentence and the database schema, and finally pre-trains the neural network model according to the first loss function and the second loss function to obtain a pre-trained Text-to-SQL model, thereby effectively improving the performance of the conversational Text-to-SQL model, thereby effectively solving the problems of low accuracy, cross-domain failure, and forgetting of conversation context in related technologies.

[0023] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0025] FIG1 is a schematic diagram of an implementation environment according to the present application;

[0026] FIG2 is a flowchart illustrating a conversational Text-to-SQL pre-training method based on fine-grained link information according to an exemplary embodiment;

[0027] FIG3 is a flow chart of step 230 in one embodiment of the embodiment corresponding to FIG2 ;

[0028] FIG4 is a flow chart of step 330 in one embodiment of the embodiment corresponding to FIG3 ;

[0029] FIG5 is a flow chart of step 250 in one embodiment of the embodiment corresponding to FIG2 ;

[0030] FIG6 is a schematic structural diagram of the Text-to-SQL model in the embodiment corresponding to FIG2 ;

[0031] FIG7 is a schematic diagram illustrating the implementation results of a conversational Text-to-SQL pre-training method based on fine-grained link information in an application scenario;

[0032] FIG8 is a block diagram of a conversational Text-to-SQL pre-training device based on fine-grained link information according to an exemplary embodiment;

[0033] FIG9 is a schematic structural diagram of a server according to an exemplary embodiment;

[0034] Fig. 10 is a block diagram of an electronic device according to an exemplary embodiment.

[0035] The above-mentioned drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0036] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.

[0037] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0038] The following is an introduction and explanation of several terms involved in this application:

[0039] Multi-Task Learning (MLM) is a machine learning method that aims to simultaneously learn to solve multiple related tasks. Traditional machine learning methods typically model and train for a single, specific task. MLM, on the other hand, uses shared underlying representations to simultaneously learn multiple tasks, thereby improving the model's generalization capabilities. In MLM, the model is designed to handle multiple distinct but related tasks, such as classification, regression, and sequence labeling. By sharing model parameters and underlying representations, MLM can extract more information from multiple tasks, helping to improve the model's performance on each task. The advantage of MLM lies in its ability to leverage the interrelationships between multiple related tasks to improve generalization. By sharing underlying representations, the model can learn more general and abstract features, resulting in better performance when handling new samples. Furthermore, MLM can address data scarcity through multi-task learning. When the number of samples for a particular task is limited, shared learning can be achieved by leveraging samples from other tasks, thereby improving model performance.

[0040] Fine-grained link information is a key concept in information extraction and natural language processing. It refers to the identification and extraction of fine-grained relationships and connections between entities in text or corpora. Traditional entity relationship extraction tasks typically focus solely on whether a relationship exists between two entities (e.g., a spouse relationship between two people). Fine-grained link information, on the other hand, aims to further refine this relationship, providing a more detailed description and classification of the specific attributes or connections between the entities. For example, in a news article, traditional entity relationship extraction might only extract that there is a "working relationship" between person A and person B. Fine-grained link information can further identify the specific attributes of this working relationship, such as "A is B's supervisor" or "A is B's colleague." Such fine-grained link information can provide more precise and detailed relationship descriptions, facilitate better understanding of text content, and support applications such as question-answering systems and information retrieval. Extracting fine-grained link information typically requires a combination of machine learning and natural language processing techniques, including entity recognition, relationship extraction, and semantic role labeling. These technologies can help identify and extract entities and relations in text, classify and categorize them, and thus obtain rich fine-grained link information.

[0041] A database schema is a blueprint or plan used to define and organize the structure and relationships of data in a database. It describes the organization and relationships of database elements such as tables, columns, keys, and constraints. A database schema defines the entities, attributes, and relationships within the database. It specifies the structure of each table, the data type and constraints of each column, and the relationships and connections between tables. A database schema can contain multiple tables, each with multiple columns, which define the types and constraints of the data stored in the table. Database schemas are typically created and defined by database administrators or developers when designing and creating a database. They use modeling tools or languages ​​(such as SQL) provided by the database management system (DBMS) to define the structure and properties of elements such as tables, columns, keys, and constraints. A database schema can be considered the "blueprint" of a database, defining the structure and organization of the data stored in the database. It provides a framework for managing and manipulating data within the database and ensuring data integrity and consistency.

[0042] With the development of neural networks, researchers have begun exploring how to combine them to solve the text-to-SQL semantic parsing task. The most representative example is the end-to-end sequence-to-sequence (Seq2seq) framework. This framework uses a neural network model to directly convert one sequence into another. The input and output sequences can be of different lengths and types. Therefore, this framework perfectly fits the task definition of text-to-SQL semantic parsing, which involves converting a text sequence in an unstructured data format into a SQL query sequence in a structured data format. This end-to-end sequence-to-sequence framework primarily consists of two modules: an encoder module, which encodes the unstructured text sequence into an intermediate representation; and a decoder module, which decodes the intermediate representation obtained by the encoder module into a structured SQL query sequence.

[0043] First, regarding the encoder module, in terms of input, some studies have concatenated the text sequence with a given database schema to obtain an input sequence in order to enable the encoder model to learn the link relationship between the text sequence and the database schema (i.e., table names and column names). This input interacts within the model, allowing the model to learn the knowledge of the link relationship. On the other hand, regarding the decoder module, since SQL query sequences have specific syntax and logic, most methods use an abstract syntax neural network model to generate SQL query sequences. This type of method first generates an abstract syntax tree that conforms to the SQL syntax, and then further generates a logical form based on this syntax tree to obtain a complete SQL query sequence. Other studies have improved the effect by constraining the decoding process, directly applying SQL syntax constraints to each search stage during the decoding beam search stage.

[0044] In conversational Text-to-SQL tasks, it is crucial to leverage contextual history information in the conversation, especially syntactic coreference relations in the context, to ensure accurate SQL statement generation. SCoRe (Yu et al., 2021) focuses on identifying semantic switches between adjacent sentences but ignores long-range semantic dependencies. STaR (Cai et al., 2022) captures sentence-level semantic dependencies through SQL similarity comparison but does not effectively model fine-grained word / token-level contextual coreference relations.

[0045] At the same time, due to the potential for SQL semantic inheritance in continuous conversations, the current SQL query may be modified from the previous one. However, when context switching occurs, this information becomes redundant, affecting the final SQL generation. Therefore, modeling a more detailed relationship between context and schema is crucial. RASAT (Qi et al., 2022) adapted its self-attention mechanism to a relational self-attention mechanism, incorporating diverse relational information to enhance encoding capabilities, but without considering information redundancy. CQR-SQL (Xiao et al., 2022) simplifies schema linking information for integration with downstream parsing models by rewriting multiple turns of conversation, but requires additional training and extensive, complex annotation. MIGA (Fu et al., 2022) integrates reference relations and schema linking information through a multi-task approach, but ignores redundancy. SCoRe (Yu et al., 2021b) predicts SQL keywords using only information from the current turn, ignoring schema linking information from previous conversations. STaR (Cai et al., 2022) tracks state at the schema level but does not consider the fine-grained relationship between contextual statements and schemas.

[0046] Therefore, the main difficulty in existing conversational text-to-SQL semantic parsing tasks lies in the need to consider complex syntactic relationships such as reference, omission, and redundancy across multiple conversations, while also building upon the single-turn text-to-SQL semantic parsing task. In the early stages of research, most approaches simply concatenated multiple conversations into a single-turn conversation during the data construction phase and then used a general single-turn text-to-SQL semantic parsing model. While this approach is convenient and fast, it is not ideal due to the syntactic complexity of multi-turn conversations.

[0047] From the above, we can see that the relevant technologies still have defects such as low accuracy, cross-domain failure, and forgetting of dialogue context.

[0048] To this end, the conversational Text-to-SQL pre-training method based on fine-grained link information provided in the present application can effectively improve the accuracy of conversational Text-to-SQL pre-training based on fine-grained link information. Accordingly, the conversational Text-to-SQL pre-training method based on fine-grained link information is applicable to a conversational Text-to-SQL pre-training device based on fine-grained link information. The conversational Text-to-SQL pre-training device based on fine-grained link information can be deployed on an electronic device configured with a von Neumann architecture. For example, the electronic device can be a desktop computer, a laptop computer, a server, and the like.

[0049] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0050] 1 is a schematic diagram of an implementation environment involved in a conversational Text-to-SQL pre-training method based on fine-grained link information, which includes a collection end 110 and a server end 130 .

[0051] Specifically, the collection terminal 110 obtains the text sentences and database schema of the conversation. The collection terminal 110 can be any electronic device with information collection function and is not limited here.

[0052] The acquisition terminal 110 and the server terminal 130 may be connected via a wired or wireless communication connection to achieve data transmission between the two. For example, the transmitted data may be text statements and database models.

[0053] The server side 130 can also be considered as the cloud, cloud platform, platform side, service side, etc. This server side 130 can be a single server, a server cluster consisting of multiple servers, or a cloud computing center consisting of multiple servers, so as to better provide backend services to the massive data collection end 110. For example, the backend service includes a conversational Text-to-SQL pre-training service based on fine-grained link information.

[0054] As the collection end 110 interacts with the server end 130, in one application scenario, for example, where the server end 130 provides a conversational text-to-SQL pre-training service based on fine-grained link information, the collection end 110 obtains text statements and database schemas and sends them to the server end 130. Server end 130 then receives the text statements and database schemas sent by the collection end 110 and provides a conversational text-to-SQL pre-training service based on these text statements and database schemas. Specifically, after obtaining the text statements and database schemas, server end 130 obtains the sentence link relationships between sentences in the conversation based on the text statements and database schemas to obtain a first loss function. It then obtains the schema link relationship between the conversation and the database based on the current sentence in the conversation and the database schema to obtain a second loss function. Finally, the neural network model is pre-trained based on the first and second loss functions to obtain a pre-trained text-to-SQL model.

[0055] Please refer to Figure 2. An embodiment of the present application provides a conversational Text-to-SQL pre-training method based on fine-grained link information. The method is applicable to an electronic device, which can be the server end 130 in the implementation environment shown in Figure 1, or a desktop computer, laptop computer, server, etc.

[0056] In the following method embodiments, for ease of description, the execution subject of each step of the method is taken as an electronic device as an example for illustration, but this does not constitute a specific limitation.

[0057] As shown in FIG2 , the method may include the following steps:

[0058] Step 210: Acquire the text sentences and database schema of the conversation; the text sentences include the current sentence and historical sentences of the conversation.

[0059] In one possible implementation, the database schema consists of multiple tables. If the current turn of the conversation is t, then the current statement of the conversation is u. t 、H history sentence is H t =[u1, u2, ..., u i ,...,u t-1 ] and a database schema S consisting of m tables = [s1, s2, ..., s j ,...,s m ], where the i-th statement is represented by u i words, which can be expressed as The jth table consists of k j columns, which can be expressed as where t jand c j Represents the table name and column name of the database schema respectively.

[0060] Through the above process, the embodiment of the present application obtains the context sentences of the conversation and the database schema corresponding to the context sentences, which is conducive to the subsequent capture of complex syntactic relationships within the context sentences, solving the problems of coreference and omission in multi-round conversations, and also conducive to obtaining a comprehensive schema link relationship between text sentences and database schema items, thereby improving the accuracy and context memory ability of conversational Text-to-SQL pre-training.

[0061] Step 230 : Acquire a sentence link relationship between sentences in the conversation based on the text sentences of the conversation and the database model, and obtain a first loss function.

[0062] Specifically, as shown in FIG3 , step 230 may include the following steps:

[0063] Step 310 : Obtain the word-level syntactic relationship between the current sentence and the text sentence in the conversation through a coreference resolution tool to obtain a sentence supervision label.

[0064] Among them, the coreference resolution tool can be a tool such as NeuralCoref, which is used to obtain the current sentence u t With all statements H t ={u1, ..., u t}, thereby obtaining the supervision label of the sentence linking task, which is not limited here.

[0065] Sentence supervision labels refer to those used in sentence linking tasks. They are typically a set of labels indicating the relationship between two or more sentences. The goal of this task is to understand the connections or associations between text segments. For example, determining whether one sentence is the cause, result, or condition of another sentence. The type of supervision label depends on the specific task settings, such as cause-effect relationships, conditional relationships, contrast relationships, sequential relationships, and explanation relationships.

[0066] When performing sentence linking tasks, in order to train the model, these relationship labels will be used together with the corresponding sentence pairs to guide the model to learn correct relationship judgments. This helps improve the model's performance in understanding text relevance, capture complex syntactic relationships within contextual sentences, and solve the problems of coreference and omission in multi-round dialogues. The specific requirements and labels can vary according to the specific application field and are not limited here.

[0067] Step 330: Obtain a sentence heuristic representation of the dialogue and the database based on the textual sentences of the dialogue and the database schema.

[0068] Specifically, as shown in FIG4 , step 330 may include the following steps:

[0069] Step 410: Get token representations of all sentences in the conversation based on the text sentences of the conversation and the database schema.

[0070] In one possible implementation, in the tth round of dialogue, the input of the sentence linking task I t As shown below: t =[{u1, ..., u t}; {s1, ..., s m}].

[0071] Here, m represents the total number of schema items (table names, column names) of all tables in the database schema.

[0072] Then, according to the input I t The output representation is:

[0073] Here, |·| represents the total number of tokens in statements and pattern items, from which the token representation of all statements is extracted:

[0074] Step 430 : Map the token representation to the sentence representation to obtain the sentence representation of all sentences in the conversation.

[0075] In one possible implementation, the subwords of all sentences in the conversation are aggregated to map tokens to word sentences to obtain sentence representations:

[0076] Step 450 : Perform matrix multiplication on the statement representation to obtain a statement heuristic representation.

[0077] In one possible implementation, computing and Matrix multiplication of is used as a heuristic to represent and predict syntactic relations.

[0078] Specifically, the calculation process is as follows:

[0079] Among them, W i and b i is a trainable parameter.

[0080] In one possible implementation, attention pooling is used to achieve subword aggregation.

[0081] Step 350: Obtain a first loss function based on the sentence supervision label and the heuristic representation.

[0082] In one possible implementation, the pre-training loss function for the sentence link prediction task is defined as the heuristic representation and word-level syntactic relation tags The cross entropy between is as follows:

[0083] Where n represents the total number of words in the entire sentence, and i and j represent the i-th and j-th words in the sentence respectively.

[0084] Step 250: Obtain a schema link relationship between the dialogue and the database based on the current statement of the dialogue and the database schema, and obtain a second loss function.

[0085] Specifically, as shown in FIG5 , step 250 may include the following steps:

[0086] Step 510: Calculate the SQL structure similarity between the current statement and the historical statements in the database schema.

[0087] In one possible implementation, to address the issue of redundant pattern link relationships in conversational Text-to-SQL tasks, refined pattern link relationships are obtained by measuring SQL structural similarity. The SQL tree edit distance is then used as the structural similarity to filter pattern link relationships. This allows the pre-trained language model to learn more accurate pattern link knowledge, further improving performance.

[0088] Step 530: Obtain a pattern supervision label based on the SQL structure similarity and a set similarity threshold.

[0089] Specifically, the current database statement SQLq in the database schema t and the historical database statements {q1, ..., q t-1} is parsed into a tree-based structure G t and {G1, ..., G t-1}, then calculate G t and G i , the SQL structure similarity of i∈[1,t-1] is as follows: fsimilarity(G t , G i )=APTED(G t , G i ).

[0090] Among them, APTED (All Path Tree Edit Distance) represents the all path tree edit distance. It is easy to understand that when the SQL structure similarity is low, it indicates that the redundancy in the pattern link relationship is high. Therefore, the entire pattern link relationship is refined according to the SQL structure similarity and the similarity threshold α to obtain the label of the pattern link prediction task.

[0091] Step 550: Obtain a schema heuristic representation of the dialogue and the database based on the textual statements of the dialogue and the database schema.

[0092] In one possible implementation, token representations of all sentences in the conversation are first obtained based on the text sentences and database patterns of the conversation, and then the token representations are mapped to sentence representations and pattern representations respectively to obtain sentence representations and pattern representations of all sentences in the conversation, and finally a pattern heuristic representation is obtained based on the sentence representations and pattern representations.

[0093] Specifically, the pattern heuristic representation is obtained based on the sentence representation and mode representation Heuristic representation between To predict the refined pattern link relationship:

[0094] Where Wi and bi are trainable parameters.

[0095] Step 570: Obtain a second loss function based on the pattern supervision label and the pattern heuristic representation.

[0096] In one possible implementation, the pre-training loss function for the pattern link prediction task is defined as the heuristic representation and mode link tags The cross entropy between is shown below:

[0097] Where n represents the total number of words in the entire sentence, k represents the number of columns in the pattern, i represents the i-th word in the sentence, and j represents the j-th column of the pattern.

[0098] In a possible implementation, the statement link relationship and the pattern link relationship are shown in Table 1.

[0099] Table 1 Supervision labels involved in sentence link relations and pattern link relations and their meanings

[0100] Step 270: Pre-train the neural network model according to the first loss function and the second loss function to obtain a pre-trained Text-to-SQL model.

[0101] In one possible implementation, the Text-to-SQL model is a trained machine learning model capable of learning the context of natural language text. During the pre-training process of the neural network model, randomly selected tokens are replaced with masked tokens. A third loss function is derived based on the masked tokens and supervisory labels, so that the third loss function participates in the pre-training; the supervisory labels are the true labels of the masked tokens.

[0102] Specifically, masked language modeling (MLM) is a pre-training task in BERT, which aims to learn the context modeling ability of natural language text. In order to enhance the generalization ability of the pre-trained language model, the embodiment of the present application retains the MLM task in the pre-training stage.

[0103] Specifically, given the t-th round dialogue input I t , the MLM task randomly selects a token and replaces it with the [MASK] token, predicts the original token of the [MASK] token based on the context, and the original 15% mask probability of BERT can be applied. The loss of the MLM task is expressed as This loss function mainly minimizes the cross entropy between the [MASK] token and the true label, and the mask probability is not limited here.

[0104] In one possible implementation, the first loss function, the second loss function, and the third loss function are summed according to their homoscedastic uncertainty to obtain a total loss function. The parameters of the neural network model are optimized according to the total loss function until the parameters meet the set conditions, thereby obtaining a pre-trained Text-to-SQL model.

[0105] The specific summation calculation formula is as follows:

[0106] in, refers to the first loss function, refers to the second loss function, Refers to the third loss function, δ1 is the weight of the first loss function, δ2 is the weight of the second loss function, and δ3 is the weight of the third loss function.

[0107] Through the above process, the embodiment of the present application first obtains the text sentences and database schema of the conversation, where the text sentences include the current sentence and historical sentences of the conversation, then obtains the sentence link relationship between the sentences in the conversation based on the text sentences and database schema to obtain a first loss function, learns the complex syntactic relationship between natural languages ​​in the conversational Text-to-SQL scenario, and then obtains the schema link relationship between the conversation and the database based on the current sentence and database schema to obtain a second loss function, captures the fine-grained link relationship between the sentence and the database schema, and finally pre-trains the neural network model based on the first loss function and the second loss function to obtain a pre-trained Text-to-SQL model, thereby effectively improving the performance of the conversational Text-to-SQL model, thereby effectively solving the problems of low accuracy, cross-domain failure, and forgetting of conversation context existing in related technologies.

[0108] In an exemplary embodiment, the conversational Text-to-SQL pre-training method based on fine-grained link information provided in an embodiment of the present application is implemented through a neural network model, which is a trained machine learning model with the ability to learn the context of natural language text and perform conversational Text-to-SQL pre-training on text statements and database schemas based on fine-grained link information.

[0109] FIG6 shows a schematic structural diagram of a Text-to-SQL model in one embodiment. In FIG6 , the Text-to-SQL model includes a statement link prediction module Utterance Linking Prediction and a schema link prediction module Schema Linking Prediction.

[0110] The following describes the pre-training process of the Text-to-SQL model in detail, based on the structure of the Text-to-SQL model in Figure 6:

[0111] In an exemplary embodiment, as shown in FIG6 , first, in the sentence link prediction module UtteranceLinking Prediction, according to the text sentences (h his h age 、h the h youngest hteacher 、h his h hometown ) and sentence supervision labels (NoMatch, Coreference, Identity) to obtain the sentence link relationship between sentences in the conversation and obtain the first loss function.

[0112] Then, in the schema linking prediction module, the schema link relationship between the dialogue and the database is obtained based on the database schema G3, G2, and G1 of the corresponding context text dialogues U3, U2, and U1. The SQL structure similarity (Similarity Calculation) between the current statement G3 and the historical statement G1 in the database schema is calculated. The second loss function is obtained based on the SQL structure similarity (Similarity Calculation) and the schema supervision label (NoMatch, ExactMatch).

[0113] Finally, the neural network model is pre-trained according to the first loss function and the second loss function to obtain a pre-trained Text-to-SQL model.

[0114] After training is completed, a trained Text-to-SQL model is obtained, which has the ability to perform conversational Text-to-SQL pre-training on text statements and database schemas based on fine-grained link information. This effectively improves the performance of the conversational Text-to-SQL model, thereby effectively solving the problems of low accuracy, cross-domain failure, and forgetting of conversation context in related technologies.

[0115] In an exemplary embodiment, the embodiment of the present application uses ELECTRA to initialize a pre-trained language model, while retaining ELECTRA's replacement token detection (RTD) task as part of the masked language modeling task to further enhance the performance of the model. The purpose of the RTD task is to detect replaced tokens in the original text, improve the language understanding ability of the pre-trained language model, and prevent misleading predictions in downstream tasks.

[0116] In the downstream stage, LGESQL was selected as the downstream inference model because it performs well in single-round Text-to-SQL semantic parsing tasks. Except for directly concatenating the current statement and historical statements as part of the input, the parameters of the downstream LGESQL model are basically the same as those of the original model. Table 2 lists some of the parameters and corresponding values ​​used during model pre-training and fine-tuning.

[0117] Table 2 Some parameters and their values ​​in the pre-training and fine-tuning stages

[0118] Through the above process, the embodiment of the present application improves the language comprehension ability of the pre-trained language model and prevents misleading predictions in downstream tasks, thereby improving the performance and accuracy of completing the Text-to-SQL semantic parsing task.

[0119] FIG7 is a schematic diagram showing the implementation results of a conversational Text-to-SQL pre-training method based on fine-grained link information in an application scenario.

[0120] In this application scenario, the embodiments of this application conducted a large number of experiments on two well-known conversational Text-to-SQL semantic parsing datasets:

[0121] (1) SParC is a cross-domain multi-turn Text-to-SQL dataset containing 4,298 conversation turns with a large corpus of approximately 12k+ natural language questions, each carefully annotated with the corresponding SQL expression in the form of question-SQL pairs.

[0122] (2) CoSQL is a conversational Text-to-SQL corpus that contains a comprehensive collection of 30k+ conversation turns and 10k+ annotated SQL queries.

[0123] Among them, CoSQL and SParC both contain 200 complex databases across 138 different domains. Compared with SParC, CoSQL is a more challenging dataset because CoSQL is more in line with practical application scenarios.

[0124] The specific experimental results are shown in Figure 7. Compared with the Text-to-SQL semantic parsing method in the prior art, the method proposed in this application has achieved significant and consistent improvements in various indicators. When considering a unified downstream model, compared with various pre-trained models, this application showed significant improvement, and also showed significant improvement compared with zero-sample ChatGPT. The best results show the powerful performance of this application in conversational Text-to-SQL semantic parsing tasks.

[0125] The following is an embodiment of the apparatus of the present application, which can be used to perform the conversational text-to-SQL pre-training method based on fine-grained link information involved in this application. For details not disclosed in the apparatus embodiment of the present application, please refer to the method embodiment of the conversational text-to-SQL pre-training method based on fine-grained link information involved in this application.

[0126] Please refer to Figure 8. In an embodiment of the present application, a conversational Text-to-SQL pre-training device 800 based on fine-grained link information is provided, including but not limited to: a data acquisition module 810, a statement linking module 830, a pattern linking module 850 and a model pre-training module 870.

[0127] The data acquisition module 810 is used to acquire the patient's text sentences and database schema to obtain the patient's first feature matrix.

[0128] The sentence linking module 830 is used to obtain the sentence linking relationship between the sentences in the conversation based on the text sentences of the conversation and the database model, and obtain a first loss function.

[0129] The pattern linking module 850 is used to obtain a pattern linking relationship between the dialogue and the database based on the current statement of the dialogue and the database pattern, and obtain a second loss function.

[0130] The model pre-training module 870 is used to pre-train the neural network model according to the first loss function and the second loss function to obtain a pre-trained Text-to-SQL model.

[0131] It should be noted that the conversational Text-to-SQL pre-training device based on fine-grained link information provided in the above embodiment only uses the division of the above-mentioned functional modules as an example when performing conversational Text-to-SQL pre-training based on fine-grained link information. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the conversational Text-to-SQL pre-training device based on fine-grained link information will be divided into different functional modules to complete all or part of the functions described above.

[0132] In addition, the conversational Text-to-SQL pre-training device based on fine-grained link information provided in the above embodiment and the conversational Text-to-SQL pre-training method based on fine-grained link information belong to the same concept, and the specific manner in which each module performs operations has been described in detail in the method embodiment and will not be repeated here.

[0133] Fig. 9 shows a schematic diagram of the structure of a server according to an exemplary embodiment. The server is applicable to the server end 130 in the implementation environment shown in Fig. 1 .

[0134] It should be noted that the server is only an example adapted for the present application and cannot be considered to provide any limitation on the scope of use of the present application. The server cannot be interpreted as needing to rely on or necessarily having one or more components in the exemplary server 2000 shown in Figure 9.

[0135] The hardware structure of the server 2000 may vary greatly due to different configurations or performances. As shown in FIG. 9 , the server 2000 includes a power supply 210 , an interface 230 , at least one memory 250 , and at least one central processing unit (CPU) 270 .

[0136] Specifically, the power supply 210 is used to provide operating voltage for each hardware device on the server 2000 .

[0137] The interface 230 includes at least one wired or wireless network interface for interacting with external devices, such as the interaction between the terminal 100 and the server 200 in the implementation environment shown in FIG1 .

[0138] Of course, in other examples adapted by this application, the interface 230 may further include at least one serial-to-parallel conversion interface 233 , at least one input / output interface 235 , and at least one USB interface 237 , as shown in FIG. 9 , which is not specifically limited here.

[0139] The memory 250 serves as a carrier for resource storage and can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon include an operating system 251, application 253 and data 255, etc. The storage method can be temporary storage or permanent storage.

[0140] Among them, the operating system 251 is used to manage and control the various hardware devices and application programs 253 on the server 200 to enable the central processing unit 270 to calculate and process the massive data 255 in the memory 250. It can be Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0141] Application 253 is a computer program that performs at least one specific task based on operating system 251. It may include at least one module (not shown in FIG9 ), each of which may include a computer program for server 2000. For example, a conversational Text-to-SQL pre-training device based on fine-grained link information may be considered an application 253 deployed on server 2000.

[0142] The data 255 may be text statements and database schemas stored on a disk, etc., stored in the memory 250 .

[0143] The central processing unit 270 may include one or more processors and is configured to communicate with the memory 250 via at least one communication bus to read computer programs stored in the memory 250, thereby performing operations and processing on the massive amount of data 255 in the memory 250. For example, the conversational Text-to-SQL pre-training method based on fine-grained link information can be implemented by the central processing unit 270 reading a series of computer programs stored in the memory 250.

[0144] In addition, the present application can also be implemented through hardware circuits or hardware circuits combined with software. Therefore, the implementation of the present application is not limited to any specific hardware circuits, software, or a combination of the two.

[0145] Please refer to Figure 10. An electronic device 4000 is provided in an embodiment of the present application. The electronic device 4000 may include: (needs to be adaptively modified according to the specific circumstances of the present application) a desktop computer, a laptop computer, a server, etc.

[0146] In FIG. 10 , the electronic device 4000 includes at least one processor 4001 , at least one communication bus 4002 , and at least one memory 4003 .

[0147] The processor 4001 and the memory 4003 are connected, for example, via a communication bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.

[0148] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0149] Communication bus 4002 may include a path for transmitting information between the aforementioned components. Communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. Communication bus 4002 may be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, FIG10 shows only one thick line, but this does not indicate that there is only one bus or only one type of bus.

[0150] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.

[0151] The memory 4003 stores a computer program, and the processor 4001 reads the computer program stored in the memory 4003 through the communication bus 4002 .

[0152] When the computer program is executed by the processor 4001 , it implements the conversational Text-to-SQL pre-training method based on fine-grained link information in the above-mentioned embodiments.

[0153] In addition, an embodiment of the present application provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the conversational Text-to-SQL pre-training method based on fine-grained link information in the above embodiments is implemented.

[0154] In one embodiment of the present application, a computer program product is provided. The computer program product includes a computer program stored in a storage medium. A processor of a computer device reads the computer program from the storage medium and executes the computer program, causing the computer device to perform the conversational text-to-SQL pre-training method based on fine-grained link information described in each of the above embodiments.

[0155] Compared with the related art, the beneficial effects of this application are:

[0156] 1. This application proposes a conversational Text-to-SQL pre-training method based on fine-grained link information. This application first obtains the text sentences and database schema of the conversation, where the text sentences include the current sentence and historical sentences of the conversation. Then, based on the text sentences and database schema of the conversation, the sentence link relationship between the sentences in the conversation is obtained to obtain a first loss function, and the complex syntactic relationship between natural languages ​​in the conversational Text-to-SQL scenario is learned. Then, based on the current sentence and database schema of the conversation, the schema link relationship between the conversation and the database is obtained to obtain a second loss function, capturing the fine-grained link relationship between the sentence and the database schema. Finally, the neural network model is pre-trained according to the first loss function and the second loss function to obtain a pre-trained Text-to-SQL model, thereby effectively improving the performance of the conversational Text-to-SQL model, thereby effectively solving the problems of low accuracy, cross-domain failure, and forgetting of conversation context in related technologies.

[0157] 2. This application introduces an innovative pre-training framework that aims to improve the conversational Text-to-SQL semantic parsing task by leveraging fine-grained link information, promoting more effective conversational Text-to-SQL semantic parsing by better representing natural language expressions and database schemas.

[0158] 3. This application optimizes the encoder and proposes two novel pre-training objectives: (1) the sentence link prediction (ULP) task, which is used to model the complex syntactic relations between natural languages ​​in conversational Text-to-SQL scenarios, and (2) the pattern link prediction (SLP) task, which focuses on capturing the fine-grained link relations between sentences and database schemas. This approach has been demonstrated on the SParC and CoSQL datasets to effectively improve the performance of conversational Text-to-SQL and has broad application scenarios.

[0159] 4. This application proposes an innovative pre-training framework that aims to address all of the above challenges by leveraging link information to improve conversational Text-to-SQL parsing tasks. It promotes more effective Text-to-SQL conversations by better representing natural language expressions and database schemas. It proposes two novel pre-training objectives: the Sentence Link Prediction (ULP) task, which is used to model complex syntactic relations between natural language sentences in conversational Text-to-SQL scenarios, and the Schema Link Prediction (SLP) task, which focuses on capturing fine-grained link relations between sentences and database schemas.

[0160] 5. This application learns the context modeling capabilities of natural language text through masked language modeling (MLM). By retaining the MLM task in the pre-training stage, the generalization ability of the pre-trained language model is enhanced. It has achieved significant and consistent improvements in various indicators and has strong performance in conversational Text-to-SQL semantic parsing tasks.

[0161] 6. This application uses ELECTRA to initialize the pre-trained language model in the pre-training stage, while retaining ELECTRA's replacement token detection (RTD) task as part of the masked language modeling task to further enhance the performance of the model. The RTD task detects the replaced tokens in the original text, improves the language understanding ability of the pre-trained language model, and prevents misleading predictions in downstream tasks.

[0162] 7. In the downstream stage, this application selects LGESQL as the downstream reasoning model because of its good performance in the single-round Text-to-SQL semantic parsing task. In addition to directly concatenating the current statement and historical statements as part of the input, the parameters of the downstream LGESQL are basically consistent with the original model.

[0163] 8. This application pioneered the sentence link prediction (ULP) task in the conversational Text-to-SQL task, which is used to explicitly model word-level coreference relationships in the context and effectively solve the complex coreference and omission problems in multi-round conversations.

[0164] 9. This application pioneered a fine-grained pattern link prediction (SLP) task in conversational Text-to-SQL tasks to ensure more accurate pattern links and enable the current utterance to focus on key pattern link information from previous utterances. After performing SQL structure similarity filtering based on tree edit distance, the model will focus on more relevant pattern link information.

[0165] 10. The results of testing this application on a commonly used dataset for conversational Text-to-SQL tasks show that the model proposed in this patent has better results than previous models. Whether in the direction of pre-trained encoding models or downstream parsing models, previous technologies all perform information fusion of pattern links in model structure, input data, and multiple tasks, but do not consider the problem of information redundancy. The pattern link prediction task based on SQL tree edit distance proposed in this application models highly fine-grained pattern link relationships in multi-round conversations, solving the problem of information redundancy.

[0166] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0167] The above content is only a preferred exemplary embodiment of the present application and is not intended to limit the implementation scheme of the present application. Ordinary technicians in this field can easily make corresponding changes or modifications based on the main ideas and spirit of the present application. Therefore, the scope of protection of the present application shall be based on the scope of protection required by the claims.

Claims

1. A conversational Text-to-SQL pre-training method based on fine-grained link information, characterized in that, the method includes: Obtain the text statements of the conversation and the database schema; the text statements include the current statement and historical statements of the conversation; Obtain the statement link relationship between statements in the conversation according to the text statements of the conversation and the database schema, and obtain the first loss function; Obtain the schema link relationship between the conversation and the database according to the current statement of the conversation and the database schema, and obtain the second loss function; Pre-train the neural network model according to the first loss function and the second loss function to obtain the pre-trained Text-to-SQL model.

2. The method according to claim 1, characterized in that, the obtaining the statement link relationship between statements in the conversation according to the text statements of the conversation and the database schema, and obtaining the first loss function includes: Obtain the word-level syntactic relationship between the current statement and the text statements in the conversation through a coreference resolution tool to obtain statement supervision labels; Obtain the statement heuristic representation of the conversation and the database according to the text statements of the conversation and the database schema; Obtain the first loss function according to the statement supervision labels and the heuristic representation.

3. The method according to claim 2, characterized in that, the obtaining the statement heuristic representation of the conversation and the database according to the text statements of the conversation and the database schema includes: Obtain the token representation of all statements in the conversation according to the text statements of the conversation and the database schema; Map the token representation to a statement representation to obtain the statement representation of all statements in the conversation; Perform matrix multiplication on the statement representations to obtain the statement heuristic representation.

4. The method according to claim 1, characterized in that, the obtaining the schema link relationship between the conversation and the database according to the current statement of the conversation and the database schema, and obtaining the second loss function includes: Calculate the SQL structure similarity between the current statement and the historical statements in the database schema; Obtain the schema supervision labels according to the SQL structure similarity and a set similarity threshold; Obtain the schema heuristic representation of the conversation and the database according to the text statements of the conversation and the database schema; Obtain the second loss function according to the schema supervision labels and the schema heuristic representation.

5. The method according to claim 4, characterized in that, the obtaining the schema heuristic representation of the conversation and the database according to the text statements of the conversation and the database schema includes: Obtain the token representation of all statements in the conversation according to the text statements of the conversation and the database schema; Map the token representation to a statement representation and a schema representation respectively to obtain the statement representation and schema representation of all statements in the conversation; Obtain the schema heuristic representation according to the statement representation and the schema representation.

6. The method according to any one of claims 1 to 5, characterized in that, the neural network model is a machine learning model that has been trained and has the ability to learn the context of natural language text; the method further includes: Replace the randomly selected tokens with masked tokens during the pre-training process; Obtain a third loss function based on the masked tokens and the supervision labels, and enable the third loss function to participate in the pre-training process of the neural network model; the supervision label is the true label of the masked token.

7. The method according to claim 6, wherein, The pre-training the neural network model according to the first loss function and the second loss function to obtain a pre-trained Text-to-SQL model includes: Calculating the sum of the first loss function, the second loss function, and the third loss function according to the homoscedastic uncertainty of the first loss function, the second loss function, and the third loss function to obtain a total loss function; Optimizing the parameters of the neural network model according to the total loss function until the parameters meet the set conditions to obtain a pre-trained Text-to-SQL model.

8. A conversational Text-to-SQL pre-training device based on fine-grained link information, wherein, The device includes: A data acquisition module for acquiring the text statements of the conversation and the database schema; the text statements include the current statement and the historical statements of the conversation; A statement link module for obtaining the statement link relationship between the statements in the conversation according to the text statements of the conversation and the database schema, and obtaining a first loss function; A schema link module for obtaining the schema link relationship between the conversation and the database according to the current statement of the conversation and the database schema, and obtaining a second loss function; A model pre-training module for pre-training the neural network model according to the first loss function and the second loss function to obtain a pre-trained Text-to-SQL model.

9. An electronic device, wherein, It includes: At least one processor and at least one memory, wherein, Program instructions or codes are stored on the memory; The program instructions or codes are loaded and executed by the processor, so that the electronic device implements the conversational Text-to-SQL pre-training method based on fine-grained link information according to any one of claims 1 to 7.

10. A storage medium, on which program instructions or codes are stored, wherein, The program instructions or codes are loaded and executed by the processor to implement the conversational Text-to-SQL pre-training method based on fine-grained link information according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Construction method and application of SQL statement generation model of natural Chinese language

    CN114020768A

  • Method for establishing pre-training language model and semantic analysis method and device

    CN114547329A

  • Pre-training model data processing method, electronic equipment and computer storage medium

    CN114579606A

  • Processing method and device for converting text into SQL statement and storage medium

    CN115563121A

  • Device and method for converting natural language query into SQL query

    US20230169074A1

Cited By

  • Dialogue state tracking method based on large language model, medium and electronic equipment

    CN121188165A

  • A dialogue state tracking method based on a large language model, a medium and an electronic device

    CN121188165B