Pre-training model data processing method, electronic device and computer storage medium
By performing specific data processing and pre-training of the pre-trained model, the problem of insufficient correlation between the model's data in the table question and answer scenario is solved, and the model's robustness and accuracy are improved.
Patent Information
- Application Number
- CN202210560697.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-05-23
AI Technical Summary
Existing pre-trained models are difficult to effectively understand the correlation between natural language query statements and database schema data in tabular Q&A scenarios, resulting in insufficient robustness and fault tolerance in practical applications.
Through the pre-processing layer of the pre-trained model, natural language query statements and database schema data are spliced into splicing vectors, and the conjunctions in the database schema data are masked to generate mask vectors. The mask vector is then pre-trained using the generator-discriminator architecture to capture the contextual relationship between natural language query statements and database schema data.
The accuracy and robustness of the model in the table question and answer scenarios are improved, so that the table question and answer system can handle user query requests more effectively and output more accurate results.
Smart Images

Figure CN114897163B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of table question answering technology, and in particular to a pre-training model data processing method, electronic device and computer storage medium. Background Art
[0002] Due to the clear data structure and easy maintenance, table / SQL databases have become the most commonly used structured data in various industries, and are also an important source of answers for intelligent dialogue systems and search engines. Traditional table queries require professional technicians to write query statements (such as SQL statements) to complete, and the high threshold has hindered the large-scale application of table queries. Table question answering technology (also known as TableQA) allows users to interact directly with table databases using natural language by directly converting natural language into SQL queries, and is becoming more and more widely used.
[0003] A table question answering system mainly consists of three parts, including natural language understanding, dialogue management and natural language generation. Among them, the natural language understanding part mainly executes the semantic parsing algorithm to convert natural language questions into corresponding executable SQL statements; the dialogue management part performs multiple rounds of state tracking and strategy optimization; the natural language generation part generates corresponding replies based on the parsed SQL statements and the execution results of SQL. For the natural language understanding part, the training output of the pre-trained model is currently used to provide functional support for the natural language understanding part of the subsequent table question answering system. The pre-trained model is an application of transfer learning, which obtains model parameters that are not related to specific tasks from large-scale data through self-supervised learning, and when supporting a new task, it only needs to use the labeled data of the task to fine-tune the pre-trained model.
[0004] However, most of the current pre-trained models focus on language understanding. In real conversation / question-and-answer scenarios, especially TableQA scenarios, natural language and tables / SQL databases are closely related. How to obtain a pre-trained model that meets the requirements of this scenario has become an urgent problem to be solved. Summary of the invention
[0005] In view of this, an embodiment of the present application provides a pre-training model data processing solution to at least partially solve the above-mentioned problems.
[0006] According to a first aspect of an embodiment of the present application, a pre-trained model data processing method is provided, comprising: generating a corresponding splicing vector according to a natural language query statement and database pattern data through a preprocessing layer of a pre-trained model; masking the associated words in the database pattern data part in the splicing vector according to information of associated words between the natural language query statement and the database pattern data to obtain a masked vector; performing mask recovery processing on the masked associated words by a generator of the pre-trained model to obtain a generated vector; and evaluating a generation result of the generator based on the generated vector using a discriminator of the pre-trained model, and training the pre-trained model according to the evaluation result.
[0007] According to the second aspect of an embodiment of the present application, another method for processing pre-trained model data is provided, comprising: obtaining model parameters of a pre-trained model to be migrated, wherein the pre-trained model is a model obtained by training according to natural language query statements and database model data, and data after masking of associated words in associated words between the natural language query statements and the database model data and associated words in the database model data part; and performing model migration from the pre-trained model to a table question and answer system.
[0008] According to the third aspect of an embodiment of the present application, there is provided an electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method described in the first aspect or the second aspect.
[0009] According to a fourth aspect of an embodiment of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect or the second aspect is implemented.
[0010] According to a fifth aspect of an embodiment of the present application, a computer program product is provided, comprising computer instructions, wherein the computer instructions instruct a computing device to perform operations corresponding to the method described in the first aspect or the second aspect.
[0011] According to the pre-trained model data processing scheme provided in the embodiment of the present application, on the one hand, based on the splicing vector generated according to the natural language query statement and the database model data, the associated words in the database model data are masked to simulate the subsequent possible changes in the natural language query statement input by the user, so that the model has better robustness and fault tolerance. On the other hand, after the pre-processing layer performs corresponding processing, the splicing vector, especially the part corresponding to the database model data in the splicing vector, is pre-trained through the generator-discriminator architecture, so that the relationship between the context can be effectively captured, the interaction between the natural language query statement and the database model data is obtained, and the accuracy of the model's judgment on the relationship between the natural language query statement and the database model data is improved. After the pre-trained model is migrated to the table question and answer system, the table question and answer system can be effectively applied to the table question and answer scenario, and more accurate results for user query requests can be output. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0013] Figure 1 A schematic diagram of an exemplary system for a pre-training model data processing method applicable to an embodiment of the present application;
[0014] Figure 2 A schematic diagram of a model structure of a pre-training model according to an embodiment of the present application;
[0015] Figure 3 A flowchart of the steps of a pre-training model data processing method according to Embodiment 1 of the present application;
[0016] Figure 4A A flowchart of a method for processing pre-trained model data according to Embodiment 2 of the present application;
[0017] Figure 4B for Figure 4A An example diagram of a scenario in the illustrated embodiment;
[0018] Figure 5 This is a schematic diagram of the structure of an electronic device according to the third embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the embodiments of the present application should fall within the scope of protection of the embodiments of the present application.
[0020] The specific implementation of the embodiment of the present application is further explained below in conjunction with the accompanying drawings of the embodiment of the present application.
[0021] Figure 1 A schematic diagram of an exemplary system for a pre-training model data processing method applicable to an embodiment of the present application is shown. Figure 1 As shown, the system 100 may include a server 102, a communication network 104 and / or one or more user devices 106. Figure 1 The example in FIG. 1 is a plurality of user devices.
[0022] The server 102 may be any suitable server for storing information, data, programs and / or any other suitable type of content. In some embodiments, the server 102 may perform any suitable function. For example, in some embodiments, a table question-and-answer system is provided in the server 102 to process a query request involving a table or database input by a user and return a query result. As an optional example, in some embodiments, a pre-trained model is also provided in the server 102, which may also be referred to as a table pre-trained model, so that after the training is completed, it can be migrated to the table question-and-answer system for use. As an optional example, in some embodiments, the pre-trained model in the server 102 adopts a pre-processing layer + (generator-discriminator) architecture, and the associated words of the database pattern data in the concatenated vector generated according to the natural language query statement and the database pattern data are masked by the pre-processing layer; then, the masked vector, i.e., the masked vector, is pre-trained by the generator-discriminator architecture, so that the relationship between the context of the data as a whole, including the natural language query statement and the database pattern data, can be effectively captured, and the interaction between the natural language query statement and the database pattern data is obtained, thereby improving the fault tolerance and robustness of the model.
[0023] In some embodiments, the communication network 104 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the following: the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN) and / or any other suitable communication network. The user device 106 can be connected to the communication network 104 via one or more communication links (e.g., communication link 112), and the communication network 104 can be linked to the server 102 via one or more communication links (e.g., communication link 114). The communication link can be any communication link suitable for transmitting data between the user device 106 and the server 102, such as a network link, a dial-up link, a wireless link, a hard-wired link, any other suitable communication link or any suitable combination of such links.
[0024] User device 106 may include any one or more user devices that have settings and interfaces for interacting with a user. In some embodiments, user device 106 may include any suitable type of device. For example, in some embodiments, user device 106 may include a mobile device, a tablet computer, a laptop computer, a desktop computer, a wearable computer, a game console, a media player, a vehicle entertainment system, and / or any other suitable type of user device.
[0025] Although server 102 is illustrated as one device, in some embodiments, any suitable number of devices may be used to perform the functions performed by server 102. For example, in some embodiments, multiple devices may be used to implement the functions performed by server 102. Alternatively, the functions of server 102 may be implemented using cloud services.
[0026] Based on the above system, an embodiment of the present application provides a pre-trained model data processing method. For the convenience of explanation, the structure of the pre-trained model used in the method is first exemplified below.
[0027] Generally speaking, the training of pre-trained models mostly adopts self-supervised learning. After the training is completed, the knowledge learned by the pre-trained model can be transferred to downstream tasks, and the functions of the corresponding downstream tasks can be realized after fine-tuning. In the embodiment of the present application, the training of the pre-trained model is mainly used for the downstream table question and answer system. Unlike traditional pre-trained models such as BERT and GPT that are mainly used for language training, the pre-trained model of the embodiment of the present application is intended to simultaneously model natural language and structured table data, and integrate the semantics of natural language into the structural content of the table in the dimension of language understanding, so as to generate fluent text based on structured data in the dimension of language generation. Based on this, the pre-trained model uses natural language query statements and database model data as input for corresponding processing and training. In one feasible method, the pre-trained model is such as Figure 2 As shown, it includes preprocessing layers, generators, and discriminators.
[0028] In the embodiment of the present application, each sample data set for training the pre-trained model includes two parts, namely, a natural language query statement for data query, and the schema data of the database queried by the query statement. Among them, the schema data of the database is also called the schema data of the database, which is a set of interrelated database objects used to represent information such as tables, table columns, column data types, indexes, foreign keys, etc. in the database. In the embodiment of the present application, the database schema data used mainly includes the table name, column name, and value data of the data table.
[0029] The preprocessing layer of the pre-trained model is used to process the input sample data, including: splicing two parts of the sample data to generate a splicing vector; then, based on the information of the associated words between the natural language query statement and the database model data obtained in advance, masking (MASK) the vector corresponding to part of the data in the database model data to obtain the mask of this part of the data; and then, combining the mask and other parts of the splicing vector except the mask to generate a mask vector.
[0030] After obtaining the mask vector, the mask vector will be input into the generator, which will encode the received mask vector as a whole and restore the mask in the mask vector through encoding. Finally, the generated vector is output, which carries the data after the mask is restored. Of course, the restored data may be the same as the original pattern data processed by the mask, or it may be similar, such as synonymous or similar in shape.
[0031] The generated vector will be further input into the discriminator to evaluate the generation result of the generator through the discriminator, mainly for the evaluation of the pattern data restored by the generator (such as the degree of difference or similarity between the restored data and the original data, etc.). In addition, the pre-trained model is trained based on the evaluation results, including but not limited to adjusting the model parameters through back propagation, etc., until the model training termination condition is reached (such as reaching a preset number of training times, or the loss value is within a preset threshold range, etc.).
[0032] Based on the above system, an embodiment of the present application provides a pre-training model data processing method, which is described in detail through multiple embodiments below.
[0033] Embodiment 1
[0034] Reference Figure 3 , shows a step flow chart of a pre-training model data processing method according to embodiment 1 of the present application.
[0035] The pre-training model data processing method of this embodiment includes the following steps:
[0036] Step S302: Generate a corresponding concatenation vector according to the natural language query statement and the database schema data through the preprocessing layer of the pretrained model; mask the associated words of the database schema data part in the concatenation vector according to the information of the associated words between the natural language query statement and the database schema data to obtain a mask vector.
[0037] As mentioned above, in the training stage of the pre-training model, natural language query statements and database schema data are two different parts of the training samples. Among them, the natural language query statement can be the data corresponding to the historical user query request obtained when the user data is authorized for use; or, it can also be a collection of the extended data generated according to certain expansion rules based on the data corresponding to some historical user query requests and the data corresponding to the said part of the historical user query requests. Correspondingly, each natural language query statement corresponds to the database schema data of the database or data table it queries. Based on this, each set of natural language query statements and their corresponding database schema data can be used as a training sample and input into the pre-training model for training.
[0038] Specifically, the preprocessing layer of the pre-trained model receives the training sample, that is, the natural language query statement and its corresponding database pattern data; then, the two parts of data are spliced to obtain the corresponding splicing vector. For the pre-trained model in the embodiment of the present application, obtaining the association between the natural language query statement and the pattern data is also called pattern linking, which is one of the important parts of training. By splicing these two parts of data, the pattern link structure can be explicitly introduced. Therefore, the pre-trained model can be used to predict which words in the natural language query statement should be linked to which items in the pattern data, and what keywords this link corresponds to in SQL, so as to obtain better query statements and pattern representations, and then after the trained model is migrated to the downstream TableQA model, it can effectively improve the performance of the downstream TableQA model.
[0039] However, it is not limited to this. In the embodiment of the present application, based on the splicing vector, part of the data in the pattern data is also masked, and the masked data is the data corresponding to the associated words that have an associated relationship with the natural language query statement. Among them, the associated words can be the same words between the natural language query statement and the database pattern data (such as "height" in the natural language query statement and "height" in the database pattern data), or they can be those words with a similarity higher than a certain similarity (such as "height" in the natural language query statement and "height" in the database pattern data, etc.). Preferably, the same words can be selected.
[0040] For example, Figure 2 As shown in , the database pattern data includes: name, height, gender, etc., then part of them can be selected for mask processing. Preferably, one of the pattern data can be selected for mask processing to make the model training more targeted. Figure 2 In the example, the height is masked. Figure 2 After obtaining the mask corresponding to some data, such as the [MASK] corresponding to "height", the mask vector is generated by combining other parts, such as Figure 2 As shown in “[S]Please tell me the names of students whose height is over 180cm[ / S]Name[ / S][MASK][ / S]Gender”. By masking some pattern data, it can be restored by the generator later to make the model more fault-tolerant and robust. However, it is not limited to this. In practical applications, multiple pattern data can also be selected for masking at the same time.
[0041] Step S304: The generator of the pre-trained model performs mask recovery processing on the mask vector for the masked associated words to obtain a generated vector.
[0042] In the embodiment of the present application, the generator can be implemented by an encoder, and the generator can be regarded as a language model, which recovers the masked associated words in the mask vector through the context (the unmasked part of the natural language query statement and the database pattern data). However, since the output of the generator is not fixed, it is possible to generate some recovered data that is different from the original masked pattern data, such as synonyms, similar words, etc.
[0043] Based on the generator's processing of the mask vector, a generated vector can be obtained, which contains the restored data corresponding to the masked associated words processed by the generator.
[0044] For example, Figure 2 As shown, after the original pattern data "height" is masked, the generator recovers the restored data "height". However, it is not limited to this, and the generator may also recover the restored data "height" which is the same as the original pattern data.
[0045] Step S306: Use the discriminator of the pre-trained model to evaluate the generation result of the generator based on the generated vector, and train the pre-trained model according to the evaluation result.
[0046] In an embodiment of the present application, corresponding to the generator, the discriminator can be implemented in the form of a decoder + classifier. The discriminator generates a corresponding decoding vector by decoding the generated vector generated by the generator. Then, by means of a classifier, if the decoded vector is consistent with the original vector, the output result of the classifier is "true", if it is inconsistent, the output result of the classifier is "false". In particular, for the pattern data part, if the output result of the classifier is "true", it means that the pre-trained model has effectively learned the pattern link between the natural language query statement and the database pattern data, and it can also effectively perform targeted deviation correction or error correction on the pattern data by processing the mask.
[0047] The more accurate the generated vector generated by the generator is, the more accurate the decoded vector obtained by decoding is, and the closer it is to the original data. Based on this, the generation result of the generator can be evaluated by the output of the discriminator. If there are more "true" ones, the generation result is better, otherwise, it is slightly worse. It should be noted that the specific implementation of the evaluation result can be implemented by those skilled in the art in a flexible manner according to actual needs, including but not limited to probability values, scores, etc. The embodiment of the present application does not limit the specific presentation method of the evaluation result.
[0048] Furthermore, based on the results obtained by the discriminator, the pre-trained model can be trained (including but not limited to the adjustment of model parameters) by conventional back propagation. The training is an iterative process until the training termination condition is reached, such as the number of training times reaches a set number, or the model loss value meets a preset threshold standard, etc.
[0049] Through the scheme of this embodiment, on the one hand, based on the splicing vector generated according to the natural language query statement and the database model data, the associated words in the database model data are masked to simulate the subsequent possible changes in the natural language query statement input by the user, so that the model has better robustness and fault tolerance. On the other hand, after the corresponding processing in the preprocessing layer, the splicing vector, especially the part corresponding to the database model data in the splicing vector, is pre-trained through the generator-discriminator architecture, so as to effectively capture the relationship between the context, obtain the interaction between the natural language query statement and the database model data, and improve the accuracy of the model's judgment on the relationship between the natural language query statement and the database model data. After the pre-trained model is migrated to the table question and answer system, the table question and answer system can be effectively applied to the table question and answer scenario, and output more accurate results for user query requests.
[0050] Embodiment 2
[0051] Reference Figure 4A , shows a step flow chart of a pre-training model data processing method according to embodiment 2 of the present application.
[0052] The pre-trained model data processing method of this embodiment exemplifies the complete process from the preliminary processing of training samples to the migration of the trained pre-trained model to the application of the downstream table question answering system. Based on this, the pre-trained model data processing method of this embodiment includes the following steps:
[0053] Step S402: performing an associative word analysis on the natural language query statement and the database schema data, and determining the associative words between the natural language query statement and the database schema data according to the analysis result.
[0054] As mentioned above, a training sample consists of two parts: a natural language query statement and its corresponding database model data. In a query based on a data table / database, the natural language query statement will eventually be converted into an SQL statement to access the data table / database. The query fields, query conditions, and other information in the SQL statement all come from the natural language query statement. The information and data related to the query fields and / or query conditions in both can be used as associated words. For example, "Please tell me Class 31 height More than 160CM"The name of the child", where "height" and "name" correspond to the fields in the data table / database, or in other words, they both correspond to the query fields in the SQL statement, and "Class 31" corresponds to the table name of the data table, and "over 160" corresponds to the query condition of the "height" field.
[0055] In some non-standard inputs, it is necessary to convert the non-standard words (words that cannot be directly mapped to fields in the database) in the natural language query statements into the final standard words, so as to obtain accurate results even when the user input is biased. Based on this, we can first perform a word analysis on the association words between the natural language query statements and the database model data to determine the association words between the two, so as to train these association words in the future and improve the fault tolerance and robustness of the model.
[0056] Among them, the specific method of associative word analysis can be implemented by technical personnel in this field in a flexible manner according to actual needs, including but not limited to: first segmenting the natural language query statement, and then calculating the similarity between the segmentation and the pattern data; or, directly comparing the pattern data with the natural language query statement; or, first determining the keywords in the natural language query statement, and then comparing the keywords with the pattern data; or, using a neural network model with an associative word analysis function, etc.
[0057] Step S404: Generate a corresponding concatenation vector according to the natural language query statement and the database schema data through the preprocessing layer of the pretrained model; mask the associated words in the database schema data part in the concatenation vector according to the information of the associated words between the natural language query statement and the database schema data to obtain a mask vector.
[0058] Among them, in one feasible way to generate the concatenation vector, the natural language query statement and the database pattern data can be concatenated through the preprocessing layer of the pre-trained model, and a separator is inserted between the concatenated natural language query statement and the database pattern data, and between adjacent pattern data of the database pattern data; according to the natural language query statement and the database pattern data after the separator is inserted, the corresponding concatenation vector is generated. By adding a separator to separate the natural language query statement and different database pattern data, it is easy to identify and process them later, and improve the speed and efficiency of model training.
[0059] Then, for the part corresponding to the database pattern data in the splicing vector, an associated word is selected from it to perform masking processing on it. There is at least one associated word in the database pattern data, usually multiple (in the embodiment of the present application, unless otherwise specified, "multiple", "multiple", etc., and the number related to "multiple" means two or more). In a preferred feasible method, one associated word can be selected at a time for masking processing to make the model processing more targeted. However, it is not limited to this. The method of masking multiple associated words at the same time is also applicable to the solution of the embodiment of the present application.
[0060] For example, Figure 2 As shown in , the preprocessing layer processes the natural language query statement (illustrated in the figure as "Please tell me the names of students whose height is over 180") and the database pattern data (illustrated in the figure as "Name, height... gender") into inputs that the pre-trained model can accept, including: first concatenating the natural language query statement and the database pattern data, and then adding a separator in the middle (illustrated in the figure as the [ / s] separator) to indicate the difference between the two; a separator is also added between each pattern data item in the database pattern data (also illustrated in the figure as [ / s]) to distinguish; in addition, it is necessary to add the [s] character at the beginning to indicate the beginning of the input. It should be noted that the above-mentioned [ / s] as the separator and [s] as the start character are only exemplary. In actual applications, those skilled in the art may adopt other forms of separators and start characters according to actual needs. The embodiments of the present application do not limit the specific implementation form of the separator. In addition, the present application also adopts a masking strategy centered on pattern data (i.e., a masking strategy). Before the preprocessing layer performs the above processing, the associated words associated with the natural language query statement and the database pattern data are obtained in advance, also known as tokens, such as Figure 2 The [height] in the natural language query statement is associated with the [height] in the database pattern data, and the [name] in the natural language query statement is associated with the [name] in the database pattern data. Then, after the concatenation vector is generated in the preprocessing layer, the corresponding part of the database pattern data is randomly masked according to these predetermined associated words (i.e., masking, changing the randomly selected associated words to [MASK]), for example Figure 2 [Height] is changed to [MASK]. It should be noted that, in this embodiment, the masking process is performed on the associated words after the splicing vector is generated in the preprocessing layer as an example, but in actual applications, the associated words may be masked first, and then spliced with other pattern data items in the database pattern data and the natural language query statement to generate a mask vector.
[0061] The mask vector containing [MASK] will be input into the generator for processing, for example, Figure 2As shown in , the mask vector Figure 2 It is expressed as “[s]Please tell me the name of the student whose height is 180cm[ / s]name[ / s][MASK]…[ / s]gender”.
[0062] Step S406: The generator of the pre-trained model performs mask recovery processing on the mask vector for the masked associated words to obtain a generated vector.
[0063] The mask vector generated by the preprocessing layer will enter the generator. In the embodiment of the present application, the generator can restore the masked associated word token, such as restoring [MASK] to [height]. The generator can be directly regarded as a language model, which performs mask recovery through the context (natural language query statements and other pattern data items in the database pattern data that have not been masked). However, since the output of the generator is not fixed, the generator may generate data that is different from the original data pattern items. For example, it may generate some synonyms, similar words, etc. For example, Figure 2 As shown in the figure, the original pattern data item "height" is masked to [MASK], and then restored by the generator to output the corresponding pattern data item "height". It can be seen that "height" and "height" are not completely consistent. However, it is precisely because of this that the subsequent discriminator can have better fault tolerance and error correction after training.
[0064] The output of the generator is the generated vector. For example, Figure 2 The generated vector is represented as "[s]Please tell me the name of the student whose height is 180cm or taller[ / s]name[ / s]height...[ / s]gender".
[0065] In addition, in a feasible manner, the generator can be specifically implemented as an encoder, including but not limited to an encoder based on a Transformer structure.
[0066] Step S408: Use the discriminator of the pre-trained model to evaluate the generation result of the generator based on the generated vector, and train the pre-trained model according to the evaluation result.
[0067] The output of the generator will be used as the input of the discriminator. In one feasible way, the discriminator can be implemented as a decoder, including but not limited to a decoder based on the Transformer structure. The discriminator can not only decode the generated vector to generate a vector form that is more similar to the original input pre-trained model, but also evaluate the generation result of the generator based on the vector form.
[0068] Based on this, in a feasible way, the concatenated vector is used as a supervision condition, and the discriminator of the pre-trained model is used to compare the generated vector and the concatenated vector, and the evaluation result is obtained according to the comparison result. For example, if the vector decoded by the discriminator is consistent with the original vector of the input preprocessing layer, the evaluation result is that the generation result of the generator is good. However, it is not limited to this. In practical applications, corresponding evaluation thresholds, such as quantity thresholds or probability thresholds, can also be set. For example, the first number of parts corresponding to each word or each word in the natural language query sentence in the decoded vector that are consistent with the vector of the original input preprocessing layer, and the second number of parts corresponding to the pattern data in the decoded vector that are consistent with the vector corresponding to the pattern data of the original input preprocessing layer can be determined. If the sum of the first number and the second number is greater than the quantity threshold, it indicates that the generation result of the generator is good. In particular, for the pattern data part, the larger the second number, the better the generation result. Of course, a higher weight can be set for the second number, and a slightly lower weight can be set for the first number, and the quality of the generator generation result can be judged based on the comprehensive result of the number and weight.
[0069] For example, Figure 2 In the figure, corresponding to the database pattern data, the result of the generator's recovery of [MASK] is "height", which is different from the original "height". Therefore, the upper right corner of the generator's processing result of the database pattern data is judged to be false for [height] (indicated by "X" in the figure), while others, such as [name] and [gender], are true (indicated by "√" in the figure). Based on this judgment, the evaluation of the generator's generation result can be considered "poor". Furthermore, based on this evaluation, the model parameters of the pre-trained model can be readjusted and training can continue.
[0070] As mentioned above, the training of the pre-trained model needs to be iterated back and forth until the model training termination condition is reached. After the termination condition is reached, the model training can be considered complete.
[0071] From a global perspective, the generator is designed to generate words that are easier to deceive the discriminator, and the discriminator is designed to better identify which words are generated by the generator. Through such an adversarial training strategy, the pre-trained model can not only capture rich contextual relationships, but also imitate the changes in query statements entered by users when making queries, making the pre-trained model more robust and fault-tolerant.
[0072] After the pre-trained model training is completed, subsequent migration applications can be performed. For ease of understanding, the migration process is further described through the following step S410 in this embodiment, but those skilled in the art should understand that the training process of the pre-trained model up to step S408 has formed a complete solution, and the following step S410 is an optional step. In actual applications, steps S408 and S410 do not need to be executed in succession. Those skilled in the art can migrate the trained pre-trained model to the table question-answering system at any time according to actual needs.
[0073] Step S410: Based on the model parameters of the discriminator in the trained pre-trained model, perform model migration from the pre-trained model to the table question answering system.
[0074] In the embodiment of the present application, after the pre-trained model is trained, only its discriminator is used to complete the downstream task. Specifically, the model migration from the pre-trained model to the table question answering system can be performed by migrating the model parameters of the discriminator in the pre-trained model to the natural language understanding part of the table question answering system.
[0075] Because the pre-trained model itself is trained for the table question answering system, the model parameters learned by the discriminator can be directly transplanted to the natural language understanding part of TableQA. With the help of the migrated model parameters, the natural language understanding part can not only perform semantic analysis for query statements input in natural language, but also has good fault tolerance and robustness. Even if the input query statement is not accurate enough or does not correspond well to the fields in the database, it can be finally converted into an accurate and executable SQL statement. Exemplarily, the natural language understanding part can be implemented as a text-to-SQL model, specifically in the form of a seq2seq neural network model, inputting a query statement and outputting a corresponding SQL statement.
[0076] After completing the model migration, the natural language understanding part of TableQA, combined with the trained dialogue management part and natural language generation part, can become a complete table question and answer system and realize the corresponding table question and answer functions.
[0077] Next, through optional step S412, combined with Figure 4B , a schematic illustration of the process of conducting form question and answering through the above-mentioned form question and answer system is given.
[0078] Step S412: receiving a natural language query request input by a user, and returning a corresponding query result through the table question answering system.
[0079] In a feasible manner, this step can be implemented as follows: the natural language understanding part of the form question answering system analyzes the natural language query request input by the user to obtain the database model data in the natural language query request; if it is determined that there is data to be corrected in the database model data, the database model data is corrected; and a database query statement corresponding to the natural language query request is generated according to the correction result. Then, a corresponding database query can be performed based on the database query statement, and the query result can be returned.
[0080] For example, Figure 4B As shown, suppose the user inputs the query request "Please tell me the names of students in Class 31 whose height exceeds 180"; the query request is input into the table question and answer system TableQA, specifically the natural language understanding part of the TableQA (such as the seq2seq model). The natural language understanding part parses the query request to obtain its corresponding database model data, including: "Class 31", "Height", and "Name". Because the model parameters of the natural language understanding part come from the pre-trained model, the pre-trained model knows through training that "height" needs to be corrected to "height". Therefore, the natural language understanding part will also follow the training results and will automatically correct the "height" in the database model data corresponding to the query request to "height". Furthermore, based on the analysis results of the query request and the correction results, the corresponding SQL statement is generated, such as Figure 4B "SELECT name FROM Class 31 WHERE height>180" as shown in the figure.
[0081] The natural language generation part of the table question answering system can access the corresponding database based on the above SQL statement, obtain the query results that meet the query conditions, and then generate a reply corresponding to the query request based on the query results, which can be fed back to the user.
[0082] As can be seen from the above, for the training part of the pre-trained model, on the one hand, based on the concatenated vector generated according to the natural language query statement and the database schema data, the associated words in the database schema data are masked to simulate the subsequent possible changes in the natural language query statement input by the user, so that the model has better robustness and fault tolerance. On the other hand, after the corresponding processing in the preprocessing layer, the concatenated vector, especially the part corresponding to the database schema data in the concatenated vector, is pre-trained through the generator-discriminator architecture, so as to effectively capture the relationship between the context, obtain the interaction between the natural language query statement and the database schema data, and improve the accuracy of the model's judgment on the relationship between the natural language query statement and the database schema data. After migrating the trained pre-trained model to the table question answering system, the table question answering system can be effectively applied to the table question answering scenario and output more accurate results for user query requests. As for the form question answering system, since its model is migrated from the pre-trained model, it can effectively handle the non-standard or irregular pattern data-related parts in the user query request, effectively improving the fault tolerance of the form question answering system, and thus protecting the accuracy of the results returned for the query request.
[0083] It should be noted that, in practical applications, the solution described in step S410 above can also form an independent model migration solution. That is, even if the pre-trained model is obtained through third-party training, as long as it has a corresponding structure and has undergone a similar training process so that the model can achieve the above functions, it can also adapt to the migration solution described in step S410 above.
[0084] In this case, the migration plan may include: obtaining model parameters of a pre-trained model to be migrated, wherein the pre-trained model is a model obtained by training based on natural language query statements and database model data, and data after masking the associated words between the natural language query statements and the database model data and the associated words in the database model data part; and performing model migration from the pre-trained model to the table question answering system.
[0085] Among them, the pre-trained model includes a preprocessing layer, a generator and a discriminator; the model migration from the pre-trained model to the table question answering system can be implemented as follows: based on the model parameters of the discriminator in the pre-trained model, the model migration from the pre-trained model to the table question answering system is performed.
[0086] If the model migration is to be migrated to the failing question-answering system, then optionally, based on the model parameters of the discriminator in the pre-trained model, the model migration from the pre-trained model to the table question-answering system can be implemented as follows: by migrating the model parameters of the discriminator in the pre-trained model to the natural language understanding part of the table question-answering system, the model migration from the pre-trained model to the table question-answering system is performed.
[0087] Further optionally, after the model migration is performed, the natural language query request input by the user can be analyzed through the natural language understanding part of the form question and answer system to obtain the database model data in the natural language query request; if it is determined that there is data to be corrected in the database model data, the database model data is corrected; and a database query statement corresponding to the natural language query request is generated based on the correction result.
[0088] The description of the above model migration process is relatively simple, and the relevant parts can refer to the relevant descriptions in the aforementioned step S410 and step S412, and have corresponding beneficial effects, which will not be repeated here.
[0089] Through model migration, the model or system that obtains the migrated data, such as the above-mentioned table question and answer system, can quickly obtain effective and suitable data, thereby accelerating the speed and efficiency of its commissioning. If the above-mentioned pre-trained model is migrated to the table question and answer system, the table question and answer system can be effectively applied to the table question and answer scenario and output more accurate results for user query requests. As for the table question and answer system, since its model is migrated from the pre-trained model, it can effectively handle the non-standard or non-standard parts of the user query request that are related to the pattern data, effectively improving the fault tolerance of the table question and answer system, and thus protecting the accuracy of the results returned for the query request.
[0090] Embodiment 3
[0091] Reference Figure 5 , shows a schematic diagram of the structure of an electronic device according to the third embodiment of the present application. The specific embodiment of the present application does not limit the specific implementation of the electronic device.
[0092] like Figure 5 As shown, the electronic device may include: a processor (processor) 502 , a communication interface (Communications Interface) 504 , a memory (memory) 506 , and a communication bus 508 .
[0093] in:
[0094] The processor 502 , the communication interface 504 , and the memory 506 communicate with each other via a communication bus 508 .
[0095] The communication interface 504 is used to communicate with other electronic devices or servers.
[0096] Processor 502 is used to execute program 510, and specifically can execute the relevant steps in the above-mentioned pre-training model data processing method embodiment.
[0097] Specifically, the program 510 may include program codes, which include computer operation instructions.
[0098] The processor 502 may be a CPU, or an application specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.
[0099] The memory 506 is used to store the program 510. The memory 506 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0100] Program 510 can be specifically used to enable processor 502 to perform operations corresponding to the pre-training model data processing method described in any of the aforementioned multiple method embodiments.
[0101] The specific implementation of each step in program 510 can refer to the corresponding description of the corresponding steps and units in the above method embodiment, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described devices and modules can refer to the corresponding process description in the above method embodiment, which will not be repeated here.
[0102] An embodiment of the present application also provides a computer program product, including computer instructions, which instruct a computing device to perform operations corresponding to any pre-trained model data processing method in the above-mentioned multiple method embodiments.
[0103] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.
[0104] The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk or magneto-optical disk), or implemented as a computer code originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded through a network and to be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a special-purpose processor or programmable or special-purpose hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a special-purpose computer for executing the method shown here.
[0105] Those of ordinary skill in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present application.
[0106] The above implementation methods are only used to illustrate the embodiments of the present application, and are not limitations on the embodiments of the present application. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application. The scope of patent protection of the embodiments of the present application should be limited by the claims.
Claims
1. A pre-training model data processing method, comprising: Generate corresponding concatenated vectors based on natural language query statements and database schema data through the preprocessing layer of the pretrained model; According to the information of the associated words between the natural language query statement and the database schema data, masking is performed on the associated words of the database schema data part in the concatenated vector to obtain a masked vector, wherein the associated words are data that are present in both the natural language query statement and the database schema data and are related to the query field and / or the query condition, and the database schema data includes the table name, column name, and value data of the data table; Performing mask recovery processing on the mask vector for the masked associated words through the generator of the pre-trained model to obtain a generated vector; The discriminator of the pre-trained model is used to evaluate the generation result of the generator based on the generation vector, and the pre-trained model is trained according to the evaluation result.
2. The method according to claim 1, wherein: The discriminator using the pre-trained model evaluates the generation result of the generator based on the generation vector, including: Taking the splicing vector as a supervision condition, using the discriminator of the pre-trained model to compare the generated vector and the splicing vector, and obtaining an evaluation result according to the comparison result.
3. The method according to claim 1, wherein: The preprocessing layer of the pretrained model generates a corresponding concatenation vector according to the natural language query statement and the database schema data, including: The natural language query statement and the database pattern data are concatenated through the preprocessing layer of the pretrained model, and separators are inserted between the concatenated natural language query statement and the database pattern data, and between adjacent pattern data of the database pattern data; Generate a corresponding concatenated vector based on the natural language query statement and database schema data after the delimiter is inserted.
4. The method according to claim 1, wherein: Before generating the corresponding concatenation vector according to the natural language query statement and the database schema data through the preprocessing layer of the pretrained model, the method further includes: An associated word analysis is performed on the natural language query statement and the database schema data, and an associated word between the natural language query statement and the database schema data is determined according to the analysis result.
5. The method according to any one of claims 1 to 4, wherein: The method further comprises: Based on the model parameters of the discriminator in the pre-trained model that has been trained, model migration from the pre-trained model to the table question answering system is performed.
6. The method according to claim 5, wherein: The model parameters of the discriminator in the pre-trained model after training are used to migrate the model from the pre-trained model to the table question answering system, including: The model migration from the pre-trained model to the table question answering system is performed by migrating the model parameters of the discriminator in the pre-trained model that has been trained to the natural language understanding part of the table question answering system.
7. The method according to claim 6, wherein: The method further comprises: Analyzing the natural language query request input by the user through the natural language understanding part to obtain database mode data in the natural language query request; If it is determined that there is data to be corrected in the database pattern data, correcting the database pattern data; A database query statement corresponding to the natural language query request is generated according to the correction result.
8. An electronic device comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method according to any one of claims 1 to 7.
9. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer program product, comprising computer instructions, wherein the computer instructions instruct a computing device to execute operations corresponding to the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Language model pre-training method and system for table pattern analysis and sequence masks
CN112559556A
Semantic structure analysis method and device, equipment, virtualization system and medium
CN113868322A