Data processing method and device, computer equipment, storage medium and program product

By preprocessing and annotating the training data of the language model, high-quality training data is determined, which solves the problem of low accuracy in SQL statement completion in language model, and realizes more efficient SQL statement generation.

CN120123458APending Publication Date: 2025-06-10BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510213501.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When the training data quality of the language model is poor, it is difficult to accurately understand and complete SQL statements that meet specific scenarios, resulting in low accuracy of output SQL statements.

Method used

By obtaining the statement code of the pending query statement, preprocessing and annotating, the training data is determined, so as to formulate processing logic for the query statement, improve the quality of the training data, and thus improve the accuracy of the query statements completed by the language model.

Benefits of technology

The accuracy of the language model for SQL statement completion is improved, and by improving the quality of training data, the model's understanding and generation ability of SQL statements in specific scenarios is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123458A_ABST
    Figure CN120123458A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data processing technologies, artificial intelligence technologies, large model technologies and large language model technologies, and discloses a data processing method and device, computer equipment, a storage medium and a program product.The data processing method comprises the steps that statement codes of query statements to be processed are obtained; preprocessing the statement code to obtain a first code; marking a to-be-complemented code in the first code to obtain a second code; training data is determined according to the first code and the second code, and the training data is used for training a language model to complement the statement code. According to the method, the statement code can be preprocessed to obtain the first code, the to-be-labeled code in the first code is labeled to obtain the second code, and the training data of the language model is determined according to the first code and the second code to formulate the processing logic for the query statement, so that the quality of the training data is improved; and the accuracy of the query statement complemented by the language model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical fields of data processing, artificial intelligence, large model technology, and large language model technology, and particularly relates to a data processing method, apparatus, computer device, storage medium, and program product. Background Art

[0002] Code completion is a crucial auxiliary function in an integrated development environment. Its core purpose is to assist developers or code writers in generating the code to be written. This function is crucial for improving the development efficiency of developers or code writers and can significantly reduce the time and resource waste during the code writing process. With the development of language models, pre-trained language models have demonstrated powerful capabilities and can write code according to user instructions.

[0003] In a specific language field, such as when completing SQL (Structured Query Language) statements, high precision is usually required because even a tiny error may lead to incorrect query results or complete failure. However, SQL statements have unique syntax, keywords, and structures, all of which require specific knowledge and understanding to be used correctly. Although language models can handle a wide range of natural language tasks, for such highly specialized code statements, often without sufficient context information, it may be difficult for language models to accurately understand and complete SQL statements that meet specific scenarios. But in related solutions for processing training data of language models, the processing logic for natural language is often used to process SQL statements, resulting in poor quality of the training data of language models, thereby affecting the precision of the SQL statements output by language models. Summary of the Invention

[0004] In view of this, the present disclosure provides a data processing method, apparatus, computer device, storage medium, and program product to solve the problem that the poor quality of the training data of language models affects the precision of the SQL statements output by language models.

[0005] In a first aspect, the present disclosure provides a data processing method, which includes:

[0006] Obtain the statement code of a query statement to be processed, where the query statement is a structured processing statement for instructing a language model to perform data query and analysis;

[0007] Preprocess the statement code to obtain a first code;

[0008] Annotate the code to be completed in the first code to obtain a second code;

[0009] Determine training data according to the first code and the second code, where the training data is used to train a language model to complete statement codes.

[0010] In a second aspect, the present disclosure provides a data processing device, which includes:

[0011] An acquisition module, configured to acquire the statement code of a query statement to be processed, where the query statement is a structured processing statement for instructing a language model to perform data query and analysis;

[0012] A preprocessing module, configured to preprocess the statement code to obtain a first code;

[0013] A marking module, configured to mark the code to be completed in the first code to obtain a second code;

[0014] A determination module, configured to determine training data according to the first code and the second code, where the training data is used to train a language model to complete statement codes.

[0015] In a third aspect, the present disclosure provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the data processing method according to the first aspect or any corresponding embodiment thereof.

[0016] In a fourth aspect, the present disclosure provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to perform the data processing method according to the first aspect or any corresponding embodiment thereof.

[0017] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to perform the data processing method according to the first aspect or any corresponding embodiment thereof.

[0018] In the embodiments of the present disclosure, first, the statement code of a query statement to be processed can be acquired, where the query statement is a structured processing statement for instructing a language model to perform data query and analysis. Next, the statement code can be preprocessed to obtain a first code, and the code to be completed in the first code can be marked to obtain a second code. Then, training data can be determined according to the first code and the second code, where the training data is used to train a language model to complete statement codes, so as to preprocess the statement code to obtain a first code, mark the code to be marked in the first code to obtain a second code, and determine the training data of the language model according to the first code and the second code, so as to formulate a processing logic for the query statement, improve the quality of the training data, and further improve the accuracy of the query statement completed by the language model. Brief Description of the Drawings

[0019] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 is a schematic flowchart of a data processing method according to an embodiment of the present disclosure;

[0021] Figure 2 is an architecture diagram of a Transformer Decoder large language model according to an embodiment of the present disclosure;

[0022] Figure 3 is a schematic flowchart of processing an SQL statement to obtain training data according to an embodiment of the present disclosure;

[0023] Figure 4 is a schematic diagram of filtering low-quality data in training data according to an embodiment of the present disclosure;

[0024] Figure 5 is a schematic flowchart of the training process of a language model according to an embodiment of the present disclosure;

[0025] Figure 6 is a structural block diagram of a data processing device according to an embodiment of the present disclosure;

[0026] Figure 7 is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present disclosure. Specific Embodiments

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present disclosure.

[0028] Combined with the application scenarios on which the execution of the data processing method depends, the application scenarios will be described herein.

[0029] Code completion is a crucial auxiliary function in an integrated development environment, whose core purpose is to assist developers or code writers in generating the code to be written. For example, the development environment can be an IDE (Integrated Development Environment), a text editor sub, or an online code editor on GitHub, etc. This function is crucial for improving the development efficiency of developers or code writers and can significantly reduce the time and resource waste in the process of writing code. With the development of language models, pre-trained language models have demonstrated powerful capabilities and can write code according to user instructions.

[0030] In a specific language domain, such as when completing SQL (Structured Query Language) statements, a high degree of precision is usually required because even a small error may lead to incorrect query results or complete failure. However, SQL statements have unique syntax, keywords, and structures, all of which require specific knowledge and understanding to be used correctly. Although language models can handle a wide range of natural language tasks, for such highly specialized code statements, without sufficient context information, language models may often have difficulty accurately understanding and completing SQL statements that conform to specific scenarios. However, in related solutions for processing training data of language models, the processing logic for natural language is often used to process SQL statements, resulting in poor quality of the training data of language models, thereby affecting the precision of the SQL statements output by the language models.

[0031] For example, the construction of training data in related solutions is often relatively single. Usually, a large number of SQL codes are cut to obtain completion training data. For example, the following text in the SQL code is intercepted to train the model's ability to complete the following text based on the above text. However, due to the large amount of code data, often reaching the million / tens of millions level, the quality of the training data is uneven, and there are a certain amount of non-derivable target code segments. After training with such data, the hallucination problem of the language model is often relatively serious, that is, when using the language model to generate results, there are phenomena such as being inconsistent with the facts, logically chaotic, or not following instructions and context.

[0032] Based on this, embodiments of the present disclosure provide a data processing method. First, the statement code of the query statement to be processed can be obtained, where the query statement is a structured processing statement used to instruct a language model to perform data query and analysis. Next, the statement code can be preprocessed to obtain a first code, and the code to be completed in the first code can be marked to obtain a second code. Then, training data can be determined based on the first code and the second code, where the training data is used to train the language model to complete the statement code, thereby preprocessing the statement code to obtain a first code, marking the code to be marked in the first code to obtain a second code, and determining the training data of the language model based on the first code and the second code to formulate the processing logic for the query statement, improving the quality of the training data, and further improving the accuracy of the query statement completed by the language model.

[0033] According to an embodiment of the present disclosure, an embodiment of a data processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0034] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0035] For example, when responding to a user's active request, a prompt message can be sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.

[0036] As an optional but non-limiting implementation manner, the way of sending a prompt message to the user in response to receiving the user's active request can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0037] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure, and other ways that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0038] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of data) should comply with the requirements of corresponding laws, regulations and related provisions.

[0039] According to an embodiment of the present disclosure, an embodiment of a data processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0040] In this embodiment, a data processing method is provided, which can be used for a server. Figure 1 It is a flowchart of the data processing method according to an embodiment of the present disclosure, as Figure 1 shown, and the process includes the following steps:

[0041] Step S101, obtain the statement code of the query statement to be processed, where the query statement is a structured processing statement used to instruct a language model to perform data query and analysis.

[0042] In an embodiment of the present disclosure, the query statement to be processed can be the above SQL statement, that is, a structured query statement that can be understood by a machine generated by the language model according to the natural language input by the user. The language model can be a large language model, such as a large language model using the Transformer Decoder architecture. Here, as Figure 2 shown is the architecture diagram of the Transformer Decoder large language model.

[0043] When obtaining the statement code of the above query statement to be processed, the statement code can be obtained based on public datasets such as the Text2SQL dataset. The public dataset contains a large number of natural language questions and corresponding SQL query statements. The purpose is to test the performance of the model on complex and unseen SQL queries and its generalization ability in new fields.

[0044] In addition, if an enterprise has its own accumulated database query logs, which contain natural language queries input by users and corresponding SQL statements, these data can be used as important resources for training the large language model to complete SQL statements. Enterprise internal data has the characteristics of strong pertinence and close combination with business scenarios, and can enable the trained language model to better adapt to the specific needs of the enterprise.

[0045] It should be understood that the statement code of the query statement to be processed can also be obtained by other means, and the present disclosure will not elaborate on the specific method of obtaining the statement code, which is subject to the usage requirements of users in the actual usage scenario.

[0046] Step S102: Preprocess the statement code to obtain the first code.

[0047] In the embodiments of the present disclosure, in some common publicly available datasets, although they provide basic data for research, the amount of data itself is relatively limited. Therefore, in order to increase the diversity and scale of training data, researchers may use these publicly available datasets multiple times, resulting in duplicate samples in the statement code.

[0048] Based on this, the statement code can be preprocessed to remove duplicates from the statement code and obtain the first code. Among them, the first code can be AR (Autoregressive Data) data, and the AR data refers to training data that predicts the subsequent text based on the previous text in chronological order.

[0049] Step S103: Label the code to be completed in the first code to obtain the second code.

[0050] In the embodiments of the present disclosure, the second code data can be FIM (Fim In Middle Data) data, and the FIM data is used to indicate training data for predicting the middle code based on the upper code and the lower code.

[0051] Specifically, the middle code in the first code can be labeled as the code to be completed to obtain the second code, so as to train the language model to predict the middle code based on the context of the middle code, thus better adapting to the characteristics that the writing of SQL statements often depends on specific database schemas, data types, and business logics, and it may be difficult for the language model to accurately understand and generate SQL statements that conform to specific scenarios without sufficient context information.

[0052] Step S104: Determine the training data according to the first code and the second code, where the training data is used to train the language model to complete the statement code.

[0053] In the embodiments of the present disclosure, the training data can include both the first code and the second code, so as to more comprehensively train the code completion ability of the language model. Specifically, the target ratio of the first code and the second code in the training data can be determined according to the specific usage scenario to generate the training data.

[0054] For example, for the scenario of completing the statement code corresponding to the above SQL statement, considering that the language model often relies on context information to predict the middle code, the proportion of the second code in the training code can be set relatively high. For example, the ratio of the first code to the second code can be set to 3:7.

[0055] As can be seen from the above description, in the embodiments of the present disclosure, first, the statement code of the query statement to be processed can be obtained, where the query statement is a structured processing statement for instructing the language model to perform data query and analysis. Next, the statement code can be preprocessed to obtain the first code, and the code to be completed in the first code can be marked to obtain the second code. Then, the training data can be determined according to the first code and the second code, where the training data is used to train the language model to complete the statement code, so as to preprocess the statement code to obtain the first code, mark the code to be marked in the first code to obtain the second code, and determine the training data of the language model according to the first code and the second code, so as to formulate the processing logic for the query statement, improve the quality of the training data, and further improve the accuracy of the query statement completed by the language model.

[0056] In some alternative embodiments, step S102 of preprocessing the statement code to obtain the first code includes:

[0057] Step S1021, obtaining the feature vectors of the statement codes corresponding to each query statement to be processed.

[0058] Step S1022, identifying the duplicate codes in the statement code based on the similarity of the feature vectors, and performing deduplication processing on the statement code based on the duplicate codes to obtain the deduplicated code.

[0059] Step S1023, intercepting the deduplicated code based on the interception position to obtain the upper text code and the lower text code.

[0060] Step S1024, marking the code to be completed in the lower text code to obtain the first code.

[0061] In the embodiments of the present disclosure, the preprocessing may include deduplication processing and marking processing, where the deduplication processing can be used to deduplicate the above statement code, and the standard processing can be used to intercept the deduplicated code to obtain the above AR data.

[0062] Here, as Figure 3 shown is a schematic flow diagram of processing the SQL statement to obtain the training data. First, the MinHash deduplication algorithm can be used to deduplicate the statement code of the SQL statement to reduce the redundant SQL code in the statement code.

[0063] Specifically, MinHash is a method for estimating the Jaccard similarity of comparison objects. For the statement code of SQL statements, it can be converted into a bag-of-words representation, and then a feature vector is generated through the MinHash algorithm. Then, the feature vector can be divided into multiple segments, and each segment is used as the key of a bucket. Only the records that meet in at least one bucket are considered candidate similar items. Based on this, redundant code in the statement code can be identified and deduplicated to obtain deduplicated code.

[0064] Then, based on the truncation position, the deduplicated code can be divided into upper-context code and lower-context code, and the lower-context code is marked as the code to be completed, so as to obtain the first code, and the ability of the language model to complete the lower-context based on the upper-context of the statement code can be trained through this first code.

[0065] In the embodiments of the present disclosure, considering that the ways to obtain training data are limited, in order to increase the diversity and scale of training data, there is often a problem of data reuse, which leads to duplicate samples in the statement code. Therefore, this application can first perform preprocessing to deduplicate the code data to improve the quality of training data.

[0066] In some optional embodiments, the above step S103 of marking the code to be completed in the first code to obtain the second code includes:

[0067] Step S1031: Based on a preset slicing position, the first code is sliced into upper-context code, lower-context code, and middle code.

[0068] Step S1032: Mark the code to be completed in the lower-context code to obtain the second code.

[0069] In the embodiments of the present disclosure, as can be seen from the above, the second code is FIM data, where the FIM data includes before code (upper-context code, that is, the code before the code segment to be completed), middle code (middle code, that is, the code to be completed), and after code (lower-context code, that is, the code after the code segment to be completed).

[0070] Specifically, the middle code can be marked based on a preset slicing position, the code before the marked middle code is marked as upper-context code, and at the same time, the code after the middle code is marked as lower-context code to obtain the second code.

[0071] In the embodiments of the present disclosure, based on the first code, the second code can be generated, thereby enhancing the diversity of training data and further improving the accuracy of the output result of the language model.

[0072] In some alternative embodiments, step S1031 of splitting the first code into an upper code, a lower code, and an intermediate code based on a preset slicing position includes:

[0073] Step a1: Obtain the abstract syntax tree corresponding to the first code, where the abstract syntax tree is used to indicate the stacking relationship between the constituent objects in the first code.

[0074] Step a2: Identify the target subtree in the abstract syntax tree and locate the preset slicing position based on the target subtree.

[0075] Step a3: Based on the preset slicing position, label the code corresponding to the target subtree as the intermediate code, and label the upper code and the lower code based on the intermediate code.

[0076] In the embodiments of the present disclosure, the abstract syntax tree (Abstract Syntax Tree, hereinafter referred to as AST) can be an abstract representation of the first code. Among them, the AST includes multiple subtrees, and each subtree usually corresponds to a main component or clause in the SQL statement.

[0077] For example, the root node represents the entire SQL statement. The SELECT subtree corresponds to the select clause in the SQL statement and contains information such as the column names and expressions to be selected. The FROM subtree corresponds to the from clause and specifies the source table or view of the data, etc. The WHERE subtree corresponds to the where clause and is used to filter the rows that meet the conditions. The GROUP BY subtree corresponds to the group by clause and is used to group the result set. The ORDER BY subtree corresponds to the order by clause and is used to sort the result set. The subquery subtree is used to represent the structure and content of the subquery statement. The JOIN subtree is used to represent the type of join, the join conditions, and the joined tables, etc.

[0078] When identifying the target subtree, a preset subtree T can be obtained. For example, a subtree with a root node of select, group by, join, etc., and the subtree in the AST that matches the preset subtree T is determined as the target subtree. Then, a starting node can be selected in the target subtree, and a preorder traversal can be performed starting from this starting node until the target subtree is traversed. Here, the traversed target subtree can be further filtered. For example, the position where the target subtree with the code layer of the target subtree meeting the slicing requirements is determined as the preset slicing position, so as to label the code corresponding to the target subtree as the intermediate code.

[0079] In the embodiments of the present disclosure, the first code may be sliced based on the abstract code tree of the first code to obtain a second code. Here, since the structure and semantic information of the first code are fully presented, by slicing the abstract code tree, the obtained slicing result can be more accurate, facilitating the model's understanding and providing more context - and semantics - compliant completion suggestions for the language model.

[0080] In some alternative embodiments, step S1031 of slicing the first code into an upper - context code, a lower - context code, and an intermediate code based on a preset slicing position further includes:

[0081] Step b1, obtaining a preset number of lines.

[0082] Step b2, marking the code of the preset number of lines after the preset slicing position in the first code as the intermediate code, and marking the upper - context code and the lower - context code based on the intermediate code.

[0083] In the embodiments of the present disclosure, the first code may be sliced using a sliding - line slicing method. Specifically, a sliding window may be generated, and the height set for the sliding window is set to a preset number of lines H. Then, any line number N in the first code may be used as the preset slicing position based on the sliding window, and at this preset slicing position, the intermediate code is sliced based on the sliding window, that is, the code in the line number range [N, N + H] in the first code is used as the intermediate code.

[0084] Next, if the number of lines of code in the first code is M, the code in the line number range [0, N] may be used as the upper - context code, and the code in the line number range [N + H, M] may be used as the lower - context code to obtain a second code containing H samples.

[0085] In the embodiments of the present disclosure, based on the first code, slicing may be performed in the first code based on a preset number of lines to generate a second code, thereby enhancing the diversity of training data and further improving the accuracy of the output result of the language model.

[0086] In some alternative embodiments, step S1031 of slicing the first code into an upper - context code, a lower - context code, and an intermediate code based on a preset slicing position further includes:

[0087] Step c1, determining a target function in the first code that matches the preset slicing position.

[0088] Step c2, marking the code corresponding to the target function as the intermediate code, and marking the upper - context code and the lower - context code based on the intermediate code.

[0089] In the embodiments of the present disclosure, the target function may be a window function. Among them, the window function may perform aggregation, sorting, and analysis operations on a subset (i.e., a window) of the query result set without changing the number of rows in the original query result.

[0090] For example, the above window functions may include: window functions such as aggregation window functions, sorting window functions, and offset window functions. Among them, the aggregation window function can be used to perform aggregation calculations on the data within the window. Common aggregation window functions include SUM(), AVG(), COUNT(), MAX(), MIN(), etc. For example, SUM() OVER() can be used to calculate the cumulative sum of sales for each department or region. The sorting window function can be used to perform sorting, ranking, and grouping calculations on the data within the window. The offset window function can be used to access the values of the previous or next few rows of the current row.

[0091] Specifically, the target function in the first code can be identified, and the position where the target function is located can be determined as the preset slicing position, so as to label the code corresponding to the target function as an intermediate function based on the preset slicing position to obtain the second code.

[0092] In the embodiments of the present disclosure, the target function may be a window function. Considering that the window function in the query statement is relatively complex and professional, and its frequency of occurrence in daily natural language expressions and conventional codes is relatively low. As a result, the number of high-quality examples of window functions encountered by the large language model during training is limited, and it is difficult to fully learn its various usage methods and patterns. Therefore, the window function can be used as intermediate code to focus on training the language model's ability to complete window functions.

[0093] In some alternative embodiments, the above preset slicing position includes a first slicing position and a second slicing position. Among them, the first slicing position includes any one of the beginning of the code line, punctuation space, and target character in the first code; the second slicing position includes any one of the punctuation space and target character after the first slicing position. The above step S1031, based on the preset slicing position, splitting the first code into upper context code, lower context code, and intermediate code, further includes:

[0094] Label the code before the preset slicing position in the first code as upper context code, label the code after the preset slicing position in the first code as lower context code, and label the intermediate code as empty.

[0095] In the embodiments of the present disclosure, the first slicing position is the starting slicing position of the intermediate code, and the second slicing position is the ending slicing position of the intermediate code. Here, the selection probability of the first slicing position can be preset. For example, the first ratios of the code line start, punctuation space, and target character can be 1:1:1 (i.e., 33%, 33%, 33%). At the same time, the selection probability of the first slicing position can be preset. For example, the second ratios of the punctuation space and the target character can be 1:1 (i.e., 50%, 50%).

[0096] When annotating the intermediate code, first, the first slicing position can be selected based on the first ratio, and after this first slicing position, the second slicing position can be arbitrarily selected based on the second ratio, and the code between the first slicing position and the second slicing position can be determined as the intermediate code to obtain the second code.

[0097] In the embodiments of the present disclosure, based on the first code, slicing can be performed in the first code based on the first slicing position and the second slicing position to generate the second code, thereby enhancing the diversity of the training data and further improving the accuracy of the output result of the language model.

[0098] In some alternative embodiments, step S1031 above, slicing the first code into the above code, the below code, and the intermediate code based on the preset slicing position, further includes:

[0099] The code before the preset slicing position in the first code is labeled as the above code, the code after the preset slicing position in the first code is labeled as the below code, and the intermediate code is labeled as empty.

[0100] In the embodiments of the present disclosure, any position in the first code can be selected as the preset slicing position. For example, the start of the Nth line can be used as the preset slicing position, or the end of the Mth line can be used as the preset slicing position.

[0101] Then, the code before the preset slicing position can be labeled as the above code, the code after the preset slicing position can be labeled as the below code, and the intermediate code can be labeled as an empty slice, thereby obtaining the second code corresponding to the first code.

[0102] In the embodiments of the present disclosure, based on the first code, slicing can be performed in the first code based on the first slicing position and the second slicing position to generate the second code. In this second code, the intermediate code is an empty slice, thereby enhancing the diversity of the training data and further improving the accuracy of the output result of the language model.

[0103] In some alternative embodiments, the above-mentioned preset slicing positions include a third slicing position and a fourth slicing position. Among them, the third slicing position is any position in the first code, and the fourth slicing position is a position after the third slicing position and the distance between the fourth slicing position and the third slicing position is less than a preset value. The above step S1031, based on the preset slicing positions, slices the first code into an upper-context code, a lower-context code, and an intermediate code, and further includes:

[0104] Label the code before the third slicing position in the first code as the upper-context code, label the code between the third slicing position and the fourth slicing position in the first code as the intermediate code, and label the code after the fourth slicing position in the first code as the lower-context code.

[0105] In the embodiments of the present disclosure, the third slicing position is the starting slicing position of the intermediate code, and the fourth slicing position is the ending slicing position of the intermediate code. The preset value may include a preset line spacing and a preset character spacing. Among them, the line spacing between the third slicing position and the fourth slicing position needs to be less than the preset line spacing dc, and the character spacing needs to be less than the preset character spacing dl.

[0106] When labeling the intermediate code, the code between the third slicing position and the fourth slicing position can be used as the intermediate code, the code between the third code can be used as the upper-context code, and the code after the fourth code can be used as the lower-context code to obtain the second code.

[0107] It should be understood that after obtaining the second code by slicing according to the embodiment corresponding to the above step S103, the second codes obtained by each slicing method can be mixed based on a preset mixing ratio to obtain training data. For example, set the ratio of generating the second code based on the AST to 50%, set the ratio of the second code labeled by the first slicing position and the second slicing position to 20%, and set the ratio of the second code determined by the remaining slicing methods to 30%.

[0108] In the embodiments of the present disclosure, based on the first code, slicing can be performed in the first code based on the third slicing position and the fourth slicing position to generate the second code, thereby improving the diversity of the training data and further improving the accuracy of the output result of the language model.

[0109] In some alternative embodiments, the above step S104, determining the training data according to the first code and the second code, includes:

[0110] Step S1041, obtain the description information of the query data table indicated by the second code.

[0111] Step S1042, splice the description information with the first code and the second code respectively to obtain spliced data.

[0112] Step S1043: Determine the training data according to the splicing data.

[0113] In the embodiments of the present disclosure, when generating training data based on the first code and the second code, it is necessary to generate the schemas of the first code and the second code, that is, the description information of the query data tables. Among them, the description information may include the data tables indicated by the SQL statements corresponding to the first code and the second code, such as data types, field information in the data tables, index information, etc.

[0114] Then, the splicing data can be obtained by splicing the schema with the first code and the second code respectively. Among them, the splicing data may include: schema, the above code, intermediate code, and the following code. Next, the ratio of the splicing data of the first code and the splicing data of the second code obtained based on the above target ratio setting can be used to obtain the training data.

[0115] In addition, a validation set can be generated based on the training data. During the model training process, the validation set can be used regularly to evaluate the performance of the model and observe the performance of the model on unseen data, so as to understand whether the model can generalize to new data and the generalization ability of the model.

[0116] In the embodiments of the present disclosure, the quality of the training data can be evaluated through the method of model data governance to filter out the low-quality data therein, and the high-quality data therein can be extracted as the validation set according to a preset ratio, such as 10:1, that is, 10% of the high-quality samples in the training data are extracted as the validation set. Here, as Figure 3 shown, the quality of the training data can be evaluated based on the pre-trained model.

[0117] Specifically, as Figure 4 shown is a schematic diagram of filtering out low-quality data in the training data. Among them, first, the training data is sampled, and a hint header for judging data quality is designed. For example, the quality judgment rules can be preset and written into the hint header, so that the pre-trained model performs quality annotation on the training data based on the hint header and trains the model through data governance. Then, the pre-trained model can be trained through the training model to obtain a classification large model for judging data quality (that is, Figure 4 the data governance model in), so as to classify the unlabeled training data through the classification large model to filter out the low-quality data therein, thereby obtaining high-quality training data.

[0118] Next, the language model can be trained according to the training data and the validation set. After a certain number of steps of training, the language model with the lowest loss on the validation set is selected as the finally output language model. The specific training process of the language model is as Figure 5As shown, the specific training steps in the present disclosure will not be elaborated.

[0119] In summary, in the embodiments of the present disclosure, first, the statement code of the query statement to be processed can be obtained, where the query statement is a structured processing statement for instructing the language model to perform data query and analysis. Next, the statement code can be preprocessed to obtain the first code, and the code to be completed in the first code can be marked to obtain the second code. Then, the training data can be determined according to the first code and the second code, where the training data is used to train the language model to complete the statement code, so as to preprocess the statement code to obtain the first code, and mark the code to be marked in the first code to obtain the second code, and determine the training data of the language model according to the first code and the second code, so as to formulate the processing logic for the query statement, improve the quality of the training data, and further improve the accuracy of the query statement completed by the language model.

[0120] In this embodiment, a data processing device is further provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be elaborated again. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0121] This embodiment provides a data processing device, as Figure 6 shown, including:

[0122] An acquisition module 601, configured to acquire the statement code of the query statement to be processed, where the query statement is a structured processing statement for instructing the language model to perform data query and analysis;

[0123] A preprocessing module 602, configured to preprocess the statement code to obtain the first code;

[0124] A marking module 603, configured to mark the code to be completed in the first code to obtain the second code;

[0125] A determination module 604, configured to determine the training data according to the first code and the second code, where the training data is used to train the language model to complete the statement code.

[0126] In some alternative implementation manners, the preprocessing module 602 is further configured to:

[0127] Acquire the feature vectors of the statement codes corresponding to each query statement to be processed;

[0128] Identify duplicate code in statement code based on the similarity of feature vectors, and perform deduplication processing on the statement code based on the duplicate code to obtain deduplicated code;

[0129] Intercept the deduplicated code based on the interception position to obtain the above code and the below code;

[0130] Mark the code to be completed in the below code to obtain the first code.

[0131] In some alternative embodiments, the marking module 603 is further configured to:

[0132] Based on a preset slicing position, slice the first code into above code, below code, and middle code;

[0133] Mark the code to be completed in the below code to obtain the second code.

[0134] In some alternative embodiments, the marking module 603 is further configured to:

[0135] Obtain the abstract syntax tree corresponding to the first code, where the abstract syntax tree is used to indicate the stacking relationship between the various constituent objects in the first code;

[0136] Identify the target subtree in the abstract syntax tree, and locate the preset slicing position based on the target subtree;

[0137] Based on the preset slicing position, mark the code corresponding to the target subtree as the middle code, and mark the above code and the below code based on the middle code.

[0138] In some alternative embodiments, the marking module 603 is further configured to:

[0139] Obtain the preset number of lines;

[0140] Mark the code of the preset number of lines after the preset slicing position in the first code as the middle code, and mark the above code and the below code based on the middle code.

[0141] In some alternative embodiments, the marking module 603 is further configured to:

[0142] Determine the target function in the first code that matches the preset slicing position;

[0143] Mark the code corresponding to the target function as the middle code, and mark the above code and the below code based on the middle code.

[0144] In some alternative embodiments, the preset slicing positions include a first slicing position and a second slicing position. Among them, the first slicing position includes any one of the beginning of a code line, a punctuation space, and a target character in the first code; the second slicing position includes any one of a punctuation space and a target character after the first slicing position; the annotation module 603 is further configured to:

[0145] Annotate the code before the first slicing position in the first code as the above text code, annotate the code between the first slicing position and the second slicing position in the first code as the middle code, and annotate the code after the second slicing position in the first code as the below text code.

[0146] In some alternative embodiments, the annotation module 603 is further configured to:

[0147] Annotate the code before the preset slicing position in the first code as the above text code, annotate the code after the preset slicing position in the first code as the below text code, and annotate the middle code as empty.

[0148] In some alternative embodiments, the preset slicing positions include a third slicing position and a fourth slicing position. Among them, the third slicing position is any position in the first code, and the fourth slicing position is a position after the third slicing position and the distance between the fourth slicing position and the third slicing position is less than a preset value; the annotation module 603 is further configured to:

[0149] Annotate the code before the third slicing position in the first code as the above text code, annotate the code between the third slicing position and the fourth slicing position in the first code as the middle code, and annotate the code after the fourth slicing position in the first code as the below text code.

[0150] In some alternative embodiments, the determination module 604 is further configured to:

[0151] Obtain the description information of the query data table indicated by the second code;

[0152] Based on the description information, splice it with the first code and the second code respectively to obtain spliced data;

[0153] Determine the training data according to the spliced data.

[0154] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding above embodiments, and will not be elaborated here.

[0155] The data processing device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0156] The embodiments of the present disclosure also provide a computer device having the above Figure 6 data processing device.

[0157] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present disclosure. As shown in Figure 7 , the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi-processor system). Figure 7 In

[0158] , a processor 10 is taken as an example.

[0159] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.

[0160] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device and the like. In addition, the memory 20 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0161] The memory 20 may include volatile memory, for example, random access memory; the memory may also include non-volatile memory, for example, flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memory.

[0162] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0163] The embodiments of the present disclosure also provide a computer-readable storage medium. The methods according to the embodiments of the present disclosure can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and to be stored in a local storage medium, so that the methods described herein can be stored in such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.

[0164] A part of the present invention can be applied as a computer program product, for example, computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the present invention through the operations of the computer. Those skilled in the art should understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.

[0165] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A data processing method, characterized in that: The method comprises: Obtaining a statement code of a query statement to be processed, wherein the query statement is a structured processing statement used to instruct a language model to perform data query and analysis; Preprocessing the statement code to obtain a first code; Marking the code to be completed in the first code to obtain a second code; Training data is determined according to the first code and the second code, wherein the training data is used to train the language model to complete the sentence code.

2. The method according to claim 1, characterized in that: The preprocessing of the statement code to obtain the first code includes: Obtaining a feature vector of a statement code corresponding to each of the query statements to be processed; Based on the similarity of the feature vectors, duplicate codes in the statement code are identified, and duplicate removal processing is performed on the statement code based on the duplicate codes to obtain duplicate removal codes; The deduplicated code is intercepted based on the interception position to obtain the above code and the below code; The code to be completed is marked in the following code to obtain the first code.

3. The method according to claim 1, characterized in that The step of marking the code to be completed in the first code to obtain the second code includes: Based on a preset slicing position, the first code is divided into an upper code, a lower code and an intermediate code; The code to be completed is marked in the following code to obtain the second code.

4. The method according to claim 3, characterized in that: The first code is divided into a previous code, a following code and an intermediate code based on a preset slice position, including: Obtaining an abstract syntax tree corresponding to the first code, wherein the abstract syntax tree is used to indicate a stacking relationship between various component objects in the first code; Identify a target subtree in the abstract syntax tree, and locate the preset slice position based on the target subtree; Based on the preset slice position, the code corresponding to the target subtree is marked as an intermediate code, and the upper code and the lower code are marked based on the intermediate code.

5. The method according to claim 3, characterized in that: The method of dividing the first code into a preceding code, a following code and an intermediate code based on a preset slicing position further includes: Get the preset number of rows; The preset number of lines of code after the preset slice position in the first code are marked as intermediate codes, and the upper code and the lower code are marked based on the intermediate codes.

6. The method according to claim 3, characterized in that The method of dividing the first code into a preceding code, a following code and an intermediate code based on a preset slicing position further includes: determining, in the first code, an objective function that matches the preset slice position; The code corresponding to the objective function is marked as an intermediate code, and the upper code and the lower code are marked based on the intermediate code.

7. The method according to claim 3, characterized in that The preset slicing position includes a first slicing position and a second slicing position, wherein the first slicing position includes any one of the beginning of a code line, a punctuation space, and a target character in the first code; the second slicing position includes any one of the punctuation space and the target character after the first slicing position; The method of dividing the first code into a preceding code, a following code and an intermediate code based on a preset slicing position further includes: The code before the first slice position in the first code is marked as the above code, the code between the first slice position and the second slice position in the first code is marked as the middle code, and the code after the second slice position in the first code is marked as the below code.

8. The method according to claim 3, characterized in that The method of dividing the first code into a preceding code, a following code and an intermediate code based on a preset slicing position further includes: The code before the preset slice position in the first code is marked as the above code, and the code after the preset slice position in the first code is marked as the below code, and the intermediate code is marked as empty.

9. The method according to claim 3, characterized in that: The preset slice position includes a third slice position and a fourth slice position, wherein the third slice position is any position in the first code, and the fourth slice position is a position after the third slice position and the distance between the third slice position and the fourth slice position is less than a preset value; The method of dividing the first code into a preceding code, a following code and an intermediate code based on a preset slicing position further includes: The code before the third slice position in the first code is marked as the above code, the code between the third slice position and the fourth slice position in the first code is marked as the middle code, and the code after the fourth slice position in the first code is marked as the below code.

10. The method according to claim 1, characterized in that The determining of training data according to the first code and the second code includes: Obtaining description information of the query data table indicated by the second code; Based on the description information, the first code and the second code are respectively spliced ​​to obtain spliced ​​data; The training data is determined according to the spliced ​​data.

11. A data processing device, characterized in that: The device comprises: An acquisition module, used to acquire a statement code of a query statement to be processed, wherein the query statement is a structured processing statement used to instruct a language model to perform data query and analysis; A preprocessing module, used for preprocessing the statement code to obtain a first code; A marking module, used for marking the code to be completed in the first code to obtain a second code; A determination module is used to determine training data based on the first code and the second code, wherein the training data is used to train the language model to complete the sentence code.

12. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data processing method according to any one of claims 1 to 10 by executing the computer instructions.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the data processing method according to any one of claims 1 to 10.

14. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the data processing method according to any one of claims 1 to 10.