Code task processing method, computing device and computer readable storage medium
By semantic analysis and training of sample code files, a pre-trained model is obtained, which solves the problem that code generation depends on labeled data in the existing technology, and achieves more efficient and accurate code generation.
Patent Information
- Application Number
- CN202410084423.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-19
- Publication Date
- 2025-07-22
AI Technical Summary
The existing technology relies on a large amount of labeled data during the code generation process, resulting in high cost of model training, insufficient performance and accuracy, and it is difficult to meet the needs of complex code generation tasks.
By obtaining the pending data of the target code task, using the pre-trained model for semantic analysis, training the structure data of the sample code file to obtain the function data, and training the initial code model based on the function data, obtaining the pre-trained model, reducing the dependence on the labeled data, and improving the model's learning ability of function-level features.
It improves the model's understanding of code structure and semantics, enhances the accuracy and efficiency of code generation, reduces dependence on labeled data, and improves the execution effect of downstream tasks.
Smart Images

Figure CN120353464A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technologies, and particularly to a method for processing code tasks, a computing device, and a computer-readable storage medium. Background Art
[0002] With the development and progress of computer technologies, in the process of software development, there is an increasing need for a deeper understanding and more accurate generation of code. Traditional code generation technologies often rely on rules or templates, and this method often fails to meet the requirements when dealing with complex code generation tasks.
[0003] However, enabling a model to learn and generate code through model training often requires a large amount of labeled data, and the acquisition cost of training samples is relatively high. Therefore, there is an urgent need for a faster and better model training method to improve the model's learning ability for code, thereby improving the accuracy of code generation. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a method for processing code tasks. One or more embodiments of this specification also relate to a method for pre-training a model applied to a cloud-side device, a method for task completion, a system for processing code tasks, a device for processing code tasks, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, there is provided a method for processing code tasks, including:
[0006] Obtain the data to be processed in the target code task;
[0007] Input the data to be processed into the code task model to obtain a target code file corresponding to the target code task, where the code task model is obtained by training a pre-trained model based on a sample set corresponding to the target code task, and the pre-trained model is obtained by training based on function data obtained by semantic analysis of sample code files.
[0008] According to a second aspect of the embodiments of this specification, there is provided a method for pre-training a model, applied to a cloud-side device, including:
[0009] Obtain a sample code set, where the sample code set includes sample code files and structural data of the sample code files;
[0010] Use an initial code model to perform semantic analysis on the sample code files based on the structural data to obtain function data, and perform prediction on the sample code files based on the function data to obtain predicted code files;
[0011] Train an initial code model based on the loss between the predicted code file and the sample code file to obtain a pre-trained model;
[0012] Send the model parameters of the pre-trained model to the edge device.
[0013] According to the third aspect of the embodiments of this specification, a code completion method is provided, including:
[0014] Receive the code to be completed sent by the front-end user;
[0015] Input the code to be completed into the code completion model to obtain a target code file, where the code completion model is trained based on a sample set corresponding to the code completion task on the pre-trained model, and the pre-trained model is trained based on function data obtained by semantic analysis of the sample code file;
[0016] Feedback the target code file to the front-end user.
[0017] According to the fourth aspect of the embodiments of this specification, a code task processing system is provided, including a client and a server:
[0018] The client is used to send the data to be processed in the target code task to the server;
[0019] The server is used to input the data to be processed into the code task model to obtain a target code file corresponding to the target code task, where the code task model is trained based on a sample set corresponding to the target code task on the pre-trained model, and the pre-trained model is trained based on function data obtained by semantic analysis of the sample code file;
[0020] The client is further used to receive the target code file corresponding to the target code task sent by the server.
[0021] According to the fifth aspect of the embodiments of this specification, a code task processing device is provided, including:
[0022] The first acquisition module is configured to acquire the data to be processed in the target code task;
[0023] The first input module is configured to input the data to be processed into the code task model to obtain a target code file corresponding to the target code task, where the code task model is trained based on a sample set corresponding to the target code task on the pre-trained model, and the pre-trained model is trained based on function data obtained by semantic analysis of the sample code file.
[0024] According to the sixth aspect of the embodiments of this specification, a model pre-training device is provided, configured in a cloud device, including:
[0025] A second acquisition module, configured to acquire a sample code set, where the sample code set includes sample code files and structure data of the sample code files;
[0026] A prediction module, configured to use an initial code model to perform semantic analysis on the sample code files based on the structure data, obtain function data, and perform prediction on the sample code files based on the function data to obtain predicted code files;
[0027] A training module, configured to train the initial code model based on the loss between the predicted code files and the sample code files to obtain a pre-trained model;
[0028] A sending module, configured to send the model parameters of the pre-trained model to the end-side device.
[0029] According to the seventh aspect of the embodiments of the present specification, there is provided a code completion device, including:
[0030] A receiving module, configured to receive the code to be completed sent by the front-end user;
[0031] A second input module, configured to input the code to be completed into a code completion model to obtain a target code file, where the code completion model is obtained by training the pre-trained model based on a sample set corresponding to the code completion task, and the pre-trained model is obtained by training based on function data obtained by performing semantic analysis on the sample code files;
[0032] A feedback module, configured to feedback the target code file to the front-end user.
[0033] According to the eighth aspect of the embodiments of the present specification, there is provided a computing device, including:
[0034] A memory and a processor;
[0035] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above method are implemented.
[0036] According to the ninth aspect of the embodiments of the present specification, there is provided a computer-readable storage medium, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above method are implemented.
[0037] According to the tenth aspect of the embodiments of the present specification, there is provided a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above method are implemented.
[0038] One embodiment of this specification realizes obtaining the data to be processed in a target code task; inputting the data to be processed into a code task model to obtain a target code file corresponding to the target code task, where the code task model is obtained by training a pre-trained model based on a sample set corresponding to the target code task, and the pre-trained model is obtained by training based on function data obtained by semantic analysis of sample code files. Training the pre-trained model with the function data obtained by semantic analysis of sample code files can, on the one hand, improve the model's learning ability and reduce the model's dependence on labeled data. On the other hand, it can improve the model's learning ability for function-level features, enabling the model to more precisely understand the structure and semantics of the code, thereby improving the model's ability to generate entire function code; training the code task model based on the sample set corresponding to the target code task can, on the basis of obtaining the pre-trained model through function-level training, train the model with the task sample set corresponding to the downstream task, enabling the model to have better execution effects for downstream tasks; furthermore, it can realize inputting the data to be processed into the code task model to obtain the target code file corresponding to the target code task, improving the generation efficiency and accuracy of the code task model for the target code file. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is an architecture diagram of a code task processing system provided by an embodiment of this specification;
[0040] Figure 2 is a flowchart of a code task processing method provided by an embodiment of this specification;
[0041] Figure 3 is an architecture diagram of a model pre-training system provided by an embodiment of this specification;
[0042] Figure 4 is a flowchart of a code completion method provided by an embodiment of this specification;
[0043] Figure 5 is a schematic diagram of the processing process of a code task processing method provided by an embodiment of this specification;
[0044] Figure 6 is a schematic diagram of post-processing of a code task processing method provided by an embodiment of this specification;
[0045] Figure 7 is a schematic structural diagram of a code task processing device provided by an embodiment of this specification;
[0046] Figure 8 is a schematic structural diagram of a model pre-training device provided by an embodiment of this specification;
[0047] Figure 9 It is a schematic structural diagram of a code completion device provided by an embodiment of this specification;
[0048] Figure 10 It is a structural block diagram of a computing device provided by an embodiment of this specification. Specific embodiments
[0049] In the following description, many specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific embodiments disclosed below.
[0050] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0051] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0052] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select to authorize or refuse.
[0053] First, the noun terms involved in one or more embodiments of this specification are explained.
[0054] AST (Abstract Syntax Tree): It is a data structure used by compilers or interpreters to understand computer programs. This tree-like data structure can follow the syntax rules of programming languages and depict the syntax structure of program code.
[0055] F1 value: It is a weighted average of the model's accuracy and recall. Its maximum value is 1 and its minimum value is 0. The larger the value, the better the model.
[0056] Currently, in the process of software development, in scenarios such as code generation, code Q&A, and code plugins, there is often a requirement for precise code generation, so it depends on a deep understanding of the code. However, traditional code generation techniques are usually based on preset rules or templates. This method often fails to meet the requirements when dealing with complex code generation tasks. With the development of deep learning technology, code generation techniques based on neural networks have begun to receive attention. However, in the process of training a model through code training techniques, it often requires a large amount of labeled data. This method usually requires a large amount of labeled data, and the acquisition cost of training samples is relatively high, resulting in the model's performance, quality, accuracy, and generalization ability often not meeting expectations.
[0057] Based on this, an embodiment of this specification realizes obtaining the data to be processed in the target code task; inputting the data to be processed into the code task model to obtain the target code file corresponding to the target code task, where the code task model is trained from a pre-trained model based on the sample set corresponding to the target code task, and the pre-trained model is trained from function data obtained by semantic analysis of the sample code file. Training the pre-trained model from the function data obtained by semantic analysis of the sample code file can, on the one hand, improve the model's learning ability and reduce the model's dependence on labeled data, and on the other hand, improve the model's learning ability for function-level features, enabling the model to more precisely understand the structure and semantics of the code, thereby improving the model's ability to generate the entire function code; training the code task model from the pre-trained model based on the sample set corresponding to the target code task can, on the basis of obtaining the pre-trained model through function-level training, train the model through the task sample set corresponding to the downstream task, enabling the model to have better performance in executing the downstream task; and then it can realize inputting the data to be processed into the code task model to obtain the target code file corresponding to the target code task, improving the generation efficiency and accuracy of the code task model for the target code file.
[0058] See Figure 1 , Figure 1 shows an architecture diagram of a code task processing system provided by an embodiment of this specification. The code task processing system may include a client 100 and a server 200;
[0059] The client 100 is used to send the data to be processed in the target code task to the server 200.
[0060] The server 200 is used to input the data to be processed into the code task model to obtain the target code file corresponding to the target code task. Among them, the code task model is trained from a pre-trained model based on a sample set corresponding to the target code task, and the pre-trained model is trained based on function data obtained by semantic analysis of the sample code file.
[0061] The client 100 is further used to receive the target code file corresponding to the target code task sent by the server 200.
[0062] By applying the code task processing system provided in the embodiments of this specification, the client can send the data to be processed in the target code task to the server. The server receives the data to be processed, can input the data to be processed into the code task model, obtain the target code file corresponding to the target code task, and send the target code file to the client. Among them, the code task model is trained from a pre-trained model based on a sample set corresponding to the target code task, and the pre-trained model is trained based on function data obtained by semantic analysis of the sample code file.
[0063] Since the pre-trained model is trained based on the function data obtained by semantic analysis of the sample code file, the pre-trained model can not only learn the code at the file level, but also learn the function data at the function level, thereby improving the model's learning ability for function-level features, enabling the model to more precisely understand the structure and semantics of the code, and further improving the model's ability to generate the entire function code. Since the code task model is trained from a pre-trained model based on a sample set corresponding to the target code task, the code task model can, on the basis of more accurately generating the entire function code, have better downstream task execution effects, improve the execution efficiency of the target code task, and improve the accuracy of the target code file. In this way, the server can more efficiently process the data to be processed in the target code task sent by the client based on the code task model and generate a more accurate target code file.
[0064] In practical applications, the code task processing system may include multiple clients 100 and a server 200. The multiple clients 100 can respectively establish communication connections with the server 200. In the code task processing scenario, the server 200 is used to obtain the data to be processed corresponding to each client 100 and provide code task processing services for each client. The multiple clients 100 can respectively act as senders or receivers and communicate through the server 200.
[0065] Among them, a connection is established between the client 100 and the server 200 through a network. The network provides a medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The data transmitted by the client 100 may need to be processed such as encoded, transcoded, compressed, etc. before being published to the server 200.
[0066] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5) application, or a light application (also known as a mini-program, a lightweight application) or a cloud application, etc. The client 100 can be based on the SDK (Software Development Kit) of the corresponding service provided by the server 200, such as developed based on the RTC (Real Time Communication) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device or certain APPs in the device to run, etc. The electronic device can, for example, have a display screen and support information browsing, etc., such as it can be a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. (device on the client side). Various other types of applications can usually be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0067] The server 200 can include servers that provide various services, such as a server that provides communication services for multiple clients, or a server for background training that supports the models used on the client, or a server that processes the data sent by the client, etc. It should be noted that the server 200 can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server (cloud-side device) of basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.
[0068] It should be noted that the code task processing method provided in the embodiments of this specification is generally executed by the server. However, in other embodiments of this specification, the client may also have a similar function as the server, so as to execute the code task processing method provided in the embodiments of this specification. In other embodiments, the code task processing method provided in the embodiments of this specification may also be jointly executed by the client and the server.
[0069] In this specification, a code task processing method is provided. This specification also relates to a model pre-training method applied to cloud-side devices, a code completion method, a code task processing device, a model pre-training device configured on cloud-side devices, a code completion device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.
[0070] See Figure 2 , Figure 2 shows a flowchart of a code task processing method provided according to an embodiment of this specification, which specifically includes the following steps.
[0071] Step 202: Obtain the data to be processed in the target code task.
[0072] In an optional embodiment of this specification, the data to be processed in the target code task can be actively obtained.
[0073] Specifically, the target code task can be understood as a task for processing the data to be processed. The data to be processed can be understood as the data used to generate the target code file. The data to be processed may have different physical meanings in different application scenarios.
[0074] Exemplarily, in the code question-and-answer scenario, the data to be processed can be the question text; in the code plug-in scenario, the data to be processed can be the source code data; in the code completion scenario, the data to be processed can be the code file to be completed; in the code generation scenario, the data to be processed can be the natural description language, which is used to describe the code to be generated, so as to provide the information required for generating the code. The language of the natural description language can be any one, such as Chinese, English, Japanese, etc.
[0075] In another optional embodiment of this specification, the target code task can be received, and the data to be processed in the target code task can be obtained from the target code task.
[0076] In practical applications, by obtaining the data to be processed in the target code task, the target code task can be processed based on the data to be processed, so as to obtain the task processing result of the target code task.
[0077] Step 204: Input the data to be processed into the code task model to obtain the target code file corresponding to the target code task. The code task model is trained from a pre-trained model based on a sample set corresponding to the target code task, and the pre-trained model is trained based on function data obtained by semantic analysis of sample code files.
[0078] In an alternative embodiment of this specification, the data to be processed can be input into the code task model to obtain the target code file corresponding to the target code task.
[0079] Specifically, the code task model can be understood as a pre-trained task processing model that can process tasks related to code. A task processing model obtained by training based on training sample sets corresponding to different target code tasks in different application scenarios can have multiple task processing functions and execute multiple specific tasks.
[0080] Exemplarily, the code task model can be applied to various application scenarios such as code generation, code completion, code question answering, and code plugins.
[0081] Specifically, the target code task can include a code generation task, a code completion task, a code question answering task, etc. The target code file can be understood as the task processing result of the code task model for the target code task. The target code file can include only at least one complete code function, or can also include other code and at least one complete code function.
[0082] Exemplarily, in the code generation scenario, the target code file can be a code file generated based on a natural description language. For example, the natural description language can be "write a bubble sort algorithm in C language", then the corresponding code file is the C language form code file of the bubble sort algorithm; in the code completion scenario, the target code file can be a complete code file generated for the code completion task, where the code completion task can include the completion of functions and the completion of files. In the function completion scenario, the target code can be function code, and in the file completion scenario, the target code can be function code and other code; in the code question answering scenario, the target code file can be the answer code file corresponding to the question text; in the code plugin scenario, the target code file can be new code generated based on the source code.
[0083] In another alternative embodiment of this specification, before inputting the data to be processed into the code task model to obtain the target code file corresponding to the target code task, the following steps can also be included:
[0084] Obtain a sample set and a pre-trained model, where the sample set includes at least one sample pair, and the sample pair includes the code sample question and the code sample answer of the target code task;
[0085] Extract sample pairs from the sample set, and input the code sample problems in the sample pairs into the pre-trained model to obtain code prediction answers;
[0086] Train the pre-trained model based on the loss between the code prediction answers and the corresponding code sample answers to obtain a code task model.
[0087] In practical applications, the code task model can be obtained by training the pre-trained model based on the sample set corresponding to the target code task.
[0088] Specifically, the sample set can be understood as the sample set corresponding to the target code task. The sample set can include at least one sample pair, and the sample pair can include the code sample problem and the code sample answer of the target code task. Among them, the code sample problem can be understood as the sample data input into the pre-trained model, and the code sample answer can be understood as the result expected to be output by the pre-trained model for the code sample problem. The code prediction answer can be understood as the actual result output by the pre-trained model for the code sample problem.
[0089] In practical applications, the sample set corresponding to the target code task and the pre-trained model that has completed training can be obtained. Extract sample pairs from the sample set, and input the code sample problems in the sample pairs into the pre-trained model to obtain the code prediction answers output by the pre-trained model for the code sample problems. Based on the code sample problems corresponding to the code prediction answers, obtain the corresponding code sample answers in the sample pairs where the code sample problems are located. The loss value can be calculated based on the code prediction answers output by the pre-trained model for the code sample problems and the code sample answers corresponding to the code sample problems in the sample pairs, that is, calculate the loss between the code prediction answers and the code sample answers corresponding to the same code sample problem. The loss between the code prediction answers and the code sample answers corresponding to the same code sample problem can be used as the model loss value of the pre-trained model. It is also possible to calculate the model loss value of the pre-trained model based on the loss values corresponding to multiple sample pairs respectively. On the basis of calculating the model loss value, the model parameters of the pre-trained model can be adjusted based on the model loss value until the preset conditions for completing training are met to obtain a code task model. Among them, the preset conditions for completing training can include: the loss between the code prediction answers output by the model and the code sample answers is lower than the preset loss threshold, or the recall rate, accuracy, etc. of the model reach the preset threshold, etc. Further, the calculation method of the model loss value can be determined according to the requirements in practical applications. For example, the loss value can be calculated through the cross-entropy loss function.
[0090] By training a pre-trained model according to a sample set corresponding to a target code task, a code task model capable of executing the target code task can be obtained, which can improve the model's execution ability for specific code tasks and the task execution effect for downstream tasks on the basis of reducing the dependence of model training on labeled data.
[0091] In an optional embodiment of the present specification, the pre-trained model can be pre-trained and directly obtained during the process of obtaining the sample set and the pre-trained model.
[0092] In another optional embodiment of the present specification, before obtaining the sample set and the pre-trained model, the following steps may further be included:
[0093] Obtain a sample code set, where the sample code set includes sample code files and structure data of the sample code files;
[0094] Use an initial code model to perform semantic analysis on the sample code files based on the structure data to obtain function data.
[0095] Specifically, the sample code set can be understood as a sample set for training an initial code model to obtain a pre-trained model. The sample code set can include sample code files and structure data of the sample code files. Among them, the structure data can be understood as data containing the structure information and semantic information of the sample code files, such as an abstract syntax tree. The initial code model can select a suitable deep learning model according to the requirements in actual applications. For example, it can be a transformer model. The function data can be understood as the function code included in the sample code files. The function data can contain the code information at the complete function level of a function code, such as the name, parameter list, return type, and function body, position, etc. of the function. These information can constitute a complete function semantic unit of a function.
[0096] In actual applications, a sample code set can be obtained, and an initial code model can be used to perform semantic analysis on the sample code files based on the structure data of the sample code files to obtain function data.
[0097] By performing semantic analysis on the sample code files based on the structure data of the sample code files, a complete function semantic unit can be extracted from the structure data, so that the model can more precisely understand the structure and semantics of the code, improve the model's learning ability for function-level features, and thus improve the model's generation ability for the entire function and the quality and accuracy of code generation.
[0098] In an optional embodiment of the present specification, after using the initial code model to perform semantic analysis on the sample code files based on the structure data to obtain function data, the following steps may further be included:
[0099] Predict the sample code file based on the function data to obtain a predicted code file;
[0100] Train the initial code model based on the loss between the predicted code file and the sample code file to obtain a pre-trained model.
[0101] Specifically, the predicted code file can be understood as the prediction result output by the initial code model for the function data in the sample code file. The predicted code file can be composed of predicted fields output by the initial code model.
[0102] In practical applications, the sample code file can be predicted based on the function data to obtain a predicted code file, and then the initial code model can be trained based on the loss between the predicted code file and the sample code file to obtain a pre-trained model. By predicting the sample code file based on the function data, the pre-trained model can learn more function-level features, so that the pre-trained model has the ability to generate the entire function, improving the accuracy and efficiency of the pre-trained model in generating function code in the code file.
[0103] In an optional embodiment of this specification, obtaining a sample code set may include the following steps:
[0104] Obtain a sample code file;
[0105] Call a parser to perform a structural analysis on the sample code file to obtain the structural data of the sample code file;
[0106] Construct a sample code set based on the sample code file and the structural data.
[0107] Specifically, the parser can be used to parse the syntax rules and semantic information of the sample code file, and specifically can be an AST parser. The structural data of the sample code file can specifically be the abstract syntax tree corresponding to the sample code file.
[0108] Optionally, obtaining the sample code file can be to obtain it from a code library or receive a code file uploaded by a user.
[0109] Optionally, calling a parser to perform a structural analysis on the sample code file to obtain the structural data of the sample code file, the language corresponding to the sample code file can be identified first, and the parser corresponding to the language can be determined, and then the parser can be called externally through an interface to parse the sample code file into an abstract syntax tree, and the sample code file and the corresponding abstract syntax tree are input into the initial code model, so that the initial code model analyzes and obtains the function data in the sample code file based on the abstract syntax tree.
[0110] Optionally, the parser can also be deployed in the initial code model. Then, the parser is called to perform structural parsing on the sample code file to obtain the structural data of the sample code file. It can first identify the language of the sample code file and determine the corresponding parser for that language, and then directly call the corresponding parser in the initial code model. The structural data corresponding to the sample code file is obtained by parsing the sample code file through the parser.
[0111] By calling the parser to perform structural parsing on the sample code file, the structural data of the sample code file can be obtained through parsing by the parser, enabling the model to analyze the syntax rules and semantic information of the sample code file based on the structural data, facilitating the extraction of function data from the sample code file according to the structural data, and enabling the model to more precisely understand the structure and semantics of the code, improving the quality and accuracy of code generation.
[0112] In an optional embodiment of this specification, the initial code model can include a preprocessing unit and a code generation unit; the function data can include function semantic information and function location information. Using the initial code model, semantic analysis is performed on the sample code file based on the structural data to obtain function data, and prediction is performed on the sample code file based on the function data to obtain a predicted code file. The steps can include the following:
[0113] Input the sample code file and the structural data into the initial code model. Use the preprocessing unit to traverse the structural data to determine the function semantic information of each function, and based on the function semantic information of each function, determine the function location information of each function in the sample code file;
[0114] Use the code generation unit to perform code prediction on the sample code file based on the function semantic information and function location information of each function to obtain a predicted code file.
[0115] Specifically, the preprocessing unit can be understood as a unit for processing the input data of the model, which may include special symbol processing and AST rule analysis. Among them, special symbol processing can be understood as preprocessing the input data with special symbols so that the model can more accurately process the preprocessed input data subsequently. Special symbol processing may include: inserting special symbols (tokens) such as start symbols, end symbols, delimiters, and placeholders into the input data. It should be noted that the special symbol processing processes corresponding to different types of models may be different. Multiple AST parsers can be integrated in the preprocessing unit, or the AST parser can be called externally through an interface to obtain the structure data of the code sample file. The code generation unit can be understood as a unit that performs code prediction based on the semantic information of functions and the function position information to generate a predicted code file. The output of the code generation unit is consistent with the input information. If the input is a sample code file, the output predicted code file is a code file; if the input is a sample function, the output predicted code file is function data.
[0116] In practical applications, when the sample code file and the structure data are input into the initial code model, the preprocessing unit can traverse the structure data to determine the function semantic information of each function in the sample code file. Optionally, traversing the structure data can be traversing each node in the abstract syntax tree to determine the node of the function definition, so as to extract the function semantic information based on the node of the function definition, and record the function position information of each function in the sample code file during the process of extracting the function semantic information. On this basis, the code generation unit can be used to perform code prediction on the sample code file based on the function semantic information and function position information of each function to obtain a predicted code file.
[0117] In an optional embodiment of this specification, performing code prediction on the sample code file based on the function semantic information and function position information of each function to obtain a predicted code file may include: extracting complete function data from the sample code file based on the function position information; generating function fields in the predicted code file based on the function semantic information and the complete function data; and obtaining the predicted code file according to the predicted function fields.
[0118] In another optional embodiment of this specification, performing code prediction on the sample code file based on the function semantic information and function position information of each function to obtain a predicted code file may further include: generating other code fields in the predicted code file based on other code feature information in the sample code file; and obtaining the predicted code file according to the function fields and the other code fields.
[0119] By using the preprocessing unit to traverse the structural data and determine the function semantic information and function location information of each function, the model can extract complete function semantic units from the sample code file, thereby more precisely understanding the structure and semantics of the code and improving the quality and accuracy of code generation. By using the code generation unit to perform code prediction on the sample code file based on the function semantic information and function location information of each function and obtain a predicted code file, the model can learn function-level features, thereby improving the model's ability to generate entire functions.
[0120] In another optional embodiment of this specification, the initial code model may further include a postprocessing unit. Through the postprocessing unit, truncation rule judgment and format adjustment can be performed on the predicted code file, thereby obtaining a code file that better meets the requirements. The postprocessing unit can be integrated into the code generation unit or can be an independent processing unit for receiving the predicted code file output by the code generation unit. Exemplarily, during the training process of the code completion task at the document level, the postprocessing unit can perform truncation processing on the code file or perform splicing processing on different code files. During the training process of the code completion task at the function level, the postprocessing unit can perform truncation processing on the function but will not perform splicing processing on different functions.
[0121] In an optional embodiment of this specification, the sample code set further includes sample functions; the code task processing method may further include the following steps:
[0122] Use the initial code model to perform code prediction on the sample function to obtain a predicted function;
[0123] Based on the loss between the predicted function and the sample function, train the initial code model to obtain a pre-trained model.
[0124] In practical applications, the sample code set used to train the initial code model may include sample code files and sample functions. Among them, the sample functions can be extracted from the sample code files by the initial code model analyzing the structural data of the sample code files, can also be directly obtained from the code library, or obtained by receiving function data uploaded by users, etc.
[0125] The initial code model can perform code prediction on the sample code file to obtain a predicted code file, or can perform code prediction on the sample function to obtain a predicted function.
[0126] It should be noted that during the process of predicting the sample code file, the initial code model can truncate the code file according to the number of fields in the code file. This processing method may cause functions to be truncated, resulting in the model being unable to learn complete function information. However, during the process of predicting the sample function, the initial code model will not truncate the sample function and can learn complete function-level information.
[0127] In the actual implementation process, the initial code model can be used to predict the code of the sample function to obtain the predicted function. Based on the loss between the predicted function and the sample function, the initial code model is trained to obtain the pre-trained model. In this way, the learning ability of the model for function data can be improved, enabling the model to pay more attention to function-level code generation and further improving the model's ability to generate the entire function. At the same time, by mixing the sample function and the sample code file for training, the training efficiency of the model can be improved, and the generalization ability of the model can be improved simultaneously.
[0128] In an optional embodiment of this specification, obtaining the sample code set may further include the following steps:
[0129] Obtain the sample function from the code library;
[0130] Weight the sample code file and the sample function based on the target weights of the sample code file and the sample function, where the target weights are determined based on the meta-information of the code library;
[0131] Construct a sample code set according to the weighted sample code file and the weighted sample function.
[0132] Specifically, the code library can be a code database in an open-source code hosting platform, or a code database within an enterprise, or a code database obtained from a specified third-party channel, etc. The code library can include at least one sample code file and can also include at least one sample function. The target weights are used to characterize the quality and importance of the sample code file or the sample function. The meta-information of the code library can reflect the quality and importance of the code library, and the meta-information can include the number of stars, the number of references, the number of contributors, the last update time, etc. The weighted sample code file can be understood as a weighted sample code file with weights based on the target weights; the weighted sample function can be understood as a weighted sample function with weights based on the target weights.
[0133] Optionally, through the meta-information of the code library, the weight information of the code library can be calculated, and thus based on the weight information of the code library, the target weights of each sample code file and each sample function in the code library can be obtained. Further, a weight function can be designed in advance, taking the meta-information of the code library as the input of the function, and thus outputting the weight information of the code library.
[0134] By obtaining sample code files and sample functions from a code library and weighting the sample code files and sample functions based on the target weights of the sample code files and sample functions, weighted sample code files and sample functions can be obtained, so that each sample code file and sample function in the sample code set carry target weights that can characterize quality and importance, enabling the model to pay more attention to high-quality code during the training process, thereby improving the generalization ability of the model.
[0135] In an alternative embodiment of this specification, the sample code set further includes sample functions. After training a pre-trained model based on the loss between the code prediction answer and the corresponding code sample answer to obtain a code task model, the following steps may further be included:
[0136] Obtain a validation set, where the validation set includes multiple validation groups, and each validation group includes a validation code question and a validation code answer for a target code task;
[0137] Extract a validation group from the validation set and input the validation code question in the validation group into the code task model to obtain a predicted code answer;
[0138] Determine the validation result of the code task model based on the validation code answer and the predicted code answer;
[0139] In the case where the validation result does not meet the preset metrics, adjust the proportion of the sample code files and sample functions in the sample code set, and return to execute the step of obtaining the sample code set.
[0140] Specifically, the validation set can be understood as the test set corresponding to the model's execution of downstream tasks. The validation set can include multiple validation groups, and each validation group can include a validation code question and a validation code answer for a target code task. Among them, the validation code question can be understood as the input data corresponding to the downstream task, that is, the target code task, and the validation code answer can be understood as the task execution result expected to be output by the model. The predicted code answer can be understood as the prediction result actually output by the model for the validation code question. The code task model can be understood as the model obtained by fine-tuning based on the downstream task. The validation result can reflect the execution ability of the code task model for downstream tasks. The preset metrics are used to define the numerical values of various metrics for the validation result that meets the expectations corresponding to the target code task.
[0141] Optionally, to determine the validation result of the code task model based on the validation code answer and the predicted code answer, the accuracy rate, recall rate, F1 value, and other metric values of the code task model can be calculated through the validation code answer and the predicted code answer.
[0142] In the actual implementation process, when the verification result does not meet the preset indicators, it indicates that the code task model has not yet achieved the expected task execution effect. The proportion of sample code files and sample functions in the sample code set can be adjusted, and the step of obtaining the sample code set can be returned to further train the pre-trained model, thereby improving the task execution effect of the code task model in executing downstream tasks. Specifically, sample functions related to the preset indicators can be increased, and sample code files less related to the preset indicators can be reduced.
[0143] Exemplarily, when the accuracy of the functions generated by the code task model is low, the proportion of sample functions in the sample code set can be increased while the proportion of sample code files is reduced, so that the pre-trained model pays more attention to the learning of function-level features and improves the accuracy of the functions generated by the model.
[0144] By obtaining a validation set, extracting validation groups from the validation set, and inputting the validation code problems in the validation groups into the code task model to obtain predicted code answers, the value of each evaluation metric can be calculated according to the performance of the model on the validation set at the end of each training cycle. According to the validation code answers and the predicted code answers, the verification result of the code task model is determined. When the verification result does not meet the preset indicators, the proportion of sample code files and sample functions in the sample code set is adjusted, and the step of obtaining the sample code set is returned, which can enable the model to better adapt to different downstream tasks and improve the performance of the model.
[0145] In an optional embodiment of this specification, after inputting the data to be processed into the code task model and obtaining the target code file corresponding to the target code task, the following steps may further be included:
[0146] Send the target code file to the front-end user;
[0147] Receive the feedback information sent by the front-end user based on the target code file;
[0148] Obtain an updated sample set based on the feedback information, and use the updated sample set to adjust the model parameters of the code task model.
[0149] Specifically, the feedback information can reflect whether the target code file is accurate or whether the quality meets the user's expectations. The feedback information may carry a standard code file uploaded by the user, and the standard code file may be a code file that meets the user's expectations. The updated sample set may include sample data re-obtained by the server based on the feedback information, or may include the standard code file carried in the feedback information.
[0150] In practical applications, based on the data to be processed in the target code task, the server outputs the corresponding target code file through the code task model and sends the target code file to the front-end user. On this basis, the server can receive the feedback information sent by the front-end user based on the target code file. In the case where the feedback information indicates that the target code file is inaccurate or the quality does not meet the user's expectations, the server can re-collect the sample data to construct an updated sample set and use the updated sample set to adjust the model parameters of the code task model. Thus, according to the feedback information of the front-end user, the iteration and optimization of the model are completed, and the quality and accuracy of the code generated by the model are improved.
[0151] In another optional embodiment of this specification, after inputting the data to be processed into the code task model to obtain the target code file corresponding to the target code task, the following steps may further be included:
[0152] Send the target code file to the front-end user;
[0153] Receive the reply feedback data sent by the front-end user for the target code file;
[0154] Generate a question-and-answer guiding message based on the reply feedback data, where the question-and-answer guiding message is used to guide the front-end user to input the target interaction data;
[0155] Use the target interaction data as the data to be processed, and return to execute the step of inputting the data to be processed into the code task model to obtain the target code file corresponding to the target code task.
[0156] Specifically, the reply feedback data can reflect whether the target code file is accurate or whether the quality meets the user's expectations. The question-and-answer guiding message can be used to guide the front-end user to send more precise task processing data, such as a more accurate problem description statement or a more refined requirement. The target interaction data can be understood as the more precise task processing data sent by the front-end user based on the question-and-answer guiding message.
[0157] In practical applications, based on the data to be processed in the target code task, the server outputs the corresponding target code file through the code task model and, on the basis of sending the target code file to the front-end user, can receive the reply feedback information sent by the front-end user based on the target code file. In the case where the reply feedback information indicates that the target code file is inaccurate or the quality does not meet the user's expectations, the server can generate a question-and-answer guiding message based on the reply feedback data, thereby guiding the user to input a more accurate question text or a more refined requirement description information, and based on the target interaction data input by the front-end user, use the target interaction data as the data to be processed and return to execute the step of inputting the data to be processed into the code task model to obtain the target code file corresponding to the target code task until a target code file that meets the expectations of the front-end user is obtained.
[0158] By generating question-and-answer guiding information based on the reply feedback data, it is possible to guide the user to input more accurate task data, thereby improving the accuracy and quality of the code generated by the model. By receiving the target interaction data input by the user and using the target interaction data as the data to be processed and inputting it into the code task model, it is possible to guide the user to input task data that better meets the model's requirements through information interaction with the user, improving the user experience.
[0159] An embodiment of this specification realizes obtaining the data to be processed in the target code task; inputting the data to be processed into the code task model to obtain the target code file corresponding to the target code task, where the code task model is trained based on a sample set corresponding to the target code task for a pre-trained model, and the pre-trained model is trained based on function data obtained by semantic analysis of the sample code file. Training the pre-trained model based on the function data obtained by semantic analysis of the sample code file can, on the one hand, improve the model's learning ability and reduce the model's dependence on labeled data. On the other hand, it can improve the model's learning ability for function-level features, enabling the model to more precisely understand the structure and semantics of the code, thereby improving the model's ability to generate entire function code; training the code task model based on a sample set corresponding to the target code task for the pre-trained model can, on the basis of obtaining the pre-trained model through function-level training, train the model through the task sample set corresponding to the downstream task, enabling the model to have better execution effects for downstream tasks; and then it can be realized to input the data to be processed into the code task model to obtain the target code file corresponding to the target code task, improving the generation efficiency and accuracy of the code task model for the target code file.
[0160] See Figure 3 , Figure 3 shows an architecture diagram of a model pre-training system provided by an embodiment of this specification. As Figure 3 shown, the model pre-training system includes a cloud-side device 302 and an edge-side device 304.
[0161] The edge-side device 304 is used to send the sample code set to the cloud-side device 302.
[0162] The cloud-side device 302 is used to obtain the sample code set, where the sample code set includes sample code files and the structure data of the sample code files; use the initial code model to perform semantic analysis on the sample code files based on the structure data to obtain function data, and perform prediction on the sample code files based on the function data to obtain predicted code files; train the initial code model based on the loss between the predicted code files and the sample code files to obtain a pre-trained model; send the model parameters of the pre-trained model to the edge-side device 304.
[0163] The edge device 304 is further configured to receive the model parameters sent by the cloud device 302.
[0164] In practical applications, after the cloud device sends the model parameters of the pre-trained model to the edge device, the edge device can construct the pre-trained model locally according to the model parameters of the pre-trained model, and further use the pre-trained model to fine-tune the downstream task, so as to obtain the code task model.
[0165] Applying the solution of the embodiments of this specification, by sending the sample code set to the cloud device through the edge device, the cloud device can train the initial code model based on the sample code set to obtain the pre-trained model, and send the model parameters of the pre-trained model to the edge device. The edge device can receive the model parameters sent by the cloud device and construct the corresponding pre-trained model locally. Since the sample code set includes the sample code file and the structure data of the sample code file, the initial code model can perform semantic analysis on the sample code file based on the structure data to obtain function data, and predict the sample code file based on the function data to obtain the predicted code file, thereby improving the model's learning ability for function-level features, enabling the model to more precisely understand the structure and semantics of the code, and improving the model's ability to generate the entire function code. In this way, the pre-trained model constructed locally by the edge device has a more accurate ability to generate the entire function code. Furthermore, the code task model obtained by fine-tuning based on the downstream task can generate the entire function code faster and better, and has a better execution effect on specific code tasks.
[0166] See Figure 4 , Figure 4 which shows the flowchart of a code completion method provided by an embodiment of this specification, specifically including the following steps.
[0167] Step 402: Receive the code to be completed sent by the front-end user.
[0168] Step 404: Input the code to be completed into the code completion model to obtain the target code file, where the code completion model is trained based on the sample set corresponding to the code completion task, and the pre-trained model is trained based on the function data obtained by performing semantic analysis on the sample code file.
[0169] Step 406: Feed back the target code file to the front-end user.
[0170] It should be noted that the specific implementation manners of steps 402 - 406 are the same as those of steps 202 - 204 above, and the embodiments of this specification will not elaborate further.
[0171] Applying the solution of the embodiments of this specification, by receiving the code to be completed sent by the front-end user, the code to be completed can be input into the code completion model to obtain the target code file, and the target code file can be fed back to the front-end user. Since the code completion model is trained from a pre-trained model based on a sample set corresponding to the code completion task, and the pre-trained model is trained from function data obtained by semantic analysis of sample code files, the pre-trained model has the ability to learn function-level features, can more precisely understand the structure and semantics of the code, and thus has the ability to generate an entire segment of function code; on this basis, the code completion model trained from the pre-trained model based on the sample set corresponding to the code completion task not only has a better execution ability for the code completion task, but also has a faster and more accurate generation ability for an entire segment of function in the code file. In this way, the front-end user can quickly obtain a more accurate target code file and a better user experience.
[0172] The following Figure 5 , taking the application of the code task processing method provided in this specification in the code completion task as an example, further illustrate the code task processing method. Among them, Figure 5 shows a schematic diagram of the processing process of a code task processing method provided by an embodiment of this specification.
[0173] Through code analysis technology, the compiler front-end receives the source code corresponding to the code completion task, performs syntax analysis and semantic analysis on the source code, and inputs the obtained abstract syntax tree into the AI model. The AI model receives the model input, performs special symbol processing and AST rule analysis through the pre-processing unit, outputs the pre-processing result and inputs it into the code model. Among them, the source code can be a function to be completed or a code file to be completed.
[0174] It should be noted that the code model is a code completion model trained from a pre-trained model based on a sample set corresponding to the code completion task. The pre-trained model is a model trained through a mixed training method at the file level and the function level based on weighted sample code files and weighted sample functions in the sample code set.
[0175] The code model outputs the predicted code data corresponding to the source code, and inputs the predicted code data into the post-processing unit. The post-processing unit processes the predicted code data according to the truncation rule, adjusts the format of the processed code data, and then sends the code data to the compiler back-end.
[0176] See Figure 6 , Figure 6The figure shows a post - processing schematic diagram of a code task processing method provided by an embodiment of this specification. Among them, in the case where the code completion task is at the document level, taking the code file as a unit, truncation of the code file or splicing of different code files can occur, and this task can be used for code completion at any position of the cursor. In the case where the code completion task is at the function level, taking the function as a unit, truncation of the function can occur but splicing of different functions will not occur.
[0177] The compiler backend can execute steps such as intermediate representation, code optimization, and generation of target code based on the code data output by the AI model. The obtained target code can be directly run in the virtual machine or can be run in other operating environments through steps such as linking and machine code.
[0178] Applying the solution of the embodiment of this specification, a pre - trained model is trained through the function data obtained by semantic analysis of the sample code file. On the one hand, it can improve the model's learning ability and reduce the model's dependence on labeled data. On the other hand, it can improve the model's learning ability for function - level features, enabling the model to more precisely understand the structure and semantics of the code, thereby improving the model's ability to generate the entire function code; by training the pre - trained model with a sample set corresponding to the target code task to obtain a code task model, on the basis of obtaining the pre - trained model through function - level training, the model can be trained with a task sample set corresponding to the downstream task, enabling the model to have a better execution effect for the downstream task; furthermore, it can realize inputting the data to be processed into the code task model to obtain the target code file corresponding to the target code task, improving the generation efficiency and accuracy of the code task model for the target code file.
[0179] Corresponding to the above - mentioned method embodiment, this specification also provides an embodiment of a code task processing device. Figure 7 The figure shows a structural schematic diagram of a code task processing device provided by an embodiment of this specification. As Figure 7 shown, the device includes:
[0180] The first acquisition module 702: configured to acquire the data to be processed in the target code task.
[0181] The first input module 704: configured to input the data to be processed into the code task model to obtain the target code file corresponding to the target code task, where the code task model is trained based on a sample set corresponding to the target code task for the pre - trained model, and the pre - trained model is trained based on the function data obtained by semantic analysis of the sample code file.
[0182] Optionally, the code task processing device further includes a pre - training module, configured to:
[0183] Obtain a sample set and a pre-trained model, where the sample set includes at least one sample pair, and the sample pair includes a code sample question and a code sample answer for a target code task;
[0184] Extract sample pairs from the sample set, and input the code sample questions in the sample pairs into the pre-trained model to obtain code prediction answers;
[0185] Train the pre-trained model based on the loss between the code prediction answers and the corresponding code sample answers to obtain a code task model.
[0186] Optionally, the pre-training module is further configured to:
[0187] Obtain a sample code set, where the sample code set includes sample code files and structural data of the sample code files;
[0188] Use the initial code model to perform semantic analysis on the sample code files based on the structural data to obtain function data, and perform prediction on the sample code files based on the function data to obtain predicted code files;
[0189] Train the initial code model based on the loss between the predicted code files and the sample code files to obtain a pre-trained model.
[0190] Optionally, the pre-training module is further configured to:
[0191] Obtain sample code files;
[0192] Call a parser to perform structural parsing on the sample code files to obtain structural data of the sample code files;
[0193] Construct a sample code set based on the sample code files and the structural data.
[0194] Optionally, the initial code model includes a pre-processing unit and a code generation unit; the function data includes function semantic information and function location information;
[0195] Optionally, the pre-training module is further configured to:
[0196] Input the sample code files and the structural data into the initial code model, use the pre-processing unit to traverse the structural data to determine the function semantic information of each function, and determine the function location information of each function in the sample code files according to the function semantic information of each function;
[0197] Use the code generation unit to perform code prediction on the sample code files based on the function semantic information and function location information of each function to obtain predicted code files.
[0198] Optionally, the sample code set further includes sample functions; the pre-training module is further configured to:
[0199] Use the initial code model to perform code prediction on the sample functions to obtain predicted functions;
[0200] Train the initial code model based on the loss between the predicted functions and the sample functions to obtain a pre-trained model.
[0201] Optionally, the pre-training module is further configured to:
[0202] Obtain sample code files and sample functions from the code library;
[0203] Weight the sample code files and sample functions based on the target weights of the sample code files and sample functions, where the target weights are determined based on the meta-information of the code library;
[0204] Construct a sample code set according to the weighted sample code files and sample functions.
[0205] Optionally, the pre-training module is further configured to:
[0206] Obtain a validation set, where the validation set includes multiple validation groups, and each validation group includes a validation code question and a validation code answer for the target code task;
[0207] Extract the validation groups from the validation set and input the validation code questions in the validation groups into the code task model to obtain predicted code answers;
[0208] Determine the validation result of the code task model according to the validation code answers and the predicted code answers;
[0209] In the case where the validation result does not meet the preset metrics, adjust the proportion of the sample code files and sample functions in the sample code set, and return to execute the step of obtaining the sample code set.
[0210] Optionally, the code task processing device further includes an adjustment module, which is configured to:
[0211] Send the target code file to the front-end user;
[0212] Receive feedback information sent by the front-end user based on the target code file;
[0213] Obtain an updated sample set based on the feedback information, and use the updated sample set to adjust the model parameters of the code task model.
[0214] Optionally, the code task processing device further includes a generation module, which is configured to:
[0215] Send the target code file to the front-end user;
[0216] Receive the reply feedback data sent by the front-end user for the target code file;
[0217] Generate Q&A guiding information based on the reply feedback data, where the Q&A guiding information is used to guide the front-end user to input target interaction data;
[0218] Take the target interaction data as the data to be processed, and return to execute the step of inputting the data to be processed into the code task model to obtain the target code file corresponding to the target code task.
[0219] Applying the solution of the embodiment of this specification, the data to be processed in the target code task is obtained through the first acquisition module, and the data to be processed is input into the code task model through the first input module to obtain the target code file corresponding to the target code task. The pre-trained model is trained with the function data obtained by semantic analysis of the sample code file. On the one hand, it can improve the model learning ability and reduce the dependence of the model on labeled data. On the other hand, it can improve the model's learning ability for function-level features, enabling the model to more precisely understand the structure and semantics of the code, thereby improving the model's generation ability for the entire function code; training the pre-trained model with the sample set corresponding to the target code task to obtain the code task model can, on the basis of obtaining the pre-trained model through function-level training, train the model with the task sample set corresponding to the downstream task, enabling the model to have a better execution effect for the downstream task; and then it can be realized to input the data to be processed into the code task model to obtain the target code file corresponding to the target code task, improving the generation efficiency and accuracy of the code task model for the target code file.
[0220] The above is a schematic solution of a code task processing device in this embodiment. It should be noted that the technical solution of this code task processing device and the technical solution of the above code task processing method belong to the same concept. For the details not described in the technical solution of this code task processing device, reference can be made to the description of the technical solution of the above code task processing method.
[0221] Corresponding to the above method embodiment, this specification also provides an embodiment of a model pre-training device configured on a cloud-side device, Figure 8 showing a structural schematic diagram of a model pre-training device provided by an embodiment of this specification. As Figure 8 shown, the device includes:
[0222] The second acquisition module 802: configured to acquire a sample code set, where the sample code set includes sample code files and the structure data of the sample code files.
[0223] Prediction module 804: Configured to perform semantic analysis on a sample code file based on structural data using an initial code model, obtain function data, and perform prediction on the sample code file based on the function data to obtain a predicted code file.
[0224] Training module 806: Configured to train the initial code model based on the loss between the predicted code file and the sample code file to obtain a pre-trained model.
[0225] Sending module 808: Configured to send the model parameters of the pre-trained model to the edge device.
[0226] Applying the solution of the embodiments of this specification, the edge device can send a sample code set to the cloud device. The cloud device can train the initial code model based on the sample code set to obtain a pre-trained model, and send the model parameters of the pre-trained model to the edge device. The edge device can receive the model parameters sent by the cloud device and build a corresponding pre-trained model locally. Since the sample code set includes sample code files and the structural data of the sample code files, the initial code model can perform semantic analysis on the sample code files based on the structural data to obtain function data, and perform prediction on the sample code files based on the function data to obtain predicted code files, thereby improving the model's learning ability for function-level features, enabling the model to more precisely understand the structure and semantics of the code, and improving the model's ability to generate entire function code segments. In this way, the pre-trained model built locally by the edge device has a more accurate ability to generate entire function code segments. Furthermore, the code task model obtained by fine-tuning based on downstream tasks can generate entire function code segments faster and better, and has a better execution effect for specific code tasks.
[0227] The above is a schematic solution of a model pre-training device configured on a cloud device in this embodiment. It should be noted that the technical solution of this model pre-training device and the technical solution of the above model pre-training method belong to the same concept. For the details not described in the technical solution of this model pre-training device, reference can be made to the description of the technical solution of the above model pre-training method.
[0228] Corresponding to the above method embodiments, this specification also provides embodiments of a code completion device. Figure 9 The structural schematic diagram of a code completion device provided by an embodiment of this specification is shown. As Figure 9 shown, this device includes:
[0229] Receiving module 902: Configured to receive the code to be completed sent by the front-end user.
[0230] Second input module 904: Configured to input the code to be completed into the code completion model to obtain a target code file, where the code completion model is obtained by training a pre-trained model based on a sample set corresponding to the code completion task, and the pre-trained model is obtained by training based on function data obtained by semantic analysis of the sample code file.
[0231] Feedback module 906: Configured to feedback the target code file to the front-end user.
[0232] Applying the solution of the embodiment of this specification, by receiving the code to be completed sent by the front-end user, the code to be completed can be input into the code completion model to obtain a target code file, and the target code file can be fed back to the front-end user. Since the code completion model is obtained by training a pre-trained model based on a sample set corresponding to the code completion task, and the pre-trained model is obtained by training based on function data obtained by semantic analysis of the sample code file, therefore, the pre-trained model has the ability to learn function-level features, can more finely understand the structure and semantics of the code, and thus has the ability to generate the entire function code; on this basis, the code completion model obtained by training the pre-trained model based on the sample set corresponding to the code completion task not only has a better execution ability for the code completion task, but also has a faster and more accurate generation ability for the entire function in the code file. In this way, the front-end user can quickly obtain a more accurate target code file and obtain a better user experience.
[0233] The above is a schematic solution of a code completion device according to an embodiment of this specification. It should be noted that the technical solution of this code completion device and the technical solution of the above code completion method belong to the same concept. For the details not described in the technical solution of this code completion device, reference can be made to the description of the technical solution of the above code completion method.
[0234] Figure 10 The structural block diagram of a computing device provided according to an embodiment of this specification is shown. The components of the computing device 1000 include but are not limited to a memory 1010 and a processor 1020. The processor 1020 is connected to the memory 1010 through a bus 1030, and the database 1050 is used to store data.
[0235] The computing device 1000 also includes an access device 1040, which enables the computing device 1000 to communicate via one or more networks 1060. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1040 may include one or more of any type of wired or wireless network interfaces (e.g., network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC).
[0236] In one embodiment of the present specification, the above components of the computing device 1000 and Figure 10 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 10 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0237] The computing device 1000 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1000 can also be a mobile or stationary server.
[0238] Wherein, the processor 1020 is configured to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above method are implemented.
[0239] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the above method.
[0240] An embodiment of this specification also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the above method are implemented.
[0241] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above method.
[0242] An embodiment of this specification also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0243] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above method belong to the same concept. For the details not described in detail in the technical solution of the computer program, reference can be made to the description of the technical solution of the above method.
[0244] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0245] The computer instructions include computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0246] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0247] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0248] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A method for processing code tasks, comprising: Obtaining the data to be processed in the target code task; Inputting the data to be processed into a code task model to obtain a target code file corresponding to the target code task, wherein the code task model is obtained by training a pre-trained model based on a sample set corresponding to the target code task, and the pre-trained model is obtained by training based on function data obtained by semantic analysis of sample code files.
2. The method according to claim 1, before inputting the data to be processed into the code task model to obtain a target code file corresponding to the target code task, further comprising: Obtaining the sample set and the pre-trained model, wherein the sample set includes at least one sample pair, and the sample pair includes a code sample problem and a code sample answer of the target code task; Extracting the sample pair from the sample set, and inputting the code sample problem in the sample pair into the pre-trained model to obtain a code prediction answer; Training the pre-trained model based on the loss between the code prediction answer and the corresponding code sample answer to obtain the code task model.
3. The method according to claim 2, before obtaining the sample set and the pre-trained model, further comprising: Obtaining a sample code set, wherein the sample code set includes the sample code file and the structural data of the sample code file; Using an initial code model to perform semantic analysis on the sample code file based on the structural data to obtain the function data.
4. The method according to claim 3, after using the initial code model to perform semantic analysis on the sample code file based on the structural data to obtain the function data, further comprising: Performing prediction on the sample code file based on the function data to obtain a predicted code file; Training the initial code model based on the loss between the predicted code file and the sample code file to obtain a pre-trained model.
5. The method according to claim 3, the obtaining of the sample code set includes: Obtaining the sample code file; Invoking a parser to perform structural parsing on the sample code file to obtain the structural data of the sample code file; Constructing the sample code set based on the sample code file and the structural data.
6. The method according to any one of claims 3-5, the initial code model includes a pre-processing unit and a code generation unit; the function data includes function semantic information and function location information; The using of the initial code model to perform semantic analysis on the sample code file based on the structural data to obtain function data, and performing prediction on the sample code file based on the function data to obtain a predicted code file includes: Inputting the sample code file and the structural data into the initial code model, using the pre-processing unit to traverse the structural data, determining the function semantic information of each function, and determining the function location information of each function in the sample code file according to the function semantic information of each function; Using the code generation unit, based on the function semantic information and function position information of each function, perform code prediction on the sample code file to obtain the predicted code file.
7. The method according to any one of claims 3-5, wherein the sample code set further includes sample functions; the method further includes: Using the initial code model, perform code prediction on the sample functions to obtain predicted functions; Based on the loss between the predicted functions and the sample functions, train the initial code model to obtain the pre-trained model.
8. The method according to claim 7, wherein obtaining the sample code set includes: Obtain the sample functions from the code library; Based on the target weights of the sample code file and the sample functions, weight the sample code file and the sample functions respectively, wherein the target weights are determined based on the meta-information of the code library; Construct a sample code set according to the weighted sample code file and the weighted sample functions.
9. The method according to claim 3, wherein the sample code set further includes sample functions; after training the pre-trained model based on the loss between the code prediction answer and the corresponding code sample answer to obtain the code task model, the method further includes: Obtain a validation set, wherein the validation set includes a plurality of validation groups, and each validation group includes a validation code question and a validation code answer of the target code task; Extract a validation group from the validation set, and input the validation code question in the validation group into the code task model to obtain a predicted code answer; Determine the validation result of the code task model according to the validation code answer and the predicted code answer; In the case that the validation result does not meet the preset index, adjust the proportion of the sample code file and the sample functions in the sample code set, and return to execute the step of obtaining the sample code set.
10. The method according to claim 1, after inputting the data to be processed into the code task model to obtain the target code file corresponding to the target code task, the method further includes: Send the target code file to the front-end user; Receive the feedback information sent by the front-end user based on the target code file; Obtain an updated sample set based on the feedback information, and use the updated sample set to adjust the model parameters of the code task model.
11. The method according to claim 1, after inputting the data to be processed into the code task model to obtain the target code file corresponding to the target code task, the method further includes: Send the target code file to the front-end user; Receive the reply feedback data sent by the front-end user for the target code file; Generate question-and-answer guiding information based on the reply feedback data, wherein the question-and-answer guiding information is used to guide the front-end user to input target interaction data; Use the target interaction data as the data to be processed, and return to execute the step of inputting the data to be processed into the code task model to obtain the target code file corresponding to the target code task.
12. A model pre-training method, applied to a cloud-side device, includes: Obtain a sample code set, where the sample code set includes sample code files and the structure data of the sample code files; Use an initial code model to perform semantic analysis on the sample code files based on the structure data to obtain function data, and perform prediction on the sample code files based on the function data to obtain predicted code files; Train the initial code model based on the loss between the predicted code files and the sample code files to obtain a pre-trained model; Send the model parameters of the pre-trained model to the edge-side device.
13. A code completion method, includes: Receive the code to be completed sent by the front-end user; Input the code to be completed into a code completion model to obtain a target code file, where the code completion model is trained based on a sample set corresponding to a code completion task, and the pre-trained model is trained based on function data obtained by performing semantic analysis on sample code files; Feed back the target code file to the front-end user.
14. A code task processing system, includes a client and a server: The client is configured to send the data to be processed in the target code task to the server; The server is used to input the data to be processed into the code task model to obtain the target code file corresponding to the target code task, where The code task model is trained based on a sample set corresponding to the target code task, and the pre-trained model is trained based on function data obtained by performing semantic analysis on the sample code files; The client is further configured to receive the target code file corresponding to the target code task sent by the server.
15. A computing device, includes: A memory and a processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 13 are implemented.
16. A computer-readable storage medium, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.
17. A computer program product, includes computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.