Method and device for model training, equipment and storage medium

By pre-training and fine-tuning the target model, using training data containing mixed text of natural language and programming languages, the problem of insufficient understanding and generation capabilities of models in the existing technology in code Q&A tasks is solved, and better code Q&A capabilities and flexibility are achieved.

CN120181256APending Publication Date: 2025-06-20FACE CUTE CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311758166.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively train models with code generation capabilities, especially in code Q&A tasks. The strict restrictions on input and output by the model and the diversity of user problem forms lead to insufficient understanding and generation capabilities.

Method used

By first pre-training the target model, using the corresponding code data of the programming language, and then fine-tuning the pre-trained model on the code question and answer tasks, and using a training data set containing sample questions and answers mixed with text in natural language and programming languages ​​to improve the model's understanding and generation ability.

Benefits of technology

It realizes better training of code Q&A tasks, improves the model's ability to handle mixed text problems in natural language and programming languages, supports free-form code Q&A, and enhances the generalization ability of the model and the flexibility of code Q&A.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181256A_ABST
    Figure CN120181256A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method, a device and equipment for model training and a storage medium. The method comprises the steps that a target model is pre-trained through a first training data set, the pre-trained target model is obtained, the target model is at least configured to execute a code question and answer task, and the first training data set comprises code data corresponding to a programming language; at least a second training data set is used for executing fine tuning on a code question and answer task on a pre-trained target model, the second training data set comprises a plurality of sample questions corresponding to model input and a plurality of sample answers corresponding to model output, and the sample questions comprise mixed texts of natural languages and programming languages; and / or the sample answer comprises a mixed text of a natural language and a programming language; and providing the fine-tuned target model for executing the code question and answer task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computer technologies, and in particular, to methods, apparatuses, devices, and computer-readable storage media for model training. Background Art

[0002] As machine learning and deep learning technologies have been widely applied in many fields. A machine learning model can be configured and trained to generate corresponding model outputs based on given model inputs. In machine learning technologies, generative models can be applied to various fields such as natural language processing, machine translation, speech synthesis, and image generation. Currently, models with code generation capabilities have also been proposed, which can assist in writing computer code. However, how to better train such models remains a challenging task. Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for model training is provided. The method includes: performing pre-training on a target model using a first training dataset to obtain a pre-trained target model, where the target model is at least configured to perform a code question-answering task, and the first training dataset includes code data corresponding to a programming language; performing fine-tuning on the pre-trained target model in the code question-answering task using at least a second training dataset, where the second training dataset includes a plurality of sample questions corresponding to model inputs and a plurality of sample answers corresponding to model outputs, the sample questions include a mixed text of natural language and programming language, and / or the sample answers include a mixed text of natural language and programming language; and providing the fine-tuned target model for performing the code question-answering task.

[0004] In a second aspect of the present disclosure, an apparatus for model training is provided. The apparatus includes: a pre-training module configured to perform pre-training on a target model using a first training dataset to obtain a pre-trained target model, where the target model is at least configured to perform a code question-answering task, and the first training dataset includes code data corresponding to a programming language; a fine-tuning module configured to perform fine-tuning on the pre-trained target model in the code question-answering task using at least a second training dataset, where the second training dataset includes a plurality of sample questions corresponding to model inputs and a plurality of sample answers corresponding to model outputs, where the sample questions include a mixed text of natural language and programming language, and / or the sample answers include a mixed text of natural language and programming language; and a model providing module configured to provide the fine-tuned target model for performing the code question-answering task.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the medium, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0007] It should be understood that the content described in this part is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In conjunction with the drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0009] Figure 1 A schematic diagram showing an example environment in which the embodiments of the present disclosure can be implemented;

[0010] Figure 2 A block diagram showing the structure of a model with code generation capabilities according to some embodiments of the present disclosure;

[0011] Figure 3 A schematic diagram showing the model training process according to some embodiments of the present disclosure;

[0012] Figure 4 A flowchart showing the process for model training according to some embodiments of the present disclosure;

[0013] Figure 5 A block diagram showing the apparatus for model training according to some embodiments of the present disclosure; and

[0014] Figure 6 An electronic device in which one or more embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0016] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.

[0017] It can be understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0018] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained through appropriate means in accordance with relevant laws and regulations.

[0019] For example, when receiving the user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested to be executed will require obtaining and using the user's personal information, so that the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.

[0020] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0021] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0022] As used herein, the term "model" can learn the correlation between corresponding inputs and outputs from training data, so that after training is completed, for a given input, a corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this document, "model" can also be referred to as "machine learning model", "learning model", "machine learning network", or "learning network", and these terms are used interchangeably herein.

[0023] A "neural network" is a machine learning network based on deep learning. A neural network can process inputs and provide corresponding outputs, and generally includes an input layer and an output layer, as well as one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications usually include many hidden layers, thereby increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also called processing nodes or neurons), and each node processes the input from the previous layer.

[0024] Generally, machine learning can roughly include three stages, namely the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously iteratively updated until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the correlation from input to output (also called the mapping from input to output) from the training data. The parameter values of the trained model are determined. In the testing stage, test inputs are applied to the trained model to test whether the model can provide correct outputs, thereby determining the performance of the model. The testing stage can sometimes be integrated into the training stage. In the application or inference stage, the trained model can be used to process actual model inputs based on the parameter values obtained from training and determine the corresponding model outputs.

[0025] Figure 1 A schematic diagram of an environment 100 in which embodiments of the present disclosure can be implemented is shown. In Figure 1 Three different stages of a model are shown in the environment 100, including a pre-training stage 102, a fine-tuning stage 104, and an application stage 106. There can also be a testing stage after the pre-training or fine-tuning stage is completed, which is not shown in the figure.

[0026] In the pre-training stage 102, the model pre-training system 110 is configured to perform pre-training of the model using the training data set 112. At the start of pre-training, the model can have initial parameter values. The pre-training process is to update the parameter values of the model to the desired values based on the training data. During the pre-training process, one or more pre-training tasks 107-1, 107-2, etc. can be designed. The pre-training tasks are used to assist in updating the parameters of the model.

[0027] In the pre-training stage 102, the model can learn powerful generalization capabilities through large-scale training data. After pre-training is completed, the parameter values of the model have been updated and have pre-trained parameter values. The pre-trained model can extract the feature representation of the image relatively accurately.

[0028] The pre-trained model can be provided to the fine-tuning stage 104 and fine-tuned by the model fine-tuning system 120 for different downstream tasks. The downstream tasks can be configured according to specific task requirements, such as downstream tasks 117-1, 117-2, etc. In some embodiments, depending on the specific downstream task, the pre-trained model can be connected to the output layer required by the downstream task to construct the corresponding downstream task model. This is because for different downstream tasks, the required outputs may be different. In the fine-tuning stage 104, the training data set 122 is further used to adjust the parameter values of the model.

[0029] During fine-tuning, the corresponding training algorithm is also used to update and adjust the parameters of the overall model. Since the model has learned a lot of knowledge from the training data in the pre-training stage, a small amount of training data can be used in the fine-tuning stage 104 to obtain a downstream task model that meets the expectations.

[0030] In the application stage 106, the obtained model 105 with trained parameter values can be provided to the model application system 130 for use. In the application stage 106, the model 105 can be used to process the corresponding target input 132 in the actual scenario and provide the corresponding target output 134.

[0031] In Figure 1 this context, the model pre-training system 110, the model fine-tuning system 120, and the model application system 130 can include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices can involve any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. Servers include, but are not limited to, mainframes, edge computing nodes, computing devices in cloud environments, and so on.

[0032] It should be understood that Figure 1 The components and arrangements in the illustrated environment 100 are merely examples, and a computing system suitable for implementing the exemplary implementations described in this disclosure may include one or more different components, other components, and / or different arrangements. For example, although shown as separate, the model pre-training system 110, the model fine-tuning system 120, and the model application system 130 may be integrated in the same system or device. For example, at least the model pre-training system 110 and the model fine-tuning system 120 may be integrated in a model training system or device. The implementations of this disclosure are not limited in this regard.

[0033] It should be understood that the structures and functions of the various elements in the environment 100 are described only for exemplary purposes and do not imply any limitation on the scope of this disclosure.

[0034] A generative model refers to a content generation technology implemented by a model. A generative model can automatically generate meaningful output content based on input content, including text, images, audio, video, etc. A generative model is sometimes also referred to as a multimodal model. Driven by large-scale pre-trained models, the capabilities of generative models have been applied in many application fields. In the field of code programming, it is also desirable to be able to assist in generating computer code by a model to help improve the code writing efficiency. Currently, some models are trained to have code generation and code completion capabilities, but these models have strict restrictions on both input and output. It is also desirable to be able to train models adapted to more flexible code question-and-answer scenarios.

[0035] In the code question-and-answer scenario, the forms of users' questions vary greatly, and may include scenarios where natural language and code are mixed, including providing code examples, etc., and it is expected that the answers obtained are not only the code itself, but also include natural language explanations of the generated code, etc. In the embodiments of this disclosure, in order to support the mixed question-and-answer ability of natural language and code, an improved model training scheme is proposed.

[0036] Figure 2 A block diagram of the structure of a target model 205 with code generation ability according to some embodiments of this disclosure is shown. The target model 205 is configured to support interactions in the form of code question-and-answer, where the model input can be considered as a question, and the model output is considered as an answer to the input question. The target model 205 may include a code large model. In an application scenario of code question-and-answer, the target model 205 is expected to be trained to support free-form code question-and-answer.

[0037] Such as Figure 2As shown, the question 210 provided to the target model 205 may include a mixed text of natural language and programming language. For example, the question 210 includes code text fragments 212 and 214 corresponding to the programming language, and the rest are natural language text fragments. This can support users to ask questions in the form of natural language mixed with code. In some cases, depending on the specific question, the answer 220 provided by the target model 205 may also include a mixed text of natural language and programming language. For example, the answer 220 includes a code text fragment 222 corresponding to the programming language, and the rest are natural language text fragments. In this way, the target model 205 can make the user understand the answer content more clearly in the form of natural language mixed with code. Such a code Q&A form is called free-form code Q&A. Of course, in some cases, the question of the target model 205 may only contain natural language or only contain code, or the answer only contains natural language or only contains code. In addition, on the basis of supporting free-form code Q&A, the target model 205 can also be trained to support traditional tasks such as code completion and natural language to code conversion.

[0038] For free-form code Q&A, due to the diverse forms of users' questions, this poses high requirements on the model's understanding ability and generalization ability. In the embodiments of the present disclosure, it is desired to enable the target model to have such capabilities during the training phase.

[0039] Figure 3 Shows a schematic diagram of a model training process 300 according to some embodiments of the present disclosure. The model training process 300 can be implemented at Figure 1 the model pre-training system 110 and the model fine-tuning system 120.

[0040] In the initial stage of the model training process 300, for the initial target model 205, the target model 205 is pre-trained using the first training dataset to obtain the pre-trained target model 205. For example, Figure 1 the model pre-training system 110 can be configured to perform the pre-training of the target model 205. The goal of pre-training is to perform unsupervised training on the target model 205 using a large amount of corpus. The first training dataset 310 includes code data corresponding to the programming language. In the pre-training stage, pre-training the target model 205 with code data can enable the target model 205 to learn and understand the corresponding programming language. In some embodiments, the first training dataset 310 may include code data of one or more programming languages (e.g., Python, C language, etc.) so that the target model 205 can understand different programming languages.

[0041] In some embodiments, before using the code data to pre-train the target model 205, the target model 205 may have been pre-trained on a third training dataset including natural language data. Alternatively, the first training dataset 310 further includes natural language data. In this way, the target model 205 can learn and understand natural language.

[0042] In some embodiments, the code data in the first training dataset 310 may include a plurality of code snippets. Considering the large amount of available code data, high-quality code data can be screened out for pre-training to improve the training effect. In some embodiments, a plurality of candidate code snippets available can be obtained from a code data source, and using a trained code quality classifier, the quality score of each of the plurality of candidate code snippets in the code data source can be determined. Then, based on the quality score of each of the plurality of candidate code snippets, a plurality of code snippets are selected from the plurality of candidate code snippets. In some embodiments, by setting a quality score threshold, a plurality of code snippets exceeding the quality score threshold can be screened out from the plurality of candidate code snippets. In some embodiments, the code quality classifier can be obtained by performing supervised training on a training dataset, and the training dataset of the code quality classifier may include sample code snippets labeled with quality scores. Using the code quality classifier, high-quality code snippets can be quickly and automatically screened out for training the target model 205. In some embodiments, alternatively or additionally, in addition to using the code quality classifier, other methods can also be used to collect the training data for pre-training the target model 205.

[0043] In some embodiments, the code data in the first training dataset 310 may additionally or alternatively include code textbook data, and the code textbook includes explanations of the use of code in a programming language. The code textbook may include defining, explaining, etc. the code in a specific programming language using natural language to help understand and apply the programming language. Based on the code textbook, the target model 205 can more quickly understand and apply the programming language.

[0044] After pre-training, at least use the second training dataset 320 to perform fine-tuning on the pre-trained target model 205 for the code question-answering task. For example, Figure 1 the model fine-tuning system 120 can be configured to perform fine-tuning on the target model 205. The goal of fine-tuning is to perform supervised fine-tuning on the target model 205 using a large number of model input-output pairs. The fine-tuning of the target model 205 is sometimes also referred to as instruction fine-tuning.

[0045] The second training dataset 320 includes a plurality of sample questions corresponding to the model input and a plurality of sample answers corresponding to the model output. The sample questions include mixed text of natural language and programming language, and / or the sample answers include mixed text of natural language and programming language. In the fine-tuning phase, for the code question-answering task, the training samples of the target model 205 are in the form of question-answer pairs. To support free-form questions, in the embodiments of the present disclosure, the question-answer pairs for fine-tuning include mixed text of natural language and programming language as questions and / or answers. From such training data, the target model 205 can be fine-tuned to learn how to understand questions containing natural language and programming language, and how to generate answers containing natural language and programming language.

[0046] In some embodiments, when constructing the second training dataset 320, the first sample questions and the first sample answers corresponding to the plurality of first sample questions can be collected from data sources related to code question-answering. Different data sources in the data sources can come from different fields, such as Front-End, Back-End, Data Science and Machine Learning (DS&ML), Mobile&Desktop, and Information Technology Operations (IT Ops) and other fields. The data sources can include data of any appropriate natural language and any appropriate programming language. Thereby, the diversity of the data sources can be ensured, and the comprehensiveness of subsequent evaluations can be improved. In some embodiments, the data sources can include forums, discussion posts, etc. related to code question-answering, in which there is a greater probability of containing question-answer pairs of mixed text of natural language and programming language.

[0047] In some embodiments, the data sources can include a plurality of questions and a plurality of answers corresponding to each question. A plurality of sample questions can be collected from the data sources. For each of the plurality of sample questions, a plurality of candidate answers corresponding to the corresponding question can also be collected from the data sources. The quality scores of the respective plurality of candidate answers can be determined, and based on the quality scores of the respective plurality of candidate answers, a sample answer matching the corresponding sample question can be selected from the plurality of candidate answers. Both the sample questions and the sample answers here can include text segments of natural language and text segments of programming language.

[0048] Specifically, for the collected first sample questions, a plurality of candidate answers corresponding to the first sample questions can be collected from the data sources; the quality scores of the respective plurality of candidate answers can be determined; and based on the quality scores of the respective plurality of candidate answers, a first sample answer matching the first sample question can be selected from the plurality of candidate answers.

[0049] In some embodiments, based on the interaction feedback information of each of multiple candidate answers in a data source, the quality score of each of the multiple candidate answers is determined. Regarding the specific manner of determining the quality score of a certain candidate answer (e.g., the first candidate answer) here, in some embodiments, the data source also includes interaction feedback information (such as like, comment, forward, view, etc. information) for different questions and answers. The quality score of the first candidate answer can be determined based on the interaction feedback information corresponding to the first candidate answer in the data source. Exemplarily, the more interactions indicated by the interaction feedback information (e.g., the higher the like count, the more forwards, the more comments, the larger the view count, etc.), the higher the quality score of the corresponding first candidate answer. For another example, the closer the interaction time corresponding to the latest interaction (e.g., the latest like, the latest comment, etc.) indicated by the interaction feedback information is to the current time, the higher the quality score of the corresponding first candidate answer. It can be understood that the quality scores of each of the multiple candidate answers can be determined based on a similar manner. In some examples, the candidate answer with the highest corresponding quality score can be determined as the sample answer matching the corresponding sample question.

[0050] In some embodiments, considering that the number of existing natural language and programming language mixed Q&A pairs in the data source is usually small, more mixed Q&A pairs can also be automatically generated through evolution and expansion, etc. For example, for the first sample question and the first sample answer collected from the data source, a prompt input for a trained generative model is constructed. The prompt input can be used to guide the generative model to output more Q&A pairs similar to the first sample question and the first sample answer. The generative model can be a text generation model based on a language model and having Q&A capabilities. The output prompt input is provided to the generative model, and based on the output of the generative model, at least one second sample question and its corresponding second sample answer in the second training dataset 320 are determined. In this way, the model can be used to generate more training data for fine-tuning the target model 205.

[0051] In some embodiments, the input of the target model 205 is in the form of a prompt. For the sample questions in the second training dataset 310, a prompt input for the target model 205 can also be generated using a prompt template and input into the target model 205 for processing.

[0052] In some embodiments, in addition to free-form question-and-answer tasks, during the fine-tuning phase of the target model 205, its training data may also include other code question-and-answer tasks, including training data corresponding to code completion tasks, natural language to code tasks, code debugging tasks, code interpretation tasks, etc., without limitation here. The training data under different tasks is related to the questions and answers under the corresponding tasks. For example, in code completion tasks and code debugging tasks, the question input to the model is code, and the answer output by the model is the completed code. In natural language to code tasks, the question input to the model is natural language text, and the answer output by the model is code. In code interpretation tasks, the question input to the model is the code to be explained, and the answer output by the model is a natural language interpretation of the code. By using task-related training data to perform instruction fine-tuning on the target model 205, the target model 205 can be made to have stronger code question-and-answer capabilities.

[0053] Provide a fine-tuned target model 205 for performing code question-and-answer tasks. For example, the model fine-tuning system 120 can provide the fine-tuned target model 205 to the model application system 130. In the model application system 130, the fine-tuned target model 205 can be used to process actual model inputs, which can be various forms of user questions. The fine-tuned target model 205 provides corresponding model outputs for providing as answers corresponding to the questions.

[0054] Figure 4 The flowchart of a process 400 for model training according to some embodiments of the present disclosure is shown. The process 400 can be implemented at Figure 1 the model pre-training system 110 and the model fine-tuning system 120.

[0055] In block 410, perform pre-training on the target model using a first training dataset to obtain a pre-trained target model, where the target model is at least configured to perform code question-and-answer tasks, and the first training dataset includes code data corresponding to a programming language.

[0056] In block 420, perform fine-tuning on the pre-trained target model in code question-and-answer tasks using at least a second training dataset, where the second training dataset includes multiple sample questions corresponding to model inputs and multiple sample answers corresponding to model outputs, and where the sample questions include a mixture of natural language and programming language text, and / or the sample answers include a mixture of natural language and programming language text.

[0057] In block 430, provide a fine-tuned target model for performing code question-and-answer tasks.

[0058] In some embodiments, the second training dataset is determined by: collecting first sample questions and first sample answers corresponding to the multiple first sample questions from a data source related to code Q&A.

[0059] In some embodiments, collecting first sample questions and first sample answers corresponding to the multiple first sample questions from a data source related to code Q&A includes: for each of the collected first sample questions, collecting multiple candidate answers corresponding to the first sample question from the data source; determining the quality score of each of the multiple candidate answers; and based on the quality scores of each of the multiple candidate answers, selecting a first sample answer that matches the first sample question from the multiple candidate answers.

[0060] In some embodiments, determining the quality score of each of the multiple candidate answers includes: determining the quality score of each of the multiple candidate answers based on the interaction feedback information of each of the multiple candidate answers in the data source.

[0061] In some embodiments, the second training dataset is further determined by: based on the first sample questions and first sample answers, constructing a prompt input for the trained generative model; providing the prompt input to the generative model; and based on the output of the generative model, determining second sample questions and second sample answers in the second training dataset.

[0062] In some embodiments, the code data in the first training dataset includes at least one of the following: multiple code snippets, code textbook data, and the code textbook includes an explanation of the use of code in a programming language.

[0063] In some embodiments, the multiple code snippets are determined by: using a trained code quality classifier to determine the quality score of each of the multiple candidate code snippets in the code data source; and based on the quality scores of each of the multiple candidate code snippets, selecting multiple code snippets from the multiple candidate code snippets.

[0064] In some embodiments, the target model has been pre-trained on a third training dataset including natural language data before pre-training.

[0065] Figure 5 The block diagram of an apparatus 500 for model training according to some embodiments of the present disclosure is shown. The apparatus 500 can be implemented as or included in Figure 1 the model pre-training system 110 and the model fine-tuning system 120. Each module / component in the apparatus 500 can be implemented by hardware, software, firmware, or any combination thereof.

[0066] As shown in the figure, the apparatus 500 includes a pre-training module 510 configured to perform pre-training on a target model using a first training dataset to obtain a pre-trained target model, where the target model is at least configured to perform a code question-answering task, and the first training dataset includes code data corresponding to a programming language. The apparatus 500 further includes a fine-tuning module 520 configured to perform fine-tuning on the pre-trained target model at least using a second training dataset on the code question-answering task. The second training dataset includes a plurality of sample questions corresponding to model inputs and a plurality of sample answers corresponding to model outputs, where the sample questions include a mixed text of natural language and programming language, and / or the sample answers include a mixed text of natural language and programming language. The apparatus 500 further includes a model providing module configured to provide the fine-tuned target model for performing the code question-answering task.

[0067] In some embodiments, the second training dataset is determined by: collecting a first sample question and first sample answers corresponding to a plurality of first sample questions from a data source related to code question-answering.

[0068] In some embodiments, collecting a first sample question and first sample answers corresponding to a plurality of first sample questions from a data source related to code question-answering includes: for the collected first sample question, collecting a plurality of candidate answers corresponding to the first sample question from the data source; determining quality scores of the plurality of candidate answers respectively; and based on the quality scores of the plurality of candidate answers respectively, selecting a first sample answer that matches the first sample question from the plurality of candidate answers.

[0069] In some embodiments, determining quality scores of the plurality of candidate answers respectively includes: determining quality scores of the plurality of candidate answers respectively based on interaction feedback information of the plurality of candidate answers in the data source.

[0070] In some embodiments, the second training dataset is further determined by: constructing a prompt input for a trained generative model based on the first sample question and the first sample answer; providing the prompt input to the generative model; and based on the output of the generative model, determining second sample questions and second sample answers in the second training dataset.

[0071] In some embodiments, the code data in the first training dataset includes at least one of the following: a plurality of code snippets, code textbook data, and the code textbook includes an explanation of the use of code in a programming language.

[0072] In some embodiments, the plurality of code snippets are determined by: using a trained code quality classifier to determine quality scores of a plurality of candidate code snippets in a code data source respectively; and based on the quality scores of the plurality of candidate code snippets respectively, selecting a plurality of code snippets from the plurality of candidate code snippets.

[0073] In some embodiments, the target model has been pre-trained on a third training dataset including natural language data before pre-training.

[0074] Figure 6 A block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that Figure 6 the illustrated electronic device 600 is merely exemplary and should not constitute any limitation on the functions and scope of the embodiments described herein. Figure 6 The illustrated electronic device 600 can be used to implement Figure 1 the model pre-training system 110, the model fine-tuning system 120, and / or the model application system 130, or Figure 5 the apparatus 500.

[0075] As Figure 6 shown, the electronic device 600 is in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 can be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 600.

[0076] The electronic device 600 generally includes multiple computer storage media. Such media can be any accessible media available to the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 can be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 can be a removable or non-removable medium and can include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 600.

[0077] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 6As shown, a disk drive for reading from and writing to a removable, non-volatile disk (such as a "floppy disk") and an optical disk drive for reading from and writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data medium interfaces. Memory 620 may include a computer program product 625 having one or more program modules configured to perform the various methods or acts of the various embodiments of the present disclosure.

[0078] Communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of electronic device 600 can be implemented by a single computing cluster or multiple computer machines that are capable of communicating via a communication connection. Thus, electronic device 600 can operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.

[0079] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a display, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) as needed via communication unit 640, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with electronic device 600, or communicate with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., a network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0080] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, where the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the methods described above.

[0081] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0082] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0083] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0084] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0085] The implementations of the present disclosure have been described above. The description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of technology in the marketplace, or to enable other ordinary skilled artisans in the art to understand the implementations disclosed herein.

Claims

1. A method for model training, comprising: Perform pre-training on a target model using a first training dataset to obtain a pre-trained target model, where the target model is at least configured to perform a code question-answering task, and the first training dataset includes code data corresponding to a programming language; Perform fine-tuning on the pre-trained target model at least using a second training dataset for the code question-answering task, where the second training dataset includes a plurality of sample questions corresponding to model inputs and a plurality of sample answers corresponding to model outputs, and where the sample questions include a mixed text of natural language and programming language, and / or the sample answers include a mixed text of natural language and programming language; And Provide the fine-tuned target model for performing the code question-answering task.

2. The method according to claim 1, wherein the second training dataset is determined by: Collecting the first sample questions and the first sample answers corresponding to the plurality of first sample questions from a data source related to code Q&A.

3. The method according to claim 2, wherein collecting the first sample questions and the first sample answers corresponding to the plurality of first sample questions from a data source related to code Q&A includes: For the first sample question collected, Collect a plurality of candidate answers corresponding to the first sample question from the data source; Determine the quality score of each of the plurality of candidate answers; And Based on the quality scores of each of the plurality of candidate answers, select a first sample answer that matches the first sample question from the plurality of candidate answers.

4. The method according to claim 3, wherein determining the quality score of each of the plurality of candidate answers includes: Determine the quality score of each of the plurality of candidate answers based on the interactive feedback information for each of the plurality of candidate answers in the data source.

5. The method according to claim 2, wherein the second training dataset is further determined by: Based on the first sample questions and the first sample answers, constructing a prompt input for the pre-trained generative model; Provide the prompt input to the generative model; And Based on the output of the generative model, determine the second sample question and the second sample answer in the second training dataset.

6. The method according to claim 1, wherein the code data in the first training dataset includes at least one of the following: A plurality of code snippets, Code textbook data, where the code textbook includes an explanation of the use of code in a programming language.

7. The method according to claim 6, wherein the plurality of code snippets are determined by: Using a pre-trained code quality classifier to determine the quality score of each of the plurality of candidate code snippets in the code data source; and Based on the quality scores of each of the plurality of candidate code snippets, selecting the plurality of code snippets from the plurality of candidate code snippets.

8. The method according to claim 1, wherein before the pre-training, the target model has been pre-trained on a third training dataset including natural language data.

9. An apparatus for model training, comprising: A pre-training module, configured to perform pre-training on a target model using a first training dataset to obtain a pre-trained target model, where the target model is at least configured to perform a code question-answering task, and the first training dataset includes code data corresponding to a programming language; A fine-tuning module, configured to perform fine-tuning on the pre-trained target model at least using a second training dataset for the code question-answering task, where the second training dataset includes a plurality of sample questions corresponding to model inputs and a plurality of sample answers corresponding to model outputs, and where the sample questions include a mixed text of natural language and programming language, and / or the sample answers include a mixed text of natural language and programming language; And A model providing module, configured to provide the fine-tuned target model for performing the code question-answering task.

10. An electronic device, comprising: At least one processing unit; And At least one memory, where the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit, and the instructions, when executed by the at least one processing unit, cause the device to perform the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Model fine tuning method and device, electronic equipment, storage medium and program product

    CN120851222A