Model code generation method and device, electronic equipment and storage medium

CN122837847APending Publication Date: 2026-09-29SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610968129.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]本申请实施例提供了一种模型代码生成方法,旨在解决现有方案存在采用固定采样策略,无法根据学生模型的实时学习进度动态调整样本难度分布,导致学生模型学习效率低、泛化能力差的问题

Benefits of technology

[0042]本发明实施例中,获取经过教师模型蒸馏的模型样本集,模型样本集包括多级难度的模型样本,每个模型样本包括一个模型实例与一个模型代码;基于学生模型的当前训练轮数,在模型样本集中采集多个不同级难度的模型样本,得到当前训练轮数的训练集;基于当前训练轮数的训练集,对学生模型进行训练,完成所有训练连轮数的训练,得到训练好的学生模型;通过训练好的学生模型对目标模型实例进行代码生成处理,得到目标模型的模型代码。本发明通过根据学生模型的当前训练轮数,在经过教师模型蒸馏的模型样本集中采集多个不同级难度的模型样本,得到当前训练轮数的训练集,并根据当前训练轮数的训练集,对学生模型进行训练,完成所有训练连轮数的训练,得到训练好的学生模型,通过训练好的学生模型对目标模型实例进行代码生成处理,得到目标模型的模型代码,解决了现有方案存在采用固定采样策略,无法根据学生模型的实时学习进度动态调整样本难度分布,导致学生模型学习效率低、泛化能力差的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122837847A_ABST
    Figure CN122837847A_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of model code generation, and provides a model code generation method, which comprises the following steps: obtaining a model sample set distilled by a teacher model, the model sample set comprising model samples of multiple levels of difficulty, each model sample comprising a model instance and a model code; collecting model samples of different levels of difficulty in the model sample set based on a current training round number of a student model, to obtain a training set of the current training round number; training the student model based on the training set of the current training round number, completing training of all training round numbers, and obtaining a trained student model; and performing code generation processing on a target model instance by using the trained student model, to obtain a model code of the target model. The application solves the problem in the prior art that a fixed sampling strategy is used, the sample difficulty distribution cannot be dynamically adjusted according to the real-time learning progress of the student model, and the student model has low learning efficiency and poor generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of model code generation technology, and particularly relates to a model code generation method, apparatus, electronic device and storage medium. Background Technology

[0002] With the rapid development of deep learning technology, various neural network models have achieved significant results in fields such as computer vision and natural language processing. Currently, existing solutions typically treat all valid samples generated by the teacher model as equivalent training data when constructing the training sample set, using random shuffling or uniform sampling methods. This lacks the ability to identify and hierarchically organize the sample difficulty. Existing methods employ a fixed sampling strategy during training, maintaining a constant sampling ratio for each type of sample in each training round. This fails to dynamically adjust the sample difficulty distribution according to the real-time learning progress of the student model. In the early stages of training, the high proportion of difficult samples can lead to gradient instability or convergence difficulties. In the later stages, the lack of sufficient high-difficulty samples hinders the improvement of the model's generalization ability. Therefore, there is an urgent need for a model code generation method that can acquire multi-level difficulty distilled sample sets and dynamically adjust the sample sampling strategy according to the training rounds. This would address the problems of existing solutions using fixed sampling strategies, failing to dynamically adjust the sample difficulty distribution according to the real-time learning progress of the student model, resulting in low learning efficiency and poor generalization ability of the student model. Summary of the Invention

[0003] This application provides a model code generation method to address the problem that existing solutions employ a fixed sampling strategy, which cannot dynamically adjust the sample difficulty distribution according to the real-time learning progress of the student model, resulting in low learning efficiency and poor generalization ability of the student model. By collecting multiple model samples of different difficulty levels from the model sample set distilled by the teacher model based on the current training epoch of the student model, a training set for the current training epoch is obtained. The student model is then trained using this training set for all training epochs to obtain a trained student model. This trained student model is then used to generate code for a target model instance, thus solving the problem of low learning efficiency and poor generalization ability in existing solutions due to the fixed sampling strategy.

[0004] In a first aspect, embodiments of the present invention provide a model code generation method, the method comprising the following steps:

[0005] Obtain a model sample set that has been distilled by the teacher model. The model sample set includes model samples of multi-level difficulty. Each model sample includes a model instance and a model code.

[0006] Based on the current training round number of the student model, multiple model samples of different difficulty levels are collected from the model sample set to obtain the training set for the current training round number.

[0007] Based on the training set of the current training round, the student model is trained to complete all training rounds and obtain a trained student model.

[0008] The target model's code is obtained by generating code from the target model instance using a trained student model.

[0009] Optionally, obtaining the model sample set after teacher model distillation includes:

[0010] Obtain the original model instance set, and use the teacher model to perform multiple inferences on the original model instance set to generate multiple initial model codes;

[0011] Based on the running time of the original model instance and the running time of the corresponding initial model code, a valid model sample is determined. The valid model sample includes the original model instance and the corresponding initial model code.

[0012] Based on the text length of the initial model code, the effective model samples are classified into difficulty levels to obtain model samples with multiple difficulty levels.

[0013] Based on the model samples with varying levels of difficulty, a model sample set is constructed.

[0014] Optionally, determining valid model samples based on the runtime of the original model instance and the runtime of the corresponding initial model code includes:

[0015] The initial model code is compiled and run to obtain the execution time of the initial model code;

[0016] Based on the baseline running time of the original model, the code speedup ratio of the initial model code is calculated;

[0017] The initial model code that compiles and runs correctly and whose speedup ratio is greater than or equal to a preset speedup ratio threshold, along with the corresponding original model instance, are determined as valid model samples.

[0018] Optionally, the step of classifying the effective model samples into difficulty levels based on the text length of the initial model code to obtain model samples with multiple difficulty levels includes:

[0019] Calculate the average text length of the initial model code corresponding to the effective model samples;

[0020] Within a preset range, the effective model sample with the smallest difference between the corresponding text length and the average value is selected as the denoising model sample;

[0021] The denoising model samples are divided into multi-level difficulty models according to the text length from shortest to longest. The longer the text, the higher the difficulty level of the denoising model sample.

[0022] Optionally, the current training round number based on the student model is obtained by collecting multiple model samples of different difficulty levels from the model sample set to obtain the training set for the current training round number, including:

[0023] Obtain the validation loss change parameters of the student model before entering the current training round;

[0024] Based on the current training round stage and the validation loss change parameters, the target sampling weights for each difficulty level are dynamically calculated.

[0025] According to the target sampling weights corresponding to each difficulty level, samples are extracted proportionally from the model samples of the multi-level difficulty to form the training set for the current training round.

[0026] Optionally, the multi-level difficulty model samples include low-difficulty, medium-difficulty, and high-difficulty model samples; the dynamic calculation of the target sampling weights for each difficulty level based on the current training round's progress stage and the validation loss change parameters includes:

[0027] The basic scheduling term is calculated by mapping the progress stage of the current training round using a preset transition activation function.

[0028] The verification loss change parameter is calculated with the preset dynamic feedback gain coefficient to obtain a dynamic feedback term that reflects the current learning progress of the model.

[0029] Based on the preset lower and upper limits of difficulty, the basic scheduling items and the dynamic feedback items are combined and calculated to output the first target sampling weight corresponding to the high difficulty level.

[0030] Optionally, after calculating the first target sampling weight corresponding to the higher difficulty level, the dynamic calculation of the target sampling weight for each difficulty level further includes:

[0031] Determine the remaining sampling space outside the first target sampling weight;

[0032] Detect whether the change in the validation loss parameter indicates that the learning state of the student model has stagnated or regressed;

[0033] If the indication stagnates or regresses, a logical buffer allocation mechanism is triggered. Specifically, the negative value of the verification loss change parameter is used to calculate the allocation of dynamic gain to increase the proportion of the medium difficulty level in the remaining sampling space, and then the second target sampling weight corresponding to the medium difficulty level is output.

[0034] The third target sampling weight corresponding to the low difficulty level is obtained by subtracting the second target sampling weight from the remaining sampling space.

[0035] Secondly, embodiments of the present invention also provide a model code generation apparatus, the model code generation apparatus comprising:

[0036] The acquisition module is used to acquire a model sample set that has been distilled by the teacher model. The model sample set includes model samples with multiple levels of difficulty, and each model sample includes a model instance and a model code.

[0037] The acquisition module is used to acquire multiple model samples of different difficulty levels from the model sample set based on the current training round number of the student model, and obtain the training set for the current training round number.

[0038] The training module is used to train the student model based on the training set of the current training round, complete the training for all training rounds, and obtain a trained student model.

[0039] The code generation module is used to generate code for the target model instance using a trained student model, thereby obtaining the model code of the target model.

[0040] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in the model code generation method provided in embodiments of the present invention.

[0041] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the model code generation method provided in the embodiments of the invention.

[0042] In this embodiment of the invention, a model sample set distilled from the teacher model is obtained. The model sample set includes model samples of multiple difficulty levels, and each model sample includes a model instance and a model code. Based on the current training round number of the student model, multiple model samples of different difficulty levels are collected from the model sample set to obtain a training set for the current training round number. Based on the training set for the current training round number, the student model is trained to complete all training rounds to obtain a trained student model. The trained student model is then used to generate code for the target model instance to obtain the model code for the target model. This invention solves the problem of existing solutions that use a fixed sampling strategy, which cannot dynamically adjust the sample difficulty distribution according to the real-time learning progress of the student model, resulting in low learning efficiency and poor generalization ability of the student model. This is achieved by collecting multiple model samples of different difficulty levels from the model sample set distilled by the teacher model, based on the current training round number of the student model. The training set is then used to train the student model, completing all training rounds to obtain a trained student model.

[0043] Other beneficial effects of this application will be described in detail in the following detailed description section. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 A flowchart illustrating a model code generation method provided in one embodiment of this application;

[0046] Figure 2 A schematic diagram of a model code generation device provided in one embodiment of this application;

[0047] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0048] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0049] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0050] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0051] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0052] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0053] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0054] like Figure 1 As shown, Figure 1This is a flowchart of a model code generation method provided in an embodiment of the present invention. The model code generation method includes the following steps:

[0055] 101. Obtain the model sample set after teacher model distillation.

[0056] In this embodiment of the invention, the above-described model code generation method can be applied to a model code generation platform. The model code generation platform can be built on a server-based or distributed architecture. The platform includes a data interface (for sensor or user uploads), a knowledge database, and a knowledge database construction program. The data interface can be used to obtain a model sample set distilled by the teacher model, and the knowledge database construction program can be used to construct the knowledge database. This knowledge database is specifically designed to provide additional relational information for the identified data entities, thereby enhancing the data recognition system's understanding of the content.

[0057] The aforementioned teacher model can be a pre-trained language model with a large parameter scale and strong reasoning ability. The teacher model can be Qwen3-235B-A22B, etc.

[0058] The aforementioned model sample set can be obtained after teacher model distillation. This set includes model samples of varying difficulty levels, each consisting of a model instance and model code. The varying difficulty levels can be attributed to the model samples within the set being categorized into different difficulty levels based on the complexity of their corresponding model code. Samples at different difficulty levels exhibit significant differences in code structure, code length, or logical complexity, including low, medium, and high difficulty. Higher difficulty indicates stronger reasoning and code structure comprehension abilities required for the student model to learn the model sample. The aforementioned model instance can be the original computational model requiring code generation, such as a PyTorch model instance. The model instance includes the model's layer type, connection method, activation function, parameter shape, etc. The aforementioned model code can be the source code string generated for the model instance and used for actual inference execution. The model code includes necessary header file references, memory allocation, computation loops, kernel function calls, etc.

[0059] The distillation process described above can be a process of generating a set of model samples to guide student model training by leveraging the reasoning capabilities of the teacher model. Specifically, it can involve inputting model instances into the teacher model, which then infers and outputs the corresponding model code, thereby forming a set of model samples that the student model can learn from.

[0060] It should be noted that by obtaining a model sample set containing multiple levels of difficulty after distillation of the teacher model, the student model can select the appropriate difficulty sample for learning based on its own ability during the training process.

[0061] 102. Based on the current training round number of the student model, collect multiple model samples of different difficulty levels from the model sample set to obtain the training set for the current training round number.

[0062] In this embodiment of the invention, the student model described above can be a lightweight language model with a small parameter size and fast inference speed.

[0063] The current training epoch number mentioned above can be the current training cycle number in the multi-round iterative training process of the student model. The current training epoch number reflects the amount of training the student model has received and its current maturity level.

[0064] The training set for the current training round can be a set of model samples of different difficulty levels collected from the model sample set based on the current training round of the student model, used to update the parameters of the student model.

[0065] It should be noted that, depending on the current stage of the training epochs, multiple model samples of different difficulty levels can be collected from the model sample set. This results in different proportions of samples of different difficulties in the training set corresponding to different training epochs. For example, when the number of training epochs is small, more low-difficulty samples are collected to help the model establish basic generation capabilities. In the middle of the training, the proportion of medium-difficulty samples is gradually increased to promote the improvement of model capabilities. When the number of training epochs is large, more high-difficulty samples are collected to challenge the generalization limit of the model. This ensures that the difficulty distribution of the training samples matches the current learning and receiving capabilities of the model.

[0066] 103. Based on the training set of the current training round, train the student model to complete all training rounds and obtain a well-trained student model.

[0067] In this embodiment of the invention, the student model can be trained according to the training set of the current training round, and the training of all training rounds can be completed to obtain a trained student model.

[0068] The above training can be performed by inputting the training set of the current training round into the student model, enabling the student model to generate prediction code, comparing the prediction code with the model code generated by the teacher model corresponding to the training set of the current training round, calculating the loss function value based on the difference between the prediction code and the model code generated by the teacher model corresponding to the training set of the current training round, and updating the network parameters of the student model through the backpropagation algorithm. The training process terminates when the number of training rounds reaches the required number of training rounds, and the training is completed, resulting in a well-trained student model.

[0069] All of the above training rounds can be the total number of complete training iterations set in advance.

[0070] It should be noted that the training set used in each training round is selected from models of different difficulty levels based on the current training round number of the student model. In the early stages of training, when the model's capabilities are relatively weak, more low-difficulty samples are tended to be allocated to help the model establish basic mapping capabilities. In the later stages of training, as the model's capabilities improve, more high-difficulty samples are gradually introduced to enhance the model's ability to generate code for complex model instances, effectively avoiding training instability or slow convergence caused by difficulty mismatch.

[0071] 104. The target model instance is processed by the trained student model to generate the model code.

[0072] In this embodiment of the invention, the target model instance can be a model instance that needs to generate executable code.

[0073] The above code generation process can be a process of generating runnable source code from a target model instance using a trained student model.

[0074] The model code of the target model can be directly used for the actual inference deployment of the target model without calling the teacher model, thereby significantly reducing inference costs and time while ensuring code quality.

[0075] It should be noted that the trained student model has learned the complex mapping relationship from the structure of model instances to the code implementation. It can perform code generation processing on the target model instance to obtain the model code of the target model. The trained student model has fast inference speed and low resource consumption, which can meet the high efficiency and real-time requirements of code generation tasks in practical applications.

[0076] In this embodiment of the invention, a model sample set distilled from the teacher model is obtained. The model sample set includes model samples of multiple difficulty levels, and each model sample includes a model instance and a model code. Based on the current training round number of the student model, multiple model samples of different difficulty levels are collected from the model sample set to obtain a training set for the current training round number. Based on the training set for the current training round number, the student model is trained to complete all training rounds to obtain a trained student model. The trained student model is then used to generate code for the target model instance to obtain the model code for the target model. This invention solves the problem of existing solutions that use a fixed sampling strategy, which cannot dynamically adjust the sample difficulty distribution according to the real-time learning progress of the student model, resulting in low learning efficiency and poor generalization ability of the student model. This is achieved by collecting multiple model samples of different difficulty levels from the model sample set distilled by the teacher model, based on the current training round number of the student model. The training set is then used to train the student model, completing all training rounds to obtain a trained student model.

[0077] It is understood that in the specific implementation of this application, sample data, instance data, code data, training data and other related data are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required. Furthermore, the collection, use and processing of related data, as well as the training, deployment and invocation of algorithm models, must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0078] Optionally, in the step of obtaining the model sample set distilled by the teacher model, the original model instance set can be obtained, and the teacher model can be used to perform multiple inferences on the original model instance set to generate multiple initial model codes; based on the running time of the original model instances and the running time of the corresponding initial model codes, the effective model samples can be determined; the effective model samples can be classified into difficulty levels according to the text length of the initial model codes to obtain multi-level difficulty model samples; and a model sample set can be constructed based on the multi-level difficulty model samples.

[0079] In this embodiment of the invention, the above-mentioned set of original model instances can be a collection of multiple original model instances, each of which includes the model's layer type, connection method, activation function, parameter shape, etc.

[0080] The aforementioned multiple inferences can be a process of using the teacher model to run inference multiple times on the original set of model instances to generate multiple initial model codes. The number of inferences can be 6, 8, etc.

[0081] The aforementioned initial model code can be the initial program code text output after the teacher model performs multiple inferences on the original model instance set.

[0082] The runtime of the original model instance can be the time consumed by the original model instance to complete a complete forward inference process under standardized hardware environment and software configuration.

[0083] The execution time of the initial model code can be the actual time consumed by the executable program to complete a complete forward inference process after the initial model code is compiled and run in the same standardized hardware environment and software configuration as the original model instance.

[0084] It should be noted that the runtime provides a clear, unified, and automated evaluation standard for judging sample validity, ensuring the usability of each sample in the model sample set.

[0085] The aforementioned valid model samples can be determined based on the runtime of the original model instance and the runtime of the corresponding initial model code. Valid model samples include the original model instance and the corresponding initial model code.

[0086] The text length of the initial model code can be either the total number of characters or the number of lines of code. Text length is used to quantify the size and structural complexity of the model code.

[0087] The above difficulty grading can be a process of dividing effective model samples into multiple difficulty levels based on the code complexity reflected by the text length of the initial model code.

[0088] The aforementioned multi-level difficulty model samples can be a set of valid model samples distributed across different difficulty levels after classifying the valid model samples according to the text length of the initial model code. Different difficulty levels correspond to different code text lengths. The multi-level difficulty can be low, medium, high, etc.

[0089] The aforementioned model sample set can be constructed based on model samples of varying difficulty levels.

[0090] It should be noted that using the teacher model to perform multiple inferences on the original model instance set can avoid the loss of high-quality samples due to accidental biases in a single inference. By classifying the effective model samples into difficulty levels according to the text length of the initial model code, a multi-level difficulty model sample set is constructed, providing a well-structured and fully labeled training dataset for subsequent training tasks, thus reducing the data loading overhead of subsequent training tasks.

[0091] Optionally, in the step of determining effective model samples based on the running time of the original model instance and the running time of the corresponding initial model code, the initial model code can be compiled and run to obtain the running time of the initial model code; the code speedup ratio of the initial model code can be calculated based on the baseline running time of the original model; and the initial model code and the corresponding original model instance that are compiled and run correctly and have a code speedup ratio greater than or equal to a preset speedup ratio threshold can be determined as effective model samples.

[0092] In this embodiment of the invention, the above-mentioned compilation and execution can be a process of parsing the syntax of the initial model code, compiling and generating intermediate representations, and actual execution.

[0093] The execution time of the initial model code mentioned above can be obtained by compiling and running the initial model code. The initial model code can be the actual time consumed by executing a complete forward inference process after compilation in a standardized hardware environment and software configuration.

[0094] The baseline running time of the original model can be the standard time consumed by an instance of the original model to complete a full forward inference process under standardized hardware and software configuration.

[0095] The code speedup of the initial model code mentioned above can be calculated based on the baseline runtime of the original model.

[0096] The aforementioned preset speedup threshold can be a pre-set speedup threshold, which is used to determine whether the initial model code has achieved an acceptable level of performance optimization relative to the original model instance.

[0097] The above-mentioned correct compilation and execution means that the initial model code does not produce any compilation errors, linking errors, or runtime exceptions during the compilation and execution process, can complete the entire process of compilation, loading, and inference execution, and returns the correct calculation results when the execution ends.

[0098] It should be noted that by calculating the code speedup ratio of the initial model code based on the baseline runtime of the original model, a standardized performance evaluation metric using the original model was established, eliminating the dimensional differences in runtime that made it incomparable between different model instances. By identifying the initial model code that compiles and runs correctly and whose code speedup ratio is greater than or equal to a preset speedup ratio threshold, along with the corresponding original model instances, as valid model samples, it is ensured that every sample in the model sample set has performance optimization value.

[0099] Optionally, in the step of classifying the effective model samples into difficulty levels according to the text length of the initial model code to obtain model samples of multiple difficulty levels, the average text length of the initial model code corresponding to the effective model samples can be calculated; within a preset range, the effective model samples with the smallest difference between the corresponding text length and the average value can be selected as denoising model samples; the denoising model samples can be divided into model samples of multiple difficulty levels in order of text length from shortest to longest, with the longer the text length, the higher the difficulty level of the denoising model sample.

[0100] In this embodiment of the invention, the above average value can be the average obtained by summing the text length values ​​of all valid model samples and dividing by the total number of samples.

[0101] The aforementioned preset range can be a pre-defined range. The range can be an offset interval centered on the average text length.

[0102] The aforementioned denoising model samples can be those selected from all valid model samples, where the difference between the text length and the average value falls within a preset range. The denoising model samples can also be the set of samples retained after excluding those with excessively long or short text lengths.

[0103] It should be noted that the average text length of the initial model code corresponding to the effective model samples can be calculated. Within a preset range, the effective model samples with the smallest difference between their text length and the average value are selected as denoising model samples, effectively removing samples that are abnormally long or short. The denoising model samples are then categorized into multiple difficulty levels from low to high according to the text length, from shortest to longest. The longer the text length, the higher the difficulty level of the denoising model sample.

[0104] Optionally, in the step of collecting multiple model samples of different difficulty levels from the model sample set based on the current training round of the student model to obtain the training set for the current training round, the validation loss change parameters of the student model before entering the current training round can be obtained; based on the progress stage of the current training round and the validation loss change parameters, the target sampling weights for each difficulty level can be dynamically calculated; according to the target sampling weights corresponding to each difficulty level, samples can be proportionally extracted from the model samples of multiple difficulty levels to form the training set for the current training round.

[0105] In this embodiment of the invention, the aforementioned validation loss change parameter can be a quantified value of the direction and magnitude of the change in the loss value calculated by the target intelligent agent on the validation set before entering the current training round, relative to the validation loss of the previous round. The validation loss change parameter is used to characterize the learning trend of the student model before the current training round.

[0106] The aforementioned progress stages can be the current training epoch's position or state category within the overall progress of all training epochs. Progress stages characterize the overall advancement of the training task.

[0107] The aforementioned dynamic calculation can be a process of calculating the target sampling weights for each difficulty level in real time based on the current training round stage and the change parameters of the validation loss.

[0108] The aforementioned target sampling weights can be sampling weights corresponding to each difficulty level. The target weight value determines the proportion of samples extracted from the model samples of the corresponding difficulty level in the training set of the current training round, relative to the total number of samples in the training set. Different difficulty levels correspond to different target sampling weights.

[0109] The above proportion can be the numerical proportion of the target sampling weight corresponding to each difficulty level.

[0110] The training set for the current training epoch can be a mixture of samples drawn proportionally from model samples across multiple difficulty levels, based on the target sampling weights corresponding to each difficulty level. This training set can also serve as the model sample set used to update the student model parameters.

[0111] Specifically, the sampling weight of the first target corresponding to a higher difficulty level can be obtained using the following formula:

[0112] D(t) = Dmin + (Dmax - Dmin) × σ(k × (t - tmid)) + λ × ΔL

[0113] Where D(t) represents the sampling weight of the first target; t represents the current training round number; tmid represents the midpoint of the course transition, usually set at 50% of the total training rounds; σ(x) represents the Sigmoid activation function, i.e., σ(x) = 1 / (1 + e^(-x)), with an output range of (0, 1), used to simulate a smooth syllabus transition; k represents the transition slope factor, used to control the steepness of the transition from simple to complex stages; Dmin represents the preset lower limit of difficulty; Dmax represents the preset upper limit of difficulty; ΔL represents the validation loss variation parameter, calculated as ΔL = L(t-1) - Lt (the loss function of the previous validation set minus the arithmetic loss of the current round), ΔL > 0 indicates that the model is progressing, ΔL ≤ 0 indicates that the model learning has stagnated or regressed; λ represents the dynamic feedback term, used to adjust the sensitivity of the model's progress speed to difficulty adjustments.

[0114] In one possible implementation, for example, the preset lower limit of difficulty is 0.1, the preset upper limit of difficulty is 0.5, the total number of training rounds is 10, that is, the course transition midpoint is 5, the transition slope factor is 1.0, and the dynamic feedback term is 0.5. In the early stage of training in scenario A, when the current training round number is 1, the model loss function decreases from 2.5 to 2.4, then the validation loss change parameter is 0.1, the basic scheduling term σ(1.0 × (1 - 5)) = σ(-4) ≈ 0.018; S(1) = 0.1 + (0.5 - 0.1) × 0.018 ≈0.107; dynamic feedback term: 0.5 × 0.1 = 0.05; final result: D(1) = 0.107 + 0.05 = 0.157; at this time, the data of the higher difficulty level only accounts for 15.7%.

[0115] In the middle of training in scenario B, when the current training round number is 5, the model loss function drops from 1.2 to 0.9, so the validation loss change parameter is 0.3; the basic scheduling term σ(1.0 × (5 - 5)) = σ(0) = 0.5, S(5) = 0.1 + (0.5 -0.1) × 0.5 = 0.3; the dynamic feedback term 0.5 × 0.3 = 0.15, D(5) = 0.3 + 0.15 = 0.45; the model improves rapidly, automatically increasing the proportion of high-difficulty data to only 45%, accelerating the model's mastery of complex GPU architecture logic.

[0116] In the later stages of training in scenario C, when the current training round number is 8, the model loss function increases from 0.4 to 0.42, so the verification loss change parameter is -0.02; the basic scheduling term σ(1.0 × (8 - 5)) = σ(3) ≈ 0.95, S(8) = 0.1 + (0.5 - 0.1) × 0.95 = 0.48; the dynamic feedback term 0.5 × (-0.02) = -0.01; D(8) = 0.48 - 0.01 = 0.47; the model falls into learning stagnation, and through the negative feedback mechanism, the proportion of data at higher difficulty levels is automatically reduced from 48% to 47%, reducing the cognitive load of the model.

[0117] It should be noted that the target sampling weights for each difficulty level can be dynamically calculated based on the current training round's progress stage and the changes in validation loss parameters of the student model before entering the current training round. This ensures that the target sampling weights for each training round not only conform to the overall plan of progressive training but also allow for fine-tuning based on the model's actual learning state in the current round. Furthermore, according to the target sampling weights corresponding to each difficulty level, samples are proportionally extracted from model samples of multiple difficulty levels to form the training set for the current training round, achieving precise control over the composition ratio of the training set samples.

[0118] Optionally, the multi-level difficulty model samples include low, medium, and high difficulty levels. In the step of dynamically calculating the target sampling weights for each difficulty level based on the current training epoch's progress stage and the validation loss change parameters, a preset transition activation function can be used to map the current training epoch's progress stage to calculate the basic scheduling term. The validation loss change parameters are then calculated with a preset dynamic feedback gain coefficient to obtain a dynamic feedback term that reflects the model's current learning progress. Based on preset lower and upper difficulty limits, the basic scheduling term and the dynamic feedback term are combined and calculated to output the first target sampling weight corresponding to the high difficulty level.

[0119] In this embodiment of the invention, the aforementioned preset transition activation function can be a pre-set transition activation function. The transition activation function is used to map the current training round's progress stage to a numerical basic scheduling term.

[0120] The aforementioned basic scheduling items can be numerical results obtained by mapping the progress stages through a preset transition activation function.

[0121] The aforementioned preset dynamic feedback gain coefficient can be a pre-set dynamic feedback gain function. The dynamic feedback gain coefficient is used to control the correction strength and influence of the verification loss variation parameters on the sampling weights of high-difficulty levels.

[0122] The aforementioned dynamic feedback term can be obtained by mathematically calculating the validation loss change parameter and a preset dynamic feedback gain coefficient, and is used to quantify the correction that should be applied to the sampling weights of higher difficulty levels based on the current model learning state. It should be noted that when the validation loss change parameter indicates that the model is learning well, the dynamic feedback term is positive and has a promoting effect; when the validation loss change parameter indicates that the model is learning poorly, the dynamic feedback term is negative and has an inhibiting effect.

[0123] The aforementioned preset lower difficulty limit can be a pre-set lower difficulty limit. The aforementioned preset upper difficulty limit can be a pre-set upper difficulty limit.

[0124] The above-mentioned superposition calculation can be the superposition of basic scheduling items and dynamic feedback items, and the superposition result is restricted to the range defined by the preset lower limit of difficulty and the preset upper limit of difficulty, so as to output the calculation process of the first target sampling weight corresponding to the high difficulty level.

[0125] It should be noted that by using a preset transition activation function to map the current training epoch's progress stage to calculate the basic scheduling term, a macro-trend planning for the smooth and non-linear increase of high-difficulty level sampling weights with training progress is achieved. The dynamic feedback term is calculated by combining the validation loss variation parameter with a preset dynamic feedback gain coefficient, allowing for fine-tuning of the high-difficulty level sampling weight calculation based on the model's real-time learning state on the validation set. By superimposing the basic scheduling term and the dynamic feedback term based on preset lower and upper difficulty limits to output the first target sampling weight, the high-difficulty level sampling weights are ensured to remain within a reasonable numerical range, effectively avoiding training imbalance caused by extreme weight values.

[0126] Optionally, after calculating the first target sampling weight corresponding to the high difficulty level, the remaining sampling space outside the first target sampling weight can be determined; whether the validation loss change parameter indicates that the learning state of the student model has stagnated or regressed; if it indicates stagnation or regression, a logical buffer allocation mechanism is triggered, specifically: using the negative value of the validation loss change parameter to calculate and allocate dynamic gain to increase the proportion of the medium difficulty level in the remaining sampling space, thereby outputting the second target sampling weight corresponding to the medium difficulty level; using the remaining sampling space to subtract the second target sampling weight to obtain the third target sampling weight corresponding to the low difficulty level.

[0127] In this embodiment of the invention, the aforementioned remaining sampling space may be the proportion of all available weights remaining after deducting the first target sampling weights already allocated to the high-difficulty level from all the sampling weights in the current training round.

[0128] The validation loss variation parameter can be compared with preset stagnation and regression thresholds. For example, if the absolute value of the validation loss variation parameter is lower than the preset stagnation threshold for multiple consecutive training epochs, it is determined as learning stagnation; if the validation loss variation parameter remains positive for multiple consecutive training epochs and cumulatively exceeds the preset regression threshold, it is determined as learning regression. The preset stagnation threshold can be a pre-set stagnation threshold. The preset regression threshold can be a pre-set regression threshold.

[0129] Furthermore, when a change in the validation loss parameter indicates a stagnation or regression in the learning state, a logical buffer allocation mechanism is triggered. The core of this mechanism is to calculate and allocate dynamic gains using the negative value of the validation loss parameter change, thereby increasing the proportion of the medium difficulty level in the remaining sampling space and outputting the second target sampling weight corresponding to the medium difficulty level. The aforementioned negative value can be taken as the input for calculating the dynamic gain when the validation loss parameter change is positive; that is, the more severe the loss increase, the larger the negative value, and the larger the dynamic gain. The aforementioned second target sampling weight can be the target sampling weight corresponding to the medium difficulty level.

[0130] Specifically, the formula for logical buffer allocation is as follows:

[0131] P2(t) = R(t) × β(t)

[0132] P1(t) = R(t) × (1 - β(t))

[0133] Where β(t) is the dynamically allocated weight for the medium difficulty level, and its calculation formula is:

[0134] β(t) = βbase + γ × max(0, -ΔL);

[0135] Where βbase is the base allocation ratio for the medium difficulty level; γ is the preset dynamic feedback gain coefficient; ΔL is the verification loss variation parameter; max(0, -ΔL) is the bottleneck trigger function, which will only output a positive value and activate the relief mechanism when ΔL < 0, and output 0 when ΔL ≥ 0.

[0136] In one possible implementation, for example, βbase = 0.3, γ = 2.0.

[0137] Scenario A corresponds to the current training round t=5. Given that the proportion of high-difficulty data is D(5) = 0.45, the remaining space R(5) = 1 - 0.45 = 0.55. At this time, ΔL = 0.3 > 0, max(0, -0.3) = 0. The dynamic weight allocation for the medium difficulty level is β(5) = 0.3 + 2.0 × 0 = 0.3. The proportion of medium difficulty is P2(5) = 0.55 × 0.3 = 16.5%; the proportion of low difficulty is P1(5) = 0.55 × (1 - 0.3) = 38.5%. At this time, the model is in good condition and can provide a small amount of medium-difficulty data to guide the logic at a regular pace, and consolidate the code with a large amount of low difficulty data.

[0138] In scenario B, corresponding to the current training round t=8, given that the proportion of high-difficulty data is D(8) = 0.47, the remaining space R(8) = 1 - 0.47 = 0.53; at this time, ΔL = -0.02 < 0, and the relief term is activated max(0, 0.02) = 0.02; the dynamic weight allocation (8) for the medium difficulty level is 0.3 + 2.0 × 0.02 = 0.34. The proportion of medium difficulty P2(8) = 0.53 × 0.34 ≈ 18.0%; the proportion of low difficulty P1(8) = 0.53 × (1 - 0.34) ≈ 35.0%.

[0139] The aforementioned third target sampling weight can be the target sampling weight corresponding to the lower difficulty level.

[0140] It should be noted that, by utilizing the remaining sampling space beyond the first target sampling weights, while ensuring that the weights already allocated to high-difficulty levels remain unchanged, medium- and low-difficulty levels are dynamically allocated within the remaining space. By triggering a logical buffer allocation mechanism when learning stagnation or regression is detected, the negative value of the validation loss change parameter is used to calculate the dynamic gain allocation, thereby increasing the proportion of medium-difficulty levels in the remaining sampling space. This achieves intelligent degradation buffering of training difficulty distribution. By subtracting the second target sampling weights from the remaining sampling space to obtain the third target sampling weights corresponding to low-difficulty levels, the normalization constraint of the sum of weights for high-difficulty, medium-difficulty, and low-difficulty levels is satisfied, ensuring the consistency of allocation.

[0141] like Figure 2 As shown, an embodiment of the present invention provides a model code generation device, which includes:

[0142] The acquisition module 201 is used to acquire a model sample set after teacher model distillation. The model sample set includes model samples of multi-level difficulty, and each model sample includes a model instance and a model code.

[0143] The acquisition module 202 is used to acquire multiple model samples of different difficulty levels in the model sample set based on the current training round number of the student model, so as to obtain the training set for the current training round number.

[0144] Training module 203 is used to train the student model based on the training set of the current training round, complete the training of all training rounds, and obtain a trained student model.

[0145] The code generation module 204 is used to generate code for the target model instance using the trained student model, thereby obtaining the model code of the target model.

[0146] Optionally, the acquisition module 201 is further configured to acquire an original model instance set, perform multiple inferences on the original model instance set using the teacher model to generate multiple initial model codes; determine valid model samples based on the running time of the original model instances and the running time of the corresponding initial model codes, wherein the valid model samples include the original model instances and the corresponding initial model codes; classify the valid model samples into difficulty levels according to the text length of the initial model codes to obtain multi-level difficulty model samples; and construct a model sample set based on the multi-level difficulty model samples.

[0147] Optionally, the acquisition module 201 is further configured to compile and run the initial model code to obtain the running time of the initial model code; calculate the code speedup ratio of the initial model code based on the baseline running time of the original model; and determine the initial model code and the corresponding original model instance that are correctly compiled and run and have a code speedup ratio greater than or equal to a preset speedup ratio threshold as valid model samples.

[0148] Optionally, the acquisition module 201 is further configured to calculate the average text length of the initial model code corresponding to the effective model sample; within a preset range, select the effective model sample with the smallest difference between the corresponding text length and the average value as the denoising model sample; and divide the denoising model sample into model samples of the multi-level difficulty according to the order of the text length from shortest to longest, wherein the longer the text length, the higher the difficulty level of the denoising model sample.

[0149] Optionally, the acquisition module 202 is further configured to acquire the validation loss change parameters of the student model before entering the current training round; dynamically calculate the target sampling weights for each difficulty level based on the progress stage of the current training round and the validation loss change parameters; and extract samples from the model samples of the multi-level difficulty according to the target sampling weights corresponding to each difficulty level to form the training set of the current training round.

[0150] Optionally, the acquisition module 202 is further configured to map the progress stage of the current training round using a preset transition activation function, and calculate the basic scheduling term; calculate the validation loss change parameter and the preset dynamic feedback gain coefficient to obtain the dynamic feedback term that reflects the current learning progress of the model; and perform superposition calculation based on the preset lower limit and upper limit of difficulty, combined with the basic scheduling term and the dynamic feedback term, to output the first target sampling weight corresponding to the high difficulty level.

[0151] Optionally, the device is further configured to determine the remaining sampling space outside the first target sampling weight; detect whether the verification loss change parameter indicates that the learning state of the student model has stagnated or regressed; if it indicates stagnation or regression, trigger a logical buffer allocation mechanism, specifically: using the negative value of the verification loss change parameter to calculate and allocate dynamic gain to increase the proportion of the medium difficulty level in the remaining sampling space, thereby outputting the second target sampling weight corresponding to the medium difficulty level; subtracting the second target sampling weight from the remaining sampling space to obtain the third target sampling weight corresponding to the low difficulty level.

[0152] like Figure 3 As shown, this embodiment of the invention also provides an electronic device, including a processor, which can execute any of the above-described model code generation methods.

[0153] Specifically, it includes a processor 301 and a memory 302, as well as a computer program stored in the memory 302 and capable of running on the processor 301, which executes the model code generation method, wherein:

[0154] The processor 301 executes the calculator program for the model code generation method stored in the memory 302, and performs the following steps:

[0155] Obtain a model sample set that has been distilled by the teacher model. The model sample set includes model samples of multi-level difficulty. Each model sample includes a model instance and a model code.

[0156] Based on the current training round number of the student model, multiple model samples of different difficulty levels are collected from the model sample set to obtain the training set for the current training round number.

[0157] Based on the training set of the current training round, the student model is trained to complete all training rounds and obtain a trained student model.

[0158] The target model's code is obtained by generating code from the target model instance using a trained student model.

[0159] Optionally, the process of obtaining the model sample set distilled by the teacher model, performed by processor 301, includes:

[0160] Obtain the original model instance set, and use the teacher model to perform multiple inferences on the original model instance set to generate multiple initial model codes;

[0161] Based on the running time of the original model instance and the running time of the corresponding initial model code, a valid model sample is determined. The valid model sample includes the original model instance and the corresponding initial model code.

[0162] Based on the text length of the initial model code, the effective model samples are classified into difficulty levels to obtain model samples with multiple difficulty levels.

[0163] Based on the model samples with varying levels of difficulty, a model sample set is constructed.

[0164] Optionally, the processor 301 determines valid model samples by calculating the execution time based on the original model instance and the execution time of the corresponding initial model code, including:

[0165] The initial model code is compiled and run to obtain the execution time of the initial model code;

[0166] Based on the baseline running time of the original model, the code speedup ratio of the initial model code is calculated;

[0167] The initial model code that compiles and runs correctly and whose speedup ratio is greater than or equal to a preset speedup ratio threshold, along with the corresponding original model instance, are determined as valid model samples.

[0168] Optionally, the processor 301 executes the step of classifying the effective model samples into difficulty levels based on the text length of the initial model code to obtain model samples with multiple difficulty levels, including:

[0169] Calculate the average text length of the initial model code corresponding to the effective model samples;

[0170] Within a preset range, the effective model sample with the smallest difference between the corresponding text length and the average value is selected as the denoising model sample;

[0171] The denoising model samples are divided into multi-level difficulty models according to the text length from shortest to longest. The longer the text, the higher the difficulty level of the denoising model sample.

[0172] Optionally, the processor 301 executes the current training round number based on the student model, collecting multiple model samples of different difficulty levels from the model sample set to obtain the training set for the current training round number, including:

[0173] Obtain the validation loss change parameters of the student model before entering the current training round;

[0174] Based on the current training round stage and the validation loss change parameters, the target sampling weights for each difficulty level are dynamically calculated.

[0175] According to the target sampling weights corresponding to each difficulty level, samples are extracted proportionally from the model samples of the multi-level difficulty to form the training set for the current training round.

[0176] Optionally, the multi-level difficulty model samples include low-difficulty, medium-difficulty, and high-difficulty model samples; the processor 301 executes the dynamic calculation of target sampling weights for each difficulty level based on the current training round stage and the validation loss change parameters, including:

[0177] The basic scheduling term is calculated by mapping the progress stage of the current training round using a preset transition activation function.

[0178] The verification loss change parameter is calculated with the preset dynamic feedback gain coefficient to obtain a dynamic feedback term that reflects the current learning progress of the model.

[0179] Based on the preset lower and upper limits of difficulty, the basic scheduling items and the dynamic feedback items are combined and calculated to output the first target sampling weight corresponding to the high difficulty level.

[0180] Optionally, after calculating the first target sampling weight corresponding to the higher difficulty level, the dynamic calculation of the target sampling weight for each difficulty level performed by the processor 301 further includes:

[0181] Determine the remaining sampling space outside the first target sampling weight;

[0182] Detect whether the change in the validation loss parameter indicates that the learning state of the student model has stagnated or regressed;

[0183] If the indication stagnates or regresses, a logical buffer allocation mechanism is triggered. Specifically, the negative value of the verification loss change parameter is used to calculate the allocation of dynamic gain to increase the proportion of the medium difficulty level in the remaining sampling space, and then the second target sampling weight corresponding to the medium difficulty level is output.

[0184] The third target sampling weight corresponding to the low difficulty level is obtained by subtracting the second target sampling weight from the remaining sampling space.

[0185] This invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the model code generation method provided in this invention and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0186] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for generating model code, characterized in that, The method includes the following steps: Obtain a model sample set that has been distilled by the teacher model. The model sample set includes model samples of multi-level difficulty. Each model sample includes a model instance and a model code. Based on the current training round number of the student model, multiple model samples of different difficulty levels are collected from the model sample set to obtain the training set for the current training round number. Based on the training set of the current training round, the student model is trained to complete all training rounds and obtain a trained student model. The target model's code is obtained by generating code from the target model instance using a trained student model.

2. The model code generation method as described in claim 1, characterized in that, The process of obtaining the model sample set after teacher model distillation includes: Obtain the original model instance set, and use the teacher model to perform multiple inferences on the original model instance set to generate multiple initial model codes; Based on the running time of the original model instance and the running time of the corresponding initial model code, a valid model sample is determined. The valid model sample includes the original model instance and the corresponding initial model code. Based on the text length of the initial model code, the effective model samples are classified into difficulty levels to obtain model samples with multiple difficulty levels. Based on the model samples with varying levels of difficulty, a model sample set is constructed.

3. The model code generation method as described in claim 2, characterized in that, The determination of valid model samples based on the runtime of the original model instance and the runtime of the corresponding initial model code includes: The initial model code is compiled and run to obtain the execution time of the initial model code; Based on the baseline runtime of the original model, the code speedup ratio of the initial model code is calculated; The initial model code that compiles and runs correctly and whose speedup ratio is greater than or equal to a preset speedup ratio threshold, along with the corresponding original model instance, are determined as valid model samples.

4. The model code generation method as described in claim 3, characterized in that, The step of classifying the effective model samples into difficulty levels based on the text length of the initial model code to obtain model samples with multiple difficulty levels includes: Calculate the average text length of the initial model code corresponding to the effective model samples; Within a preset range, the effective model sample with the smallest difference between the corresponding text length and the average value is selected as the denoising model sample; The denoising model samples are divided into multi-level difficulty models according to the text length from shortest to longest. The longer the text, the higher the difficulty level of the denoising model sample.

5. The model code generation method according to any one of claims 1 to 4, characterized in that, The current training epoch based on the student model is obtained by collecting model samples of different difficulty levels from the model sample set to obtain the training set for the current training epoch, including: Obtain the validation loss change parameters of the student model before entering the current training round; Based on the current training round stage and the validation loss change parameters, the target sampling weights for each difficulty level are dynamically calculated. According to the target sampling weights corresponding to each difficulty level, samples are extracted proportionally from the model samples of the multi-level difficulty to form the training set for the current training round.

6. The model code generation method as described in claim 5, characterized in that, The multi-level difficulty model samples include low-difficulty, medium-difficulty, and high-difficulty model samples; the dynamic calculation of target sampling weights for each difficulty level based on the current training epoch progress stage and the validation loss change parameters includes: The basic scheduling term is calculated by mapping the progress stage of the current training round using a preset transition activation function. The verification loss change parameter is calculated with the preset dynamic feedback gain coefficient to obtain a dynamic feedback term that reflects the current learning progress of the model. Based on the preset lower and upper limits of difficulty, the basic scheduling items and the dynamic feedback items are combined and calculated to output the first target sampling weight corresponding to the high difficulty level.

7. The model code generation method as described in claim 6, characterized in that, After calculating the first target sampling weight corresponding to the high difficulty level, the dynamic calculation of the target sampling weight for each difficulty level further includes: Determine the remaining sampling space outside the first target sampling weight; Detect whether the change in the validation loss parameter indicates that the learning state of the student model has stagnated or regressed; If the indication stagnates or regresses, a logical buffer allocation mechanism is triggered. Specifically, the negative value of the verification loss change parameter is used to calculate the allocation of dynamic gain to increase the proportion of the medium difficulty level in the remaining sampling space, and then the second target sampling weight corresponding to the medium difficulty level is output. The third target sampling weight corresponding to the low difficulty level is obtained by subtracting the second target sampling weight from the remaining sampling space.

8. A model code generation device, characterized in that, The model code generation device includes: The acquisition module is used to acquire a model sample set that has been distilled by the teacher model. The model sample set includes model samples with multiple levels of difficulty, and each model sample includes a model instance and a model code. The acquisition module is used to acquire multiple model samples of different difficulty levels from the model sample set based on the current training round number of the student model, and obtain the training set for the current training round number. The training module is used to train the student model based on the training set of the current training round, complete the training for all training rounds, and obtain a trained student model. The code generation module is used to generate code for the target model instance using a trained student model, thereby obtaining the model code of the target model.

9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the model code generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the model code generation method as described in any one of claims 1 to 7.