Code generation method and device and computing equipment

Through the hybrid programming language training sample data training method of general coding models and selecting lightweight expert adapters, the accuracy and speed problems in multi-programming language code generation are solved, and efficient code generation is achieved, suitable for intelligent programming auxiliary tools and automated software development platforms.

CN120469674APending Publication Date: 2025-08-12DIGITAL TRADING SCI & TECH (BEIJING) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510559626.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient accuracy of generating code and high inference delay in code generation in multi-programming languages, which cannot meet the needs of real-time scenarios.

Method used

The general coding model is trained using the training sample data of a hybrid programming language, and the context representation process is performed by selecting the target expert adapter, and fusing it with the language detailed features to generate code, and using the lightweight expert adapter to avoid repeated calculations.

Benefits of technology

It improves the accuracy and speed of code generation, reduces computing consumption, supports real-time high concurrent request scenarios, and is suitable for intelligent programming auxiliary tools and automated software development platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469674A_ABST
    Figure CN120469674A_ABST
Patent Text Reader

Abstract

The invention discloses a code generation method and device and computing equipment, and the method comprises the steps: processing a received input content through a universal coding model, and obtaining a context representation; the universal coding model is obtained by training the training sample data of multiple programming languages; selecting a target expert adapter from a plurality of expert adapters according to the context representation; different expert adapters are obtained by training the training sample data of different programming languages; processing the context representation by using a target expert adapter to obtain language detailed features; and fusing the context representation and the language detailed features, and generating a target code according to features obtained through fusion. By means of the mode, the universal features and the unique features of the specific programming language are fused for code generation, the correctness of code generation can be improved, part of expert adapters are lightly activated instead of all expert adapters, the code generation speed can be increased, and calculation consumption can also be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and specifically to a code generation method, apparatus, computing device, computer storage medium, and computer program product. Background Art

[0002] With the rapid development of large language models, they are increasingly used in various fields because they can understand natural language input by humans and generate logically rigorous results. Especially in the field of code generation, code generation models are widely used in application scenarios where high-quality code is generated in real time, providing auxiliary tools for programming work.

[0003] Current technology has shifted toward multimodal collaborative frameworks, using datasets from multiple programming languages to train models and support code generation for these languages. However, existing technologies typically use a single model for code generation, which has significant limitations. These include difficulty balancing the grammatical characteristics of multiple languages, resulting in insufficient accuracy in generated code, and high inference latency, making it incapable of meeting the demands of real-time scenarios. These drawbacks limit the practicality and sustainability of existing technologies in complex multilingual environments. Summary of the Invention

[0004] In view of the above problems, the present application is proposed to provide a code generation method, apparatus, computing device, computer storage medium and computer program product that overcome the above problems or at least partially solve the above problems.

[0005] According to one aspect of the present application, a code generation method is provided, comprising:

[0006] Processing the received input content using a universal encoding model to obtain a contextual representation; wherein the universal encoding model is trained using training sample data of multiple programming languages;

[0007] A target expert adapter is selected from a plurality of expert adapters according to the context representation; wherein different expert adapters are trained using training sample data of different programming languages;

[0008] Utilize the target expert adapter to process the context representation and obtain detailed language features;

[0009] The context representation and detailed language features are fused, and the target code is generated based on the fused features.

[0010] Optionally, selecting a target expert adapter from a plurality of expert adapters according to the context representation further comprises:

[0011] Determine the weight scores of various programming languages based on context representation;

[0012] A target programming language is determined from various programming languages according to the weight scores, and an expert adapter corresponding to the target programming language is determined as a target expert adapter.

[0013] Optionally, the method further comprises:

[0014] For any programming language, the initial universal coding layer is trained according to the training sample data of the programming language to obtain a first model;

[0015] Training an initial second model according to the training sample data corresponding to the programming language to obtain a second model;

[0016] The initial second model includes a third model consistent with the first model and an initial expert adapter, and during the training of the initial second model, parameters of the third model are frozen while parameters of the initial expert adapter are adjusted;

[0017] The initial expert adapter after parameter adjustment included in the second model is determined as the expert adapter corresponding to the programming language.

[0018] Optionally, training the initial second model according to the training sample data corresponding to the programming language further includes:

[0019] Inputting the training input data in the training sample data into the first model and the initial second model for processing respectively;

[0020] Calculate the output layer loss based on the output probability distribution of the first model and the output probability distribution of the initial second model;

[0021] Obtain a first feature of the intermediate layer output of the first model and a second feature of the corresponding intermediate layer output in the initial second model, and calculate the intermediate layer loss based on the first feature and the second feature;

[0022] Calculate the task loss based on the predicted labels output by the initial second model and the true labels corresponding to the training input data;

[0023] The total loss is calculated based on the output layer loss, the intermediate layer loss, and the task loss, and the parameters of the initial expert adapter are adjusted based on the total loss.

[0024] Optionally, if the number of target expert adapters is at least two, processing the context representation using the target expert adapter to obtain detailed language features further includes:

[0025] The initial language detailed features obtained by processing at least two target expert adapters are weightedly summed to obtain the language detailed features.

[0026] Optionally, determining the weight scores of various programming languages according to the context representation further includes:

[0027] The context representation is input into a dynamic routing selector for processing to obtain weight scores of various programming languages; wherein the loss function used for training the dynamic routing selector is determined according to the cross entropy loss function.

[0028] According to another aspect of the present application, there is provided a code generating device, comprising:

[0029] a first feature extraction module adapted to process received input content using a universal coding model to obtain a contextual representation; wherein the universal coding model is trained using training sample data of multiple programming languages;

[0030] a selection module adapted to select a target expert adapter from a plurality of expert adapters according to the context representation; wherein the different expert adapters are trained using training sample data of different programming languages;

[0031] a second feature extraction module adapted to process the context representation using a target expert adapter to obtain detailed language features;

[0032] The generation module is suitable for fusing the context representation and the detailed language features and generating the target code based on the fused features.

[0033] Optionally, the selection module is further adapted to:

[0034] Determine the weight scores of various programming languages based on context representation;

[0035] A target programming language is determined from various programming languages according to the weight scores, and an expert adapter corresponding to the target programming language is determined as a target expert adapter.

[0036] Optionally, the device further comprises:

[0037] A training module is suitable for any programming language, and trains an initial universal coding layer according to the training sample data of the programming language to obtain a first model; trains an initial second model according to the training sample data corresponding to the programming language to obtain a second model; wherein the initial second model includes a third model consistent with the first model and an initial expert adapter, and during the training of the initial second model, the parameters of the third model are frozen and the parameters of the initial expert adapter are adjusted; the initial expert adapter after the parameters are adjusted contained in the second model is determined as the expert adapter corresponding to the programming language.

[0038] Optionally, the training module is further adapted to:

[0039] Inputting the training input data in the training sample data into the first model and the initial second model for processing respectively;

[0040] Calculate the output layer loss based on the output probability distribution of the first model and the output probability distribution of the initial second model;

[0041] Obtain a first feature of the intermediate layer output of the first model and a second feature of the corresponding intermediate layer output in the initial second model, and calculate the intermediate layer loss based on the first feature and the second feature;

[0042] Calculate the task loss based on the predicted labels output by the initial second model and the true labels corresponding to the training input data;

[0043] The total loss is calculated based on the output layer loss, the intermediate layer loss, and the task loss, and the parameters of the initial expert adapter are adjusted based on the total loss.

[0044] Optionally, the second feature extraction module is further adapted to:

[0045] The initial language detailed features obtained by processing at least two target expert adapters are weightedly summed to obtain the language detailed features.

[0046] Optionally, the selection module is further adapted to:

[0047] The context representation is input into a dynamic routing selector for processing to obtain weight scores of various programming languages; wherein the loss function used for training the dynamic routing selector is determined according to the cross entropy loss function.

[0048] According to another aspect of the present application, a computing device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus;

[0049] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute operations corresponding to the above-mentioned code generation method.

[0050] According to another aspect of the present application, a computer storage medium is provided, wherein the storage medium stores at least one executable instruction, and the executable instruction enables a processor to perform operations corresponding to the above-mentioned code generation method.

[0051] According to another aspect of the present application, a computer program product is provided, comprising at least one executable instruction, wherein the executable instruction enables a processor to perform operations corresponding to the above-mentioned code generation method.

[0052] According to the code generation method, apparatus, computing device, computer storage medium and computer program product provided in the embodiments of the present application, a universal coding model is used to process the received input content to obtain a context representation; wherein the universal coding model is trained using training sample data of multiple programming languages; a target expert adapter is selected from multiple expert adapters based on the context representation; wherein different expert adapters are trained using training sample data of different programming languages; the context representation is processed using the target expert adapter to obtain language detailed features; the context representation and the language detailed features are fused, and the target code is generated based on the fused features. In the above manner, a universal context representation is obtained using a universal coding model to provide a cross-language semantic representation basis for the expert adapter, and the expert adapters of different programming languages share the context representation of the universal coding model, which can avoid repeated calculations; by lightweight activation of some expert adapters instead of all expert adapters, the code generation speed can be improved and the computing consumption can be reduced; the code generation is performed by fusing the universal features between programming languages and the unique features of a specific programming language, which can improve the correctness of the generated code.

[0053] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0055] Figure 1 A flowchart of a code generation method provided by an embodiment of the present application is shown;

[0056] Figure 2 A flowchart of a code generation method provided by another embodiment of the present application is shown;

[0057] Figure 3 A flowchart of a code generation method provided by another embodiment of the present application is shown;

[0058] Figure 4 A schematic diagram showing the functional structure of a code generation device provided in an embodiment of the present application is shown;

[0059] Figure 5 A schematic diagram of the structure of a computing device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0060] The following describes exemplary embodiments of the present application in more detail with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0061] Figure 1 FIG. 1 shows a flow chart of a code generation method provided by an embodiment of the present application. Figure 1 As shown, the method includes the following steps:

[0062] Step S110 : Process the received input content using a general coding model to obtain a context representation.

[0063] Among them, the universal coding model is trained using training sample data of multiple programming languages. Specifically, the initial universal coding layer is trained using training sample data of mixed programming languages to obtain the universal coding model. The initial universal coding layer can adopt a Transformer structure. The pre-trained universal coding model is used to learn common knowledge across programming languages, including syntax, semantic features, code logic, data structure, control flow, algorithm logic, etc., programming languages ​​such as Java, Python, C++, etc.

[0064] Input content refers to the content entered by the user to represent their code requirements. It can be a natural language description or a code snippet. The input content is input into the universal coding model, and the high-dimensional feature vector obtained by processing the universal coding model is extracted, namely the context representation. The context representation refers to the feature vector used to directly generate tokens. A token is the smallest semantic unit for text segmentation and encoding. The universal coding model generates code based on the token.

[0065] Step S120 : selecting a target expert adapter from a plurality of expert adapters according to the context representation.

[0066] Different expert adapters are trained using training sample data from different programming languages. Each programming language is equipped with an expert adapter. This means that the expert adapter is trained based on the corresponding language's training sample data. The expert adapter learns the language's unique knowledge, including grammatical rules, keywords, and coding style, and is responsible for generating grammatical details. Expert adapters are lightweight modules for specific programming languages and can use LoRa architectures or small FFN networks. Their parameter size is smaller than that of general-purpose coding models.

[0067] According to the context representation obtained by processing the general coding model, the target programming language related to the context representation is determined, and the expert adapter corresponding to the target programming language is determined as the target expert adapter. The number of target expert adapters may be one or more.

[0068] Step S130 : Process the context representation using the target expert adapter to obtain detailed language features.

[0069] Activate the selected target expert adapter, input the context representation into the target expert adapter for processing, extract the initial language detailed features obtained by the target expert adapter from the context representation, and determine the language detailed features based on the initial language detailed features corresponding to the target expert adapter. Similarly, the initial language detailed features obtained by the target expert adapter from the context representation are the features used to directly generate tokens.

[0070] Step S140 , fusing the context representation and the detailed language features, and generating target code based on the fused features.

[0071] The context representation and language detailed features are fused together, and based on the fused features, the target code is generated in an autoregressive manner. In the method of the embodiment of the present application, the language detailed features are extracted through the target expert adapter, and the detailed features of the specific programming language are injected into the general context representation, thereby improving the grammatical correctness of the generated code.

[0072] In the existing technology, large language models for mixed programming languages have certain limitations when processing different programming languages. Specifically: on the one hand, when a single large language model generates code for different programming languages, there are large differences in the accuracy of the generated code, especially in language syntax and detail processing, and consistency and high quality cannot be guaranteed; on the other hand, large language models with strong code generation capabilities usually require a higher number of parameters, resulting in slower inference speed and poor real-time performance, which cannot meet the needs of scenarios with high real-time requirements, and also consumes a lot of computing resources.

[0073] In summary, according to the code generation method provided in this embodiment, a code generation model architecture based on hybrid experts is provided, which solves the pain points of existing large language models in code generation through a dynamic adaptation architecture, builds corresponding independent lightweight expert adapters for each programming language, decouples them from the shared universal coding model, and adopts a hybrid architecture of trunk plus plug-ins. The shared trunk, i.e., the universal coding model, learns cross-language programming logic, while the pluggable plug-ins, i.e., the expert adapters, are specially optimized for specific programming languages. The universal coding model is pre-trained through large-scale training sample data of mixed programming languages, and the common features of multiple programming languages are learned. The received input content is processed using the universal coding model to obtain a universal Context representation, so as to provide a cross-language semantic representation basis for expert adapters; select a target expert adapter from multiple expert adapters according to the context representation, and only load the selected target expert adapter. By lightweight activation of some expert adapters instead of all expert adapters, the code generation speed can be improved and the computing consumption can be reduced; the general context representation extracted by the general coding model and the unique refined features extracted by the expert adapter are fused to generate the target code, and the general features between programming languages and the unique features of a specific programming language are fused for code generation, which can improve the correctness of the generated code. Expert adapters of different programming languages share the context representation of the general coding model, which can avoid repeated calculations. In summary, the embodiment of the present application provides an optimization method for a large language model in cross-programming language code generation, which is suitable for application scenarios such as intelligent programming auxiliary tools and automated software development platforms that require real-time generation of high-quality multi-programming language codes.

[0074] Figure 2 FIG. 1 shows a flow chart of a code generation method provided by another embodiment of the present application. Figure 2 As shown, the method includes the following steps:

[0075] Step S210 : For any programming language, the initial universal coding layer is trained according to the training sample data of the programming language to obtain a first model.

[0076] The training of the expert adapter is divided into two steps. The first step is to train the first model (also called the teacher model), and the second step is to train the second model (also called the student model).

[0077] The training sample data includes code generation requirement descriptions and their sample labels, i.e., code. For any programming language, the initial general coding layer is trained using the corresponding training sample data to obtain the first model. The input is the code generation requirement description, and the output includes code and probability. The initial general coding layer can adopt a Transformer structure.

[0078] Specifically, all parameters of the initial general encoding layer are fine-tuned, and a complete first model is trained for various programming languages. The first model can generate high-quality code, but has a large number of parameters and high inference cost.

[0079] Step S220 , training an initial second model based on the training sample data corresponding to the programming language to obtain a second model; and determining the initial expert adapter after parameter adjustment included in the second model as the expert adapter corresponding to the programming language.

[0080] The initial second model includes a third model consistent with the first model and an initial expert adapter. During the training of the initial second model, the parameters of the third model are frozen and the parameters of the initial expert adapter are adjusted. The initial second model refers to the model in the training process, and the initial expert adapter is also an adapter in the training process. When the training meets the iteration termination condition, the model parameters at this time are solidified to obtain the second model, and the initial expert adapter after parameter adjustment is the expert adapter used for code generation. In the method of the embodiment of the present application, the structure and parameters of the first model are reused to obtain the third model. During the training of the initial second model, the parameters of the third model are kept unchanged and the parameters of the initial expert adapter are adjusted.

[0081] The initial expert adapter is a lightweight module with only 1%-5% of the parameters of the first model. The third model is connected to the initial expert adapter. The features extracted by the third model for direct token generation serve as the input of the initial expert adapter. The output of the initial expert adapter includes the generated code and probability.

[0082] In an optional manner, training the initial second model according to the training sample data corresponding to the programming language specifically includes the following steps:

[0083] Step S1: input the training input data in the training sample data into the first model and the initial second model for processing.

[0084] The training input data is essentially a description of code requirements, and the ground truth labels corresponding to the training input data are essentially code. The training input data is fed into the first model and the initial second model, respectively. The loss is calculated using the processing results of the first and initial second models, and the parameters of the initial expert adapter in the initial second model are adjusted based on the loss.

[0085] Step S2: Calculate the output layer loss based on the output probability distribution of the first model and the output probability distribution of the initial second model.

[0086] The output probability distribution refers to the model's quantification of the likelihood of each code result. By minimizing the output layer loss function, the initial second model learns the knowledge of the first model so that the outputs of the two are aligned. The output layer loss is calculated as follows:

[0087]

[0088] in, represents the output layer loss, D KL represents KL divergence, p teacher (y|x) represents the output probability distribution of the first model, p student (y|x) represents the output probability distribution of the initial second model.

[0089] Step S3: Obtain the first feature of the intermediate layer output of the first model and the second feature of the corresponding intermediate layer output in the initial second model, and calculate the intermediate layer loss based on the first feature and the second feature.

[0090] By minimizing the intermediate layer loss function, the intermediate features of the first model are aligned with the intermediate features of the initial second model to enhance the synergy between the two. The calculation method of the intermediate layer loss is as follows:

[0091]

[0092] in, represents the intermediate layer loss, ||2 represents the L2 norm, h teacher represents the first feature of the intermediate layer output of the first model, h student Represents the second feature output by the corresponding intermediate layer in the initial second model. Optionally, the intermediate layer of the first model is specifically an attention head output layer, and the intermediate layer in the initial second model corresponding to the intermediate layer of the first model can be selected according to actual needs.

[0093] Step S4: Calculate the task loss based on the predicted labels output by the initial second model and the true labels corresponding to the training input data.

[0094] The task loss function is specifically a cross-entropy loss function. The cross-entropy loss is calculated based on the predicted label, i.e., code and probability, output by the initial second model and the true label, i.e., label code, of the training input data to obtain the task loss.

[0095] Step S5: Calculate the total loss based on the output layer loss, the intermediate layer loss, and the task loss, and adjust the parameters of the initial expert adapter based on the total loss.

[0096] A total loss is calculated based on the output layer loss, the intermediate layer loss, and the task loss, and the total loss is used to adjust the parameters of the initial expert adapter.

[0097] In an optional approach, the total loss is obtained by calculating the weighted sum of the output layer loss, the intermediate layer loss, and the task loss. The specific calculation method is as follows:

[0098]

[0099] Among them, α, β, and γ are weights, and their specific values can be set according to actual needs. represents the total loss, Indicates mission loss.

[0100] The method of the embodiment of the present application further includes: when the programming language has a grammar update, using the training sample data corresponding to the updated grammar of the programming language to fine-tune and update the expert adapter corresponding to the programming language.

[0101] Step S230 , pre-training the initial universal coding layer using training sample data of multiple programming languages to obtain a universal coding model.

[0102] Specifically, the universal encoding model is obtained by training the initial universal encoding layer using training sample data from mixed programming languages. The goal of training the initial universal encoding layer is to learn universal knowledge across programming languages. The training data used is a mixed multi-language code base, and the training objectives include masked language modeling, code completion, and cross-language translation. The initial universal encoding layer can adopt a Transformer structure, which includes a multi-head attention network and a feedforward network.

[0103] Preferably, after the universal encoding model is pre-trained once, the parameters are fixed and remain unchanged to avoid catastrophic forgetting due to differences between programming languages.

[0104] Step S240: Process the received input content using a general coding model to obtain a context representation.

[0105] Input content refers to the content entered by the user to represent their code requirements. It can be a natural language description or a code snippet. The input content is input into the general coding model, and the high-dimensional feature vector obtained by processing the general code model is extracted, namely the context representation. The context representation refers to the feature vector used to directly generate tokens.

[0106] Step S250 , determining weight scores of various programming languages according to the context representation; determining a target programming language from the various programming languages according to the weight scores, and determining an expert adapter corresponding to the target programming language as a target expert adapter.

[0107] Specifically, the context representation is input into the dynamic routing selector for processing to obtain the weight scores of various programming languages. Afterwards, the various programming languages are sorted in descending order, and the programming languages with the top k (i.e., top-k, k≥1) weight scores are determined as target programming languages. The expert adapter corresponding to the target programming language is the target expert adapter, which realizes sparse activation of the top-k expert adapters based on the input content, avoiding the problem of high computational overhead caused by activating all expert adapters.

[0108] Optionally, the dynamic routing selector uses a small neural network model, which also needs to be pre-trained before it can be used. Specifically, during the process of training expert adapters corresponding to various programming languages, the training input data is input into the initial second model for processing. The third model processes the sample context representation (i.e., the features used to directly generate tokens). The sample context representation is determined as the input of the initial dynamic routing selector. The loss function is calculated based on the output of the initial dynamic routing selector. The parameters of the initial dynamic routing selector are adjusted based on the calculated loss. The above steps are repeated until the iteration termination condition is reached, and the model parameters at this time are solidified to obtain the dynamic routing selector.

[0109] In addition, when a new programming language is added, the dynamic routing selector is updated using the training sample data corresponding to the new programming language, and the output dimension of the dynamic router is expanded.

[0110] In an optional manner, the loss function used to train the dynamic router is determined according to a cross entropy loss function.

[0111] In another alternative approach, the loss function used to train the dynamic router is determined by a cross-entropy loss function and a load balancing loss function. The training goal is to balance the usage frequency of expert adapters and improve routing accuracy. The cross-entropy loss function is used to maximize the weight score of the most relevant programming languages, while the load balancing loss function is used to penalize the overuse or inactivity of certain expert adapters, thereby preventing imbalanced usage of these adapters.

[0112] Step S260 : Process the context representation using the target expert adapter to obtain detailed language features.

[0113] If the number of target expert adapters is 1, the initial language detailed features obtained by processing the target expert adapter are determined as language detailed features, wherein the initial language detailed features obtained by processing the target expert adapter are features used to directly generate tokens; if the number of target expert adapters is at least two, at least two target expert adapters are calculated in parallel, and the initial language detailed features obtained by processing the at least two target expert adapters are weighted and summed to obtain language detailed features, and the weight used in the weighted summation is the weight score of the corresponding programming language calculated in the above steps; based on this, the method of the embodiment of the present application can support mixed code output of multiple programming languages or code output of a single programming language.

[0114] Step S270: fuse the context representation and the detailed language features, and generate target code based on the fused features.

[0115] The context representation and detailed language features are fused by addition, weighted sum calculation, or a specific fusion model, thereby injecting the unique features of the target programming language into the general context representation, such as grammatical rules (Python indentation, Java type declarations), etc.

[0116] The fused features are processed by the decoder to generate the target code. Specifically, the fused features are input to the encoder, which generates tokens based on the fused features and then generates the target code token by token. The decoder can use the standard Transformer decoder architecture, including a masked multi-head attention mechanism and a feedforward network, supporting autoregressive generation. The encoder can generate the target code through greedy search, beam search, or sampling.

[0117] In an optional manner, when generating the target code, the decoder calls a programming language-specific syntax constraint module (such as an AST parser) to assist in generating code, and ensures the grammatical correctness of the output code through syntax tree verification, indentation rules, etc.

[0118] An embodiment of the present application provides a cross-programming language code generation method based on dynamic routing and lightweight distillation, which uses a general coding model to process the received input content to obtain a general context representation to provide a cross-language semantic representation basis for the expert adapter; selects a target expert adapter from multiple expert adapters based on the context representation, and only loads the selected target expert adapter; extracts detailed features of the programming language through the target expert adapter, and injects unique features of the programming language into the context representation, which can improve the grammatical correctness of the generated code. At the same time, the lightweight expert adapter and sparse activation strategy can improve the code generation speed and reduce computing consumption. Under the same hardware conditions, it can support high-concurrency request scenarios, such as achieving second-level response of online programming assistants; in summary, the method of the embodiment of the present application can ensure a balance between code quality and real-time performance.

[0119] In addition, in the prior art, the entire model needs to be retrained to adapt to the language version update, and the economic cost is too high. The industry attempts to update model parameters through knowledge distillation compression model or incremental training, but the former sacrifices the logical integrity of the code, and the latter still faces the risk of parameter drift. The code generation method of the embodiment of the present application, through the knowledge distillation technology, through the framework of the teacher model and the student model, migrates the capabilities of the large model (first model) with full parameter fine-tuning to a lightweight expert adapter. By minimizing the output distribution difference, the adapter approaches the performance of the first model with a very small amount of parameters, improves the code generation quality of the trained expert adapter, and maintains the efficiency of the parameters. Moreover, the expert adapter only requires a small amount of labeled data to complete the training, which significantly reduces the data collection cost. When there is a new programming language, the expert adapter corresponding to the new programming language is trained in accordance with the above-mentioned training expert adapter method, and the dynamic routing selector is trained and updated and the relevant output dimensions are expanded. This avoids full model reconstruction and can effectively reduce the model training cost.

[0120] At present, programming languages are constantly updated and changed, and code generation models need to be retrained regularly to adapt to new syntax and features. This not only increases the training cost, but also makes the update cycle of the model longer, making it difficult to quickly respond to changes in programming languages, and the update and maintenance costs are high. The method of the embodiment of the present application can efficiently support the dynamic update of programming languages. When there is a syntax update in a programming language, the newly added training sample data of the programming language is collected and the expert adapter is fine-tuned. There is no need to retrain the corresponding expert adapter, and there is no need to adjust the general coding model and the expert adapters of other programming languages. The high cost of retraining the entire model in the traditional solution is avoided, which can effectively reduce the cost of model updates and shorten the time spent on training updates. At the same time, the fixed parameters of the general coding model ensure that the functions of other programming languages are not affected during the update process, eliminate the risk of version conflicts, and ensure compatibility.

[0121] Furthermore, the method of the embodiment of the present application can also solve the coverage problem of long-tail programming languages, has small sample adaptation capabilities, and only requires a small amount of training data to train expert adapters for niche or emerging programming languages. Compared with traditional large models of multiple programming languages, the generation accuracy is greatly improved, and it can also avoid model deviations caused by differences in sample data volume, ensuring balanced generation quality for mainstream programming languages and niche programming languages.

[0122] Furthermore, the method of the embodiment of the present application can also significantly reduce training and deployment costs. The general encoding model is pre-trained once and then the parameters are fixed. New languages only train lightweight expert adapters, which effectively reduces training costs. In addition, only some expert adapters are activated during inference. Compared with dense models with the same parameter scale, the code generation speed is significantly improved.

[0123] Furthermore, the method of the embodiment of the present application also has good ecological scalability. First, it can open language support. Developers can quickly expand the supported programming languages by defining new expert adapters without having to master the full model training technology; second, it has industrial implementation advantages. The expert adapter supports dynamic loading and hot replacement, and adapts to the continuous integration requirements of enterprise-level code generation platforms.

[0124] Figure 3 FIG. 4 shows a flow chart of a code generation method provided by another embodiment of the present application, such as Figure 3 As shown, the general process of the method is as follows: the input content is processed by a general encoding model to obtain a context representation; the dynamic routing selector processes the context representation, determines the target programming language based on the output weight score, and activates the expert adapter corresponding to the target programming language; the context representation enters the activated expert adapter for processing, and if at least two expert adapters are activated, the at least two activated expert adapters are calculated in parallel, and the initial language detailed features output by the at least two expert adapters are weightedly fused to obtain the language detailed features; after the context representation and the language detailed features are fused, they are processed by an autoregressive decoder, which first generates tokens and then generates codes token by token, and finally outputs the target code.

[0125] The method of the embodiment of the present application can be executed by a code generation model, which includes a general coding model, multiple independent expert adapters corresponding to programming languages, and a dynamic routing selector. In addition, the code generation model is configured with a front-end interface for content display and reception. The user enters natural language instructions or code snippets in the front-end interface, and the code model executes the solution of the embodiment of the present application to generate target code, which is displayed in the front-end interface. For example, if the user enters: "Write a quick sort function", the code generation model executes the following process: (1) The general coding model generates a context representation; (2) The dynamic routing selector predicts the target programming language, such as Python with a weight of 0.9 and Java with a weight of 0.1; (3) The Top-1 expert adapter (the expert adapter corresponding to Python) is activated and the generated features are refined; (4) Python code is generated based on the fused features and displayed in the front-end interface. In addition, if code generation for a new programming language (such as Rust) needs to be supported, the newly added update module and the newly added training sample data can be used to add the expert adapter corresponding to Rust and update the dynamic routing selector without retraining the general coding model.

[0126] Figure 4 FIG. 1 shows a functional structure diagram of a code generating device provided in an embodiment of the present application. Figure 4 As shown, the device includes:

[0127] A first feature extraction module 41 is adapted to process received input content using a universal coding model to obtain a contextual representation; wherein the universal coding model is trained using training sample data of multiple programming languages;

[0128] A selection module 42 is adapted to select a target expert adapter from a plurality of expert adapters according to the context representation; wherein the different expert adapters are trained using training sample data of different programming languages;

[0129] A second feature extraction module 43 is adapted to process the context representation using a target expert adapter to obtain detailed language features;

[0130] The generation module 44 is adapted to fuse the context representation and the detailed language features, and generate target code according to the fused features.

[0131] In an optional manner, the selection module 42 is further adapted to:

[0132] Determine the weight scores of various programming languages based on context representation;

[0133] A target programming language is determined from various programming languages according to the weight scores, and an expert adapter corresponding to the target programming language is determined as a target expert adapter.

[0134] In an optional manner, the device further comprises:

[0135] A training module is suitable for any programming language, and trains an initial universal coding layer according to training sample data of the programming language to obtain a first model; trains an initial second model according to training sample data corresponding to the programming language to obtain a second model; wherein the initial second model includes a third model consistent with the first model and an initial expert adapter, and during the training of the initial second model, the parameters of the third model are frozen while the parameters of the initial expert adapter are adjusted; and the initial expert adapter after parameter adjustment contained in the second model is determined as the expert adapter corresponding to the programming language.

[0136] In an optional manner, the training module is further adapted to:

[0137] Inputting the training input data in the training sample data into the first model and the initial second model for processing respectively;

[0138] Calculate the output layer loss based on the output probability distribution of the first model and the output probability distribution of the initial second model;

[0139] Obtain a first feature of the intermediate layer output of the first model and a second feature of the corresponding intermediate layer output in the initial second model, and calculate the intermediate layer loss based on the first feature and the second feature;

[0140] Calculate the task loss based on the predicted labels output by the initial second model and the true labels corresponding to the training input data;

[0141] The total loss is calculated based on the output layer loss, the intermediate layer loss, and the task loss, and the parameters of the initial expert adapter are adjusted based on the total loss.

[0142] In an optional manner, the second feature extraction module 43 is further adapted to:

[0143] The initial language detailed features obtained by processing at least two target expert adapters are weightedly summed to obtain the language detailed features.

[0144] In an optional manner, the selection module 42 is further adapted to:

[0145] The context representation is input into a dynamic routing selector for processing to obtain weight scores of various programming languages; wherein the loss function used for training the dynamic routing selector is determined according to the cross entropy loss function.

[0146] In summary, according to the code generation device provided by this embodiment, a general coding model is pre-trained through large-scale mixed programming language training sample data, common features of multiple programming languages are learned, and the received input content is processed using the general coding model to obtain a general context representation, so as to provide a cross-language semantic representation basis for the expert adapter. Expert adapters of different programming languages share the context representation of the general coding model, and repeated calculations can be avoided; a target expert adapter is selected from multiple expert adapters according to the context representation, and only the selected target expert adapter is loaded. By lightweight activation of some expert adapters instead of all expert adapters, the code generation speed can be improved and the computing consumption can be reduced; the general context representation extracted by the general coding model and the unique refined features extracted by the expert adapter are fused to generate the target code, and the general features between programming languages and the unique features of a specific programming language are fused for code generation, which can improve the correctness of the generated code.

[0147] An embodiment of the present application provides a non-volatile computer storage medium, which stores at least one executable instruction or computer program, and the executable instruction or computer program can enable a processor to perform operations corresponding to the code generation method in any of the above method embodiments.

[0148] An embodiment of the present application provides a computer program product, which includes at least one executable instruction or computer program, and the executable instruction or computer program can enable a processor to perform operations corresponding to the code generation method in any of the above method embodiments.

[0149] Figure 5 A schematic structural diagram of an embodiment of a computing device of the present application is shown. The specific embodiment of the present application does not limit the specific implementation of the computing device.

[0150] like Figure 5 As shown, the computing device may include: a processor (processor) 502 , a communications interface (Communications Interface) 504 , a memory (memory) 506 , and a communication bus 508 .

[0151] Processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508. Communication interface 504 is used to communicate with other devices, such as clients or other server network elements. Processor 502 is used to execute program 510, which may specifically perform the steps described in the aforementioned embodiment of the code generation method for a computing device.

[0152] Specifically, the program 510 may include program codes, which include computer operation instructions.

[0153] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the computing device may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.

[0154] The memory 506 is used to store the program 510. The memory 506 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0155] Program 510 can be specifically used to cause processor 502 to execute the code generation method in any of the above-mentioned method embodiments. The specific implementation of each step in program 510 can refer to the corresponding description of the corresponding steps and units in the code generation embodiment, and will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the above-mentioned devices and modules can refer to the corresponding process description in the above-mentioned method embodiments, and will not be repeated here.

[0156] The algorithm or demonstration provided here are not inherently relevant to any particular computer, virtual system or other equipment. Various general purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious that the structure required for constructing this type of system. In addition, the present application embodiment is not directed to any specific programming language yet. It should be understood that various programming languages can be utilized to realize the content of the present application described here, and the above description of specific languages is for the purpose of disclosing the best mode of implementation of the present application.

[0157] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.

[0158] Similarly, it should be understood that in order to streamline the present application and aid in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, various features of the embodiments of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, this disclosed method should not be interpreted as reflecting an intention that the claimed application requires more features than are expressly recited in each claim. Rather, as reflected in the claims, inventive aspects lie in less than all the features of the individual embodiments disclosed above. Accordingly, the claims that follow the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present application.

[0159] Those skilled in the art will appreciate that the modules in the devices in the embodiments may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments may be combined into one module or unit or component, and in addition may be divided into multiple submodules or subunits or subcomponents. All features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed herein may be combined in any combination, except that at least some of such features and / or processes or units are mutually exclusive. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature providing the same, equivalent or similar purpose.

[0160] Furthermore, those skilled in the art will appreciate that although some embodiments herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of this application and to form different embodiments. For example, in the claims, any of the claimed embodiments may be used in any combination.

[0161] The various component embodiments of the present application can be implemented in hardware, or in a software module running on one or more processors, or in a combination thereof. Those skilled in the art will appreciate that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of the present application. The application can also be implemented as a device or apparatus program (e.g., computer program and computer program product) for performing a part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0162] It should be noted that the above embodiments illustrate rather than limit the present application, and that a person skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of appropriately programmed computers. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names. The steps in the above embodiments should not be understood as limiting the order of execution unless otherwise specified.

Claims

1. A code generation method, comprising: Processing the received input content using a universal coding model to obtain a context representation; wherein the universal coding model is trained using training sample data of multiple programming languages; Selecting a target expert adapter from a plurality of expert adapters according to the context representation; wherein different expert adapters are trained using training sample data of different programming languages; Processing the context representation using the target expert adapter to obtain detailed language features; The context representation and the language detailed features are fused, and target code is generated based on the fused features.

2. The method according to claim 1, wherein The selecting a target expert adapter from a plurality of expert adapters according to the context representation further comprises: Determining weight scores for various programming languages based on the context representation; A target programming language is determined from various programming languages according to the weight scores, and an expert adapter corresponding to the target programming language is determined as a target expert adapter.

3. The method according to claim 1 or 2, wherein: The method further comprises: For any programming language, the initial universal coding layer is trained according to the training sample data of the programming language to obtain a first model; Training an initial second model according to the training sample data corresponding to the programming language to obtain a second model; The initial second model includes a third model consistent with the first model and an initial expert adapter, and during the training of the initial second model, parameters of the third model are frozen while parameters of the initial expert adapter are adjusted; The initial expert adapter after parameter adjustment included in the second model is determined as the expert adapter corresponding to the programming language.

4. The method according to claim 3, wherein: The training of the initial second model according to the training sample data corresponding to the programming language further includes: Inputting the training input data in the training sample data into the first model and the initial second model for processing respectively; Calculating an output layer loss based on the output probability distribution of the first model and the output probability distribution of the initial second model; Obtaining a first feature of an intermediate layer output of the first model and a second feature of a corresponding intermediate layer output in the initial second model, and calculating an intermediate layer loss based on the first feature and the second feature; Calculating the task loss based on the predicted label output by the initial second model and the true label corresponding to the training input data; A total loss is calculated according to the output layer loss, the intermediate layer loss, and the task loss, and parameters of the initial expert adapter are adjusted according to the total loss.

5. The method according to claim 1, wherein If the number of the target expert adapters is at least two, the processing the context representation using the target expert adapter to obtain detailed language features further includes: The initial language detailed features obtained by processing at least two target expert adapters are weightedly summed to obtain the language detailed features.

6. The method according to claim 2, wherein: Determining the weight scores of various programming languages according to the context representation further includes: The context representation is input into a dynamic routing selector for processing to obtain weight scores of various programming languages; wherein the loss function used in training the dynamic routing selector is determined according to a cross entropy loss function.

7. A code generation device comprising: a first feature extraction module adapted to process received input content using a universal coding model to obtain a contextual representation; wherein the universal coding model is trained using training sample data of multiple programming languages; a selection module adapted to select a target expert adapter from a plurality of expert adapters according to the context representation; wherein different expert adapters are trained using training sample data of different programming languages; a second feature extraction module adapted to process the context representation using the target expert adapter to obtain detailed language features; The generation module is adapted to fuse the context representation and the language detailed features, and generate target code according to the fused features.

8. A computing device comprising: A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the code generation method according to any one of claims 1 to 6.

9. A computer storage medium, wherein at least one executable instruction is stored in the storage medium, and wherein the executable instruction enables a processor to execute an operation corresponding to the code generation method according to any one of claims 1 to 6.

10. A computer program product, comprising at least one executable instruction, wherein the executable instruction enables a processor to execute operations corresponding to the code generation method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Code generation method and device, equipment and storage medium

    CN120743244A