Code representation model training method and device based on multi-task learning

By employing a multi-task learning-based code representation model training method, which combines a Transformer encoder and decoder with dynamic weight optimization, the problem of poor performance of existing models in multi-task processing is solved, achieving more efficient code representation and task transfer capabilities.

CN121094010AActive Publication Date: 2025-12-09SHENZHEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511650065.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2025-12-09
Estimated Expiration
2045-11-12

AI Technical Summary

Technical Problem

Existing code representation models perform poorly in multi-task processing, especially in task transfer and complex code understanding and generation tasks. Furthermore, the static weight mechanism leads to training imbalance, which affects the model's generalization ability.

Method used

A code representation model training method based on multi-task learning is adopted. The preprocessing module generates code tag sequences, linearized abstract syntax tree structure sequences, and natural language tag sequences. The Transformer encoder and decoder are used for feature extraction and decoding. The multi-task loss is optimized by dynamic weight adjustment to achieve balanced optimization of the model across different tasks.

Benefits of technology

It improves the efficiency and generalization ability of the code representation model in multi-task processing, ensures the balanced optimization of the model across different tasks, and enhances the performance of code understanding and generation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121094010A_ABST
    Figure CN121094010A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of artificial intelligence, and provides a code representation model training method and device based on multi-task learning, and the code representation model comprises a preprocessing module, a code task encoder and a code task decoder. The method comprises the steps that code samples of a plurality of target code tasks are serialized through a preprocessing module, code marking sequences, linearized AST structure sequences and NL marking sequences of the code samples are obtained, and code representation sequences are generated; performing feature extraction on the code representation sequence by using a code task encoder to obtain deep semantic representation; decoding the deep semantic representation by using a code task decoder to obtain a corresponding decoding sequence; calculating the task learning loss and the dynamic weight of each target code task, further calculating the total learning loss of the current training, updating the parameters of the code representation model according to the total learning loss, and continuing to train the code representation model until the training is finished.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, and in particular relates to a training method and apparatus for a code representation model based on multi-task learning. Background Technology

[0002] With the continuous advancement of computer technology and the ever-expanding application scenarios of software, the scale of software projects is growing explosively, and system structures are becoming increasingly complex. Modern software not only boasts numerous functions but also places higher demands on performance, reliability, and security, posing more severe challenges to developers in the process of code understanding, writing, and maintenance. Especially in large-scale projects, the sheer volume of code, the complex dependencies between modules, and the abstract and variable nature of the code further increase the difficulty of development and maintenance. Therefore, how to efficiently and accurately achieve code understanding and generation has become a core problem that urgently needs to be addressed in the fields of software engineering and intelligent development.

[0003] The rapid development of deep learning technology, especially the successful application of Transformer models and their variants in natural language processing, has further driven innovation in the field of code intelligence. Researchers have gradually introduced deep learning methods into code representation and generation tasks, utilizing models to automatically learn latent patterns in code sequences, abstract syntax trees (ASTs), and natural language annotations, thus achieving significant results in various code understanding and generation tasks. However, most current mainstream Transformer models adopt a single-task learning paradigm, that is, training and optimizing the model for a single task. While this approach can achieve good results on specific tasks, the model often exhibits significant performance degradation when transferring to other tasks or facing new related tasks. Even code representation models that perform well on a single task may experience a significant performance drop when transferred to related tasks, revealing the inherent limitations of single-task training in terms of representation generality and transferability.

[0004] On the other hand, while code representation models for multi-task learning have begun to attempt joint optimization of multiple tasks, they generally rely on static weight mechanisms to allocate task losses. This approach cannot fully consider the differences in convergence speed and optimization difficulty among different tasks during training. As a result, tasks with fast convergence often dominate the training process, while the optimization of slow-converging tasks is suppressed, leading to unbalanced model training and difficulty in achieving true multi-task collaborative optimization. This not only limits the model's performance in multi-task environments but also affects the generalization ability of code representation, causing the model to underperform when faced with complex code understanding and generation tasks. Summary of the Invention

[0005] The purpose of this invention is to provide a training method and apparatus for a code representation model based on multi-task learning, which aims to solve the problem of poor multi-task processing performance of existing code representation models.

[0006] In a first aspect, the present invention provides a training method for a code representation model based on multi-task learning. The code representation model includes a preprocessing module, a code task encoder, and a code task decoder. The code task encoder is a Transformer encoder, and the code task decoder is a Transformer decoder. The method includes the following steps: The preprocessing module is used to serialize code samples from multiple target code tasks to obtain code tag sequences, linearized abstract syntax tree structure sequences, and natural language tag sequences. Based on the code tag sequences, linearized abstract syntax tree structure sequences, and natural language tag sequences, a code representation sequence for each target code task is generated. The code representation sequence of each target code task is input into the corresponding code task encoder, and the code task encoder is used to extract features from the code representation sequence to obtain deep semantic representation. The deep semantic representation output by each code task encoder is input into the corresponding code task decoder, and the deep semantic representation is decoded by the code task decoder to obtain the corresponding decoding sequence. Based on the code samples and target code of the multiple target code tasks, the task learning loss and dynamic weights of each target code task are calculated. According to the task learning loss and dynamic weights of each target code task, the total learning loss of the current training is calculated. According to the total learning loss, the parameters of the code task encoder and code task decoder are updated, and the code representation model is trained again until the training ends.

[0007] In some embodiments, the step of using the preprocessing module to serialize code samples from multiple target code tasks to obtain a code tag sequence, a linearized abstract syntax tree structure sequence, and a natural language tag sequence includes: Lexical analysis is performed on the code sample to obtain the code tag sequence of the code sample; Obtain the abstract syntax tree and natural language tag sequence of the code sample, and perform linearization on the abstract syntax tree to obtain the linearized abstract syntax tree structure sequence of the code sample.

[0008] In some embodiments, the number of code task encoders and code task decoders is the same, and all code task encoders share parameters.

[0009] In some embodiments, the step of inputting the deep semantic representation output by each code task encoder into the corresponding code task decoder, and using the code task decoder to decode the deep semantic representation to obtain the corresponding decoded sequence includes: The deep semantic representation is input into the code task decoder of the target code task associated with the deep semantic representation, and the code task decoder is used to decode the deep semantic representation to obtain the corresponding decoded sequence.

[0010] In some embodiments, the step of calculating the task learning loss and dynamic weights for each target code task based on code samples and target code of the plurality of target code tasks includes: Based on the code samples and target code of the multiple target code tasks, calculate the original task learning loss for each target code task; The original task learning loss for each target code task is smoothed, and the smoothed original task learning loss is set as the task learning loss for each target code task. Based on the task learning loss for each target code task, the learning speed for each target code task is calculated. Then, based on the calculated learning speed and the total number of target code tasks, the dynamic weight for each target code task is calculated.

[0011] In some embodiments, the following formula is used: Calculate the learning speed for each target code task, where, This represents the learning speed of the target code task k at time t. This represents the task learning loss of target code task k at time t. Let represent the task learning loss of target code task k at time t in the previous time step. This represents a pre-defined minimum constant; Using the following formula: Calculate the dynamic weights for each target code task, where, This represents the dynamic weight of the target code task k at time t. Indicates the total number of target code tasks. This indicates the preset temperature coefficient.

[0012] Secondly, the present invention provides a method for processing target code tasks, the method comprising the following steps: Receive the source code to be processed and the corresponding target code task from the user; The source code to be processed and the target code task are input into the trained code representation model to obtain the target code of the source code to be processed. The code representation model is a code representation model trained by any of the training methods described above.

[0013] Thirdly, the present invention provides a training device for a code representation model based on multi-task learning, wherein the code representation model includes a preprocessing module, a code task encoder, and a code task decoder, wherein the code task encoder is a Transformer encoder, and the code task decoder is a Transformer decoder, and the training device includes: The preprocessing unit is used to serialize code samples of multiple target code tasks using the preprocessing module to obtain code tag sequences, linearized abstract syntax tree structure sequences, and natural language tag sequences of the code samples. Based on the code tag sequences, linearized abstract syntax tree structure sequences, and natural language tag sequences, a code representation sequence for each target code task is generated. The feature extraction unit is used to input the code representation sequence of each target code task into the corresponding code task encoder, and use the code task encoder to extract features from the code representation sequence to obtain deep semantic representation. A sequence decoding unit is used to input the deep semantic representation output by each code task encoder into the corresponding code task decoder, and use the code task decoder to decode the deep semantic representation to obtain the corresponding decoded sequence; and The parameter update unit is used to calculate the task learning loss and dynamic weights for each target code task based on the code samples and target code of the multiple target code tasks, calculate the total learning loss of the current training based on the task learning loss and dynamic weights of each target code task, update the parameters of the code task encoder and code task decoder based on the total learning loss, and continue to train the code representation model until the training ends.

[0014] Fourthly, the present invention also provides a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described above.

[0015] Fifthly, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described above.

[0016] In this embodiment of the invention, after obtaining the code representation sequence for each target code task, the code representation sequence is input into the corresponding code task encoder. The code task encoder extracts features from the code representation sequence to obtain a deep semantic representation. The code task decoder, which processes the same task, decodes the corresponding deep semantic representation to obtain a decoded sequence. Based on code samples and target code from multiple target code tasks, the task learning loss and dynamic weights for each target code task are calculated. Based on the task learning loss and dynamic weights for each target code task, the total learning loss for the current training is calculated. Based on the total learning loss, the parameters of the code task encoder and code task decoder are updated. Then, the code representation model continues to be trained, and finally a trained code representation model is obtained. Thus, multiple code tasks are optimized simultaneously during the training process of the code representation model, and the weight ratio of each code task in the total learning loss is dynamically adjusted through dynamic weights, so that the model maintains an optimal balance among different tasks, thereby improving the multi-task processing efficiency of the code representation model. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the training method for a code representation model based on multi-task learning provided in Embodiment 1 of the present invention. Figure 2 This is a flowchart illustrating step S104 in Embodiment 1 provided in Embodiment 5 of the present invention; Figure 3 This is a flowchart illustrating the training method for a code representation model based on multi-task learning provided in Embodiment Six of the present invention. Figure 4 This is a flowchart illustrating the method for processing target code tasks provided in Embodiment 7 of the present invention; Figure 5 This is a schematic diagram of the structure of the training device for the code representation model based on multi-task learning provided in Embodiment 8 of the present invention; Figure 6 This is a schematic diagram of the structure of the computing device provided in Embodiment 9 of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0019] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described feature, integral, step, operation, element, and / or component, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof. Furthermore, the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. The terms "first," "second," and similar words do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Words such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Above," "below," "left," "right," etc., are used only to indicate relative positional relationships, which may change accordingly when the absolute position of the described object changes.

[0020] To keep the following description of the embodiments of the present invention clear and concise, detailed descriptions of some known functions and known components are omitted in this specification.

[0021] The specific implementation of the present invention will be described in detail below with reference to specific embodiments: Example 1: This invention provides a training method for a code representation model based on multi-task learning. This code representation model can handle multiple code tasks, such as code completion, code translation, and code vulnerability repair. The model includes a preprocessing module, a code task encoder, and a code task decoder, and may also include other processing modules. The code task encoder is a Transformer encoder, and the code task decoder is a Transformer decoder.

[0022] As shown in the figure Figure 1 The implementation flow of the training method for the code representation model is shown. For ease of explanation, only the parts related to the embodiments of the present invention are shown, and are described in detail below.

[0023] In step S101, the preprocessing module is used to serialize the code samples of multiple target code tasks to obtain the code tag sequence, the linearized abstract syntax tree structure sequence, and the natural language tag sequence of the code samples. Based on the code tag sequence, the linearized abstract syntax tree structure sequence, and the natural language tag sequence, the code representation sequence of each target code task is generated.

[0024] In this embodiment of the invention, the target code task is the operation performed on the input source code, such as code completion, code translation, and code vulnerability repair. The code sample includes multiple source codes, each source code corresponding to a target code task, so as to realize multi-task learning of the code representation model through the code sample.

[0025] In step S102, the code representation sequence of each target code task is input into the corresponding code task encoder, and the code task encoder is used to extract features from the code representation sequence to obtain deep semantic representation.

[0026] In step S103, the deep semantic representation output by each code task encoder is input into the corresponding code task decoder, and the deep semantic representation is decoded by the code task decoder to obtain the corresponding decoded sequence.

[0027] In this embodiment of the invention, the number of code task encoders and code task decoders is the same, and all code task encoders share parameters. Each code task encoder corresponds to one code task, used to extract features from the input code representation sequence to obtain deep semantic representations. Similarly, each code task decoder also corresponds to one code task, used to decode the deep semantic representations. One code task encoder is associated with one code task decoder, and both process the same code task.

[0028] In step S104, based on code samples and target code of multiple target code tasks, the task learning loss and dynamic weights of each target code task are calculated. Based on the task learning loss and dynamic weights of each target code task, the total learning loss of the current training is calculated. Based on the total learning loss, the parameters of the code task encoder and code task decoder are updated, and the code representation model is trained again until the training ends.

[0029] In this embodiment of the invention, the target code is the target code (answer) obtained after the code sample (problem) has been processed by the code task. Using the code samples and target code of the multiple target code tasks, the task learning loss of each target code task can be calculated. Based on the calculated task learning loss of each target code task, the dynamic weight of each target code task can be further calculated, so that the weight of the code task that converges quickly will be reduced, while the weight of the code task that converges slowly will be increased, thereby preventing individual code tasks from dominating the parameter update of the code representation model and ensuring the training stability of the code representation model.

[0030] Example 2 In this embodiment of the invention, in step S101 of embodiment one, when using the preprocessing module to serialize code samples of multiple target code tasks, lexical analysis can be performed on the code samples to obtain the code token sequence, the abstract syntax tree (Linearized AST) and natural language token sequence (NL token) of the code samples can be obtained, the abstract syntax tree can be linearized to obtain the linearized abstract syntax tree structure sequence of the code samples, and finally the code token sequence, the linearized abstract syntax tree structure sequence and the natural language token sequence of the code samples can be obtained.

[0031] Specifically, when obtaining the code tag sequence of code samples, a tokenizer performs lexical analysis on code samples from multiple target code tasks to extract the code tag sequence, reflecting the basic syntactic composition of the code. When obtaining the linearized abstract syntax tree (AST) structure sequence of code samples, an AST parser is used to parse the code samples, obtaining the AST. Then, based on an XML structure traversal method, the AST is linearized to obtain the linearized AST structure sequence, which characterizes the structural information of the code. When obtaining the natural language tag sequence of code samples, named entity recognition and invocation expression recognition are used to generate the natural language tag sequence of the code samples, expressing the semantic intent contained in the code.

[0032] After obtaining the code tag sequence, the linearized abstract syntax tree structure sequence, and the natural language tag sequence, a code representation sequence for each target code task is generated based on these sequences. Specifically, the three types of sequences can be concatenated using a special delimiter to obtain the code representation sequence for the target code task, thereby achieving comprehensive modeling of the code sample features of multiple target code tasks.

[0033] Example 3 In this embodiment of the invention, step S102 in Embodiment 1 can be implemented in the following way: (1) The input code representation sequence is vectorized so that discrete code tags are mapped into continuous low-dimensional semantic vectors through word embedding matrix, and positional encoding is introduced to preserve the order information of code statements, so as to obtain the initial code representation with semantic and positional information; (2) Context dependency modeling is performed on the initial code representation. By calculating the attention weights between different sequences, the semantic relationships in the global scope are captured, thereby obtaining the code feature representation containing structural and contextual information; (3) Perform nonlinear mapping and dimensional transformation on the obtained code feature representation to further refine high-order semantic information, enhance the model's ability to abstract complex code semantics, and obtain intermediate feature representation after feature transformation; (4) Normalize the intermediate feature representation and use residual connections to retain the original feature information to prevent gradient vanishing and improve training stability, so as to obtain high-dimensional context features, which is the output of the code task encoder (deep semantic representation).

[0034] Example 4 In this embodiment of the invention, step S103 in embodiment one can be implemented in the following way: inputting the deep semantic representation into the code task decoder of the target code task associated with the deep semantic representation, and using the code task decoder to decode the deep semantic representation to obtain the corresponding decoding sequence.

[0035] Specifically, the deep semantic representation is input into the code task decoder of the target code task associated with the deep semantic representation, and the deep semantic representation is decoded using the code task decoder, including: (1) The target code of the code sample is vectorized, and discrete symbols are mapped into continuous semantic vectors through the embedding matrix. The positional encoding is combined to preserve the sequence position information, so as to obtain the initial input representation that can be used for subsequent decoding. (2) Perform internal dependency modeling on the initial input representation. By calculating the attention relationship between each sequence in the initial input representation, capture the semantic association within the generation context and achieve effective modeling of the target code; (3) Interact the deep semantic representation output by the encoder with the current state of the decoder, calculate the attention weight between the decoder query vector and the encoder key vector, and perform semantic alignment and information fusion between the code sample and the target code to obtain the fused context features, thereby guiding the decoder to generate output that conforms to semantic logic. (4) Perform nonlinear mapping and semantic transformation on the fused context features to obtain the intermediate feature representation after feature extraction; (5) Normalize the intermediate feature representation and use residual connections to maintain gradient flow and semantic continuity to obtain a high-dimensional vector representation, which is the output of the code task decoder (decoding sequence).

[0036] Example 5 In embodiments of the present invention, such as Figure 2 As shown in the figure, the implementation method of step S104 in Embodiment 1 is illustrated, and is described in detail below: In step S201, based on code samples and target code of multiple target code tasks, the original task learning loss of each target code task is calculated; In step S202, the original task learning loss of each target code task is smoothed, and the smoothed original task learning loss is set as the task learning loss of each target code task. In step S203, based on the task learning loss of each target code task, the learning speed of each target code task is calculated, and the dynamic weight of each target code task is calculated according to the calculated learning speed and the total number of target code tasks. In step S204, the total learning loss of the current training is calculated based on the task learning loss and dynamic weights of each target code task. Based on the total learning loss, the parameters of the code task encoder and code task decoder are updated, and the code representation model continues to be trained until the training ends.

[0037] In this embodiment of the invention, when calculating the total learning loss of the code representation model, based on code samples and target code of multiple target code tasks, the task learning loss and dynamic weight of each target code task are calculated. Based on the task learning loss and dynamic weight of each target code task, the total learning loss of the current training is calculated. Based on the total learning loss, the weight ratio of each code task in the total learning loss is dynamically adjusted through dynamic weights, so that the model maintains an optimal balance among different tasks and improves the multi-task processing efficiency of the code representation model.

[0038] In some embodiments, the original task learning loss for each target code task is smoothed using the following formula: in, Represents the target code task At any moment The learning loss of the original task, This is a smoothing coefficient used to control the decay rate of historical losses. Represents the target code task At any moment The learning loss of the original task, Representation (target code task) At any moment The original task learning loss after smoothing.

[0039] In this embodiment of the invention, by smoothing the original task learning loss of each target code task, the noise impact during the training process of the code representation model is reduced, and the training efficiency of the code representation model is improved.

[0040] In some embodiments, the following formula is used: Calculate the learning speed for each target code task, where, This represents the learning speed of the target code task k at time t. This represents the task learning loss of target code task k at time t. Let represent the task learning loss of target code task k at time t in the previous time step. This represents a preset minimum constant.

[0041] In this embodiment of the invention, the above formula can accurately characterize the learning speed or difficulty of each target code task, thereby improving the accuracy of subsequent dynamic weight adjustment.

[0042] In some embodiments, the following formula is used: Calculate the dynamic weights for each target code task, where, This represents the dynamic weight of the target code task k at time t. Indicates the total number of target code tasks. This indicates the preset temperature coefficient.

[0043] In this embodiment of the invention, the above formula enables dynamic adjustment of the weight ratio of each code task in the total loss function, so that the model maintains an optimal balance among different tasks and improves the multi-task processing efficiency of the code representation model.

[0044] In some embodiments, the task learning loss for each code task is obtained. and their corresponding weights Then, the total learning loss for the current training can be calculated using the following formula: The total learning loss serves as the optimization objective for the code representation model at each training step, guiding the update process of the code representation model parameters and thus enabling collaborative learning across multiple code tasks.

[0045] When determining whether to terminate the training of the code representation model, specifically, the model training iterates a maximum of 50 times and employs an early stopping strategy: if the performance on the validation set does not improve within 20 consecutive times, training is terminated early to reduce computational overhead and avoid overfitting. After each training iteration, the total learning loss is calculated, and the training termination condition is determined. Training ends when the iteration limit is reached or the early stopping condition is met, and the trained code representation model is obtained; otherwise, training of the code representation model continues.

[0046] Example 6 In an embodiment of the present invention, as an example, Figure 3 The flowchart of the training method for a code representation model based on multi-task learning is shown.

[0047] As shown in the figure, in the preprocessing stage, code samples from multiple target code tasks are serialized to obtain code tag sequences, linearized abstract syntax tree (LST) structure sequences, and natural language tag sequences. Then, a concatenation operation is performed on these sequences to generate a code representation sequence for each target code task. In the encoding stage, the code representation sequence for each target code task is input into the corresponding code task encoder (e.g., first, second, and third task encoders). The encoders extract features from the code representation sequences to obtain deep semantic representations. In the decoding stage, the deep semantic representations output by each code task encoder are input into code task decoders (e.g., first, second, and third task decoders) that process the same task. The decoders decode the deep semantic representations to obtain the corresponding decoded sequences. In the learning loss calculation stage, based on the code samples and target code from multiple target code tasks, the task learning loss for each target code task is calculated (e.g., ...). L 1. L 2. L 3) and dynamic weights (e.g., , , Based on the task learning loss and dynamic weights for each target code task, the total learning loss for the current training is calculated. Based on the total learning loss, the parameters of the code task encoder and code task decoder are updated through backpropagation, and the code representation model continues to be trained until the training ends.

[0048] In this embodiment of the invention, the first, second, and third task decoders are all Transformer encoders, and the parameters of the first, second, and third task decoders are the same.

[0049] Example 7 Figure 4 The implementation flow of the target code task processing method provided by the embodiment of the present invention is shown. For ease of explanation, only the parts related to the embodiment of the present invention are shown, and are described in detail below.

[0050] In step S401, the user inputs the source code to be processed and the corresponding target code task; In step S402, the source code to be processed and the target code task are input into the trained code representation model to obtain the target code of the source code to be processed. The code representation model is a code representation model trained according to the training method described in any of the above embodiments.

[0051] Example 8: Figure 5 The structure of the training device for the code representation model based on multi-task learning provided in Embodiment 8 of the present invention is shown. For ease of explanation, only the parts related to the embodiments of the present invention are shown.

[0052] In this embodiment of the invention, the code representation model includes a preprocessing module, a code task encoder, and a code task decoder. The code task encoder is a Transformer encoder, and the code task decoder is a Transformer decoder. The training device includes: Preprocessing unit 51 is used to serialize code samples of multiple target code tasks using the preprocessing module to obtain code tag sequence, linearized abstract syntax tree structure sequence and natural language tag sequence of code samples, and generate code representation sequence for each target code task based on code tag sequence, linearized abstract syntax tree structure sequence and natural language tag sequence. The feature extraction unit 52 is used to input the code representation sequence of each target code task into the corresponding code task encoder, and use the code task encoder to extract features from the code representation sequence to obtain deep semantic representation. Sequence decoding unit 53 is used to input the deep semantic representation output by each code task encoder into the corresponding code task decoder, and use the code task decoder to decode the deep semantic representation to obtain the corresponding decoded sequence; and The parameter update unit 54 is used to calculate the task learning loss and dynamic weights for each target code task based on code samples and target code from multiple target code tasks. Based on the task learning loss and dynamic weights for each target code task, the total learning loss for the current training is calculated. Based on the total learning loss, the parameters of the code task encoder and code task decoder are updated, and the code representation model is trained again until the training ends.

[0053] In this embodiment of the invention, for the sake of convenience and brevity, only the division of the above-described functional units and modules is used as an example. In practical applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to achieve all or part of the functions described above. Each unit and module of the device can be implemented by corresponding hardware or software units. Each unit and module can be an independent hardware or software unit, or it can be integrated into a single hardware or software unit, which is not intended to limit the invention. In addition, the specific names of each functional unit and module are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the device can be referred to the corresponding description in the foregoing method embodiments, and will not be repeated here.

[0054] Example 9 Figure 6 The structure of the computing device provided in Embodiment 9 of the present invention is shown. For ease of explanation, only the parts related to the embodiments of the present invention are shown.

[0055] The computing device 6 of this embodiment includes a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60. When the processor 60 executes the computer program 62, it implements the steps in the training method embodiments of the various code representation models described above, for example... Figure 1 Steps S101 to S104 are shown. Alternatively, when the processor 60 executes the computer program 62, it implements the functions of each unit in the above-described device embodiment, for example... Figure 5 The functions of units 51 to 54 are shown.

[0056] The computing device 6 in this embodiment of the invention can be a personal computer, a server, or the like. The steps implemented by the processor 60 in this computing device 6 when executing the computer program 62 to implement the training method for the code representation model can be referred to the description of the foregoing method embodiments, and will not be repeated here.

[0057] Example 10 In this embodiment of the invention, a computer-readable storage medium is provided, which stores a computer program. When executed by a processor, the computer program implements the steps in the above-described training method embodiment for the code representation model. For example... Figure 1 Steps S101 to S104 are shown. Alternatively, when the computer program is executed by a processor, it implements the functions of each unit in the above-described device embodiment, for example... Figure 5 The functions of units 51 to 54 are shown.

[0058] The computer-readable storage medium of this invention can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EEPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0059] The above embodiments are merely illustrative of the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the scope of disclosure involved in the above embodiments is not limited to technical solutions formed by specific combinations of the above technical features, but should also cover other technical solutions formed by arbitrary combinations of the above technical features or their equivalent features without departing from the above-disclosed concept. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0060] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the invention. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

Claims

1. A training method for a code representation model based on multi-task learning, characterized in that, The code representation model includes a preprocessing module, a code task encoder, and a code task decoder. The code task encoder is a Transformer encoder, and the code task decoder is a Transformer decoder. The method includes the following steps: The preprocessing module is used to serialize code samples from multiple target code tasks to obtain code tag sequences, linearized abstract syntax tree structure sequences, and natural language tag sequences. Based on the code tag sequences, linearized abstract syntax tree structure sequences, and natural language tag sequences, a code representation sequence for each target code task is generated. The code representation sequence of each target code task is input into the corresponding code task encoder, and the code task encoder is used to extract features from the code representation sequence to obtain deep semantic representation. The deep semantic representation output by each code task encoder is input into the corresponding code task decoder, and the deep semantic representation is decoded by the code task decoder to obtain the corresponding decoding sequence. Based on the code samples and target code of the multiple target code tasks, the task learning loss and dynamic weights of each target code task are calculated. According to the task learning loss and dynamic weights of each target code task, the total learning loss of the current training is calculated. According to the total learning loss, the parameters of the code task encoder and code task decoder are updated, and the code representation model is trained again until the training ends.

2. The method as described in claim 1, characterized in that, The steps of using the preprocessing module to serialize code samples from multiple target code tasks to obtain a code tag sequence, a linearized abstract syntax tree structure sequence, and a natural language tag sequence include: Lexical analysis is performed on the code sample to obtain the code tag sequence of the code sample; Obtain the abstract syntax tree and natural language tag sequence of the code sample, and perform linearization on the abstract syntax tree to obtain the linearized abstract syntax tree structure sequence of the code sample.

3. The method as described in claim 1, characterized in that, The number of code task encoders and code task decoders is the same, and all code task encoders share parameters.

4. The method as described in claim 1, characterized in that, The steps of inputting the deep semantic representation output by each code task encoder into the corresponding code task decoder, and using the code task decoder to decode the deep semantic representation to obtain the corresponding decoded sequence include: The deep semantic representation is input into the code task decoder of the target code task associated with the deep semantic representation, and the code task decoder is used to decode the deep semantic representation to obtain the corresponding decoded sequence.

5. The method as described in claim 1, characterized in that, Based on the code samples and target code of the multiple target code tasks, the steps for calculating the task learning loss and dynamic weights for each target code task include: Based on the code samples and target code of the multiple target code tasks, calculate the original task learning loss for each target code task; The original task learning loss for each target code task is smoothed, and the smoothed original task learning loss is set as the task learning loss for each target code task. Based on the task learning loss for each target code task, the learning speed for each target code task is calculated. Then, based on the calculated learning speed and the total number of target code tasks, the dynamic weight for each target code task is calculated.

6. The method as described in claim 5, characterized in that, Using the following formula: Calculate the learning speed for each target code task, where, This represents the learning speed of the target code task k at time t. This represents the task learning loss of target code task k at time t. Let represent the task learning loss of target code task k at time t in the previous time step. This represents a pre-defined minimum constant; Using the following formula: Calculate the dynamic weights for each target code task, where, This represents the dynamic weight of the target code task k at time t. Indicates the total number of target code tasks. This indicates the preset temperature coefficient.

7. A method for processing target code tasks, characterized in that, The method includes the following steps: Receive the source code to be processed and the corresponding target code task from the user; The source code to be processed and the target code task are input into the trained code representation model to obtain the target code of the source code to be processed. The code representation model is a code representation model trained according to any one of the training methods according to claims 1 to 6.

8. A training device for a code representation model based on multi-task learning, characterized in that, The code representation model includes a preprocessing module, a code task encoder, and a code task decoder. The code task encoder is a Transformer encoder, and the code task decoder is a Transformer decoder. The training device includes: The preprocessing unit is used to serialize code samples of multiple target code tasks using the preprocessing module to obtain code tag sequences, linearized abstract syntax tree structure sequences, and natural language tag sequences of the code samples. Based on the code tag sequences, linearized abstract syntax tree structure sequences, and natural language tag sequences, a code representation sequence for each target code task is generated. The feature extraction unit is used to input the code representation sequence of each target code task into the corresponding code task encoder, and use the code task encoder to extract features from the code representation sequence to obtain deep semantic representation. A sequence decoding unit is used to input the deep semantic representation output by each code task encoder into the corresponding code task decoder, and use the code task decoder to decode the deep semantic representation to obtain the corresponding decoded sequence; and The parameter update unit is used to calculate the task learning loss and dynamic weights for each target code task based on the code samples and target code of the multiple target code tasks, calculate the total learning loss of the current training based on the task learning loss and dynamic weights of each target code task, update the parameters of the code task encoder and code task decoder based on the total learning loss, and continue to train the code representation model until the training ends.

9. A computing device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Code abstract automatic generation method and device

    CN111651198A

  • Combination optimization automatic modeling-oriented training data generation method and system

    CN119294403A

  • Multi-modal feature fusion software supply chain vulnerability intelligent positioning method

    CN120068095A

  • Video-based webpage generation method, system, equipment and medium

    CN120469686A