Code generation method and system fusing two-dimensional code representation and dependent coding
By modeling code snippets into two-dimensional structures and using sparse autoencoder and self-attention mechanisms, the problem of insufficient generalization ability in code generation is solved, and more accurate code generation and long sequence understanding is achieved.
Patent Information
- Application Number
- CN202510295492.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The existing code generation model lacks generalization ability when dealing with code structures and long sequences, ignoring the two-dimensional structure of the code and the dependencies between lines, resulting in inaccurate generation results.
The code snippet is modeled into a two-dimensional structure, and the sparse autoencoder is used for inter-line dependency encoding, combining the self-attention mechanism and multi-layer perceptron to predict the intermediate representation of characters, improving the generalization ability and structural understanding ability of the code generation model.
It improves the modeling ability of the code generation model to model code structure and relationships, enhances the long context processing ability, and improves the accuracy and rationality of generated code.
Smart Images

Figure CN120255874A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a code generation method and system that integrates two-dimensional code representation and dependency encoding. Background Art
[0002] Existing code generation models, such as Codex, CodeLlama, etc., are usually based on traditional natural language processing (NLP) models and are obtained through incremental training on large-scale code datasets; these models regard code snippets as one-dimensional byte sequences and capture the sequential information in the sequence through positional encoding (such as RoPE, ALiBi, etc.); however, this method also has some obvious deficiencies:
[0003] 1. Poor generalization of positional encoding: Although the positional encoding index can effectively capture the relative positional relationship between characters, it ignores the potential translational invariance of the code itself in some aspects; specifically, functions, classes, or modules defined in many code snippets, although in different positions in the code, are equivalent in semantics and function; therefore, changing the positions of these code snippets does not affect the execution logic of the program; however, the existing positional encoding methods can only learn the distribution patterns between fixed positions, resulting in limited generalization ability of the model when faced with changes in code structure or order.
[0004] 2. Lack of code structure information: Existing code generation models usually regard code as a one-dimensional linear sequence and ignore the essential structural features of the code; code is not just a linear string composed of individual characters or tokens, but has an obvious hierarchical structure, including functions, classes, conditional statements, etc., and the logical relationships between these elements constitute the function and execution flow of the program; traditional serialization processing methods fail to effectively capture the inter-line logical flow and intra-line operations in the code, making it difficult for the model to understand the overall structure of the code, thereby affecting the correctness and rationality of the generation results.
[0005] 3. Insufficient long sequence processing ability: The performance of traditional positional encoding methods usually drops significantly when processing long sequences that exceed the training length; this is mainly because traditional positional encoding methods usually represent positions based on fixed-length indexes, and when the length of the input sequence exceeds the training length, the model cannot model the positional relationships between characters beyond the training length; therefore, for tasks that need to process long code snippets or complex code structures, such as warehouse-level code generation tasks, the effect of traditional positional encoding methods is relatively limited. Summary of the Invention
[0006] The object of the present invention is to overcome the deficiencies in the prior art, and provide a code generation method and system that integrates two-dimensional code representation and dependency encoding. By modeling code snippets as two-dimensional structures and combining sparse autoencoders (SAEs) for inter-line dependency encoding, the generalization ability, structural understanding ability, and long-context processing ability of the code generation model are improved, which can effectively enhance the model's ability to model code structures and relationships, thereby improving the model's performance.
[0007] To solve the above technical problems, the present invention is implemented by adopting the following technical solutions:
[0008] In a first aspect, the present invention provides a code generation method that integrates two-dimensional code representation and dependency encoding, including:
[0009] Model the obtained code data as a two-dimensional structure to obtain a two-dimensional code;
[0010] Input the two-dimensional code into a trained code generation model: perform word embedding on the two-dimensional code through an embedding layer to obtain word embedding vectors; perform inter-line masking on the two-dimensional code through an intra-line attention mask to obtain a mask matrix; perform dependency modeling on the word embedding vectors through a sparse autoencoder to obtain a dependency encoding of the code inter-line dependency relationship; perform attention calculation using the self-attention mechanism based on the word embedding vectors, mask matrix, and dependency encoding to obtain an intermediate representation of each character; perform model prediction through a multi-layer perceptron and a softmax layer based on the intermediate representation of each character to obtain a generated code.
[0011] Optionally, modeling the obtained code data as a two-dimensional structure includes:
[0012] Obtain the positions of each line break character in the code data to obtain line break position encodings , where represents the line break serial number;
[0013] Based on the line break position encodings , for any line index , if , then it is considered that the character belongs to the same line, where represents the line index, represents the position of the character in the code data, and represent the positions of two adjacent code lines.
[0014] Optionally, the mask matrix M is calculated by the following formula:
[0015]
[0016] Among them, represents the accessibility between character and character . By setting a value as small as possible, the attention score between the two characters after softmax calculation is 0.
[0017] Optionally, the data processing process of the dependency modeling includes:
[0018] Obtain the word embedding of the tail character of the code line as the semantic anchor of the code line , where the characters of the code line include ; ;
[0019] According to the semantic anchor , calculate the dependency encoding of the code line through the sparse autoencoder SAE. The specific calculation is as follows:
[0020]
[0021] Among them, is the learnable parameter in the sparse autoencoder SAE, represents the encoder in the sparse autoencoder SAE, which is used to convert the input into a feature activation value, represents the feature dictionary, and activates the features in the dictionary according to the feature activation value, represents the bias value, represents the activation function, represents the code line activation value of the dictionary feature of, represents the code line dependency encoding of.
[0022] Optionally, the dependency modeling includes distance-based activation value enhancement, which is achieved through the following formula:
[0023]
[0024] Among them, represents the activation value of the dictionary feature of the code line , represents the clipping operation, taking the larger value of and , represents the input length of the model, represents the position of the semantic anchor .
[0025] Optionally, the data processing process of the attention calculation includes:
[0026] The attention scores between characters are calculated according to the word embedding vector, dependency encoding and mask matrix M, specifically through the following formula:
[0027]
[0028] in, Representing characters word embedding, Representing characters The dependency code of the line of code, Represents characters after relying on encoding word embedding, Representing characters word embedding, Representing characters The dependency code of the line of code, Represents characters after relying on encoding word embedding, represents the key vector in the self-attention mechanism, represents the query mapping matrix, represents the value vector in the self-attention mechanism, represents the key mapping matrix, Representing characters and characters The attention score between Representing characters and characters accessibility between;
[0029] According to the attention scores between the characters, the intermediate representation of each character is calculated, specifically through the following formula:
[0030]
[0031] in, Representing characters The intermediate representation generated by the self-attention mechanism, represents the value vector in the self-attention mechanism, represents the value mapping matrix, Represents the normalization function.
[0032] Optionally, the self-attention mechanism alleviates the problem of distraction through MLP, specifically through the following formula:
[0033]
[0034] in, Represents a character and a character The attention score between them Represents a set of multi-layer perceptrons Represents a dimensionality concatenation operation
[0035] Optionally, the model prediction by the multi-layer perceptron and the softmax layer is achieved through the following formula
[0036]
[0037] where Represents the probability distribution of the final prediction result of the model Represents a normalization function that converts the predicted values into a probability distribution Represents a set of multi-layer perceptrons Represents a probability mapping matrix that converts the model output into an unnormalized score in the vocabulary
[0038] Optionally, the code generation model adopts the following loss function
[0039]
[0040] where Represents the total loss value of the model Represents the cross-entropy loss of the model, which is used to optimize the gap between the model prediction result and the true label Represents the sparsity loss of the model, which is used to control the sparsity of the activation vectors in the sparse coding so that the model can focus on key dependencies Represents a hyperparameter of the model, which is used to control the activation sparsity in the sparse autoencoder Represents the activation value of the features in the sparse autoencoder
[0041] In a second aspect, the present invention provides a code generation system that fuses two-dimensional code representation and dependency coding, including
[0042] A two-dimensional structure modeling module for modeling the obtained code data into a two-dimensional structure to obtain a two-dimensional code
[0043] A word embedding module for performing word embedding on the two-dimensional code through an embedding layer to obtain word embedding vectors
[0044] An inter-line masking module for performing inter-line masking on the two-dimensional code through an intra-line attention mask to obtain a mask matrix
[0045] A dependency modeling module for performing dependency modeling on the word embedding vectors through a sparse autoencoder to obtain a dependency coding of the code inter-line dependencies
[0046] An attention calculation module, configured to: perform attention calculation by using a self-attention mechanism according to the word embedding vectors, the mask matrix, and the dependency encoding, so as to obtain an intermediate representation of each character;
[0047] A model prediction module, configured to: perform model prediction through a multi-layer perceptron and a softmax layer according to the intermediate representation of each character, so as to obtain generated code.
[0048] Compared with the prior art, the beneficial effects achieved by the present invention:
[0049] 1. The code generation method that fuses two-dimensional code representation and dependency encoding provided by the present invention is based on a Transformer model. The code is regarded as a two-dimensional structure, and a dependency encoding is proposed to model the dependency relationship of this two-dimensional structure, which can effectively improve the model's ability to model the code structure and relationship, thereby improving the performance of the model; the code snippet is modeled as a two-dimensional structure, where the vertical dimension represents the logical flow between code lines, and the horizontal dimension represents the fine-grained meta-operations of characters within a line; in the shallow layer of the Transformer model, the semantic representation of the functionality of each line of code is extracted by setting an in-line attention mask; in the deep layer of the Transformer model, a sparse autoencoder (SAE) is used to extract features of the inter-line dependency relationship and generate corresponding dependency encoding to realize the embedding of the inter-line dependency relationship; the dependency embedding is fused with the character embedding, the attention scores between characters are calculated, and attention fusion and position-aware feature enhancement are performed on the attention scores to ensure the performance of the method in long-context scenarios;
[0050] 2. The code generation system that fuses two-dimensional code representation and dependency encoding provided by the present invention is significantly superior to the existing system in tasks such as code modeling, long-sequence understanding, functional correctness, and context retrieval through the settings of a two-dimensional structure modeling module, a word embedding module, an inter-line mask module, a dependency modeling module, an attention calculation module, and a model prediction module, and has practical significance and good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a flowchart of a code generation method that fuses two-dimensional code representation and dependency encoding according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] The technical solution of the present invention will be described in detail below through the drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present application and the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.
[0053] It should be noted that the term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this article generally represents an "or" relationship between the preceding and following associated objects.
[0054] Embodiment 1:
[0055] The embodiment of the present invention discloses a code generation method that combines two-dimensional code representation and dependency encoding. Referring to Figure 1 as shown, the specific steps are as follows:
[0056] S1, model the obtained code data into a two-dimensional structure to obtain a two-dimensional code;
[0057] S2, input the two-dimensional code into a trained code generation model: perform word embedding on the two-dimensional code through an embedding layer to obtain a word embedding vector; perform inter-line masking on the two-dimensional code through an in-line attention mask to obtain a mask matrix; perform dependency modeling on the word embedding vector through a sparse autoencoder to obtain a dependency encoding of the code inter-line dependency relationship; perform attention calculation using the self-attention mechanism based on the word embedding vector, mask matrix, and dependency encoding to obtain an intermediate representation of each character; perform model prediction through a multi-layer perceptron and a softmax layer based on the intermediate representation of each character to obtain a generated code.
[0058] Specifically,
[0059] In step S1, the code snippet is modeled as a two-dimensional structure, where the vertical dimension represents the logical flow between code lines, and the horizontal dimension represents the fine-grained meta-operations of characters within a line; current code large models mainly regard code snippets as ordinary text. Although this method facilitates the direct application of techniques in natural language processing to code processing, it fundamentally ignores the hierarchical and modular structures in source code; during the actual programming process, developers pay more attention to the dependencies between code lines when writing or understanding code, rather than the specific positions of individual characters in the input sequence; from this perspective, representing code as a two-dimensional structure is more important than positioning character positions as a one-dimensional sequence; in this embodiment, the code is represented as a two-dimensional structure and organized by code lines (vertical dimension) and characters in each line (horizontal dimension); specifically, the vertical dimension encodes the logical flow across program lines, covering elements such as control structures, function declarations, and dependencies between lines; while the horizontal dimension captures the fine-grained data operations within each line; by reflecting the spatial organizational structure of the code, this representation method is more in line with the natural way for developers to read, understand, and write code, embodying the structural organization of the source code; to represent the source code as a two-dimensional structure, a simple approach is to represent it based on a grid similar to an image; however, due to the usually large variation in the lengths of source code lines, this method will introduce a large number of padding placeholders, thus reducing the computational efficiency; therefore, this embodiment adopts a one-dimensional sequence representation method to process the source code, but highlights its line structure, specifically:
[0060] Given an input sequence , the source code is divided into different lines by the newline character \n; obtain the positions of each newline character in the code data to get the newline position encoding (assuming the boundary condition is ), where represents the newline sequence number;
[0061] According to the newline position encoding , for any line index , such that , then it is considered that the character belongs to the same line, where represents the line index, represents the position of the character in the code data, and represent the positions of two adjacent code lines.
[0062] In step S2, an intra-line attention mask is set to extract the semantic representation of the code lines. At the shallow layer of the model, the purpose is to extract the semantic information of each code line so as to perform dependency analysis on these code lines at the deep layer of the model. For this purpose, in this embodiment, the attention is restricted to the operations within each line, and the mask matrix M is used to achieve this goal. By decoupling the intra-line dynamics from the inter-line relationships, the model can achieve a more structured understanding of each line of code, thereby being able to model the translational invariance between codes. The mask matrix M is calculated by the following formula:
[0063]
[0064] where represents the accessibility between character and character .
[0065] Many code fragments, such as functions, classes, or independent modules, exhibit semantic translational invariance, that is, rearranging these elements usually does not change their basic logic. However, it is difficult for code language models that rely on fixed-position encoding to capture this invariance, which makes it difficult for them to accurately understand the semantics of the code. Nevertheless, there are also sequential relationships in the source code (for example, variables must be defined before they can be used), which indicates that completely abandoning position encoding is not an ideal solution. Instead, a more flexible, structured, and coarse-grained way must be adopted to view the relationships between characters, so as to balance the global invariance and follow the local sequential constraints. In addition, the positional information lacking a semantic basis itself is unreliable. For example, source code usually contains non-functional elements such as comments, indents, and line breaks, which enhance readability but do not directly affect the program logic in most cases. These elements may interfere with the position index, making the position encoding more complex and unreliable.
[0066] For this purpose, this embodiment proposes a dependency encoding to capture the hierarchical relationships between code lines, using the newline character \n as an anchor point for extracting the dependencies of subsequent code lines, since it is the end of each line of code and the start of the next line of code. The data processing process of dependency modeling includes:
[0067] According to the word embedding vector, obtain the word embedding of the tail character of the code line , as the semantic anchor point of the code line , where the characters of the code line include .
[0068] According to the semantic anchor point , calculate the code line The dependency encoding is calculated as follows:
[0069]
[0070] Among them, is the learnable parameter in the sparse autoencoder SAE, represents the encoder in the sparse autoencoder SAE, represents the feature dictionary, and each column represents a feature, represents the bias value, represents the activation function, represents the code line of the activation value of the dictionary feature, represents the code line of the dependency encoding, which is fused with the original word embedding in the way of additive positional encoding.
[0071] This embodiment also additionally introduces a distance-based activation enhancement technique. Compared with the position-based attention attenuation, in the context of the length of the context, this mechanism dynamically adjusts the activation value of the activation features in the above-mentioned sparse autoencoder , so that the model can focus on the key dependency relationships in a longer context, rather than only emphasizing local information; it is specifically implemented through the following formula:
[0072]
[0073] Among them, represents the activation value of the dictionary feature of the code line , represents the clipping operation, taking the larger value of and , represents the input length of the model, represents the semantic anchor position.
[0074] As the context length increases (such as in the warehouse-level code generation task), the attention of the model will be dispersed, resulting in a decline in the performance of the model; traditional distance-based attention attenuation methods can avoid attention dispersion. However, such methods inevitably lose long-distance dependency relationships; in this example, unified dependency encoding is performed on the characters in the same code line. As the dependency relationship grows, the number of characters covered by attention will also increase; considering two embeddings and , let and respectively represent the dependency relationships extracted by SAE corresponding to them, and the result output by the model attention layer can be obtained through and After performing linear mapping, calculate the dot product attention scores, and then obtain the intermediate representation of each character by weighting according to the attention scores; specifically as follows:
[0075] Calculate the attention scores between characters according to the word embedding vectors, dependency encodings, and mask matrix M, specifically through the following formula:
[0076]
[0077] Among them, represents the word embedding of character , represents the dependency encoding of the code line where character is located, represents the word embedding of character after being encoded by the dependency encoding, represents the word embedding of character , represents the word embedding of character represents the dependency encoding of the code line where character is located, represents the word embedding of character after being encoded by the dependency encoding, represents the key vector in the self-attention mechanism, represents the query mapping matrix, represents the value vector in the self-attention mechanism, represents the key mapping matrix, and character between the attention scores, represents the word embedding of character and character between the accessibility;
[0078] According to the attention scores between the characters, calculate the intermediate representation of each character, specifically through the following formula:
[0079]
[0080] Among them, represents the intermediate representation of character generated by the self-attention mechanism, represents the value vector in the self-attention mechanism, represents the value mapping matrix, represents the normalization function.
[0081] Intuitively, the first term in the expansion represents semantic attention, capturing and Semantic interaction between; the second and third items represent cross-attention between semantic content and encoding; the fourth item represents dependence attention, which is a direct interaction representing the inter-line dependence relationship; although the fourth item directly reflects the dependence relationship between code lines, however, in the setting of this embodiment, the characters within the same code line share the same dependence information, which may cause the attention scores between characters to uniformly increase between lines with strong dependence relationships, thereby inadvertently increasing the model's attention to more characters and resulting in model attention dispersion; to solve this problem, this embodiment suppresses the direct positional attention (i.e., the fourth item) and processes its corresponding attention scores as independent features; specifically, the attention scores are flattened along the head dimension, and the scores of the fourth item are concatenated from each head. These concatenated features are then processed by an MLP, alleviating the attention dispersion problem; specifically through the following formula:
[0082]
[0083] Wherein, represents the attention score between character and character , represents a set of multi-layer perceptrons, represents the dimension concatenation operation.
[0084] In this embodiment, model prediction is performed through a multi-layer perceptron and a softmax layer, specifically implemented through the following formula:
[0085]
[0086] Wherein, represents the probability distribution of the final prediction result of the model, represents the normalization function, represents a set of multi-layer perceptrons, represents the probability mapping matrix.
[0087] In this embodiment, the code generation model adopts the following loss function:
[0088]
[0089] Wherein, represents the total loss value of the model, represents the cross-entropy loss of the model, represents the sparsity loss of the model, represents the hyperparameter of the model, represents the activation value of the feature in the sparse autoencoder.
[0090] This loss function not only encourages the model to focus on modeling key relationships, enhancing the model's generalization ability, but also enhances the interpretability of the learned features due to the sparsity constraint.
[0091] In this embodiment, a code modeling task is used to verify the effectiveness of the code generation method that fuses two-dimensional code representation and dependency encoding; the code modeling task is to predict the next character in a programming-related context; and perplexity is used to reflect the likelihood degree of the model prediction sequence and the target sequence, and at the same time, the accuracy rate is used to evaluate the proportion of characters correctly predicted by the model; the results are shown in Table 1:
[0092] Table 1: Comparison of perplexity and accuracy rate between this method and the baseline method
[0093]
[0094] Experiments show that this method has stronger long text understanding and generation capabilities compared to the baseline method in code modeling.
[0095] In summary, the code generation method that fuses two-dimensional code representation and dependency encoding proposed in this embodiment solves the problems of poor generalization, poor code structure perception, and insufficient extrapolation ability existing in the traditional position encoding method during the code generation process; for the first time, it proposes to model code fragments as a two-dimensional structure, decompose the code from the one-dimensional modeling method in the traditional language model into a logical process in the vertical dimension and meta-operations in the horizontal dimension, and use a sparse autoencoder (SAE) to capture the semantic dependency relationships between code lines, replacing the traditional position encoding method based on absolute position numbers; the steps are as follows: design of a hierarchical Transformer model architecture, the shallow layer of the model extracts meta-operation semantic information by setting an attention mask to avoid the influence of irrelevant contexts; the deep layer integrates dependency encoding, extracts a feature relationship dictionary between code lines by setting a sparse autoencoder, and models the dependency relationships between code lines and assigns corresponding dependency encodings; experiments show that this method is significantly superior to existing methods in tasks such as code modeling, long sequence understanding, functional correctness, and context retrieval.
[0096] Embodiment 2:
[0097] Based on the same inventive concept as Embodiment 1, this embodiment of the present invention discloses a code generation system that fuses two-dimensional code representation and dependency encoding, including:
[0098] Two-dimensional structure modeling module, used for: modeling the obtained code data as a two-dimensional structure to obtain two-dimensional code;
[0099] Word embedding module, used for: performing word embedding on the two-dimensional code through an embedding layer to obtain word embedding vectors;
[0100] The line-by-line masking module is used to: perform line-by-line masking on the two-dimensional code through an in-line attention mask to obtain a mask matrix;
[0101] The dependency modeling module is used to: perform dependency modeling on the word embedding vectors through a sparse autoencoder to obtain a dependency encoding of the line-by-line dependencies of the code;
[0102] The attention calculation module is used to: perform attention calculation using the self-attention mechanism based on the word embedding vectors, the mask matrix, and the dependency encoding to obtain an intermediate representation of each character;
[0103] The model prediction module is used to: perform model prediction through a multi-layer perceptron and a softmax layer based on the intermediate representation of each character to obtain the generated code.
[0104] The specific functional implementations of the above modules refer to the relevant content in the method of Embodiment 1 and will not be elaborated here.
[0105] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0106] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0107] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or one block or a plurality of blocks in the flow Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps of the functions specified in one block or a plurality of blocks.
[0109] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.
Claims
1. A code generation method that integrates two-dimensional code representation and dependency encoding, characterized in that, Including: Model the obtained code data into a two-dimensional structure to obtain a two-dimensional code; Input the two-dimensional code into the trained code generation model: perform word embedding on the two-dimensional code through the embedding layer to obtain a word embedding vector; perform inter-line masking on the two-dimensional code through the in-line attention mask to obtain a mask matrix; Perform dependency modeling on the word embedding vector through a sparse autoencoder to obtain a dependency encoding of the code inter-line dependency relationship; According to the word embedding vector, mask matrix, and dependency encoding, use the self-attention mechanism to perform attention calculation to obtain the intermediate representation of each character; according to the intermediate representation of each character, perform model prediction through a multi-layer perceptron and a softmax layer to obtain the generated code.
2. The method for generating a code by fusing a two-dimensional code representation and a dependency encoding according to claim 1, wherein Modeling the obtained code data into a two-dimensional structure includes: Obtain the positions of each line break character in the code data to get line break position codes , where represents the line break sequence number; Encoding according to the newline character position , for any line index , such that , then the character is considered to be on the same line, where represents the line index, represents the position of the character in the code data, and represent the positions of two adjacent code lines.
3. The method for generating a code integrating a two-dimensional code representation and a dependency encoding according to claim 2, wherein The mask matrix M is calculated by the following formula: Among them, represents the accessibility between character and character.
4. The method for generating a code that fuses two-dimensional code representation and dependency encoding according to claim 2, wherein The data processing process of the dependency modeling includes: Based on the word embedding vectors, obtain the tail character of the code line and the word embedding of as the semantic anchor of the code line , where the characters of the code line include ; According to the semantic anchor , calculate the dependency encoding of the code line through the sparse autoencoder SAE, and the specific calculation is as follows: Among them, is a learnable parameter in the sparse autoencoder SAE, represents the encoder in the sparse autoencoder SAE, represents the feature dictionary, represents the bias value, represents the activation function, represents the code line activation value of the dictionary feature, represents the code line dependency encoding.
5. The method for generating a code by fusing a two-dimensional code representation and a dependency encoding according to claim 4, wherein The dependency modeling includes distance-based activation value enhancement, which is achieved by the following formula: Among them, represents the activation value of the dictionary feature of the code line . represents a clipping operation, taking the larger value between and represents the input length of the model, represents the semantic anchor position.
6. The method for generating a code by fusing a two-dimensional code representation and a dependency encoding according to claim 4, wherein The data processing process of the attention calculation includes: Calculate the attention scores between characters according to the word embedding vector, dependency encoding, and mask matrix M, specifically through the following formula: Among them, represents the word embedding of the character . represents the dependency encoding of the code line where the character is located. represents the word embedding of the character after being encoded by the dependency. represents the word embedding of the character . represents the word embedding of the character and the dependency encoding of the code line where it is located. represents the word embedding of the character after being encoded by the dependency. represents the key vector in the self-attention mechanism. represents the query mapping matrix. represents the value vector in the self-attention mechanism. represents the key mapping matrix. represents the attention score between the character and the character . represents the accessibility between the character and the character . Calculate the intermediate representation of each character according to the attention scores between characters, specifically through the following formula: Among them, represents a character the intermediate representation generated by the self-attention mechanism, represents the value vector in the self-attention mechanism, represents the value mapping matrix, represents the normalization function.
7. The method for generating a code by integrating a two-dimensional code representation and a dependency encoding according to claim 6, wherein The self-attention mechanism alleviates the problem of attention dispersion through an MLP, specifically through the following formula: Among them, represents the attention score between character and character represents a set of multi-layer perceptrons, represents the dimension concatenation operation.
8. The method for generating a code by fusing a two-dimensional code representation and a dependency encoding according to claim 6, wherein The model prediction through the multi-layer perceptron and the softmax layer is achieved by the following formula: Among them, represents the probability distribution of the final prediction result of the model, represents the normalization function, represents a group of multi-layer perceptrons, represents the probability mapping matrix.
9. The method for generating a code that fuses two-dimensional code representation and dependency encoding according to claim 1, wherein The code generation model adopts the following loss function: Among them, represents the total loss value of the model, represents the cross-entropy loss of the model, represents the sparsity loss of the model, represents the hyperparameter of the model, represents the activation value of the feature in the sparse autoencoder.
10. A code generation system integrating two-dimensional code representation and dependency encoding, characterized in that, Including: A two-dimensional structure modeling module, which is used to: model the obtained code data into a two-dimensional structure to obtain a two-dimensional code; A word embedding module, which is used to: perform word embedding on the two-dimensional code through the embedding layer to obtain a word embedding vector; An inter-line masking module, which is used to: perform inter-line masking on the two-dimensional code through the in-line attention mask to obtain a mask matrix; A dependency modeling module, which is used to: perform dependency modeling on the word embedding vector through a sparse autoencoder to obtain a dependency encoding of the code inter-line dependency relationship; An attention calculation module, which is used to: according to the word embedding vector, mask matrix, and dependency encoding, use the self-attention mechanism to perform attention calculation to obtain the intermediate representation of each character; A model prediction module, which is used to: according to the intermediate representation of each character, perform model prediction through a multi-layer perceptron and a softmax layer to obtain the generated code.
Citation Information
Patent Citations
Code generation method and device, storage medium and electronic equipment
CN116166271A
Space-time prediction method and device based on space-time mask auto-encoder, and electronic equipment
CN118171773A
Code generation method and apparatus, storage medium and electronic device
WO2024174911A1
Cited By
Code generation model training method, device and equipment
CN120873595A