A hierarchical adaptive code generation method, system and medium
Through a hierarchical adaptive code generation method based on decoding layer information, dynamic selection of model layers and using classification decoding strategies, the problem of low code quality in existing code generation methods is solved, and higher quality and reliable code generation is achieved.
Patent Information
- Application Number
- CN202411775766.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-05
AI Technical Summary
The existing code generation methods have hallucinations when generating code, insufficient knowledge utilization, and lack targeted decoding strategies for code characteristics, resulting in low quality of generated code.
A hierarchical adaptive code generation method based on decoding layer information is proposed. Through the code token type prediction module and the decoding layer adaptive selection algorithm, the appropriate model layer is dynamically selected for output prediction, and three different classification decoding strategies are used to generate tokens of basic structure, code logic and high-level semantic content.
Improves the reliability of LLMs in code generation tasks, reduces structural or semantic errors in generated code, ensures the logic and executability of generated code, and does not require the introduction of external knowledge bases.
Smart Images

Figure CN119248289B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of software engineering automation technology, and in particular to a hierarchical adaptive code generation method, system and medium based on decoding layer information. Background Art
[0002] In the field of software engineering automation, code generation has become an important part of the automated development process. Code generation technology aims to automatically generate code that conforms to grammatical rules, is logically correct, and is executable. In recent years, deep learning models, especially Transformer-based models, have been able to automatically generate high-quality text through multi-layer encoder and decoder structures. Large language models (LLMs) can learn the syntax, structure, and patterns of code after pre-training with large-scale code data, realize automatic code generation, and show good results in some scenarios. However, these models still face challenges when generating code, and the generated code may not conform to grammatical rules or function requirements.
[0003] Most current decoding strategies are derived from text generation in the field of natural language processing and lack targeted optimization for code generation, which limits the quality and practicality of model-generated code.
[0004] Specifically, existing code generation methods usually generate code step by step based on the language model's understanding of the input description. However, these methods have the following defects:
[0005] 1. Hallucination problem: During the code generation process, LLMs often generate code snippets with syntax errors or incorrect logic. This phenomenon is called hallucination, which causes the generated results to fail to execute correctly.
[0006] 2. Insufficient knowledge utilization: LLMs learn a lot of code-related knowledge during pre-training, but cannot effectively call this knowledge when generating code. In particular, the model is prone to errors when predicting complex logic or calling specific library functions, and cannot reliably generate expected code.
[0007] 3. Lack of targeting of code features: Current technologies mainly rely on the general architecture of text generation models, adding a large amount of code data for pre-training to achieve code generation. The single text generation decoding strategy fails to fully consider the particularity of the code. The lack of specialized decoding strategies for code syntax, local logic, and global logic results in low quality of generated code.
[0008] 4. Dependence on external knowledge base: Some methods correct model errors by retrieving external knowledge bases, but this increases the complexity of implementation and has unstable performance in the absence of external resources. Summary of the invention
[0009] The main purpose of the present invention is to propose a hierarchical adaptive code generation method, system and medium based on decoding layer information, aiming to improve the reliability of LLMs in code generation tasks, enable the model to more effectively utilize its inherent knowledge at all levels, reduce structural or semantic errors in the generated code, and ensure the logic and executableness of the generated code.
[0010] To achieve the above object, the present invention provides a hierarchical adaptive code generation method based on decoding layer information, the method comprising the following steps:
[0011] Step S10, analyzing the context of the code to be generated based on the code token type prediction module, identifying the basic type of the next token to be generated, wherein the basic type includes basic structure, code logic and high-level semantic content;
[0012] Step S20, automatically selecting an appropriate model layer for output prediction based on a decoding layer adaptive selection algorithm;
[0013] Step S30, using three different classification decoding strategies to generate tokens belonging to basic structure, code logic and high-level semantic content respectively.
[0014] A further technical solution of the present invention is that step S10 comprises:
[0015] Step S101, in the autoregressive decoding process, the code-aware attention mechanism with distance decay is used to identify the basic type of the next token to be generated based on the given existing previous output. The type set is recorded as , .
[0016] A further technical solution of the present invention is that step S101 comprises:
[0017] Based on formula (1), the probability distribution of three types of tokens is calculated:
[0018] Formula (1);
[0019] in, Indicates that given a sequence of preceding code tokens In the case of , generate the i-th token, The probability distribution of belonging to a class, is the hidden state calculated from the previously generated token sequence, , and are the parameters of the classifier, The context vector, and the calculated output probability distribution represents which of the three categories the current token belongs to.
[0020] A further technical solution of the present invention is that the context vector is obtained by weighted summation of the hidden states of the previous token sequence. To ensure that the context information closer to the current generation position has a greater weight and the influence of the more distant context is weakened, a distance attenuation coefficient is introduced. This coefficient exponentially decays according to the distance between the token and the current generation position. The specific calculation formula is:
[0021] Formula (2);
[0022] where represents the weight of the j-th previous token; i represents the position of the current generated token; e is the natural constant, approximately equal to 2.718, which is used for exponential decay calculation so that the distance decay follows the characteristics of the exponential function; is the distance attenuation coefficient, which is used to control the attenuation speed of the weight; k represents the index of all previous token positions, and the range is 1 ≤ k < i; the context vector is obtained by weighted summation of the previous hidden states with the attenuated weights:
[0023] Formula (3).
[0024] A further technical solution of the present invention is that the step S20 includes:
[0025] Step S201, assuming that the model has a total of N layers, the structure token, the logic token, and the semantic token will respectively select appropriate decoding layers from the intervals of [1 - N / 3], [N / 3 - 2N / 3], and [2N / 3 - N];
[0026] Step S202, based on the logits entropy value and the confidence level selection rule, adaptively select which specific layer should be used for decoding the current token.
[0027] A further technical solution of the present invention is that the step S202 includes:
[0028] By calculating the entropy value of the logits of different layers , while measuring the confidence level of this layer for the current generation step , thus dynamically select the optimal layer for decoding. The entropy value is used to measure the uncertainty of the probability distribution output by the i-th layer, and the calculation method is:
[0029] Formula (4);
[0030] Where N is the size of the vocabulary, It is the predicted probability of the kth word in the vocabulary of the i-th layer as the next token. The lower the entropy value, the higher the certainty of the generation result of the layer and the more concentrated the confidence of the output.
[0031] Confidence It is used to reflect the maximum possibility of the output of the i-th layer, which is the maximum probability value of generating the next token:
[0032] Formula (5);
[0033] A layer with a high confidence level indicates that the layer has a clearer prediction for generating the token. The final layer selection is performed using the following formula:
[0034] Formula (6);
[0035] in, is the optimal layer finally selected, It is the hierarchical interval corresponding to the current token type. It is an adjustment parameter used to balance the influence of entropy and confidence.
[0036] A further technical solution of the present invention is that step S30 comprises:
[0037] For structural tokens, the basic structure of the low-level decoding generated code is adopted. After the decoding layer is determined using the decoding layer adaptive selection algorithm, the probability distribution of the next token is calculated using formula (7):
[0038] Formula (7);
[0039] Among them, low represents the decoding layer selected from the low interval by the decoding layer adaptive selection algorithm. Logits is the real-valued vector output by the model decoding layer, which represents the score corresponding to each word. The dimension of this vector is equal to the size of the vocabulary. These logits values are normalized by the softmax function and converted into probability distribution;
[0040] For logical tokens, the middle layer is used for decoding to generate the control logic structure in the code. The probability distribution is calculated using formula (8) by combining the structural information of the middle layer and the lower layer:
[0041] Formula (8);
[0042] in, is the weight coefficient, which is used to balance the weight of structural information and logical information; and They are the logits of the layers selected from the low-layer interval and the middle-layer interval using the decoding layer adaptive selection algorithm;
[0043] For semantic tokens, high-level decoding is used to generate specific semantic details, and tokens are selected based on the difference between high-level and low-level logits.
[0044] The step of selecting a token by the difference between high-level and low-level logits specifically includes:
[0045] First, use formula (9) to calculate the output probability distribution:
[0046] Formula (9),
[0047] and They are the logits of the layers selected from the high-level interval and the low-level interval using the decoding layer adaptive selection algorithm;
[0048] Then select the token with the highest probability.
[0049] A further technical solution of the present invention is that after step S30, the following steps are further included:
[0050] Step S40, repeating steps S20 and S30 until an end symbol is generated;
[0051] Step S50, output: obtaining the code generated by the model and verifying the generated code;
[0052] Before step S10, the following steps are also included:
[0053] Step S00, input: obtaining a partial code input by the user or code generation prompt information.
[0054] To achieve the above-mentioned objectives, the present invention also proposes a hierarchical adaptive code generation system based on decoding layer information, the system comprising a memory, a processor, and a hierarchical adaptive code generation program based on decoding layer information stored on the processor, the hierarchical adaptive code generation program based on decoding layer information executing the steps of the method described above when run by the processor.
[0055] To achieve the above objectives, the present invention also proposes a computer-readable storage medium, which stores a hierarchical adaptive code generation program based on decoding layer information, and the hierarchical adaptive code generation program based on decoding layer information executes the steps of the method described above when executed by a processor.
[0056] The beneficial effects of the hierarchical adaptive code generation method, system and medium based on decoding layer information of the present invention are:
[0057] 1. This paper proposes a hierarchical decoding strategy, which divides the tokens generated by the code into three types, representing the basic structure, code logic, and high-level semantic content. This hierarchical decoding strategy effectively utilizes the knowledge differences of each layer of the large model during the code generation process, which is more in line with the characteristics of the code language.
[0058] 2. The present invention introduces a code token type prediction module and a decoding layer adaptive selection algorithm to guide the subsequent decoding process by predicting different types of tokens. The decoding layer adaptive selection algorithm flexibly selects the decoding layer through technical means such as entropy calculation, making full use of the content of different layers within the model.
[0059] 3. The present invention does not need to introduce an external knowledge base, does not need to fine-tune the entire code to generate a large model, and only needs to dynamically utilize the information at different levels of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.
[0061] Figure 1 It is a flow chart of a preferred embodiment of the hierarchical adaptive code generation method based on decoding layer information of the present invention;
[0062] Figure 2 It is a system architecture diagram of the hierarchical adaptive code generation system based on decoding layer information of the present invention.
[0063] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0064] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0065] The present invention proposes a hierarchical adaptive code generation method based on decoding layer information, such as Figure 1 As shown, a preferred embodiment of the hierarchical adaptive code generation method based on decoding layer information of the present invention includes the following steps:
[0066] Step S10, analyzing the context of the code to be generated based on the code token type prediction module, identifying the basic type of the next token to be generated, wherein the basic type includes basic structure, code logic and high-level semantic content.
[0067] Step S20, based on the decoding layer adaptive selection algorithm, automatically select the appropriate model layer for output prediction.
[0068] Step S30, using three different classification decoding strategies to generate tokens belonging to basic structure, code logic and high-level semantic content respectively.
[0069] Different layers of a large language model have different information storage capabilities: lower layers mainly store structural information, while higher layers tend to capture semantic information. In order to mine the knowledge stored inside the model and reduce the hallucination phenomenon caused by large-scale language models (LLMs) in code generation tasks, this paper proposes a hierarchical adaptive decoding strategy based on decoding layer information for code generation tasks, aiming to dynamically use the information of different decoding layers of the model to generate high-quality code. The specific operations are as follows:
[0070] Use the low-level code to generate basic structures and ensure the syntactic correctness of the code. It is mainly responsible for generating basic parts such as variable declarations and function definitions.
[0071] Use the middle layer to generate code logic and refine the code logic, such as conditional branches, loops and other code control structures.
[0072] Use high-level semantic content generation to focus on capturing complex function calls, library dependencies, and implementation details, thereby ensuring the execution correctness and contextual consistency of the generated code.
[0073] The overall process of the present invention mainly consists of three stages: 1) Code token type prediction module, by analyzing the context of the code to be generated, automatically identifies the basic type of the next token to be generated, and determines whether it belongs to the basic structure, code logic or high-level semantic content, providing guidance for the subsequent hierarchical decoding strategy; 2) Decoding layer adaptive selection algorithm, automatically selects the appropriate model layer for output prediction. The algorithm dynamically locates the layer used when generating code by calculating the entropy value of logits of each layer; 3) Classification decoding strategy, using three different decoding strategies to generate tokens belonging to basic structure, code logic and high-level semantic content respectively. This decoding strategy fully utilizes the parameter knowledge stored in the large model through the collaborative work of the initial layer, the middle layer and the high layer, so that the generated code has a reasonable grammatical structure, rigorous logic and correct functional execution.
[0074] Specifically, the step S10 includes:
[0075] Step S101, in the autoregressive decoding process, the code-aware attention mechanism with distance decay is used to identify the basic type of the next token to be generated based on the given existing previous output. The type set is recorded as , .
[0076] In this embodiment, in step S10, a multi-layer neural network is trained as a classifier using a code token type prediction module. In the autoregressive decoding process, in each step of generation, given the existing pre-order output, the module will determine which type the next token to be generated should belong to. The type set is , In order to more accurately capture the contextual information of the generated code, this module introduces a code-aware attention mechanism with distance decay. This mechanism focuses on the code information at a closer distance and uses the weight distribution with distance decay to adjust the influence of code information at different distances.
[0077] In this embodiment, step S101 specifically includes:
[0078] The classifier calculates the probability distribution of three token types based on formula (1):
[0079] Formula (1);
[0080] in, Indicates that given a sequence of preceding code tokens In the case of , generate the i-th token, The probability distribution of belonging to a class, is the hidden state calculated from the previously generated token sequence, , and are the parameters of the classifier, represents the context vector, and the calculated output probability distribution represents which of the three categories the current token belongs to.
[0081] Among them, the context vector is obtained by weighted summation of the hidden states of the previous token sequence. To ensure that the context information closer to the current generation position has a greater weight and the influence of the more distant context is weakened, a distance decay coefficient is introduced. This coefficient exponentially decays according to the distance between the token and the current generation position. The specific calculation formula is:
[0082] Formula (2);
[0083] Among them, represents the weight of the j-th previous token; i represents the position of the current generated token; e is the natural constant, approximately equal to 2.718, which is used for exponential decay calculation so that the distance decay follows the characteristics of the exponential function; is the distance decay coefficient, which is used to control the decay speed of the weight; k represents the index of all previous token positions, and the range is 1 ≤ k < i; the context vector is obtained by weighted summation of the previous hidden states with the decay weights:
[0084] Formula (3).
[0085] In this embodiment, the step S20 includes:
[0086] Step S201, assuming that the model has a total of N layers, the structural tokens, logical tokens, and semantic tokens will respectively select appropriate decoding layers from the intervals [1 - N / 3], [N / 3 - 2N / 3], and [2N / 3 - N].
[0087] Step S202, based on the logits entropy value and the confidence selection rule, adaptively select which specific layer should be used for decoding the current token.
[0088] Step S20 The decoding layer adaptive selection algorithm is intended to select the model decoding layer for outputting the result. First, the model hierarchy is divided into three intervals, representing the low layer, the middle layer, and the high layer. Assuming that the model has a total of N layers, the structural token, the logical token, and the semantic token will select the appropriate decoding layer from the intervals [1-N / 3], [N / 3-2N / 3], and [2N / 3-N], respectively. After the interval is selected, the present invention adaptively selects which specific layer the current token should use for decoding based on the logits entropy value and the confidence selection rule.
[0089] The step S202 specifically includes:
[0090] By calculating the entropy of logits at different layers , while measuring the confidence of the layer in the current generation step , thereby dynamically selecting the optimal layer for decoding, the entropy value It is used to measure the uncertainty of the probability distribution of the output of the i-th layer, and is calculated as:
[0091] Formula (4);
[0092] Where N is the size of the vocabulary, It is the predicted probability of the kth word in the vocabulary of the i-th layer as the next token. The lower the entropy value, the higher the certainty of the generation result of the layer and the more concentrated the confidence of the output.
[0093] Confidence It is used to reflect the maximum possibility of the output of the i-th layer, which is the maximum probability value of generating the next token:
[0094] Formula (5);
[0095] A layer with a high confidence level indicates that the layer has a clearer prediction for generating the token. The final layer selection is performed using the following formula:
[0096] Formula (6);
[0097] in, is the optimal layer finally selected, It is the hierarchical interval corresponding to the current token type. It is an adjustment parameter used to balance the influence of entropy and confidence.
[0098] In this embodiment, in the classification decoding strategy of step S30, different decoding methods are used for three different types of tokens. Step S30 specifically includes:
[0099] For structural tokens, low-level decoding is used to generate the basic structure of the code. When using low-level decoding, the focus is on generating the basic structure of the code. This information is usually captured by the low-level model. After the decoding layer is determined using the decoding layer adaptive selection algorithm in this embodiment, the probability distribution of the next token is calculated using formula (7):
[0100] Formula (7);
[0101] Among them, low represents the decoding layer selected from the low interval by the decoding layer adaptive selection algorithm. Logits is the real-valued vector output by the model decoding layer, which represents the score corresponding to each word. The dimension of this vector is equal to the size of the vocabulary. These logits values are normalized by the softmax function and converted into probability distribution. In this stage, the token with the highest probability is selected as the generated result of this step.
[0102] For logical tokens, the middle layer decoding is used to generate the control logic structure in the code. The middle layer decoding focuses on the control logic structure in the generated code, such as conditional statements, loops, etc. These structures are relatively complex and require both basic grammatical correctness and logical clarity. This embodiment combines the structural information of the middle layer and the lower layer and uses formula (8) to calculate the probability distribution:
[0103] Formula (8);
[0104] in, is the weight coefficient, which is used to balance the weight of structural information and logical information; and They are the logits of the layers selected from the low-layer interval and the middle-layer interval using the decoding layer adaptive selection algorithm.
[0105] For semantic tokens, high-level decoding is used to generate specific semantic details. When using high-level decoding, the focus is on generating specific semantic details, such as function calls, library dependencies, and complex algorithm implementations. At this time, it is necessary to rely on the high-level semantic representation of the model. In order to strengthen the use of semantic information, this embodiment proposes a strategy for selecting tokens by the difference between high-level and low-level logits.
[0106] The step of selecting a token by the difference between high-level and low-level logits specifically includes:
[0107] First, use formula (9) to calculate the output probability distribution:
[0108] Formula (9);
[0109] Then select the token with the highest probability.
[0110] and The logits of the layers are selected from the high-level interval and the low-level interval using the decoding layer adaptive selection algorithm. Tokens with large difference are usually tokens that express high-level semantic information, so selecting this token during high-level decoding can improve the semantic accuracy of the generated code.
[0111] Furthermore, in this embodiment, after step S30, the following steps are further included:
[0112] Step S40, repeating steps S20 and S30 until an end symbol is generated;
[0113] Step S50, output: obtain the code generated by the model and verify the generated code.
[0114] Before step S10, the following steps are also included:
[0115] Step S00, input: obtaining partial code input by the user or code generation prompt information, such as function definition, requirement description, etc.
[0116] In this embodiment, step S20 mainly calls the trained code generation model. The model first uses the type classifier to predict the type of the next token to be generated, and then automatically selects the corresponding decoding layer through the decoding layer adaptive selection algorithm; step S30 is that the model uses the corresponding decoding strategy to predict the next token.
[0117] The beneficial effects of the hierarchical adaptive code generation method based on decoding layer information of the present invention are:
[0118] 1. This paper proposes a hierarchical decoding strategy, which divides the tokens generated by the code into three types, representing the basic structure, code logic, and high-level semantic content. This hierarchical decoding strategy effectively utilizes the knowledge differences of each layer of the large model during the code generation process, which is more in line with the characteristics of the code language.
[0119] 2. The present invention introduces a code token type prediction module and a decoding layer adaptive selection algorithm to guide the subsequent decoding process by predicting different types of tokens. The decoding layer adaptive selection algorithm flexibly selects the decoding layer through technical means such as entropy calculation, making full use of the content of different layers within the model.
[0120] 3. The present invention does not need to introduce an external knowledge base, does not need to fine-tune the entire code to generate a large model, and only needs to dynamically utilize the information at different levels of the model.
[0121] To achieve the above object, the present invention also proposes a hierarchical adaptive code generation system based on decoding layer information, such as Figure 2 As shown, the system includes a processor 1001, a CPU, a network interface 1004, a user interface 1003, a memory 1005, a communication bus 1002, and a hierarchical adaptive code generation program based on decoding layer information stored on the processor, wherein the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0122] Those skilled in the art will understand that Figure 2 The system structure shown in the figure does not constitute a limitation of the system, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0123] like Figure 2 As shown, the memory 1005 as a computer storage medium may include an operating device, a network communication module, a user interface module, and a hierarchical adaptive code generation program based on decoding layer information.
[0124] exist Figure 2 In the system shown, the network interface 1004 is mainly used to connect to the network server and communicate data with the network server; the user interface 1003 is mainly used to interact with the user terminal and receive instructions input by the user; and the processor 1001 can be used to call the hierarchical adaptive code generation program based on decoding layer information stored in the memory 1005.
[0125] To achieve the above-mentioned purpose, the present invention also proposes a computer-readable storage medium, which stores a hierarchical adaptive code generation program based on decoding layer information. When the hierarchical adaptive code generation program based on decoding layer information is executed by a processor, the steps of the method described above are executed, which will not be repeated here.
[0126] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. All equivalent structural changes made by using the contents of the present invention specification and drawings under the concept of the present invention, or directly / indirectly applied in other related technical fields are included in the patent protection scope of the present invention.
Claims
1. A hierarchical adaptive code generation method, characterized in that: The method comprises the following steps: Step S10, analyzing the context of the code to be generated based on the code token type prediction module, identifying the basic type of the next token to be generated, wherein the basic type includes basic structure, code logic and high-level semantic content; Step S20, automatically selecting an appropriate model layer for output prediction based on a decoding layer adaptive selection algorithm; Step S30, using three different classification decoding strategies to generate tokens belonging to basic structure, code logic and high-level semantic content respectively; The step S10 comprises: Step S101, in the autoregressive decoding process, a code-aware attention mechanism with distance decay is used to identify the basic type of the next token to be generated based on the given existing previous output, and the type set is denoted as T, T = {structural token, logical token, semantic token}; The step S101 includes: Based on formula (1), the probability distribution of three types of tokens is calculated: P(t i |t1, t2, ..., t i-1 )=softmax(W a h i-1 +W b c t +b) Formula (1); Among them, P(t i |t1, t2, ..., t i-1 ) means that given the preceding code token sequence t1, t2, ..., t i-1 In the case of , generate the i-th token, t i The probability distribution of belonging to a certain category, h i-1 is the hidden state calculated from the previously generated token sequence, W a , W b and b are the parameters of the classifier, c t represents the context vector, and the calculated output probability distribution indicates which of the three categories the current token belongs to; Context vector c t It is obtained by weighted summing the hidden states of the previous token sequence. To ensure that the context information closer to the current generation position has a greater weight, while the influence of the farther context is weakened, the distance decay coefficient α is introduced j , the coefficient α j The distance between the token and the current generated position decays exponentially. The specific calculation formula is: Among them, α j represents the weight of the j-th previous token; i represents the position of the currently generated token; e is the natural constant used for exponential decay calculation, such that the distance decay follows the characteristics of the exponential function; λ is the distance decay coefficient used to control the decay speed of the weight; k represents the index of all previous token positions, with the range 1 ≤ k < i; the context vector c t is obtained by weighted summation of the previous hidden states with decaying weights:
2. The hierarchical adaptive code generation method according to claim 1, characterized in that: The step S20 comprises: Step S201, assuming that the model has N layers in total, the structural token, logical token and semantic token will select appropriate decoding layers from the intervals of [1, N / 3], (N / 3, 2N / 3] and (2N / 3, N] respectively; Step S202: based on the logits entropy value H(L i ) and confidence C(L i ) selection rules to adaptively select which specific layer the current token should use for decoding.
3. The hierarchical adaptive code generation method according to claim 2, characterized in that: The step S202 includes: By calculating the entropy value H(L i ), and at the same time measures the confidence C(L i ), thereby dynamically selecting the optimal layer for decoding, the entropy value H(L i ) is used to measure the uncertainty of the probability distribution of the output of the i-th layer, and is calculated as: Where N is the size of the vocabulary, P i,k It is the predicted probability of the kth word in the vocabulary of the i-th layer as the next token. The lower the entropy value, the higher the certainty of the generation result of the layer and the more concentrated the confidence of the output. Confidence C(L i ) is used to reflect the maximum possibility of the output of the i-th layer, which is the maximum probability value of generating the next token: C(L i )=max P(t i |t1, t2, ..., t i-1 ) formula (5); A layer with a high confidence level indicates that the layer has a clearer prediction for generating the token. The final layer selection is performed using the following formula: Among them, L * is the optimal layer finally selected, L k is the level interval corresponding to the current token type, and α is the adjustment parameter used to balance the influence of entropy and confidence.
4. The hierarchical adaptive code generation method according to claim 3, characterized in that: The step S30 comprises: For structural tokens, the basic structure of the low-level decoding generated code is adopted. After the decoding layer is determined using the decoding layer adaptive selection algorithm, the probability distribution of the next token is calculated using formula (7): P(t i |t1, t2, ..., t i-1 )=softmax(logits low ) formula (7); Among them, low represents the decoding layer selected from the low-level interval by the decoding layer adaptive selection algorithm, logits is the real-valued vector output by the model decoding layer, which represents the score corresponding to each word. The dimension of this vector is equal to the size of the vocabulary. These logits values are normalized by the softmax function and converted into probability distribution. The token with the largest probability is selected as the generation result of this step; For the logic token, the middle layer decoding is used to generate the control logic structure in the code. The probability distribution is calculated using formula (8) by combining the structure information of the middle layer and the lower layer: P(t i |t1, t2,..., t i-1 ) = θsoftmax(logits low ) + (1 - θ)softmax(logits middle ) Equation (8); Among them, θ is the weight coefficient, which is used to balance the weight of structural information and logical information; logits low and logits middle They are the logits of the layers selected from the low-layer interval and the middle-layer interval using the decoding layer adaptive selection algorithm; For semantic tokens, high-level decoding is used to generate semantic details, and tokens are selected by the difference between high-level and low-level logits.
5. The hierarchical adaptive code generation method according to any one of claims 1 to 4, characterized in that: After step S30, the following steps are further included: Step S40, repeating steps S20 and S30 until an end symbol is generated; Step S50, output: obtaining the code generated by the model and verifying the generated code; Before step S10, the following steps are also included: Step S00, input: obtaining a partial code input by the user or code generation prompt information.
6. A hierarchical adaptive code generation system based on decoding layer information, characterized in that: The system includes a memory, a processor, and a hierarchical adaptive code generation program stored on the processor. When the hierarchical adaptive code generation program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are performed.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a hierarchical adaptive code generation program, and the hierarchical adaptive code generation program executes the steps of the method according to any one of claims 1 to 5 when executed by a processor.
Citation Information
Patent Citations
Multi-language machine translation method and device, electronic equipment and storage medium
CN113239710A
Code automatic completion method, system and device based on deep reinforcement learning
CN118519681A