Code dynamic completion method for large language model

By constructing a Transformer model architecture with mixed training data sets and initializing the fusion probability spatial topological transformation, combining spherical probability mapping and taboo table mechanisms, the problem that traditional code completion technology is difficult to adapt to code diversity and variability is solved, and higher code completion accuracy and syntax correctness are achieved.

CN120085873AActive Publication Date: 2025-06-03SHENZHEN HAIYUNAN NETWORK SECURITY TECH CO LTD

Patent Information

Application Number
CN202510564227.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Traditional code completion techniques are difficult to adapt to the diversity and variation of code, cannot handle complex or unseen code patterns, and deep learning-based methods are difficult to ensure syntactic rules and semantic constraints of generated code.

Method used

A dynamic completion method for coding for large language models is proposed. By constructing a hybrid training data set, including multilingual code snippets, syntax constraint annotation and domain pre-training corpus, the Transformer model architecture of fusion probability spatial topological transformation is initialized, and the probability distribution generated is combined with spherical probability mapping and taboo table mechanism.

Benefits of technology

It improves the accuracy and syntax correctness of code completion, enhances the model's satisfaction with syntax and semantic constraints, and significantly improves the compilation pass rate of long code snippets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120085873A_ABST
    Figure CN120085873A_ABST
Patent Text Reader

Abstract

The invention provides a code dynamic completion method for a large language model. The method comprises the following steps: constructing a mixed training data set; the method comprises the following steps of: initializing a Transform model architecture fused with probability space topological transformation; based on the mixed training data set, executing model training of collaborative optimization of the main target and the auxiliary target; in each step of decoding, the model calculates the probability distribution of the next token according to the current context; the generated probability distribution is restrained in an effective probability subspace through spherical probability mapping, and grammar error candidates are filtered in real time and the probability distribution is adjusted in combination with a tabu table mechanism; and dynamically selecting a sampling strategy according to the current decoding depth, sampling from the adjusted probability distribution to obtain a next code snippet, and finally generating a code completion suggestion conforming to abstract syntax tree rules and semantic constraints. According to the method, the probability distribution constraint and the taboo table mechanism are combined, the model can dynamically adjust the generated probability distribution, the more appropriate candidate token can obtain the higher probability, and therefore the accuracy of code completion is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of code completion, and particularly to a coding dynamic completion method for large language models. Background Art

[0002] In the process of software development, code completion technology can help programmers quickly and accurately complete code snippets, so as to improve programming efficiency and reduce programming errors, thereby accelerating the development cycle and reducing the error rate.

[0003] Traditional code completion technologies often rely on predefined code rules or templates, and generate completion suggestions by matching the current context with the patterns in the rule library. However, such completion methods have great limitations: due to the fixity of the rules or templates, it is difficult to adapt to the diversity and variability in the code, and it is impossible to handle complex or unseen code patterns; as programming languages and frameworks are constantly updated, the rule library needs to be frequently updated and maintained, increasing the complexity and cost of technical implementation.

[0004] Currently, some code completion methods based on deep learning have also emerged, such as RNN, LSTM, Transformer, etc. These methods use deep learning models to learn complex patterns and dependencies in the code and can generate more accurate code completion suggestions. But there are also some limitations as follows: it is often difficult to ensure compliance with syntax rules and semantic constraints when generating code, resulting in syntax errors or semantic inconsistencies in the generated code; during the decoding process, the probability distribution generated by the model may contain a large number of candidate tokens that do not conform to syntax or semantics, lacking an effective mechanism to constrain and optimize its probability distribution, reducing the accuracy of code completion. Summary of the Invention

[0005] Based on this, the purpose of the present invention is to propose a coding dynamic completion method for large language models to solve the above-mentioned problems.

[0006] A coding dynamic completion method for large language models according to the present invention, the method includes: Construct a mixed training data set, including multi-language code snippets, syntax constraint annotations, and domain pre-training corpora; Initialize the Transformer model architecture with fused probability space topological transformation, and define the hypersphere dimension and dynamic effective region parameters; Based on the mixed training data set, perform model training with collaborative optimization of the main objective and the auxiliary objective, where the main objective is the code completion accuracy rate, and the auxiliary objectives include syntax tree validity prediction and semantic constraint satisfaction discrimination; At each decoding step, the model calculates the probability distribution of the next token according to the current context; During the model decoding stage, the generated probability distribution is constrained within the effective probability subspace through spherical probability mapping. Combining with the taboo table mechanism, syntax error candidates are filtered in real-time and the probability distribution is adjusted; Dynamically select a sampling strategy according to the current decoding depth. Based on the selected sampling strategy, sample the next code snippet from the adjusted probability distribution, and finally generate code completion suggestions that conform to the abstract syntax tree rules and semantic constraints.

[0007] Furthermore, the steps of initializing the Transformer model architecture that fuses probability space topological transformation include: Insert a probability space transformation module into the standard Transformer architecture; Define the parameters of the hypersphere probability space; Set the dynamic effective region parameters and constrain the generated probability distribution through the spherical mapping function; Load domain pre-trained weights when initializing the model parameters.

[0008] Furthermore, the steps of the model calculating the probability distribution of the next token according to the currently generated code sequence at each decoding step include: Convert the currently generated code sequence into an embedding vector; Capture the internal dependencies of the currently generated code sequence through the masked self-attention mechanism; Use the output of the self-attention layer as the query, and the output of the encoder as the key and value; Integrate the output of the encoder and the information of the currently generated code sequence, perform a non-linear transformation on the integrated information to generate the final output representation; Map the final output representation to the space of the vocabulary size through a linear layer; Use the Softmax function to generate the predicted probability distribution of each candidate token in the output sequence as the next token, and use the set of all candidate tokens in the output sequence as the candidate sequence.

[0009] Furthermore, the steps of constraining the generated probability distribution within the effective probability subspace through spherical probability mapping, combining with the taboo table mechanism, filtering syntax error candidates in real-time and adjusting the probability distribution include: Map the generated probability distribution to the unit sphere, and achieve uniform constraint of the probability distribution through octahedral parameterization; Create a taboo table to record the error patterns and their probability characteristics generated during the historical decoding process; At each decoding step, query the taboo table according to the current candidate sequence to filter known error patterns, where the candidate sequence is the sequence composed of all possible tokens generated by the model for the next token position; Recalculate the spherical probability for the filtered candidate sequences to improve the generation probability of valid candidate tokens; Dynamically adjust the spherical mapping parameters and the taboo table update strategy according to the decoding progress.

[0010] Furthermore, the step of mapping the generated probability distribution to the unit sphere and realizing the uniform constraint of the probability distribution through octahedral parameterization includes: Convert the high-dimensional vector representation of each generated candidate token into spherical coordinates through the octahedral mapping formula. The mapping formula is: |x| + |y| + |z| = 1, where x, y, and z are the coordinates of the generated candidate token in three-dimensional space respectively; Normalize the mapped coordinates to ensure that all candidate tokens are located on the unit sphere; Calculate the probability value of each candidate token through the spherical coordinate system so that the probability distribution satisfies the uniformity and effectiveness of the spherical space.

[0011] Furthermore, the step of querying the taboo table according to the current candidate sequence and filtering known error patterns at each decoding step includes: Perform a syntax check on the currently generated candidate sequence and extract the error pattern; Query the taboo table. If the candidate sequence contains the error pattern in the taboo table, filter it out; Update the filtered candidate sequence.

[0012] Furthermore, the step of recalculating the spherical probability for the filtered candidate sequences to improve the generation probability of valid candidate tokens includes: Resample the filtered candidate sequences through the spherical probability mapping method; Calculate the spherical probability value of each valid candidate token and adjust the probability distribution according to the spherical probability value so that the adjusted probability distribution is more concentrated on the valid candidates.

[0013] Furthermore, the step of dynamically adjusting the spherical mapping parameters and the taboo table update strategy according to the decoding progress includes: Monitor the code generation quality and syntax validity during the decoding process and adjust the spherical mapping parameters in real time; Dynamically update the taboo length and capacity of the taboo table according to the filling situation of the taboo table and the frequency of error patterns; Through the reinforcement learning mechanism, give higher taboo priorities to the candidates that lead to error patterns and optimize the self-update strategy of the taboo table.

[0014] Furthermore, the step of dynamically updating the taboo length and capacity of the taboo table according to the filling situation of the taboo table and the frequency of error patterns includes: Set the optimal interval of the taboo length, and randomly select a taboo length from the optimal interval of the taboo length as the initial taboo length; If the optimal solution formed by combining the optimal candidate token in the current candidate sequence with the currently generated code sequence is superior to the currently recorded optimal solution in terms of the evaluation criterion, increase the taboo length; Otherwise, decrease the taboo length; Dynamically adjust the capacity of the taboo table according to the frequency of the error pattern to ensure that the taboo table can effectively store and update the error pattern.

[0015] Furthermore, the step of dynamically selecting a sampling strategy according to the current decoding depth, sampling the next code snippet from the adjusted probability distribution based on the selected sampling strategy, and finally generating a code completion suggestion that conforms to the abstract syntax tree rules and semantic constraints includes: During the code completion process, monitor the length of the generated code sequence in real time to determine the current decoding depth; When the current decoding depth is relatively shallow, adopt the Top-k sampling strategy and select the k code snippets with the highest probabilities in the probability distribution as candidates; When the current decoding depth is relatively deep, adopt the temperature sampling strategy and control the randomness of the sampling process by adjusting the temperature parameter; Based on the selected sampling strategy, sample the next code snippet from the adjusted probability distribution; Add the sampled code snippet to the generated code sequence to form a code completion suggestion that conforms to the abstract syntax tree rules and semantic constraints.

[0016] In summary, the encoding dynamic completion method for large language models of the present invention integrates multi - language code snippets, explicit syntax constraint annotations, and domain - specific pre - trained corpora to construct a hybrid training set. The syntax constraint annotations provide the model with clear syntax rule information, enabling the model to better understand the syntax structure of the code during the learning process. The domain pre - trained corpora help the model master the semantic knowledge of specific domains, so that more domain - semantic - compliant code can be generated during code completion. By introducing hyper - sphere dimension constraints and dynamic effective region parameters in Transformer, the generated probability distribution is restricted within the syntax - valid subspace, which can more accurately constrain the generated probability distribution and improve the accuracy of code completion. A collaborative optimization framework of the main objective (i.e., completion accuracy) and the auxiliary objectives (including syntax tree validity prediction and semantic constraint satisfaction discrimination) is established to guide the model to not only focus on code completion itself during training, but also pay attention to whether the generated code conforms to syntax rules and semantic constraints, thus enhancing the model's satisfaction with syntax and semantic constraints. At each decoding step, the model aggregates the semantic vectors of the current context and generates a precise probability distribution through the self - attention mechanism. And after generating the probability distribution at each step, by converting the discrete token probability into a continuous probability flow optimization problem through spherical probability mapping and combining with the taboo list mechanism, the model can dynamically adjust the generated probability distribution, making more suitable candidate tokens obtain higher probabilities, thereby improving the accuracy of code completion. According to the decoding depth, the sampling strategy is dynamically adjusted, enabling the model to flexibly select the most appropriate sampling strategy according to different decoding stages, thus effectively balancing the problem of generation efficiency and quality and significantly improving the compilation passing rate of long code snippets.

[0017] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The above - mentioned and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which: Figure 1 It is a flowchart of a method for encoding dynamic completion of large language models according to Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.

[0020] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected to" another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used herein in the description of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items. Embodiment

[0022] Please refer to Figure 1 , the present invention provides a coding dynamic completion method for large language models, and the method includes steps S101 to S106: S101, construct a mixed training dataset, including multilingual code snippets, syntax constraint annotations, and domain pre-training corpora.

[0023] Fuse multilingual code snippets, explicit syntax constraint annotations, and domain-specific pre-training corpora to construct a mixed training set. This dataset enables the model to have the ability to understand multilingual paradigms and at the same time master domain-specific syntax rules (such as Python indentation, Java type declarations), which can significantly improve the adaptability of the model in different programming languages, thereby solving the semantic drift problem of traditional single-modal models during cross-domain completion.

[0024] Further optionally, the step of constructing the mixed training dataset includes: Collect open-source code libraries in different languages from multiple sources; Use a code parsing tool or library to parse the collected code into an abstract syntax tree structure; Utilize a static code analysis tool to annotate the parsed abstract syntax tree, extract syntax constraint features, including variable types, function parameters, syntax rule compliance, and potential problems; Fuse the code pattern data in the domain pre-training corpus with the annotated code; Generate mixed training samples containing syntax tags and semantic features according to the annotated syntax features and the fused code pattern data.

[0025] Specifically, code in different languages can be collected from multiple open-source code platforms. For example, collect the project code of the Python web development framework Django, the project code of the Java enterprise application development framework Spring Boot, the project code of the JavaScript front-end framework React, etc. from GitHub. These projects cover applications in different fields and of different scales, with rich code examples.

[0026] Use Python's Abstract Syntax Tree library to parse Python code. After parsing, its abstract syntax tree structure can be obtained, including function definition nodes, parameter nodes, return statement nodes, etc.

[0027] For the parsed abstract syntax tree, use a static code analysis tool for annotation. Extract syntactic constraint features from the output of the static code analysis tool. For example, for the following add function code, the extracted features include the function name add, the types of parameters a and b (assumed to be integers by inferring from comments or context), the return value type of the function as an integer, and the compliance with syntax rules (correct indentation), etc. Code example: def add(a, b): return a + b Extract code pattern data from the pre-trained corpus. For example, extract the code pattern for handling GET requests: from flask import Flask, request app = Flask(__name__) @app.route(' / get_data', methods=['GET']) def get_data(): data = request.args.get('param') return data Adopt the method of fine-tuning the pre-trained model to fuse the annotated code with the code pattern data extracted from the domain pre-trained corpus. That is, use the annotated code and the code pattern data as input to fine-tune the pre-trained model. During the fine-tuning process, let the model learn information such as syntactic features and domain-specific code patterns in the code. Make the fused data fully reflect the code features within the domain.

[0028] Generate mixed training samples containing syntactic labels and semantic features based on the labeled syntactic features and integrated code pattern data. For example, for the add function and the code handling GET requests, generate a training sample that includes syntactic labels of the function definition (such as function name, parameter types, return value type), semantic features of the code (such as the function's function is to add, and the function of handling GET requests is to obtain parameters and return), etc. The sample is as follows: { "code": "def add(a, b):\n return a + b", "language": "Python", "syntax_labels": { "function_name": "add", "parameters": ["a", "b"], "return_type": "int" }, "semantic_features": { "function_purpose": "Add two numbers" } }, { "code": "from flask import Flask, request\n\napp = Flask(__name__)\n\n@app.route(' / get_data', methods=['GET'])\ndef get_data():\ndata = request.args.get('param')\n return data", "language": "Python", "syntax_labels": { "framework": "Flask", "route": " / get_data", "method": "GET" }, "semantic_features": { "function_purpose": "Handle GET request and return parametervalue" } } Thus, a multi-modal hybrid training dataset containing multi-language code snippets, syntax constraint annotations, and domain pre-trained corpora is constructed.

[0029] S102. Initialize the Transformer model architecture for fusing probability space topological transformation, and define the hypersphere dimension and dynamic effective region parameters.

[0030] Introduce hypersphere probability space constraints in the Transformer to limit the generation process to the syntactically valid submanifold, making the code generated by the model more compliant with syntax rules. By dynamically adjusting the effective region parameters, the model can maintain an exploratory nature (wide probability distribution) at the initial stage of generation, enabling it to explore more code generation possibilities and increasing the diversity of the generated code. In the later stage, it gradually shrinks to legal candidates (narrow probability distribution) to avoid generating code that does not conform to syntax, thereby improving the syntactic correctness of the generated code.

[0031] Further optionally, the steps of initializing the Transformer model architecture for fusing probability space topological transformation include: Insert a probability space transformation module into the standard Transformer architecture; Define hypersphere probability space parameters; Set the dynamic effective region parameters and constrain the generation probability distribution through a spherical mapping function; Load domain pre-trained weights when initializing the model parameters.

[0032] Specifically, insert a probability space transformation module after the decoder layer of the Transformer. The probability space transformation module includes a spherical mapping layer and a dynamic effective region adjustment mechanism; Set the hypersphere dimension N of the probability space and define the radius of the standardized hypersphere surface; Randomly sample an initial center vector on the hypersphere surface, and use orthogonality constraints to ensure ; Set the initial effective radius and define the effective region as a hypersphere cap region centered at with a radius of ; Define the effective region expansion coefficient α ∈ [0.9, 1.1], and set the update rule for the effective region radius as to achieve dynamic adjustment of the effective region. Construct a spherical mapping function. When v≠0, ; when v = 0, ; Map the probability vector output by the model to the hypersphere through the spherical mapping function to achieve probability space transformation; Load the pre-trained parameters from the standard GPT model and freeze the parameters of its first L - 1 layers, where L is the total number of layers of the GPT model; Perform Xavier initialization on the newly added parameters while keeping the GPT parameters unchanged. The newly added parameters include the center vector c 0 and the effective radius , and the GPT parameters include the weight matrix W and the bias term b; Implement parameter constraints after each parameter update to ensure that the center vector and the effective radius are within a reasonable range.

[0033] S103. Based on the mixed training dataset, perform model training with collaborative optimization of the main objective and the auxiliary objectives. The main objective is the code completion accuracy, and the auxiliary objectives include the prediction of the validity of the syntax tree and the discrimination of the satisfaction of semantic constraints; Adopt a multi-task learning framework, where the main task (completion accuracy) and the auxiliary tasks (prediction of the validity of the syntax tree, discrimination of semantic constraints) form complementary supervision. For example, when the model generates candidates with high probability but syntax errors, the auxiliary tasks will backpropagate constraint signals, forcing the model to recalibrate the probability distribution, and finally the passing rate of the generated code under the verification of the abstract syntax tree parser is significantly improved.

[0034] Further optionally, the steps of the model training with collaborative optimization of the main objective and the auxiliary objectives based on the mixed training dataset include: Define a joint loss function, which is composed of the weighted cross-entropy loss of the code completion task, the binary cross-entropy loss of the syntax tree validity prediction task, and the binary cross-entropy loss of the semantic constraint satisfaction discrimination task; Through the backpropagation algorithm and the adaptive optimizer, iteratively optimize the model parameters on the mixed training dataset to achieve the collaborative optimization of the main objective and the auxiliary objectives.

[0035] Specifically, when training the model, first preprocess the mixed training dataset, including operations such as data cleaning, tokenization, and encoding. The mixed training dataset contains multi-language code snippets, syntax constraint annotations, and domain pre-trained corpora.

[0036] Then, calculate the cross-entropy loss between the code completion results predicted by the model and the true code, which is used to measure the prediction accuracy of the model in the code completion task and is the core optimization goal of the code completion scheme. For the syntax tree validity prediction task, calculate the binary cross-entropy loss between the syntax tree validity labels predicted by the model and the true labels, which is used to evaluate the model's ability to judge the validity of the syntax tree and assist the code completion scheme in generating code that conforms to the syntax rules. For the semantic constraint satisfaction discrimination task, calculate the binary cross-entropy loss between the semantic constraint satisfaction labels predicted by the model and the true labels to measure the model's discrimination ability for the degree of satisfaction of semantic constraints and ensure that the code generated by the code completion scheme is semantically reasonable and consistent. Combine the loss functions of the above three tasks with a certain proportion of weighting to form a joint loss function. During the training process, dynamically adjust the weights of each task in the joint loss function according to the performance of the model on each task. For example, when the performance of the model improves rapidly on a certain task, the weight of this task can be appropriately reduced to encourage the model to achieve better performance on other tasks.

[0037] Input the preprocessed mixed training dataset into the initialized model for forward propagation calculation to obtain the prediction results on each task (including code completion results, syntax tree validity labels, semantic constraint satisfaction labels). And calculate the total loss of the model on the current batch of data according to the joint loss function. Then, through the backpropagation algorithm, calculate the gradient of the loss function with respect to the model parameters. Use an adaptive optimizer to update the model parameters according to the calculated gradient to minimize the joint loss function.

[0038] Repeat the above steps of forward propagation, calculating the joint loss, backpropagation, and parameter update until the preset number of training epochs is reached or other stopping conditions are met.

[0039] S104. At each decoding step, the model calculates the probability distribution of the next token based on the current context.

[0040] In the decoding stage, use the trained model to gradually calculate the probability distribution of the next token at each decoding step based on the current context. Specifically, the model will receive the currently generated token sequence (i.e., the context) and predict the next possible token based on the language patterns and probability distributions learned internally. This prediction process will iterate continuously until a complete text sequence is generated or the preset stopping condition is reached.

[0041] The probability distribution generated by the model is the basis for subsequent spherical probability mapping and taboo list mechanism operations. Subsequently, the generated probability distribution can be constrained within the effective probability subspace through spherical probability mapping, and combined with the taboo list mechanism to filter syntax error candidates in real time and adjust the probability distribution, further improving the quality and syntax correctness of the generated code.

[0042] Further optionally, the step of the model calculating the probability distribution of the next token according to the currently generated code sequence at each decoding step includes: Convert the currently generated code sequence into an embedding vector; Through the masked self-attention mechanism, capture the internal dependencies of the currently generated code sequence; Use the output of the self-attention layer as the query, and the output of the encoder as the key and value; Integrate the output of the encoder and the information of the currently generated code sequence, perform a non-linear transformation on the integrated information, and generate the final output representation; Map the final output representation to a space of the vocabulary size through a linear layer to associate the output of the model with all possible code tokens; Use the Softmax function to generate the predicted probability distribution of each candidate token in the output sequence as the next token, and use the set of all candidate tokens in the output sequence as the candidate sequence.

[0043] Specifically, each token of the currently generated code sequence is converted into its corresponding word embedding vector. In this process, each token is represented as a high-dimensional vector.

[0044] Through the self-attention mechanism, the model is allowed to consider all tokens in the currently generated sequence when generating the next token. The masked self-attention mechanism can ensure that when the model generates the t-th token, it can only access the information of the 1st to the t-1-th tokens. In this process, the model calculates the attention weights between each token and other tokens, and these weights reflect the dependencies between the tokens.

[0045] In the decoder of the Transformer model, use the output of the self-attention layer as the query, and the output of the encoder as the key and value, in order to combine the information of the encoder (key and value) and the output of the decoder self-attention layer (query). Enable the model to integrate information from the encoder and decoder, and understand the input sequence and the generated sequence more comprehensively.

[0046] Then, integrate the output of the encoder and the information of the currently generated code sequence, perform a non-linear transformation on the integrated information to generate the final output representation, that is, a final high-dimensional vector representation, which captures the complex relationship between the input sequence and the generated sequence.

[0047] Map the final output representation to a space of the size of the vocabulary through a linear layer, that is, map this high-dimensional vector representation of the final output to a vector space with the same size as the vocabulary. Each dimension corresponds to a token in the vocabulary, and each value in the mapped vector represents the score or probability of the corresponding token as the next token.

[0048] Then use the Softmax function to generate the predicted probability distribution of each candidate token in the output sequence as the next token, making the sum of the values of all dimensions equal to 1, and each value of each dimension represents the predicted probability of the corresponding token as the next token. And take the set of all candidate tokens in the output sequence as the candidate sequence: S105, in the model decoding stage, constrain the generated probability distribution within the valid probability subspace through spherical probability mapping, and combine the taboo table mechanism to filter out syntax error candidates in real time and adjust the probability distribution.

[0049] At each step of model decoding, based on the probability distribution generated by the model, constrain the generated probability distribution within the valid probability subspace through spherical probability mapping, and combine the taboo table mechanism to filter out syntax error candidates in real time and adjust the probability distribution, further improving the quality and syntax correctness of the generated code.

[0050] Further optionally, the step of constraining the generated probability distribution within the valid probability subspace through spherical probability mapping, combining the taboo table mechanism, filtering out syntax error candidates in real time and adjusting the probability distribution includes: Map the generated probability distribution to the unit sphere and achieve uniform constraint of the probability distribution through octahedral parameterization; Create a taboo table to record the error patterns and their probability characteristics generated during the historical decoding process; At each step of decoding, query the taboo table according to the current candidate sequence to filter out known error patterns; Recalculate the spherical probability for the filtered candidate sequence to increase the generation probability of valid candidate tokens; Dynamically adjust the spherical mapping parameters and the taboo table update strategy according to the decoding progress.

[0051] It is understandable that the primary task of the code completion system is to generate syntactically correct code. Traditional models generate code snippets with syntax errors during the decoding process, such as unclosed parentheses, type mismatches, etc., resulting in compilation failures or runtime errors. By means of spherical probability mapping, the generated probability distribution is restricted within a syntactically valid subspace, which can significantly reduce the probability of generating code with syntax errors. Specifically, the probability distribution can be mapped to the unit sphere through spherical probability mapping, and uniform constraints can be achieved using methods such as octahedral parameterization, making the model more inclined to select syntactically correct candidate tokens. This kind of constraint can ensure the syntactic correctness of the generated code and improve the compilability and executability of the code.

[0052] During the decoding process, dynamically identify and filter out candidate tokens that may cause syntax errors, thereby further improving the quality of the generated code. The taboo list mechanism is used to record the error patterns and their probability characteristics that occur in the historical decoding process. At each step of decoding, the model queries the taboo list to filter out those candidate sequences that are known to cause syntax errors. This real-time filtering mechanism can effectively avoid repeating the same syntax errors and improve the accuracy and reliability of the generated code.

[0053] After filtering out the syntax error candidates, readjust the probability distribution so that the remaining valid candidate tokens have a higher generation probability, thereby optimizing the generation result. Specifically, the spherical probability of the filtered candidate sequences can be recalculated, and the probability distribution can be adjusted according to the spherical probability values. This adjustment can ensure that the model is more inclined to select those candidate tokens that are syntactically correct and semantically reasonable when generating code, improving the quality and practicality of the generated code.

[0054] While ensuring the syntactic correctness of the generated code, maintain the diversity of the generation and avoid over-constraint resulting in overly single or lack of innovation in the generated code. Specifically, by dynamically adjusting the spherical mapping parameters and the taboo list update strategy, the model can adjust the strictness of the constraints in real time according to the decoding progress and the quality of the generated code. For example, a relatively loose constraint can be maintained at the initial stage of decoding to encourage generation diversity; while at the later stage of decoding, the constraint can be gradually tightened to ensure the syntactic correctness and semantic consistency of the generated code. This dynamic adjustment mechanism enables the model to find a balance between syntactic constraints and generation diversity.

[0055] Further optionally, the step of mapping the generated probability distribution to the unit sphere and achieving uniform constraint of the probability distribution through octahedral parameterization includes: Convert the high-dimensional vector representation of each generated candidate token into spherical coordinates through the octahedral mapping formula. The mapping formula is: |x| + |y| + |z| = 1, where x, y, and z are the coordinates of the generated candidate token in three-dimensional space respectively; Normalize the mapped coordinates to ensure that all candidate tokens lie on the unit sphere; Calculate the probability values of each candidate token through the spherical coordinate system, such that the probability distribution satisfies the uniformity and validity of the spherical space.

[0056] In this embodiment, the high-dimensional vector representations of the generated candidate tokens are converted into spherical coordinates through the octahedron mapping formula, and the probability values of each candidate token are calculated, thereby realizing the uniform constraint of the probability distribution. The uniform constraint can ensure that the generated probability distribution is smoother and more reasonable, thereby improving the quality of the generated code.

[0057] The following example is used for illustration: Suppose that in a certain decoding step, the model generates three candidate tokens for the next token: A, B, C, and the original probability distribution is: 0.2, 0.3, 0.5. Each token is represented as a high-dimensional vector inside the model. For example, Token A: the vector representation is (0.5, 0.5, 0.5), Token B: the vector representation is (-0.5, 0.5, 0), Token C: the vector representation is (0, -0.5, 0.5).

[0058] Map the above three-dimensional vectors to the surface of the unit octahedron through octahedron mapping and approximately satisfy the spherical coordinates. The octahedron mapping formula is |x| + |y| + |z| = 1. Direct application of this formula may require scaling and adjustment of the vectors.

[0059] Token A is mapped to (0.267, 0.267, 0.466) through scaling and adjustment (satisfying |0.267| + |0.267| + |0.466| ≈ 1); Token B is mapped to (-0.267, 0.267, 0.466); Token C is mapped to (0, -0.267, 0.733).

[0060] The mapped coordinates approximately satisfy the surface of the unit octahedron, but to ensure that they strictly lie on the unit sphere (the modulus length is 1), normalization is performed. The normalization formula is: normalized coordinate = mapped coordinate / ||mapped coordinate||, where ||mapped coordinate|| is the modulus length of the mapped coordinate.

[0061] In the spherical coordinate system, calculate their probability values according to the positions of the candidate tokens. It can be calculated based on the area ratio occupied by the candidate tokens on the sphere, and the probability of Token A is 0.4, the probability of Token B is 0.3, and the probability of Token C is 0.3.

[0062] Through the above steps, the high-dimensional vector representation of the generated candidate tokens is converted into spherical coordinates by the octahedron mapping formula, and the probability value of each candidate token is calculated, thereby realizing the uniform constraint of the probability distribution.

[0063] Further optionally, the step of querying the taboo table according to the current candidate sequence and filtering known error patterns at each decoding step includes: Perform a syntax check on the currently generated candidate sequence and extract the error pattern; Query the taboo table. If the candidate sequence contains an error pattern in the taboo table, filter it out; Update the filtered candidate sequence.

[0064] It can be understood that at each step of decoding, it is first necessary to identify possible syntax errors in the currently generated candidate sequence. Specifically, each candidate token in the candidate sequence can be combined with the currently generated code sequence in turn to simulate the possible state of the code sequence if the candidate token is selected, which can more comprehensively evaluate the rationality of the candidate token. For each candidate token in the candidate sequence, add it to the next token position (the token position to be generated) of the currently generated code sequence in turn to form a new code sequence combination.

[0065] The taboo table stores known error patterns that occurred in the historical decoding process. By querying the taboo table, those code sequence combinations containing error patterns can be identified. Specifically, each new code sequence combination can be matched with the patterns in the taboo table. If a combination matches an error pattern in the taboo table, mark the candidate token as likely to cause an error and filter it.

[0066] Remove those candidate tokens marked as likely to cause errors from the candidate sequence. In this way, in subsequent decoding steps, these tokens will no longer be considered, thus avoiding potential errors. By filtering out those candidate tokens that may cause errors, it can be ensured that only those safe and effective candidate tokens are considered in the subsequent decoding process.

[0067] At each step of decoding, repeat the above filtering steps until a complete code sequence is generated.

[0068] Further optionally, the step of recalculating the spherical probability for the filtered candidate sequence to increase the generation probability of effective candidate tokens includes: Resample the filtered candidate sequence through the spherical probability mapping method; Calculate the spherical probability values for each valid candidate token, and adjust the probability distribution according to the spherical probability values so that the adjusted probability distribution is more concentrated on the valid candidates.

[0069] Understandably, using a spherical probability mapping method (such as octahedral parameterization), the filtered candidate sequence is remapped onto the unit sphere. Resampling is to re-evaluate the distribution of the remaining candidate sequences based on the spherical probability mapping method after filtering out the invalid candidates, ensuring more accurate subsequent probability calculations.

[0070] In the spherical coordinate system, according to the positions of the valid candidate tokens, calculate the spherical probability values for each filtered valid candidate token. These probability values quantify the distribution density or importance of the valid candidate tokens in the spherical space.

[0071] According to the calculated spherical probability values, adjust the originally generated probability distribution. The probability values of the valid candidate tokens can be increased, while the probability values of the invalid candidate tokens are decreased (or kept at zero). The adjusted probability distribution is more concentrated on the valid candidates, thereby improving the quality of the generated code.

[0072] When adjusting the probability distribution, it is necessary to ensure that the adjusted distribution still maintains the uniformity of the spherical space (i.e., the probability density of any region on the sphere is proportional to the area of that region) and validity (i.e., the sum of all probability values is 1, and each probability value is between 0 and 1).

[0073] In this embodiment, by recalculating the spherical probability and adjusting the probability distribution, the generation probability of the valid candidate tokens can be significantly improved, thereby improving the quality of the generated code. By concentrating on the valid candidates, the model can converge to the optimal solution faster, thus accelerating the decoding process.

[0074] Further optionally, the step of dynamically adjusting the spherical mapping parameters and the taboo table update strategy according to the decoding progress includes: Monitor the code generation quality and syntax validity during the decoding process, and adjust the spherical mapping parameters in real time; According to the filling situation of the taboo table and the frequency of error patterns, dynamically update the taboo length and capacity of the taboo table; Through a reinforcement learning mechanism, give higher taboo priorities to the candidates that lead to error patterns, and optimize the self-update strategy of the taboo table.

[0075] It is understandable that in the code completion or generation task, in order to improve the quality and efficiency of the generated code, the spherical mapping parameters and the taboo table update strategy can be dynamically adjusted according to the decoding progress. Specifically, during the decoding process, the quality and syntax validity of the generated code are monitored in real time to promptly detect and correct potential problems. And according to the monitoring results, the spherical mapping parameters are adjusted in real time, such as the coefficients of the mapping function, the dimension of the spherical space, etc., to optimize the probability distribution of the generated code and improve the generation quality. Through monitoring feedback, the spherical mapping parameters are continuously iteratively optimized until a satisfactory generation effect is achieved.

[0076] The taboo table is used to store the error patterns that occurred during the historical decoding process to avoid making the same mistakes repeatedly. By dynamically updating the taboo length and capacity of the taboo table according to the filling situation of the taboo table and the frequency of the error patterns, the effectiveness and efficiency of the taboo table can be improved.

[0077] The reinforcement learning mechanism can optimize the decision-making strategy according to the environmental feedback. In the code generation task, the reinforcement learning mechanism can give higher taboo priorities to the candidates that lead to error patterns, thereby optimizing the self-update strategy of the taboo table and improving the quality of the generated code.

[0078] Specifically, a reinforcement learning model is designed, and this reinforcement learning model can adjust its decision-making strategy according to feedback signals such as the quality of the generated code and syntax validity. During the decoding process, when a certain candidate token causes a syntax error, the reinforcement learning model gives a higher taboo priority to this candidate token, that is, it is more likely to add it to the taboo table. By continuously iteratively training the reinforcement learning model and optimizing the self-update strategy of the taboo table, the model can more effectively avoid making the same mistakes repeatedly in the subsequent decoding process.

[0079] In this embodiment, by dynamically adjusting the spherical mapping parameters and the taboo table update strategy, the quality and syntax validity of the generated code can be significantly improved. The dynamic adjustment strategy enables the model to adapt to different decoding progress and generation environments, improving the robustness and flexibility of the model. And by optimizing the self-update strategy of the taboo table, the model can converge to the optimal solution faster, reducing the unnecessary exploration and trial-and-error process.

[0080] Further optionally, the step of dynamically updating the taboo length and capacity of the taboo table according to the filling situation of the taboo table and the frequency of the error patterns includes: Set an optimal interval for the taboo length, and randomly select a taboo length from the optimal interval for the taboo length as the initial taboo length; If the optimal solution formed by combining the optimal candidate token in the current candidate sequence with the currently generated code sequence is better than the currently recorded optimal solution in terms of the evaluation criteria, increase the taboo length; Otherwise, decrease the taboo length; Dynamically adjust the capacity of the taboo table according to the frequency of error patterns to ensure that the taboo table can effectively store and update error patterns.

[0081] Understandably, the taboo table can be used to store and avoid repeated access to those token patterns (or so-called bad candidate patterns) that are known to be non-optimal or will cause error codes during the decoding process. To make more efficient use of the taboo table, the taboo length and capacity of the taboo table can be dynamically adjusted according to its filling situation and the occurrence frequency of error token patterns.

[0082] Specifically, first set an optimal interval for the taboo length, and randomly select a value from this interval as the initial taboo length. This taboo length determines how long a certain error token pattern (or bad candidate pattern) will be retained in the taboo table, and thus affects the diversity and convergence speed of the generation process. For example, the initial taboo length of a dynamic language (such as Python) can be set to 5 - 8 steps. The initial taboo length of a static language (such as Java) can be set to 8 - 12 steps.

[0083] During the generation process, if the optimal candidate token in the current candidate sequence (according to the probability value) forms an optimal solution when combined with the currently generated code sequence, and this optimal solution is better than the currently recorded optimal solution (i.e., the optimal solution found in the previous round or previous generation process) in terms of the evaluation criteria, then increase the taboo length. This can make this pattern (or a similar pattern) be retained for a longer time in the subsequent code generation process, avoiding returning to sub-optimal token combinations prematurely, and thus promoting the generation process towards a better token combination. On the contrary, if no solution containing a better token combination is found, then decrease the taboo length. This can increase the diversity of the generation, enabling the model to have the opportunity to explore token patterns (or bad candidate patterns) that were previously tabooed, and thus may discover new and better token combinations.

[0084] In addition, the capacity of the taboo table can also be dynamically adjusted according to the frequency of error token patterns, ensuring that the taboo table can store enough error token patterns while not causing low storage and update efficiency due to excessive capacity. The capacity of the taboo table directly determines the number of error token patterns it can store.

[0085] This embodiment makes the generation process more efficient by dynamically adjusting the taboo length and capacity. This adjustment mechanism can not only avoid premature convergence to sub-optimal token combinations but also maintain the diversity of the generation, thus being able to better adapt to the characteristics of different code completion tasks and generation stages.

[0086] S106. Dynamically select a sampling strategy according to the current decoding depth. Based on the selected sampling strategy, sample the next code snippet from the adjusted probability distribution, and finally generate a code completion suggestion that conforms to the abstract syntax tree rules and semantic constraints.

[0087] Dynamically select an appropriate sampling strategy according to the current decoding depth, sample the next code snippet from the probability distribution adjusted by the spherical probability mapping and taboo list mechanism, and finally generate a code completion suggestion that conforms to both the abstract syntax tree rules and semantic constraints. Its core goal is to establish a dynamic balance between syntax rules and generation freedom during the code completion process, improve the compilability and practicality of code completion, and at the same time balance the contradiction between exploration and exploitation to improve the compilation pass rate of long code snippets.

[0088] The steps of dynamically selecting a sampling strategy according to the current decoding depth, sampling the next code snippet from the adjusted probability distribution based on the selected sampling strategy, and finally generating a code completion suggestion that conforms to the abstract syntax tree rules and semantic constraints include: During the code completion process, monitor the length of the generated code sequence in real time to determine the current decoding depth. When the current decoding depth is relatively shallow, adopt the Top-k sampling strategy, and select the k code snippets with the highest probabilities in the probability distribution as candidates to enhance the diversity of the generated code completion suggestions. When the current decoding depth is relatively deep, adopt the temperature sampling strategy, and control the randomness of the sampling process by adjusting the temperature parameter to ensure the semantic consistency of the generated code completion suggestions. Based on the selected sampling strategy, sample the next code snippet from the adjusted probability distribution. Add the sampled code snippet to the generated code sequence to form a code completion suggestion that conforms to the abstract syntax tree rules and semantic constraints.

[0089] Specifically, during the code completion process, monitor the length of the generated code sequence in real time to determine the current decoding depth. When the current decoding depth is relatively shallow (e.g., <5 steps), adopt the Top-k sampling strategy. This strategy selects the k code snippets with the highest probabilities in the probability distribution as candidates. That is, at the initial stage of decoding, when the code generation is in the exploration stage, adopting this strategy can enhance the diversity of the generated code completion suggestions, encourage the model to generate more innovative code snippets, and avoid falling into local optima prematurely.

[0090] When the current decoding depth is relatively shallow (e.g., ≥5 steps), a temperature sampling strategy is adopted. The randomness of the sampling process is controlled by adjusting the temperature parameter. That is, in the later stage of decoding, code generation needs to pay more attention to semantic consistency to ensure that the generated code can be correctly compiled and run. The temperature sampling strategy can reduce randomness to a certain extent, making the model more inclined to select code fragments with higher probabilities, thus ensuring the semantic consistency of the generated code completion suggestions.

[0091] Based on the selected sampling strategy, the next code fragment is sampled from the probability distribution adjusted by the spherical probability mapping and the taboo list mechanism. Among them, the spherical probability mapping constrains the generated probability distribution within the effective probability subspace to ensure that the generated code is syntactically correct; the taboo list mechanism filters out syntax error candidates in real time, further improving the quality of the generated code.

[0092] The sampled code fragment is added to the generated code sequence to form a code completion suggestion that conforms to the abstract syntax tree rules and semantic constraints. In this way, a complete code completion result is gradually constructed to ensure that the generated code not only meets the syntax requirements but also has reasonable semantics.

[0093] In this embodiment, by dynamically selecting the sampling strategy, it is ensured that at different decoding depths, the model can find a balance between exploration and exploitation and generate high-quality code.

[0094] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.

Claims

1. A method for dynamic completion of encoding for a large language model, characterized in that: The method comprises: Build a hybrid training dataset that includes multi-language code snippets, grammatical constraint annotations, and domain pre-training corpus; Initialize the Transformer model architecture that integrates the topological transformation of the probability space, and define the hypersphere dimension and dynamic effective area parameters; Based on the mixed training data set, the model training is performed with the main goal and the auxiliary goal of co-optimization. The main goal is the code completion accuracy, and the auxiliary goals include the prediction of syntax tree validity and the judgment of semantic constraint satisfaction. At each decoding step, the model calculates the probability distribution of the next token based on the current context; In the model decoding stage, the generated probability distribution is constrained within the valid probability subspace through spherical probability mapping, and combined with the taboo table mechanism, grammatical error candidates are filtered in real time and the probability distribution is adjusted; The sampling strategy is dynamically selected according to the current decoding depth. Based on the selected sampling strategy, the next code snippet is sampled from the adjusted probability distribution, and finally code completion suggestions are generated that comply with the abstract syntax tree rules and semantic constraints.

2. The method for dynamic encoding completion of a large language model according to claim 1, characterized in that: The step of initializing the Transformer model architecture of the fusion probability space topology transformation includes: Insert a probability space transformation module into the standard Transformer architecture; Define the hypersphere probability space parameters; Set dynamic effective area parameters and generate probability distribution through spherical mapping function constraints; Load domain pre-trained weights when initializing model parameters.

3. The method for dynamic encoding completion of a large language model according to claim 1, characterized in that: At each decoding step, the model calculates the probability distribution of the next token based on the currently generated code sequence, including: Convert the current generated code sequence into an embedding vector; Through the masked self-attention mechanism, the internal dependencies of the currently generated code sequence are captured; The output of the self-attention layer is used as the query, and the output of the encoder is used as the key and value; Integrate the encoder output and the currently generated code sequence information, perform nonlinear transformation on the integrated information, and generate the final output representation; Map the final output representation to the space of vocabulary size through a linear layer; Use the Softmax function to generate the predicted probability distribution of each candidate token in the output sequence as the next token, and take the set of all candidate tokens in the output sequence as the candidate sequence.

4. The method for dynamic encoding completion of a large language model according to claim 1, characterized in that: The steps of constraining the generated probability distribution within the valid probability subspace by spherical probability mapping, combining the taboo table mechanism, filtering grammatical error candidates in real time and adjusting the probability distribution include: The generated probability distribution is mapped to the unit sphere, and the uniform constraint of the probability distribution is achieved through octahedral parameterization; Create a taboo table to record the error patterns and their probability characteristics generated during the historical decoding process; At each decoding step, the taboo table is queried based on the current candidate sequence to filter out known error patterns, where the candidate sequence is a sequence of all possible tokens generated by the model for the next token position; Recalculate the spherical probability of the filtered candidate sequence to improve the probability of generating valid candidate tokens; The spherical mapping parameters and taboo table update strategy are dynamically adjusted according to the decoding progress.

5. The method for dynamic encoding completion of a large language model according to claim 4, characterized in that: The step of mapping the generated probability distribution to the unit sphere and implementing uniform constraints on the probability distribution through octahedral parameterization includes: The high-dimensional vector representation of each generated candidate token is converted into spherical coordinates through the octahedral mapping formula. The mapping formula is: |x|+|y|+|z|=1, where x, y, and z are the coordinates of the generated candidate token in three-dimensional space respectively; Normalize the mapped coordinates to ensure that all candidate tokens are located on the unit sphere; The probability value of each candidate token is calculated through the spherical coordinate system, so that the probability distribution satisfies the uniformity and effectiveness of the spherical space.

6. The method for dynamic encoding completion of a large language model according to claim 4, characterized in that: The step of searching the taboo table according to the current candidate sequence and filtering known error patterns at each decoding step includes: Perform syntax check on the currently generated candidate sequence and extract error patterns; Query the taboo table, and if the candidate sequence contains an error pattern in the taboo table, filter it out; The filtered candidate sequences are updated.

7. The method for dynamic completion of encoding for a large language model according to claim 4, characterized in that: The step of recalculating the spherical probability of the filtered candidate sequence to improve the probability of generating a valid candidate token includes: The filtered candidate sequences are resampled by spherical probability mapping method; Calculate the spherical probability value of each valid candidate token, and adjust the probability distribution according to the spherical probability value, so that the adjusted probability distribution is more concentrated on the valid candidates.

8. The method for dynamic encoding completion of a large language model according to claim 4, characterized in that: The step of dynamically adjusting the spherical mapping parameters and the taboo table update strategy according to the decoding progress comprises: Monitor the quality of code generation and grammatical validity during decoding, and adjust spherical mapping parameters in real time; Dynamically update the taboo length and capacity of the taboo table based on the filling status of the taboo table and the frequency of the error pattern; Through the reinforcement learning mechanism, higher taboo priority is given to candidates that cause error patterns, and the self-update strategy of the taboo table is optimized.

9. The method for dynamic encoding completion of a large language model according to claim 8, characterized in that: The step of dynamically updating the taboo length and capacity of the taboo table according to the filling condition of the taboo table and the frequency of the error mode comprises: Set the optimal interval of taboo length, and randomly select a taboo length from the optimal interval of taboo length as the initial taboo length; If the optimal solution formed by combining the best candidate token in the current candidate sequence with the currently generated code sequence is better than the currently recorded optimal solution in terms of the evaluation criteria, the taboo length is increased; Otherwise, reduce the taboo length; According to the frequency of the error pattern, the capacity of the taboo table is dynamically adjusted to ensure that the taboo table can effectively store and update the error pattern.

10. The method for dynamic completion of encoding for a large language model according to claim 1, characterized in that: The steps of dynamically selecting a sampling strategy according to the current decoding depth, sampling the next code snippet from the adjusted probability distribution based on the selected sampling strategy, and finally generating a code completion suggestion that complies with the abstract syntax tree rules and semantic constraints include: During the code completion process, the length of the generated code sequence is monitored in real time to determine the current decoding depth; When the current decoding depth is shallow, the Top-k sampling strategy is used to select the k code snippets with the highest probability in the probability distribution as candidates; When the current decoding depth is deep, the temperature sampling strategy is adopted to control the randomness of the sampling process by adjusting the temperature parameters; Based on the selected sampling strategy, sample the next code snippet from the adjusted probability distribution; The sampled code snippets are added to the generated code sequence to form code completion suggestions that comply with the abstract syntax tree rules and semantic constraints.

Citation Information

Patent Citations

  • Code completion method and system based on large language model

    CN119356686A

  • System and method for performing code completion in an integrated development environment

    US20050015747A1

Cited By

  • Software quality assessment method based on dependency chain quality conduction

    CN120994244A

  • A software quality evaluation method based on dependency chain quality conduction

    CN120994244B

  • RAG financial credit decision-making method based on policy mask constraint decoding

    CN121234927A

  • Large language model structured preference alignment method and device, electronic equipment and medium

    CN121787541A

  • Large language model structured preference alignment method and device, electronic equipment and medium

    CN121787541B