Method for optimizing generation of code thinking chain in small model
By adopting novel data set construction and partition training strategies in small models, the COTTON_lite model was developed, which solved the problem of low quality of code thinking chain generation in resource-constrained environments, and achieved efficient and reliable high-quality thinking chain generation.
Patent Information
- Application Number
- CN202510274280.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The prior art is difficult to efficiently generate high-quality code thinking chains in resource-constrained environments, resulting in a decline in output quality.
A method to optimize the generation of code thinking chains in small models is proposed. Through novel data set construction methods and partition training strategies, a COTTON_lite model is developed. This model contains only 0.38B parameters and can efficiently generate high-quality thinking chains.
It realizes efficient generation of high-quality code thinking chains in resource-constrained environments, improves the reliability and interpretability of code generation tasks, and reduces resource consumption.
Smart Images

Figure CN120218176A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method for optimizing the generation of code thought chains in small models. Background Art
[0002] In recent years, with the emergence of large language models, significant progress has been made in code generation. These advancements have been mainly driven by large-scale models such as GPT4 and DeepSeek-v3, which typically contain over 100 billion parameters. However, applying these huge models to code generation tasks poses significant challenges in terms of time, computation, and financial costs, especially in resource-constrained environments such as limited LLM API access, data security issues, or limited GPU availability.
[0003] Through model architecture innovation and knowledge distillation, researchers have developed lightweight models such as Qwen2.5-Coder 7B and DeepSeek-Coder 6.7B, which can run on consumer-grade GPUs. Although these models show some potential, they still face difficulties in complex code generation tasks, often resulting in a decline in output quality. To address these issues, the thought chain technique has been widely adopted by researchers because it can improve model performance without the need for retraining or fine-tuning. Essentially, a thought chain consists of a series of intermediate natural language reasoning steps that guide the final output, enabling large language models to provide more reliable answers through deliberation and explanation.
[0004] Although these CoT-based methods show some promise, they typically rely on large models such as GPT4 to generate high-quality CoT, which introduces significant computational and financial overhead. This dependence makes them impractical in resource-constrained scenarios. To address this limitation, Yang et al.
[12] proposed COTTON, a lightweight alternative that uses a 7B-parameter model for CoT generation. Therefore, there is a need to explore a solution for efficiently generating code thought chains in resource-constrained environments.
[0005] How to solve the above technical problems becomes the subject of the present invention. Summary of the Invention
[0006] The object of the present invention is to provide a method for optimizing the generation of code thought chains in small models, which can automatically generate comments based on Bash code.
[0007] The idea of the present invention is: the present invention proposes a method for optimizing the generation of code thought chains in small models to improve the reliability and interpretability of code generation tasks. Traditional methods rely on large language models and are difficult to apply in resource-constrained environments. The COTTON_lite model developed by the present invention contains only 0.38B parameters and efficiently generates high-quality thought chains. Its innovations include: 1. Proposing a novel data set construction method, multiple large language models collaborate to generate high-quality thought chain samples, and the quality is guaranteed by a special evaluation model; 2. Designing a divide-and-conquer training strategy, the data set is divided into three difficulty subsets, and the model is trained in a targeted manner to improve performance through merging and fine-tuning. The present invention provides an efficient solution for code generation in resource-constrained environments.
[0008] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is specifically: a method for optimizing the generation of code thinking chains in a small model, which includes the following steps:
[0009] (1) Using three high-performance large-scale language models as teacher models, based on the CodeHarmony dataset, we use prompt strategies to generate corresponding thought chains, including the following steps:
[0010] (1-1) Three high-performance large-scale language models are selected as teacher models, namely GPT-4, DeepSeek-V2.5 and QwenCoder32B-Instruct;
[0011] (1-2) Based on the CodeHarmony dataset, which contains more than 15,000 code examples and their corresponding test cases;
[0012] (1-3) For each input data, the three large language models use the prompt strategy to generate the corresponding thought chain Cji;
[0013] (2) In order to ensure the semantic consistency and validity of the generated thought chain, a strict quality assessment framework was implemented based on functional correctness. After screening, the final data set was obtained. The data set format was set to <code segment, thought chain, code description>, which specifically included the following steps:
[0014] (2-1) For each generated thought chain Cji, the evaluation model E first uses the input data Xi and the thought chain Cji to synthesize the corresponding code Yi;
[0015] (2-2) Use the test case Ti to evaluate the functional correctness of the code Yi and calculate the thinking chain Cji score score(Cji). Where compiler(Y i,t) indicates compiling and validating the code Yi with the test case t. If the code passes the test case, it returns 0; otherwise, it returns a non-zero value. If the generated code Yi passes all the test cases, then score(Cji) = 1; otherwise, score(Cji) = 0;
[0016] (2-3) When score(Cji) is 1, that is, when Yi passes all the test cases, the corresponding chain-of-thought sample is selected as the candidate chain of thought; if score(Cji) is 0, then this sample is excluded from the dataset to construct the final dataset D;
[0017] (3) Select Qwen2.5-Coder-0.5Binstruct as the base model and adopt lexical pruning as the parameter reduction strategy to compress the model size to 0.38B. The specific steps are as follows:
[0018] (3-1) Construct the base large language model Qwen2.5-Coder-0.5Binstruct;
[0019] (3-2) According to the frequency of tokens in the training corpus, a frequency-based pruning function P(Vorig) is defined, P(Vorig) = {v ∈ Vorig | f(v) > τ}, where Vorig is the original vocabulary, v is a token in the vocabulary, f(v) is the frequency of the token v in the training corpus, and τ is an empirically determined frequency threshold;
[0020] (3-3) Retain the tokens in the original vocabulary Vorig with frequencies higher than the threshold τ to obtain the compressed vocabulary Vpruned;
[0021] (3-4) Set τ to 1, the original vocabulary size N to 151,936, and the pruned vocabulary size M to 23,954, reducing the total parameters from 0.5B to 0.38B while retaining the basic capabilities of the model;
[0022] (4) Divide the dataset into three subsets according to the training difficulty: easy, medium, and hard. The specific steps are as follows:
[0023] (4-1) Apply supervised fine-tuning on the base model and generate the chain of thought Csi for the entire training set to evaluate the inherent difficulty of each training sample;
[0024] (4-2) Divide the dataset D into three subsets: easy (Deasy), medium (Dmedium), and hard (Dhard) according to the BLEU score distribution between the generated chain of thought Csi and the candidate chain of thought Cji;
[0025] (5) Train three sub-models on three training subsets respectively using the divide-and-conquer training strategy, and merge and fine-tune the three sub-models by combining DARE with task arithmetic to obtain the final code chain of thought generation model COTTON_lite, which specifically includes the following steps:
[0026] (5-1) For each difficulty level k ∈ {easy, medium, hard}, train a dedicated model Mk respectively;
[0027] (5-2) Calculate the difference between the dedicated model and the base model of Mk, and then perform weighted combination according to the importance weight α determined by the validation performance;
[0028] (5-3) Apply dropout at each layer through the rescaling mechanism of DARE to prevent interference between different tasks;
[0029] (5-4) Use the complete dataset to comprehensively fine-tune the merged model to obtain the final model COTTON_lite;
[0030] (6) Deploy COTTON_lite on a single GPU device to generate corresponding chain of thought prompts during the code generation stage to guide code generation.
[0031] As a further optimization solution for the method of optimizing the code chain of thought generation in a small model provided by the present invention, in the steps (1) and (2), a brand-new high-quality dataset D is constructed; in the step (3), the vocabulary pruning technology is used to greatly reduce the number of model parameters; in the step (5), the divide-and-conquer idea is adopted to train and merge the models, providing an efficient solution for code generation in resource-constrained environments.
[0032] The optimal parameter settings of the method for optimizing the code chain of thought generation in a small model are as shown in Table 1 below:
[0033] Table 1
[0034]
[0035] Compared with the prior art, the beneficial effects of the present invention are as follows: The method for optimizing the code chain of thought generation in a small model proposed by the present invention proposes a novel dataset construction method, uses multiple large language models to collaborate to generate the chain of thought, and ensures the quality through a dedicated evaluation model; introduces a divide-and-conquer training strategy, divides the dataset into simple, medium and difficult subsets, trains the corresponding models respectively, and realizes progressive performance improvement through model merging and fine-tuning; proposes COTTON_lite, which is a CoT generation model with significantly reduced parameters and superior performance, and it can be seamlessly integrated with existing lightweight code generation models to provide enhanced generation performance while reducing resource consumption. Brief Description of the Drawings
[0036] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention.
[0037] Figure 1 It is a system framework diagram of a method for optimizing the generation of code thought chains in a small model provided by the present invention. Detailed Embodiments
[0038] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the drawings and embodiments. Of course, the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0039] Embodiment 1
[0040] See Figure 1 As shown, this embodiment provides a method for optimizing the generation of code thought chains in a small model, specifically including the following content:
[0041] (1) Using three high-performance large language models as teacher models, based on the CodeHarmony dataset, respectively using the prompting strategy to generate corresponding thought chains, specifically including the following steps:
[0042] (1-1) Selecting three high-performance large language models as teacher models, namely GPT-4, DeepSeek-V2.5 and QwenCoder32B-Instruct;
[0043] (1-2) Based on the CodeHarmony dataset, which contains more than 15,000 code examples and their corresponding test cases;
[0044] (1-3) For each input data, the three large language models use the prompting strategy to generate the corresponding thought chain Cji;
[0045] (2) In order to ensure the semantic consistency and effectiveness of the generated thought chains, a strict quality assessment framework is implemented based on functional correctness, and the final dataset is obtained after screening. The dataset format is set as <code segment, thought chain, code description>, specifically including the following steps:
[0046] (2-1) For each generated thought chain Cji, first evaluate that the model E synthesizes the corresponding code Yi using the input data Xi and the thought chain Cji;
[0047] (2-2) Use the test case Ti to evaluate the functional correctness of the code Yi, and calculate the score score(Cji) of the chain of thought Cji. where compiler(Y i , t) means to compile and verify the code Yi with the test case t. If the code passes the test case, return 0; otherwise, return a non-zero value. If the generated code Yi passes all test cases, then score(Cji) = 1; otherwise, score(Cji) = 0.
[0048] (2-3) When score(Cji) is 1, that is, Yi passes all test cases, the corresponding chain of thought sample is selected as a candidate chain of thought; if score(Cji) is 0, then the sample is excluded from the dataset to construct the final dataset D.
[0049] (3) Select Qwen2.5-Coder-0.5Binstruct as the base model, and use vocabulary pruning as a parameter reduction strategy to compress the model size to 0.38B. The specific steps are as follows:
[0050] (3-1) Construct the base large language model Qwen2.5-Coder-0.5Binstruct.
[0051] (3-2) According to the frequency of tokens in the training corpus, define a frequency-based pruning function P(Vorig), P(Vorig) = {v ∈ Vorig | f(v) > τ}, where Vorig is the original vocabulary, v is a token in the vocabulary, f(v) is the frequency of the token v in the training corpus, and τ is an empirically determined frequency threshold.
[0052] (3-3) Retain the tokens in the original vocabulary Vorig with frequencies higher than the threshold τ to obtain the compressed vocabulary Vpruned.
[0053] (3-4) Set τ to 1, the original vocabulary size N to 151,936, and the pruned vocabulary size M to 23,954. Reduce the total parameters from 0.5B to 0.38B while retaining the basic capabilities of the model.
[0054] (4) Divide the dataset into three subsets according to the training difficulty: easy, medium, and hard. The specific steps are as follows:
[0055] (4-1) Apply supervised fine-tuning on the base model and generate the chain of thought Csi for the entire training set to evaluate the inherent difficulty of each training sample.
[0056] (4-2) Divide the dataset D into three subsets: easy (Deasy), medium (Dmedium), and hard (Dhard) according to the BLEU score distribution between the generated chain of thought Csi and the candidate chain of thought Cji;
[0057] (5) Use the divide-and-conquer training strategy to train three sub-models on the three training subsets respectively, and adopt the method of combining DARE with task arithmetic to merge and fine-tune the three sub-models to obtain the final chain of thought generation model for code COTTON_lite, which specifically includes the following steps:
[0058] (5-1) For each difficulty level k ∈ {easy, medium, hard}, train a dedicated model Mk respectively;
[0059] (5-2) Calculate the difference between the dedicated model and the base model of Mk, and then perform weighted combination according to the importance weight α determined by the validation performance;
[0060] (5-3) Apply dropout in each layer through the rescaling mechanism of DARE to prevent interference between different tasks;
[0061] (5-4) Use the complete dataset to comprehensively fine-tune the merged model to obtain the final model COTTON_lite;
[0062] (6) Deploy COTTON_lite on a single GPU device to generate corresponding chain of thought prompts during the code generation stage to guide code generation;
[0063] (7) Evaluate the method of this embodiment and the existing chain of thought generation method on the same dataset, and use three performance metrics in the field of neural machine translation (i.e., BLEU, METEOR, and ROUGE-L) to automatically evaluate the quality of the generated chain of thought:
[0064] Table 2 Comparison table of the results of the method of this embodiment and other methods (%)
[0065]
[0066]
[0067] After experiments, Table 2 shows that the method for generating code thought chains in the optimized small model proposed in this embodiment can generate thought chains of higher quality compared to the baseline method. Specifically, in terms of automatic evaluation metrics, COTTON_lite continuously outperforms all baseline models and achieves the highest scores in the BLEU, METEOR, and ROUGE-L metrics. For example, the performance of COTTON_lite in the BLEU, METEOR, and ROUGE-L metrics has increased by at least 12.28%, 4.24%, and 4.88% respectively.
[0068] In addition, compared with models of similar size (such as OPT, Pythia, SmolLM2, PolyCoder), COTTON_lite shows significant improvement, with a relative increase in the BLEU score of 51.38% - 131.90%. Even compared with large models with more parameters (such as Qwen2.5-Coder, with 500 million parameters), COTTON_lite (with 380 million parameters) still achieves better performance while using fewer parameters, which demonstrates the efficiency of our method and the strong competitiveness of the method proposed in this embodiment.
[0069] Example 2
[0070] In Example 2, the specific improvement effect of the thought chains generated by the COTTON_lite method of the present invention on the performance of the code generation model is further verified. Specifically, it includes the following contents:
[0071] (1) Select three high-performance large language models as evaluation models, namely: DeepSeek-Coder, Qwen2.5-Coder, and Yi-Coder;
[0072] (2) Select three public data sets as validation sets, namely:
[0073] (2-1) HumanEval: Contains 164 Python programming problems, with an average of 7.8 test cases per problem;
[0074] (2-2) OpenEval: Contains 178 programming problems, with an average of 5 test cases per problem;
[0075] (2-3) CodeHarmony: Contains 153 programming problems, with an average of 3 test cases per problem;
[0076] (3) Evaluate the improvement effect of the method in this embodiment on the code generation model on the data set, and use the performance metric Pass@1 in the code generation field to automatically evaluate the quality of the generated thought chains:
[0077] Table 3 Comparison Table of the Improvement Effects of the Method in This Embodiment on Three Code Generation Models under Three Datasets (%)
[0078]
[0079] Through experiments, Table 3 shows the improvement effects of the method in this embodiment on three code generation models (DeepSeek-Coder, Qwen2.5-Coder, Yi-Coder) on three datasets (HumanEval, OpenEval, CodeHarmony). The results show that the method in this embodiment significantly improves the code generation performance of the models on different models and datasets.
[0080] On the HumanEval dataset, the zero-shot Pass@1 of DeepSeek-Coder increased from 32.93% to 53.66%, an increase of 20.73 percentage points; the zero-shot Pass@1 of Yi-Coder increased from 41.46% to 59.15%, an increase of 17.69 percentage points. This indicates that this method significantly enhances the zero-shot performance of the models on the HumanEval dataset, especially with the most prominent effect on DeepSeek-Coder.
[0081] On the OpenEval dataset, the zero-shot Pass@1 of DeepSeek-Coder increased from 26.40% to 41.01%, an increase of 14.61 percentage points; the zero-shot Pass@1 of Yi-Coder increased from 24.72% to 41.57%, an increase of 16.85 percentage points. This shows that this method improves the model performance more evenly on the OpenEval dataset, especially with better performance on Yi-Coder.
[0082] On the CodeHarmony dataset, the zero-shot Pass@1 of DeepSeek-Coder increased from 51.63% to 69.28%, an increase of 17.65 percentage points; the zero-shot Pass@1 of Yi-Coder increased from 48.37% to 66.67%, an increase of 18.30 percentage points. This indicates that this method has a significant improvement effect on the models on the CodeHarmony dataset, especially with a large increase on DeepSeek-Coder and Yi-Coder.
[0083] Generally speaking, the method in this embodiment significantly improves the performance of the three code generation models on different datasets, especially with more prominent effects on the HumanEval and CodeHarmony datasets. This indicates that this method can effectively enhance the zero-shot performance of the code generation models and has high practicality and universality.
[0084] Example 3
[0085] In Example 3, we further verified the applicability and superiority of the method of the present invention in different programming languages. The specific contents are as follows:
[0086] (1) Select a code generation benchmark dataset HumanEval-XL containing multiple programming languages, which covers 11 programming languages, including Java, C#, Kotlin, JavaScript, PHP, Python, Ruby, Perl, Go, Scala, and TypeScript;
[0087] (2) Select three high-performance large language models as evaluation models, namely: DeepSeek-Coder, Qwen2.5-Coder, and Yi-Coder;
[0088] (3) Evaluate the improvement effect of the method of this embodiment on different programming languages on the dataset HumanEval-XL, and use the performance metric Pass@1 in the field of code generation to automatically evaluate the quality of the generated chain of thought:
[0089] Table 4 Comparison table of the improvement effect of the method of this embodiment on different programming languages (%)
[0090]
[0091] Through experiments, Table 4 details the performance differences between the method of this embodiment and zero-shot learning on different programming languages. Through three different code generation models (DeepSeek-Coder, Qwen2.5-Coder, Yi-Coder), it can be seen that the method of this embodiment can significantly improve the performance of the model on multiple programming languages.
[0092] Taking the DeepSeek-Coder model as an example, the accuracy rate of the Java language has increased from 60.00% to 73.75%, showing a significant performance improvement. For the Qwen2.5-Coder model, the accuracy rate of the JavaScript language has increased from 80.00% to 83.75%. Although the improvement range is small, it still proves the effectiveness of this method. The Yi-Coder model performs particularly well in the C# language, with the accuracy rate increasing significantly from 46.25% to 65.00%, an increase of 18.75 percentage points.
[0093] In addition, for a language like Scala that performs poorly in zero-shot learning (only 16.25% under the Yi-Coder model), the method of this embodiment can also significantly improve its performance, with the accuracy increasing to 31.25%. This shows that the method of this embodiment is not only effective for mainstream programming languages, but also has an improvement effect on some niche languages.
[0094] In summary, the method of this embodiment shows good generalization ability and improvement effect on different programming languages and models, especially having significant advantages in improving the code generation performance of niche languages.
[0095] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for optimizing the generation of code thought chains in small models, characterized in that: The following steps are involved: S1, using three open source large language models as teacher models, based on the CodeHarmony dataset, and using prompt strategies to generate corresponding thought chains; S2. Ensure the semantic consistency and validity of the generated thought chain, obtain a quality assessment framework based on functional correctness, obtain the final data set after screening, and set the data set format to <code segment, thought chain, code description>; S3, select Qwen2.5-Coder-0.5Binstruct as the basic model, use vocabulary pruning as the parameter reduction strategy, and compress the model size to 0.38B; S4, divide the dataset into three subsets according to the training difficulty: easy, medium and hard; S5. Use the divide-and-conquer training strategy to train three sub-models on three training subsets respectively, and use DARE combined with task arithmetic to merge and fine-tune the three sub-models to obtain the final code thinking chain generation model COTTON_lite; S6. Deploy COTTON_lite on a single GPU device and generate corresponding thought chain prompts in the code generation phase to guide code generation.
2. The method for optimizing the generation of code thought chains in small models according to claim 1, characterized in that: The step S1 comprises the following steps: S11: Three high-performance large-scale language models are selected as teacher models, namely GPT-4, DeepSeek-V2.5 and QwenCoder32B-Instruct; S12: Based on the CodeHarmony dataset, which contains more than 15,000 code examples and their corresponding test cases; S13: For each input data, the three large language models use the prompt strategy to generate the corresponding thought chain Cji.
3. The method for optimizing the generation of code thought chains in small models according to claim 1, characterized in that: The step S2 comprises the following steps: S21: For each generated thought chain Cji, firstly, the evaluation model E uses the input data Xi and the thought chain Cji to synthesize the corresponding code Yi; S22: Use the test case Ti to evaluate the functional correctness of the code Yi and calculate the thinking chain Cji score score(Cji). Where compiler(Y i ,t) means compiling and verifying the code Yi with the test case t. If the code passes the test case, it returns 0, otherwise it returns a non-zero value. If the generated code Yi passes all test cases, score(Cji)=1, otherwise score(Cji)=0; S23: When score(Cji) is 1, that is, Yi passes all test cases, the corresponding thinking chain sample is selected as a candidate thinking chain; if score(Cji) is 0, the sample is excluded from the data set and the final data set D is constructed.
4. The method for optimizing the generation of code thought chains in small models according to claim 1, characterized in that: The S3 comprises the following steps: S31: Build the basic large language model Qwen2.5-Coder-0.5Binstruct; S32: According to the frequency of the token in the training corpus, a frequency-based pruning function P(Vorig) is defined, P(Vorig) = {v∈Vorig|f(v)>τ}, where Vorig is the original vocabulary, v is a token in the vocabulary, f(v) is the frequency of token v in the training corpus, and τ is an empirically determined frequency threshold; S33: retain the tokens in the original vocabulary Vorig whose frequency is higher than the threshold τ, and obtain the compressed vocabulary Vpruned; S34: Set τ to 1, the original vocabulary size N to 151,936, and the pruned vocabulary size M to 23,954, reducing the total parameters from 0.5B to 0.38B while retaining the basic capabilities of the model.
5. The method for optimizing the generation of code thought chains in small models according to claim 1, characterized in that: The S4 comprises the following steps: S41: Apply supervised fine-tuning on the base model and generate thought chains Csi for the entire training set to evaluate the inherent difficulty of each training sample; S42: According to the BLEU score distribution between the generated thinking chain Csi and the candidate thinking chain Cji, the dataset D is divided into three subsets: simple Deasy, medium Dmedium and difficult Dhard.
6. The method for optimizing the generation of code thought chains in small models according to claim 1, characterized in that: The step S5 comprises the following steps: S51. For each difficulty level k∈{easy, medium, hard}, train a dedicated model Mk respectively; S52, calculating the difference between the specialized model and the Mk base model, and then performing a weighted combination according to the importance weight α determined by the verification performance; S53, apply dropout at each layer through DARE’s rescaling mechanism to prevent interference between different tasks; S54. Use the complete dataset to fully fine-tune the merged model to obtain the final model COTTON_lite.
7. The method for optimizing the generation of code thought chains in small models according to claim 1, characterized in that: In the S6, COTTON_lite is deployed on a single GPU device, and corresponding thought chain prompts are generated in the code generation stage to guide code generation.
Citation Information
Patent Citations
Depth fine tuning method for generating Chinese text logical reasoning thinking chain
CN117669536A
Generation method and device of thinking chain data, electronic equipment and storage medium
CN117933391A
Flexible thinking chain learning method and device based on large language model and medium
CN118333153A
Event extraction method for training by generating thinking chain interpretation based on large language model
CN118467737A
Chinese and cross-language query expansion method based on retrieval enhancement and knowledge distillation
CN118673095A
Cited By
Enterprise inference model construction method, device and equipment based on large model distillation
CN120525027A