Automatic deep thinking model selection training method and system for large language model
By adding special tags and dialogue templates to large language models, combining reinforcement learning, and dynamically selecting mode switching, the problems of resource waste and accuracy imbalance in existing models are solved, and flexible response and efficient training are achieved.
Patent Information
- Application Number
- CN202510823870.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-26
AI Technical Summary
Existing large language models are unable to dynamically adjust reasoning strategies according to the complexity of the problem, resulting in waste of computing resources and imbalance in answer accuracy. They also lack training designs for different response modes, making it difficult to effectively switch between normal mode and deep thinking mode.
Add special tokens by extending the word breaker vocabulary
The model can flexibly switch response strategies according to the difficulty of the questions, optimize resource utilization, improve computing efficiency and answer accuracy, and enhance the model's adaptability in diverse scenarios.
Smart Images

Figure CN120706561A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to an automated deep thinking pattern selection training method and system for large language models. Background Art
[0002] Current large language models (LLMs) generally use a single response mode to generate answers and are unable to dynamically adjust reasoning strategies based on the complexity of the question. For simple questions, existing models may waste computing resources due to excessive reasoning; for complex questions, relying solely on fast response modes may reduce the accuracy of answers due to a lack of in-depth analysis. In addition, traditional training methods do not design discriminative semantic structures for different response modes, making it difficult for the model to effectively switch between normal mode (instant response) and deep thinking mode (multi-step reasoning). There is no comprehensive solution in the existing technology that combines multi-modal training, dynamic mode selection, and reinforcement learning reward mechanisms, and the following major defects exist:
[0003] Single mode: The model cannot flexibly switch response strategies according to demand, resulting in an imbalance between efficiency and accuracy; Waste of resources: Using simple modes for complex problems is prone to errors, while using complex modes for simple problems increases computing overhead; Training limitations: The lack of balanced data design for different modes causes the model to be biased towards a single capability; Insufficient decision optimization: The dynamic selection mechanism based on problem difficulty and the reinforcement learning guidance strategy are not introduced.
[0004] Therefore, we propose an automated deep thinking pattern selection training method and system for large language models to solve the above problems. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides an automated deep thinking pattern selection training method and system for a large language model to solve the problems raised in the above-mentioned background technology.
[0006] To achieve the above objectives, the present invention provides the following technical solution: an automated deep thinking pattern selection training method for a large language model, comprising the following methods:
[0007] Extend the word breaker vocabulary and add special tokens <think>and< / think> , used to identify the content structure of deep thinking patterns;
[0008] Design dialogue templates to distinguish the input and output formats of normal mode and deep thinking mode, where deep thinking mode includes <think>Marker-guided reasoning process;
[0009] Construct training datasets of the same size for both normal and deep thinking modes, and jointly train the model to enable it to respond to both modes simultaneously.
[0010] Dynamically select the mode based on the question difficulty function. When the accuracy of the normal mode answer exceeds the preset threshold, the normal mode will be called first, otherwise the deep thinking mode will be triggered;
[0011] Through the mode reward mechanism in reinforcement learning, the reward weight of the common mode is adjusted to guide the model to optimize the mode selection strategy.
[0012] In a preferred embodiment, the problem difficulty function is defined as:
[0013] Generate n candidate answers through the normal mode, count the number of correct answers c, when the accuracy When it is judged as a simple problem, normal mode is called first.
[0014] In a preferred embodiment, in the mode reward mechanism, the reward weight for the correct answer generated in the normal mode is:
[0015]
[0016] β is an adjustable hyperparameter, and the reward for the answer sampled by DeepThinking mode is not adjusted.
[0017] In a preferred embodiment, when constructing training data, samples of the normal mode (direct answer) and the deep thinking mode (step-by-step reasoning) are generated for the same question, and the data volume of the two modes is kept balanced.
[0018] In a preferred embodiment, the automatic mode selection process includes:
[0019] After entering the user question, call the normal mode to generate n candidate answers;
[0020] Count the number of correct answers c, if Use normal mode to output; otherwise trigger deep thinking mode to generate the answer.
[0021] In a preferred embodiment, in reinforcement learning optimization, the β value is dynamically adjusted according to the actual application scenario to balance the usage frequency of the normal mode and the deep thinking mode.
[0022] In a preferred embodiment, an automated deep thinking pattern selection training system for a large language model includes:
[0023] A deep thinking module enables the model to generate answers that include detailed thinking processes through step-by-step reasoning;
[0024] Normal modules enable the model to directly output concise and immediate answers;
[0025] The mode selection module automatically selects whether to respond in deep thinking mode or normal mode based on the difficulty of the input question.
[0026] In a preferred embodiment, the system implements mode differentiation through the following steps:
[0027] Add special tokens to the word breaker's vocabulary <think> and< / think> , used to identify semantic structures in deep thinking mode;
[0028] Design a dialogue template that includes normal mode and deep thinking mode, and trigger the corresponding mode through different starting markers.
[0029] In a preferred embodiment, training data of the same size is generated for the normal mode and the deep thinking mode respectively;
[0030] In the data of deep thinking mode, through <think> and< / think> Markers wrap the step-by-step reasoning and provide the final answer at the end.
[0031] In a preferred embodiment, the triggering mode of the mode selection module is:
[0032] Normal mode is triggered when the input dialog contains the <|im_start|>assistant tag;
[0033] When the input dialog contains the <|im_start|>think tag, deep think mode is triggered.
[0034] The difficulty assessment method is:
[0035] Defining the difficulty function Where n is the total number of sampled answers in the normal mode, and c is the number of correct answers;
[0036] Set the accuracy threshold α. When D < α, the deep thinking mode is automatically selected, otherwise the normal mode is selected.
[0037] Technical effects and advantages of the present invention:
[0038] Flexible mode switching:
[0039] By adding special markers ( <think> and< / think> ) and customized dialogue templates to achieve semantic distinction between normal mode and deep thinking mode, enabling the model to accurately switch response strategies according to needs.
[0040] Resource efficiency optimization:
[0041] Dynamically select the mode based on the question difficulty function (such as the normal mode answer accuracy threshold α): simple questions prioritize calling the normal mode for quick response, and complex questions trigger the deep thinking mode to ensure answer quality, effectively balancing computing resource consumption and result reliability.
[0042] Improved training balance:
[0043] Construct training data of the same scale for normal mode and deep thinking mode to prevent the model from being biased towards a single mode and enhance its adaptability in diverse scenarios.
[0044] Reinforcement learning guidance strategy:
[0045] A mode reward mechanism is introduced to prioritize the normal mode for handling simple problems by adjusting the reward weight of the normal mode, while retaining the basic rewards of the deep thinking mode to cope with complex needs.
[0046] Technical uniqueness:
[0047] By integrating semantic tag design, dynamic mode selection, and reinforcement learning optimization, a technical solution is formed that is different from the traditional single-mode model. It is suitable for scenarios such as intelligent customer service, educational assistance, and complex decision support, and is both efficient and accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 The model training steps in the present invention are shown in the flowchart;
[0049] Figure 2 This is the thinking logic diagram of the deep thinking mode in the present invention. DETAILED DESCRIPTION
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0051] Reference Figure 1 ,An automated deep thinking mode selection training method and system for large language models,Phase 1: Learning deep thinking mode and normal mode
[0052] Follow these steps:
[0053] 1. Expand the vocabulary of the tokenizer and add custom special tags ( <think> and< / think> ) to give the model the ability to process specific semantic structures.
[0054] 2. Design chat_template, based on the original (normal mode:
[0055] <|im_start|>assistant\n);
[0056] Added a model for deep thinking mode, use: <|im_start|>assistant\n <think>\n.
[0057] The Deep Thinking model follows this pattern when generating answers:
[0058] <|im_start|>assistant\n <think> \nDeep Thinking Content\n< / think> \nFinal answer.
[0059] 3. For the normal mode and deep thinking mode, construct training data of the same size and train the model so that the model has the ability to answer both the normal mode and the deep thinking mode.
[0060] Deep Thinking Mode: In Deep Thinking Mode, the model reasoned step by step, giving a final answer after careful consideration. This approach is well-suited for complex problems that require deep thought.
[0061] Normal mode: a non-deep-thinking mode in which the model provides quick and immediate responses. It is suitable for simple problems where speed is a high requirement.
[0062] Training data examples and format description
[0063] Take the question 12*12=? as an example:
[0064] Normal mode:
[0065] 12*12=144;
[0066] Corresponding training data:
[0067] "<|im_start|>system\nYouareahelpfulassistant.<|im_end|>\n<|im_start|>user\n12*12=?<|im_end|>\n<|im_start|>assistant\n12*12=144<|im_end|>"
[0068] The last part is the standard answer for model training.
[0069] Deep thinking mode
[0070] Refer to the instruction manual Figure 2 ;
[0071] Corresponding training data:
[0072] "<|im_start|>system\nYouareahelpfulassistant.<|im_end|>\n<|im_start|>user\n12*12=?<|im_end|>\n<|im_start|>think\n <think>\n;
[0073] Well, what I need to calculate now is 12 times 12. Let me think about how to solve this problem. First, I remember that the basic concept of multiplication is repeated addition, but since both numbers are two-digit numbers, it might be faster to just multiply them directly.
[0074] \n\nFirst, I could try to decompose the two numbers to see if there's an easier way. For example, split one of the numbers into parts that are easier to calculate.
[0075] For example, I know that multiplying 10 by any number is relatively simple, so perhaps I could split one of the 12s into 10 plus 2 and then use the distributive property. This would convert the original equation to (10 + 2) multiplied by another 12. Now, according to the distributive property, this should equal 10 times 12 plus 2 times 12, right? So, let's start with the first part: what is 10 times 12? This should be easy, because multiplying any number by 10 just adds a zero, so the result is 120. The second part is 2 times 12, which I remember clearly: 24. Then, adding these two results together—120 plus 24—should give 144. So, is the answer 144? However, I'm still a little unsure. Maybe I'm going to go wrong. Are there other ways to verify this? For example, I could try another method, like vertical multiplication. Let's try this: First, align the two numbers: 12 × 12. First, multiply the 2 in the ones place by 12. This gives us 24. Note the carry here: you should write the ones-digit result first, then move the carry to the tens place. However, a more accurate way to multiply each digit separately and then add them together is probably more accurate. Specifically, the second multiplier is 12, so we can break it down into two steps: first, multiply each digit of the first number by the 2 in the ones place, and second, multiply the first digit by the 1 in the tens place, shifting the result left one place, and then adding them together. So, step one: 2 times 12 equals 24, which is written down as 24. Step two: multiply the 1 in the tens place by 12. This should add a trailing zero to the result, making it 120. Then, add these two results: 24 plus 120, which is indeed 144. It seems that the two methods produce the same result, so they should be correct. \n\nIn addition, I can also recall my knowledge of square numbers. For example, I know some common square numbers, such as 10 squared is 100, 11 squared is 121, so is 12 squared a larger number than these?
[0076] For example, we can use the formula (a+b) 2 =a 2 +2ab+b 2 To calculate, where a=10, b=2, then:\n\n(10+2) 2 =10 2 +2×10×2+2 2 =100+40+4=144. This also gives the same result.
[0077] \n\nIt seems that no matter which method is used, the result is 144, so I think this answer is correct.
[0078] \n< / think> \n12 multiplied by 12 can be verified in several ways:\n\n1.
[0079] Distributive law decomposition method:
[0080] \nSplit one of the factors into \\(10+2\\),\n\\[\n(10+2)\times 12=10\times 12+2\times 12=120+24=144.\n\\]\n\n2.
[0081] Vertical multiplication verification:\n
[0082] Align the numbers and perform digit-by-digit calculations: \n\\[\n\begin{array}{r}\n\\quad 12\\\n\times\\quad 12\\\n\\hline\n\\quad 24\\,\\,\\,(\text{units digit}\times12)\\\n+120\\,\\,(\text{tens digit}\times 12, shift one bit left)\\\n\\hline\n\\quad 144\\\n\\end{array}\n\\]\n\n3.
[0083] Square formula expansion method:\n
[0084] Using the square formula \\((a+b)^2=a^2+2ab+b^2\\),\n\\[\n(10+2)^2=10^2+2\times10\times 2+2^2=100+40+4=144.
[0085] \n\\]\n\nIn summary, all three methods produce the same result, confirming that they are correct. \n\nAnswer: \boxed{144}<|im_end|>".
[0086] The above part is the standard answer for model training. <think> and< / think> The part in between is the content for deep thinking.
[0087] Instructions for triggering different modes after training
[0088] When the following information (assistant) is input to the model, the model will enter the normal mode to answer:
[0089] <|im_start|>system\nYouareahelpfulassistant.<|im_end|>\n<|im_start|>user\n12*12=? <|im_end|>\n<|im_start|>assistant
[0090] When the model is given the following information (think), it will enter deep thinking mode to answer:
[0091] <|im_start|>system\nYouareahelpfulassistant.<|im_end|>\n<|im_start|>user\n12*12=? <|im_end|>\n<|im_start|>think
[0092] In this way, during the second stage of intensive training, we can control the use of two different models to answer the same question and sample answers from ordinary and deep thinking.
[0093] Stage 2: Learning to automatically choose between Normal Mode and Deep Thinking Mode
[0094] Stage 1 has learned two modes of answering questions:
[0095] The second stage is to let the model learn to automatically choose whether to use deep thinking mode based on the difficulty of the user's questions:
[0096] (After the model is trained in stage 1, after the input `<|im_start|>system\nYou are a helpful assistant.<|im_end|>\n<|im_start|>user\n12*12=?<|im_end|>\n<|im_start|>`, it is uncertain whether the model will use the normal mode or the deep thinking mode for subsequent output. We expect the model to select the appropriate mode based on the difficulty of the input question. For simple questions, we hope to use the normal mode, as this can save a lot of time and computing resources, provided the answer is correct. For complex questions, the normal mode may not be able to answer correctly, so we hope the model will use the deep thinking mode.
[0097] Definition of Difficulty:
[0098]
[0099] Assumptions:
[0100] -D stands for difficulty function.
[0101] -n represents the number of answers sampled in normal mode.
[0102] -c represents the number of correct answers among these n answers.
[0103] -α represents the hyperparameter threshold of accuracy.
[0104] The mode reward part is calculated according to the difficulty definition.
[0105] Here are the reward answers for normal mode sampling:
[0106]
[0107] β is an adjustable hyperparameter. The mode reward is mainly used to allow the model to give more rewards for questions that can also be answered by the normal mode through reinforcement learning.
[0108] The rewards for answers sampled in Deep Thinking mode will not be adjusted.
[0109] An automated deep thinking pattern selection training system for large language models, comprising:
[0110] Deep thinking mode, which enables the model to generate answers that include detailed thinking processes through step-by-step reasoning;
[0111] Normal mode, which enables the model to directly output concise and immediate answers;
[0112] The mode selection module automatically selects whether to respond in deep thinking mode or normal mode based on the difficulty of the input question.
[0113] The system achieves mode differentiation through the following steps:
[0114] Add special tokens to the tokenizer's vocabulary <think> and< / think> , used to identify semantic structures in deep thinking mode;
[0115] Design a conversation template (chat_template) that includes normal mode and deep thinking mode, and trigger the corresponding mode through different start tags.
[0116] The methods for constructing training data include:
[0117] Generate training data of the same size for both normal mode and deep thinking mode;
[0118] In the data of deep thinking mode, through <think> and< / think> Markers wrap the step-by-step reasoning and provide the final answer at the end.
[0119] The triggering method of the mode selection module is:
[0120] Normal mode is triggered when the input dialog contains the <|im_start|>assistant tag;
[0121] When the input dialog contains the <|im_start|>think tag, deep think mode is triggered.
[0122] The difficulty assessment method is:
[0123] Defining the difficulty function Where n is the total number of sampled answers in normal mode, and cc is the number of correct answers;
[0124] Set the accuracy threshold α. When D < α, the deep thinking mode is automatically selected, otherwise the normal mode is selected.
[0125] The system optimizes mode selection through reinforcement learning, including:
[0126] In normal mode, set the dynamic reward weight β and adjust the reward value according to the accuracy of the answer;
[0127] Deep Thinking Mode maintains the base rewards without additional weight adjustments.
[0128] The training process includes the following stages:
[0129] Phase 1: Through dual-mode training data, the model is able to master the response capabilities of both normal mode and deep thinking mode;
[0130] Phase 2: Introduce the mode selection module and optimize the mode switching strategy through difficulty assessment and reinforcement learning mechanism.
[0131] The data format of deep thinking mode is:
[0132] Starting from the question input, it includes the reasoning process, verification method and final answer in sequence;
[0133] The reasoning process is achieved through mathematical formula decomposition, logical chain deduction or multi-method cross-validation.
[0134] The system is suitable for the following scenarios:
[0135] Multi-step reasoning and verification of complex problems;
[0136] Dynamically balance response speed and answer depth based on user needs in real-time interactions.
[0137] An automated deep thinking pattern selection training method for a large language model, comprising the following steps:
[0138] Extend the word breaker vocabulary and add special tokens <think> and< / think> , used to identify the reasoning content of deep thinking patterns;
[0139] Design a conversation template (chat_template) to distinguish the input and output formats of the normal mode and the deep thinking mode. The deep thinking mode is <think>Marker-guided multi-step reasoning;
[0140] Construct training datasets of the same size for both the normal mode (direct answers) and the deep thinking mode (step-by-step reasoning), and jointly train the model to enable it to respond to both modes simultaneously.
[0141] Dynamically select the mode based on the question difficulty function. When the accuracy of the answer generated by the normal mode exceeds the preset threshold α, the normal mode will be called first, otherwise the deep thinking mode will be triggered;
[0142] Through the mode reward mechanism in reinforcement learning, the reward weight of the normal mode is adjusted to encourage the model to give priority to the normal mode to handle simple problems.
[0143] The difficulty function of the question is defined as: generate n candidate answers through the normal mode, count the number of correct answers c, and when the accuracy When it is judged as a simple problem, normal mode is called first.
[0144] In the mode reward mechanism, the reward weight for the correct answer generated by the normal mode is R 普通 =R 基础 +β, where β is an adjustable hyperparameter, and the deep thinking mode maintains the basic reward weight R 基础 .
[0145] The normal mode format of the conversation template is: <|im_start|>assistant\n<answer>
[0146] The format of deep thinking mode is: <|im_start|>assistant\n <think> \n<Reasoning Process>\n< / think> \n<Answer>.
[0147] When constructing the training dataset, samples of the normal mode (direct answer) and the deep thinking mode (step-by-step reasoning) are generated for the same question, and the data volume of the two modes is kept in a 1:1 ratio;
[0148] The automatic mode selection process includes:
[0149] After entering the user question, call the normal mode to generate n candidate answers;
[0150] Count the number of correct answers c, if Directly output the normal mode answer; otherwise, trigger the deep thinking mode to generate the answer.
[0151] Special Marking <think> and< / think> It is defined as an independent semantic unit in the word segmenter and is used by the model to distinguish the reasoning process from the generation logic of the final answer.
[0152] In reinforcement learning optimization, the hyperparameter β is dynamically adjusted according to the actual application scenario to balance the usage frequency of normal mode and deep thinking mode.
[0153] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.< / think> < / think> < / think>
Claims
1. An automated deep thinking pattern selection training method for large language models, characterized by; Includes the following methods: Extend the word breaker vocabulary and add special tokens <think> and< / think> , used to identify the content structure of deep thinking patterns; Design dialogue templates to distinguish the input and output formats of normal mode and deep thinking mode, where deep thinking mode includes <think> Marker-guided reasoning process;< / think> Construct training datasets of the same size for both normal and deep thinking modes, and jointly train the model to enable it to respond to both modes simultaneously. Dynamically select the mode based on the question difficulty function. When the accuracy of the normal mode answer exceeds the preset threshold, the normal mode will be called first, otherwise the deep thinking mode will be triggered; Through the mode reward mechanism in reinforcement learning, the reward weight of the common mode is adjusted to guide the model to optimize the mode selection strategy.
2. The method for automatic deep thinking mode selection and training for a large language model according to claim 1, characterized in that: The problem difficulty function is defined as: Generate n candidate answers through the normal mode, count the number of correct answers c, when the accuracy When it is judged as a simple problem, normal mode is called first.
3. The method and system for automatic deep thinking mode selection training for a large language model according to claim 1, characterized in that: In the mode reward mechanism, the reward weight for the correct answer generated by the normal mode is: β is an adjustable hyperparameter, and the reward for the answer sampled by DeepThinking mode is not adjusted.
4. The method and system for automatic deep thinking mode selection training for a large language model according to claim 1, characterized in that: When constructing training data, samples of normal mode (direct answers) and deep thinking mode (step-by-step reasoning) are generated for the same question, and the amount of data in the two modes is kept balanced.
5. The method and system for automatic deep thinking mode selection training for a large language model according to claim 1, characterized in that: The automatic mode selection process includes: After entering the user question, call the normal mode to generate n candidate answers; Count the number of correct answers c, if Use normal mode to output; otherwise trigger deep thinking mode to generate the answer.
6. The method and system for automatic deep thinking mode selection training for a large language model according to claim 1, characterized in that: In reinforcement learning optimization, the β value is dynamically adjusted according to the actual application scenario to balance the usage frequency of normal mode and deep thinking mode.
7. An automated deep thinking pattern selection and training system for a large language model according to any one of claims 1 to 6, characterized in that: include: A deep thinking module enables the model to generate answers that include detailed thinking processes through step-by-step reasoning; Normal modules enable the model to directly output concise and immediate answers; The mode selection module automatically selects whether to respond in deep thinking mode or normal mode based on the difficulty of the input question.
8. The automated deep thinking pattern selection training system for a large language model according to claim 7, characterized in that: The system achieves mode differentiation through the following steps: Add special tokens to the word breaker's vocabulary <think> and< / think> , used to identify semantic structures in deep thinking mode; Design a dialogue template that includes normal mode and deep thinking mode, and trigger the corresponding mode through different starting markers.
9. The automated deep thinking pattern selection training system for a large language model according to claim 7, characterized in that: Generate training data of the same size for both normal mode and deep thinking mode; In the data of deep thinking mode, through <think> and< / think> Markers wrap the step-by-step reasoning and provide the final answer at the end.
10. The automated deep thinking pattern selection training system for a large language model according to claim 7, characterized in that: The triggering method of the mode selection module is: Normal mode is triggered when the input dialog contains the <|im_start|>assistant tag; When the input dialog contains the <|im_start|>think tag, deep think mode is triggered. The difficulty assessment method is: Defining the difficulty function Where n is the total number of sampled answers in the normal mode, and c is the number of correct answers; Set the accuracy threshold α. When D < α, the deep thinking mode is automatically selected, otherwise the normal mode is selected.
Citation Information
Cited By
Model deep thinking control method and device
CN121210519A
Media data processing method and device, equipment, storage medium and program product
CN121562789A
Session processing method and electronic equipment
CN121597798A