Large model reply method and device, electronic equipment and medium
By combining intuitive responses with deep thinking methods in the target large model, the output of the general large model is optimized, solving the problem of low response accuracy in the general large model, achieving more efficient and accurate responses, and improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-04-07
AI Technical Summary
General-purpose large models have low accuracy in answering mathematical and logical reasoning problems and lack logical reasoning ability, which limits their application scenarios.
By receiving questions to be answered and pre-saved prompts, the target large model intuitively responds and retrieves relevant information from the database for in-depth thinking. Combined with supervised fine-tuning and reinforcement learning training, the intuitive output is optimized to improve accuracy and efficiency.
It achieves improved output efficiency while increasing response accuracy, reducing user waiting time, enhancing interaction response speed and transparency, adapting to more complex tasks, and improving user experience.
Smart Images

Figure CN121809667A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and natural language processing, and in particular to a large model response method, device, electronic device, and medium. Background Technology
[0002] In the field of artificial intelligence, especially in Natural Language Processing (NLP), General Large Language Models (GMLs) have seen rapid development in recent years and have become a hot topic in research and application. GMLs typically refer to large neural network models that have been trained on large-scale data and are capable of handling multiple tasks. Multimodal GMLs, such as the ChatGPT series and the Qwen series, have also emerged, capable of understanding data from multiple modalities.
[0003] However, while the general big model can quickly output responses that are perceived as good by users, it lacks training in logical reasoning ability, which makes it unable to respond well to mathematical and logical reasoning problems, thus limiting its application. In addition, the general big model responds only based on intuition, resulting in a low accuracy rate for its responses. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, and medium for large model recovery, which addresses the problem of low accuracy in general large model recovery in related technologies.
[0005] In a first aspect, embodiments of this application provide a large model response method, the method comprising: Receive questions to be answered, and input the questions to be answered and pre-saved prompts into the target large model; wherein, the prompts are used to prompt the target large model to output intuitive and deep thinking responses; Receive the target intuitive response output by the target large model based on intuition when processing the question to be answered; It also receives the target deep thinking response output by the target big model; wherein, the target big model calls the database based on the target intuitive response and the question to be answered, obtains target information related to the keywords in the target intuitive response and the keywords in the question to be answered, and the target big model supplements and adjusts the target intuitive response by slow thinking according to the target information to obtain the target deep thinking response.
[0006] Secondly, embodiments of this application also provide a large model recovery device, the device comprising: The input receiving module is used to receive questions to be answered and input the questions to be answered and pre-saved prompts into the target large model; wherein, the prompts are used to prompt the target large model to output intuitive and deep thinking responses; The processing module is used to receive the target intuitive response output by the target big model based on intuition when processing the question to be answered; and to receive the target deep thinking response output by the target big model; wherein, the target big model calls the database based on the target intuitive response and the question to be answered to obtain target information related to the keywords in the target intuitive response and the keywords in the question to be answered, and the target big model supplements and adjusts the target intuitive response by slow thinking based on the target information to obtain the target deep thinking response.
[0007] Thirdly, embodiments of this application also provide an electronic device, including: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the steps included in any of the above-described methods for large model recovery according to the obtained program instructions.
[0008] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the large model recovery method as described in any of the preceding claims.
[0009] In this application embodiment, a novel target big model is proposed. The electronic device inputs the question to be answered and pre-saved prompts into the target big model. The target big model first provides a target intuitive response based on intuition, and simultaneously performs deep thinking. During the output of the target intuitive response, the target big model calls a database based on the question to be answered and the intuitive response to obtain relevant information from the database. Based on the obtained relevant information and the intuitive response, it performs deep thinking to obtain a target deep thinking response. This target big model can both provide intuitive output and perform deep thinking, resulting in a more accurate response. It improves both output efficiency and output quality. Compared to general large models, this application optimizes intuitive output through inference training, enabling initial responses to possess good intuition. Compared to inference large models, this application outputs intuitive responses first, reducing user waiting time. This application uses a training method combining supervised fine-tuning and reinforcement learning to specifically improve the training speed and accuracy of intuitive output of the target large model. Compared to traditional large model training methods, it can more efficiently improve model performance and adapt to more complex tasks. Furthermore, this application uses special labeling guidance to clearly show users the process and results of fast and slow thinking, enhancing the responsiveness, transparency, and interpretability of the interaction, thus improving the user experience. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram illustrating a large model response method provided in an embodiment of this application. Figure 2 This is a schematic diagram illustrating a cold start dataset construction process provided in an embodiment of this application. Figure 3 A structural diagram of a large model recovery device provided in an embodiment of this application; Figure 4 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art are within the scope of protection of this application.
[0013] Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without any inventive effort. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, any changes to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.
[0014] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0015] The terms "connection," "linked," and "coupled" used in this application are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Multiple" in this application refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. The terms "first," "second," and "third" used in this application are merely to distinguish similar objects and do not represent a specific ordering of objects.
[0016] Furthermore, the technical solutions of the various embodiments can be combined with each other, but only if they are based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0017] The terms mentioned in the embodiments of this application are explained as follows: General-purpose large models: These are large-scale artificial intelligence models designed to handle multiple tasks and data types. They are typically trained on large amounts of multimodal data, such as text, images, and audio, to be able to understand and generate content in various formats.
[0018] Large-scale reasoning models: Models trained on large-scale pre-trained models for task reasoning (such as DeepSeek-R1), which are good at complex tasks such as logical reasoning and mathematical proof; General-purpose large-scale models: Models trained on large-scale data that can handle multiple tasks (such as GPT-4 and Qwen-2.5), which are good at quickly generating intuitive responses.
[0019] Quick Think: The model generates answers quickly based on user instructions, without the need for complex reasoning.
[0020] Slow thinking: The process by which the model deeply analyzes a problem through multi-step reasoning and chain-of-thought.
[0021] Cold start: In the early stages of supervised fine-tuning of a large model, training is conducted to enable the model to quickly master a fixed output paradigm (such as a special labeling separation format).
[0022] To address the issue of low accuracy in large model recovery in related technologies, embodiments of this application provide a large model recovery method, apparatus, electronic device, and medium.
[0023] The large-scale model response method includes: receiving a question to be answered; inputting the question and pre-saved prompts into the target large-scale model; wherein, the prompts are used to prompt the target large-scale model to output an intuitive response and a deep thinking response; receiving the target intuitive response output by the target large-scale model based on intuition; and receiving the target deep thinking response output by the target large-scale model; wherein, the target large-scale model calls the database based on the target intuitive response and the question to be answered, obtains target information related to the keywords in the target intuitive response and the keywords in the question to be answered, and the target large-scale model supplements and adjusts the target intuitive response based on the target information to obtain the target deep thinking response.
[0024] Example 1: Figure 1 This application provides a schematic diagram of a large model response method, which includes the following steps: S101: Receive the question to be answered, and input the question to be answered and the pre-saved prompt words into the target large model; wherein, the prompt words are used to prompt the target large model to output intuitive and deep thinking responses.
[0025] The large model response method provided in this application is applied to an electronic device, which can be a smart device such as a PC or a server.
[0026] To improve the accuracy of responses from large-scale models, this application discloses a novel target large-scale model for responding to questions. Specifically, the electronic device can receive the question to be answered. In one example, a user can enter the question on a preset webpage on their device and click a preset button, at which point the electronic device can receive the question. Furthermore, if the user requires improved accuracy, they can click a button on the preset page that indicates improved accuracy, such as the "Accurate Thinking" button.
[0027] After receiving a question to be answered, the electronic device can input the question and a pre-saved prompt into the target large model. The prompt is used to prompt the target large model to output an intuitive response and a deep thinking response.
[0028] S102: Receive the target intuitive response output by the target large model based on intuition to process the question to be answered.
[0029] To improve response efficiency, after receiving a question to be answered, the target large model can process the question based on intuition, obtain a target intuitive response, and output the target intuitive response, which can be received by electronic devices.
[0030] For example, Table 1 illustrates a question to be answered and a target intuitive response provided in an embodiment of this application:
[0031] Table 1
[0032] Table 1 shows that when the input question is "A cuboid and a cube have the same perimeter and height on their bases, then their volumes are also equal. ( )", the target model outputs the target intuition response as "×\n\n1. Let the length of the rectangular base of the cuboid be \(a\) and the width be \(b\), and the side length of the base of the cube be \(c\).\n - Given that the perimeters of the bases of the cuboid and the cube are equal, then \(2(a + b) = 4c\), which simplifies to \(a + b = 2c\).\n2. Calculate the base areas of the cuboid and the cube respectively:\n - Base area of the cuboid \(S_{length} = ab\).\n - Base area of the cube \(S_{cube} = c^{2}\).\n - According to the arithmetic mean inequality \(a + b\geqslant 2\\sqrt{ab}\\) (equality holds if and only if\\(a =b\\)), from\\(a + b = 2c\\) we get\\(2c\\geqslant 2\\sqrt{ab}\\), that is\\(c\\geqslant\\sqrt{ab}\\), squaring both sides gives\\(c^{2}\\geqslant ab\\), which is the base area of the cube\\(S_{rectangular}\\geqslant S_{length}\\) (equality holds if and only if\\(a = b\\)). \n3. Calculate the volume of the cuboid and cube: \n - Volume of the cuboid\\(V_{length}=S_{length}h = abh\\) (\\(h\\) is the height). \n - Volume of the cube\\(V_{cubic}=S_{cubic}h = c^{2}h\\). \n - Because S_{rectangular} geqslant S_{length} and their heights (h) are equal, V_{rectangular} geqslant V_{length} , their volumes are only equal when a = b. Generally, the volumes of a rectangular prism and a cube are not equal. Therefore, the statement "If a rectangular prism and a cube have the same base perimeter and height, then their volumes are also equal" is incorrect.
[0033] S103: And receive the target deep thinking response output by the target big model; wherein, the target big model calls the database based on the target intuitive response and the question to be answered, obtains target information related to the keywords in the target intuitive response and the keywords in the question to be answered in the database, and the target big model supplements and adjusts the target intuitive response by slow thinking according to the target information to obtain the target deep thinking response.
[0034] To improve the accuracy of responses, the target-oriented big data model can access a database based on the target intuitive response and the question to be answered. In one example, it can retrieve keywords from the target intuitive response and keywords from the question to be answered, obtain target information related to the keywords from the database, perform slow thinking based on the obtained target information, supplement and adjust the target intuitive response, and obtain a target-oriented deep thinking response, which is then output. Electronic devices can then receive the target-oriented deep thinking response output by the target-oriented big data model.
[0035] Among them, the target big model can respond immediately after receiving a question to be answered, and output the target intuitive response. In the process of outputting the target intuitive response, the target big model conducts in-depth thinking and outputs the target deep thinking response. The target deep thinking response is more accurate and can improve the accuracy of the response.
[0036] For example, Table 2 illustrates a response to a question to be answered and a target in-depth thinking response provided in an embodiment of this application:
[0037] Table 2
[0038] Table 2 shows that when the input question is "A cuboid and a cube have the same base perimeter and height, then their volumes are also equal. ( )", the target depth thinking response output by the target large model is " <think>First, assume the length and width of the base of the cuboid are *a* and *b*, respectively, and the side length of the base of the cube is *c*. The perimeter of the base of the cuboid is 2(a + b), and the perimeter of the base of the cube is 4c. Since the perimeters are equal, we have 2(a + b) = 4c, which simplifies to a + b = 2c. Assume the height of both the cuboid and the cube is *h*. The volume of the cuboid is... = a b h. The volume of the cube is... = c³ h.\n\nTo make = It needs to satisfy a b = c³. However, according to a + b = 2c, it cannot be guaranteed that a... b must equal c³, unless a = b = c, which is equivalent to the volume of a cube itself. Therefore, the volume of a cuboid is usually smaller than the volume of a cube. In conclusion, the volumes of a cuboid and a cube are not necessarily equal.< / think> Question: If a cuboid and a cube have the same base perimeter and height, then their volumes are also the same. ( ) answer: Let the base of the cuboid have length and width of (a) and (b) respectively, and the base of the cube have side length of (c). Both have height of (h). According to the problem: - The perimeter of the cuboid's base is (2(a + b)). - The perimeter of the cube's base is (4c). - Since the perimeters are the same, we have: \\[\n 2(a + b) = 4c \\quad \\Rightarrow \\quad a + b = 2c\n \\]\n We need to note that although the perimeters are equal, the base areas of the cuboid and the cube are not necessarily equal. The volume expression is as follows:\n- The volume of the cuboid\\( = a \\times b \\times h\\)\n- Volume of the cube\\( = c^3 \\times h\\)\n\nIf you want\\( = If \\), then it needs to satisfy: \n\\[\na \\times b = c^3\n\\]\nHowever, since \\(a + b = 2c\\), this does not guarantee \\(a \\times b = c^3\\). In fact, only when \\(a =b = c\\), that is, when the cuboid becomes a cube, does \\( = In general, the base area of a cuboid (a × b) is usually smaller than the base area of a cube (c³). For example: Let the base length of a cuboid be 3 and its width be 1. Then its perimeter is 2(3 + 1) = 8, and the corresponding base side length of the cube is 2. The volume of the cuboid is 3 × 1 × h = 3h. The volume of the cube is 2³ × h = 8h. Clearly, 3h ≠ 8h, meaning their volumes are not equal. Therefore, The volumes are not necessarily equal Answer: \boxed{\text{Error}}”.
[0039] In one possible implementation, the special token (special_token) serves to distinguish between the target's intuitive response and the target's deep thinking response during the reasoning phase. <intuition>The content within `<head>` can be directly displayed to the user. This content contains detailed reasoning processes and background information, providing users with a deeper understanding. This content can also be selectively displayed, depending on the user's needs or system settings. For example, this content can be shown when the user needs more details, and hidden when only brief information is required. Selective display is possible. By using these special tags, different types of responses can be flexibly managed and displayed, satisfying both the user's need for quick information and providing in-depth explanations and analysis.
[0040] This application provides a method for optimizing the intuitive output of a large model based on inference training.
[0041] Among them, large-scale reasoning models such as DeepSeek-R1 are trained on the basis of pre-trained large-scale models to improve the model's thought process and logical reasoning ability in order to cope with more complex task requirements.
[0042] General-purpose large-scale models, represented by GPT-4, PaLM-2, Claude, Qwen-2.5, and DeepSeek-V3, are based on the Transformer architecture. They capture semantic and grammatical patterns from text data through self-supervised learning, excelling in tasks such as text generation, question answering, and summarization. Based on pre-trained large-scale models, they undergo fine-tuning training using human preference feedback, resulting in a human-like intelligent assistant dialogue experience. In practical applications, these general-purpose large-scale models receive user input, utilize learned language patterns and knowledge reserves, and autoregressively generate corresponding answers through attention mechanisms. For example, in a knowledge-based question-answering scenario, if a user asks about the time of a historical event, the general-purpose large-scale model can quickly extract relevant information from its knowledge system and provide an answer. In text generation tasks, such as writing news reports, it can quickly organize language and generate text that meets the requirements. Furthermore, building upon large language models, by fusing multiple modalities such as text, images, and audio for joint training, cross-modal semantic alignment is achieved, enabling the large-scale models to understand multi-modal data and support visual understanding and image-text dialogue capabilities, such as GPT-4V.
[0043] Large-scale reasoning models: DeepSeek-R1 is a leader in this category. Building upon pre-trained large-scale models, large-scale reasoning models are trained on thought processes and logical reasoning abilities. For example, they undergo supervised fine-tuning using carefully prepared human-prepared reasoning data, and reinforcement learning training is used with output format requirements and the correctness of the final answer as training objectives. When receiving a user question, the model first enters a thinking mode, extensively considering the question, such as breaking it down, guessing the user's intent, etc., then reasoning step by step, and finally providing a final answer based on the thought process. For example, when solving mathematical proofs, it can deduce step by step, demonstrating a detailed proof process; when dealing with logical reasoning puzzles, it can also systematically analyze various conditions and arrive at the correct conclusion. Large-scale reasoning models can handle mathematical, programming, and logical reasoning scenarios well. However, because it needs to start thinking from the user's original input each time, it lacks an intuitive initial answer. This often means that large-scale reasoning models need to reason from scratch and spend a lot of time thinking before responding to the user with a final answer. The response time is slow, and users have to wait a long time for an answer. Furthermore, large-scale reasoning models are also adept at solving scenarios requiring spatial reasoning, such as embodied intelligence in robots and autonomous driving. However, because robots or driverless cars often need to make immediate decisions about their next move, the external spatial environment can change significantly while waiting for a considerable amount of time to process the information.
[0044] However, the output of general-purpose large models lacks logical thinking and can be considered intuitive. While they are fast and can handle common tasks, they often produce inaccurate and imprecise answers when faced with complex problems requiring deep logical reasoning. This is because their training mechanisms are not specifically designed for such problems, and they lack a deep understanding and analytical ability of complex logical relationships. For example, they are prone to giving incorrect or one-sided answers when solving complex mathematical word problems or logically interpreting legal provisions. Similarly, they cannot provide accurate answers in multimodal scenarios requiring spatial reasoning, such as embodied intelligence in robots and autonomous driving. Although large reasoning models perform well in logical reasoning tasks, the lengthy thinking process, starting from scratch each time rather than building on a good intuitive initial answer, leads to excessively long processing times. This forces users to wait a long time for the final result when interacting with the model, severely impacting the user experience, especially in scenarios with high real-time requirements, such as rapid online customer service Q&A, instant decision assistance, and robot spatial reasoning, where they struggle to meet real-time demands.
[0045] This application provides a model that possesses both the logical capabilities of a large-scale inference model and the rapid response capabilities of a general-purpose large-scale model, enabling wider application in various scenarios and providing a more user-friendly experience in terms of waiting time.
[0046] To address the problems inherent in general-purpose target models and inference-based target models, this invention proposes a method and system for optimizing the intuitive output of a target model based on inference training, aiming to integrate the advantages of both. The rapid response process of a general-purpose target model can be likened to fast thinking, which can handle common problems quickly and provide preliminary answers; the rigorous reasoning process of an inference-based target model can be likened to slow thinking, which can deeply analyze complex problems. This application proposes using inference training to optimize intuitive output, enabling the trained target model to possess both the rapid response of the general-purpose target model's fast thinking and the logical reasoning ability of the inference-based target model's slow thinking. After training according to this application, when responding to user input, the target model first enters fast thinking mode, quickly providing a basic preliminary answer based on intuition, meeting the user's need for timeliness. After fast thinking ends, it automatically enters slow thinking mode, re-evaluating and deeply considering the user's question and the preliminary answer provided by fast thinking. Utilizing the logical reasoning ability of slow thinking, it verifies, corrects, and refines the preliminary answer output by fast thinking, ensuring the accuracy and rigor of the response. This autoregressive generation method, which prioritizes fast thinking followed by slow thinking, ensures rapid interaction in general scenarios while providing high-quality solutions for complex problems, thus comprehensively improving user experience and model usability.
[0047] In practical applications, users can choose to output either an intuitive output or a deep-thinking output from the target model, as needed. If the application scenario has high real-time requirements and does not require extensive deep thinking, the target model can end its response after providing a preliminary answer based on the optimized intuitive output. If the application scenario requires extensive deep thinking and logical reasoning, the target model will first provide a preliminary answer, then conduct deep thinking based on the preliminary answer, and finally output the final answer.
[0048] In this application embodiment, a novel target big model is proposed. The electronic device inputs the question to be answered and pre-saved prompts into the target big model. The target big model first provides a target intuitive response based on intuition, and simultaneously performs deep thinking. During the output of the target intuitive response, the target big model calls a database based on the question to be answered and the intuitive response to obtain relevant information from the database. Based on the obtained relevant information and the intuitive response, it performs deep thinking to obtain a target deep thinking response. This target big model can both provide intuitive output and perform deep thinking, resulting in a more accurate response. It improves both output efficiency and output quality. Compared to general large models, this application optimizes intuitive output through inference training, enabling initial responses to possess good intuition. Compared to inference large models, this application outputs intuitive responses first, reducing user waiting time. This application uses a training method combining supervised fine-tuning and reinforcement learning to specifically improve the training speed and accuracy of intuitive output of the target large model. Compared to traditional large model training methods, it can more efficiently improve model performance and adapt to more complex tasks. Furthermore, this application uses special labeling guidance to clearly show users the process and results of fast and slow thinking, enhancing the responsiveness, transparency, and interpretability of the interaction, thus improving the user experience.
[0049] Example 2: To improve the accuracy of the response, based on the above embodiments, in this embodiment of the application, the target large model is trained in the following way: Obtain any sample question from a pre-saved cold start dataset, and the labeled responses saved for that sample question, wherein the labeled responses include labeled intuitive responses and labeled deep thinking responses; wherein the cold start dataset is used for supervised fine-tuning training of the target large model; The sample question and the labeled response are input into the original target large model to obtain the output response output by the original target large model; wherein, the output response includes an intuitive response and a deep thinking response; The original target large model is trained based on the deviation between the labeled response and the output response.
[0050] To improve the accuracy of responses, electronic devices can train large target models.
[0051] Specifically, the electronic device locally stores a cold-start dataset. This dataset is used to perform supervised fine-tuning training on the target large-scale model, ensuring that the output paradigm of the target large-scale model conforms to a specified paradigm. The electronic device acquires any sample question from the pre-saved cold-start dataset, along with a saved labeled response to that sample question. This labeled response includes labeled intuitive responses and labeled deep-thinking responses. After acquiring the sample question and the saved labeled response, the electronic device can input the sample question and labeled response into the original target large-scale model to obtain the output response from the original target large-scale model. This output response includes output intuitive responses and output deep-thinking responses.
[0052] Electronic devices can determine the target deviation between labeled responses and output responses. For example, they can determine a first deviation between labeled intuitive responses and output intuitive responses, and a second deviation between labeled deep thinking responses and output deep thinking responses. Based on the first deviation between labeled intuitive responses and output intuitive responses, and the second deviation between labeled deep thinking responses and output deep thinking responses, the target deviation between labeled responses and output responses is determined. Based on this target deviation, the original target model is trained so that the trained target model can output both intuitive responses and deep thinking responses, and the output conforms to the standard paradigm.
[0053] The method provided in this application embodiment can be referred to as cold start training. Using the method provided in this application embodiment, supervised fine-tuning training of all parameters of the model is performed using a cold start dataset on the basis of a pre-trained large model. After cold start training, the output of the target large model will follow: |<special_token1> |<intuition_answer> |<special_token1> |<special_token2> |<reasoning_process> |<special_token2> | <summary>The `intuition_answer` is the intuitive output, `reasoning_process` is the output based on the user's thought process, and `summary` summarizes all the preceding content and finally answers the user's question. `special_token1` and `special_token2` separate the different parts of the output. Cold start training helps the model master the output format and accelerates the convergence speed of subsequent reinforcement learning training.
[0054] Example 3: To improve the accuracy of the responses, based on the above embodiments, in this embodiment, the cold start dataset is generated in the following way: Obtain any pre-saved input question, input the input question into a general model, and obtain the first response output by the general model, wherein the first response is the response output by the general model based on intuition; then input the input question into a reasoning model, and obtain the second response output by the reasoning model, wherein the second response is the response output by the reasoning model based on deep thinking; concatenate the first response and the second response, and add a preset mark between the first response and the second response to obtain the labeled response; The labeled response is added to the cold start dataset along with the input question.
[0055] To improve the accuracy of the response, the electronic device can first generate a cold start dataset. Specifically, the electronic device can first obtain any pre-saved input question, input the input question into the general large model, and the general large model performs quick thinking based on intuition to obtain a response. For easy distinction, this response can be called the first response, and the general large model outputs the first response.
[0056] Table 3 illustrates an input question and a first response provided in an embodiment of this application:
[0057] Table 3
[0058] Table 3 shows that when the input question is "A cuboid and a cube have the same perimeter and height on their bases, then their volumes are also equal. ( )", the first response output by the general large model is "×\n\n1. Let the length of the rectangular base of the cuboid be \(a\) and the width be \(b\), and the side length of the base of the cube be \(c\).\n - Given that the perimeters of the bases of the cuboid and the cube are equal, then \(2(a + b) = 4c\), which simplifies to \(a + b = 2c\).\n2. Calculate the base areas of the cuboid and the cube respectively:\n - Base area of the cuboid \(S_{length} = ab\).\n - Base area of the cube \(S_{cube} = c^{2}\).\n - According to the AM-GM inequality \(a + b \geqslant 2 \sqrt{ab}\) (if and only if \(a When =b\\), equality holds. From \\(a + b = 2c\\), we get \\(2c\\geqslant 2\\sqrt{ab}\\), which is \\(c\\geqslant\\sqrt{ab}\\). Squaring both sides, we get \\(c^{2}\\geqslant ab\\), which is the base area of the cube \\(S_{rectangular}\\geqslant S_{length}\\) (equality holds if and only if \\(a = b\\). \n3. Calculate the volume of the cuboid and cube: \n - Volume of the cuboid \\(V_{length}=S_{length}h = abh\\) (\\(h\\) is the height). \n - Volume of the cube \\(V_{cubic}=S_{cubic}h = c^{2}h\\). \n - Because \\(S_{cubic}\\geqslant Since S_{length}\\) and height\\(h\\) are equal, V_{cubic}\\geqslant V_{length}\\). The volumes of a cuboid and a cube are equal only when \\(a = b\\). Generally, the volumes of a cuboid and a cube are not equal. Therefore, the statement "If the perimeter of the base and the height of a cuboid and a cube are equal, then their volumes are also equal" is incorrect.
[0059] The electronic device also inputs the input question into the reasoning model, which responds based on deep thinking. For easy distinction, this response can be called the second response. The reasoning model outputs the second response, and the electronic device can then obtain it.
[0060] Table 4 illustrates an input question and a second response provided in an embodiment of this application:
[0061] Table 4
[0062] Table 4 shows that when the input question is "A cuboid and a cube have the same base perimeter and height, then their volumes are also equal. ( )", the second response output by the reasoning model is " <think>First, assume the length and width of the base of the cuboid are *a* and *b*, respectively, and the side length of the base of the cube is *c*. The perimeter of the base of the cuboid is 2(a + b), and the perimeter of the base of the cube is 4c. Since the perimeters are equal, we have 2(a + b) = 4c, which simplifies to a + b = 2c. Assume the height of both the cuboid and the cube is *h*. The volume of the cuboid is... = a b h. The volume of the cube is... = c³ h.\n\nTo make = It needs to satisfy a b = c³. However, according to a + b = 2c, it cannot be guaranteed that a... b must equal c³, unless a = b = c, which is equivalent to the volume of a cube itself. Therefore, the volume of a cuboid is usually smaller than the volume of a cube. In conclusion, the volumes of a cuboid and a cube are not necessarily equal.< / think> Question: If a cuboid and a cube have the same base perimeter and height, then their volumes are also the same. ( ) answer: Let the base of the cuboid have length and width of (a) and (b) respectively, and the base of the cube have side length of (c). Both have height of (h). According to the problem: - The perimeter of the cuboid's base is (2(a + b)). - The perimeter of the cube's base is (4c). - Since the perimeters are the same, we have: \\[\n 2(a + b) = 4c\\quad \\Rightarrow \\quad a + b = 2c\n \\]\n\nIt is important to note that although the perimeters are equal, the base areas of the cuboid and the cube are not necessarily equal. The volume expression is as follows:\n- The volume of the cuboid\\( = a \\times b \\times h\\)\n- Volume of the cube\\( = c^3 \\times h\\)\n\nIf you want\\( = If \\), then it needs to satisfy: \n\\[\na \\times b = c^3\n\\]\nHowever, since \\(a +b = 2c\\), this does not guarantee \\(a \\times b = c^3\\). In fact, only when \\(a = b = c\\), that is, when the cuboid becomes a cube, does \\( = In general, the base area of a cuboid (a × b) is usually smaller than the base area of a cube (c³). For example: Let the base length of a cuboid be 3 and its width be 1. Then its perimeter is 2(3 + 1) = 8, and the corresponding base side length of the cube is 2. The volume of the cuboid is 3 × 1 × h = 3h. The volume of the cube is 2³ × h = 8h. Clearly, 3h ≠ 8h, meaning their volumes are not equal. Therefore, The volumes are not necessarily equal Answer: \boxed{\text{Error}}”.
[0063] After obtaining the first and second responses, the electronic device can use the first response as the labeled intuitive response to the input question and the second response as the labeled deep thinking response to the input question. The first and second responses are then concatenated using preset special markers to obtain the labeled response. For example, additional markers can be added... <intuition> 、``、 <reasoning>Alternatively, use other markers to separate the first and second responses, and add the marked response along with the input question to the cold start dataset.
[0064] In one example, an electronic device can first construct an intuition dataset and an inference dataset. Based on these datasets, a corresponding cold-start dataset, also known as a cold-start training dataset, is then built. The intuition data in the intuition dataset is essentially the output data of the general-purpose large model. There are two approaches to constructing the intuition dataset: the first is to directly use the dataset used in the supervised fine-tuning phase of the general-purpose large model's training process; the second is to directly use a general-purpose large model, such as DeepSeek-V3, to generate the corresponding output for pre-prepared user input data. Combining the data from these two approaches yields the intuition dataset. Each data point in the intuition dataset is a pair of user input and general-purpose large model output. The inference data in the inference dataset is essentially the output data of the inference large model. When constructing the inference dataset, an inference large model, such as DeepSeek-R1, is used to respond to each user input in the intuition output dataset. Each data point in the inference dataset is a pair of user input and inference large model output. The construction of the cold-start training dataset involves concatenating the intuition dataset and the inference dataset. Each data point in the cold-start training dataset is a concatenation of user input, intuition output, and inference output. A special token (special_token) is added to separate intuitive data from inference data.
[0065] Table 5 illustrates an example of an input question and response provided in an embodiment of this application:
[0066] Table 5
[0067] Table 5 shows that when the input question is "A cuboid and a cube have the same base perimeter and height, then their volumes are also equal. ( )", the first and second responses can be " <intuition> 1. Let the length of the rectangle on the base of the cuboid be \\(a\\) and the width be \\(b\\). Let the side length of the base of the cube be \\(c\\). - Given that the perimeters of the bases of the cuboid and the cube are equal, then \\(2(a + b) = 4c\\), which simplifies to \\(a + b = 2c\\). 2. Calculate the base areas of the cuboid and the cube respectively: - Base area of the cuboid \\(S_{length} = ab\\). - Base area of the cube \\(S_{cube} = c^{2}\\). - According to the AM-GM inequality \\(a + b\\geqslant 2\\sqrt{ab}\\) (equality holds if and only if \\(a = b\\), from \\(a + b = 2c\\), we get \\(2c\\geqslant 2\\sqrt{ab}\\), which is \\(c\\geqslant\\sqrt{ab}\\). Squaring both sides, we get \\(c^{2}\\geqslant ab\\), which is the base area of the cube \\(S_{cube}\\geqslant S_{length}\\) (equality holds if and only if \\(a = b\\). 3. Calculate the volume of the cuboid and cube: - The volume of the cuboid \\(V_{length}=S_{length}h = abh\\) (\\(h\\) is the height). - The volume of a cube is V_{rectangular} = S_{rectangular}h = c^{2}h. - Because S_{rectangular} ∈ S_{length} and the heights (h) are equal, V_{rectangular} ∈ V_{length} is only equal in volume when a = b. Generally, the volumes of a cuboid and a cube are not equal. Therefore, the statement "If a cuboid and a cube have the same base perimeter and height, then their volumes are also equal" is incorrect.< / intuition> <think>First, assume the length and width of the base of the cuboid are *a* and *b*, respectively, and the side length of the base of the cube is *c*. The perimeter of the base of the cuboid is 2(a + b), and the perimeter of the base of the cube is 4c. Since the perimeters are equal, we have 2(a + b) = 4c, which simplifies to a + b = 2c. Assume the height of both the cuboid and the cube is *h*. The volume of the cuboid is... = a b h. The volume of the cube is... = c³ h.\n\nTo make = It needs to satisfy a b = c³. However, according to a + b = 2c, it cannot be guaranteed that a... b must equal c³, unless a = b = c, which is equivalent to the volume of a cube itself. Therefore, the volume of a cuboid is usually smaller than the volume of a cube. In conclusion, the volumes of a cuboid and a cube are not necessarily equal.< / think> Question: If a cuboid and a cube have the same base perimeter and height, then their volumes are also the same. ( ) answer: Let the base of the cuboid have length and width of (a) and (b) respectively, and the base of the cube have side length of (c). Both have height of (h). According to the problem: - The perimeter of the cuboid's base is (2(a + b)). - The perimeter of the cube's base is (4c). - Since the perimeters are the same, we have: \\[\n 2(a + b) = 4c \\quad \\Rightarrow \\quad a + b = 2c\n \\]\n We need to note that although the perimeters are equal, the base areas of the cuboid and the cube are not necessarily equal. The volume expression is as follows:\n- The volume of the cuboid\\( = a \\times b \\times h\\)\n-volume of the cube\\( = c^3 \\times h\\)\n\nIf you want\\( = If \\), then it needs to satisfy: \n\\[\na \\times b =c^3\n\\]\nHowever, since \\(a + b = 2c\\), this does not guarantee \\(a \\times b = c^3\\). In fact, only when \\(a = b = c\\), that is, when the cuboid becomes a cube, does \\( = In general, the base area of a cuboid (a × b) is usually smaller than the base area of a cube (c³). For example: Let the base length of a cuboid be 3 and its width be 1. Then its perimeter is 2(3 + 1) = 8, and the corresponding base side length of the cube is 2. The volume of the cuboid is 3 × 1 × h = 3h. The volume of the cube is 2³ × h = 8h. Clearly, 3h ≠ 8h, meaning their volumes are not equal. Therefore, The volumes are not necessarily equal Answer: \boxed{\text{Error}}”.
[0068] The purpose of constructing cold-start data provided in this application embodiment is to enable the target large model to learn an output paradigm that first provides intuitive output, then engages in deep thinking, and finally summarizes the final answer. In this way, during subsequent reinforcement learning training, the large model can optimize its intuitive output based on the content of deep thinking.
[0069] Figure 2 This application provides a schematic diagram of a cold start dataset construction process, which includes the following steps: Figure 2 This will be illustrated by taking the example of obtaining the first reply first and then the second reply.
[0070] S201: Obtain any pre-saved input question.
[0071] S202: Input the input question into the general large model and obtain the first response output by the general large model.
[0072] S203: Input the input question into the large inference model and obtain the second response output by the large inference model.
[0073] S204: Take the first reply as the labeled intuitive reply corresponding to the input question and the second reply as the labeled deep thinking reply corresponding to the input question. Combine the first reply and the second reply and add a preset special mark to obtain the labeled reply.
[0074] S205: Add the labeled response along with the input question to the cold start dataset.
[0075] Example 4: To improve the accuracy of the response, based on the above embodiments, the method in this application embodiment further includes: The target large model is trained based on the Group Relative Policy Optimization (GRPO) algorithm in mathematical, programming, logical reasoning, or other scenarios.
[0076] In mathematics, programming, logical reasoning, or other scenarios, the GRPO algorithm can be used to train large target models. Specifically, a suitable reward function is designed to optimize thought chain reasoning, thereby correcting intuitive output and evaluating the model's performance in different scenarios. The reward function can be defined based on the complexity of the problem, the accuracy of the answer, and the rationality of the reasoning process. For example, in mathematical problems, the reward can be based on the correctness of the answer and the simplicity of the solution steps; in programming problems, the reward can be based on the correctness and efficiency of the code. The GRPO algorithm is used for training. The GRPO algorithm is a reinforcement learning method that maximizes cumulative rewards by optimizing the policy. In each training step, the model generates an action (i.e., the output answer or reasoning process) based on the current state (i.e., the input problem). The reward for the current action is calculated according to the reward function, and the model's parameters are updated to optimize the policy.
[0077] To improve the accuracy of the response, based on the above embodiments, in this embodiment, training the target large model based on the GRPO algorithm includes: Based on pre-saved recognition rules, the target answer corresponding to the sample question is determined. A consistency score is determined based on whether the output answer in the summary of the output response is consistent with the target answer. A format score is determined based on whether the format of the output response conforms to the pre-saved target format. A language score is determined based on whether the language corresponding to the output intuitive response and the output deep thinking response contained in the output response is consistent with the pre-saved target language. The target score is determined based on the consistency score, the format score, the language score, and their respective weights. The step of training the original target large model based on the deviation between the labeled response and the output response includes: The original target large model is trained based on the deviation between the labeled response and the output response and the target score.
[0078] To improve the accuracy of responses, electronic devices can train large target models using reinforcement learning. The GRPO algorithm is used to train models in scenarios such as mathematics, programming, science, and logical reasoning question answering. These scenarios are characterized by clear answers and ease of scoring and verification.
[0079] Specifically, the electronic device can determine the target answer corresponding to the sample question based on pre-stored recognition rules. The sample question refers to the specific question or query that needs to be answered; it can be a simple question, such as "What's the weather like today?", or a complex question requiring multi-step reasoning or analysis. Furthermore, the electronic device can obtain the output answer from both the intuitive response and the deep thinking response. For example, it can obtain the answer from the output response... <summary>The system outputs the answer and determines a consistency score based on whether the output answer matches the target answer. In one example, if the output answer matches the target answer, the consistency score is set to a first preset score; otherwise, a second preset score is set, where the first preset score is greater than the second preset score. In other words, the electronic device uses a rule-based method to analyze the target large model... <summary>The answer in the question is judged as correct or incorrect.
[0080] The electronic device also determines a format score based on whether the format of the output intuitive response and the output deep thinking response conforms to a pre-saved target format. In one example, a regular expression method can be used to determine whether the model's output follows a certain format.<special_token1> |<intuition_answer> |<special_token1> |<special_token2> |<reasoning_process> |<special_token2> | <summary>The format and order. Specifically, it involves determining the target large model within...<special_token1>< / special_token1> An intuitive response was given between them.<special_token2>< / special_token2> The responses generated demonstrated deep thinking. Specifically, the format score was higher when the format conformed to the pre-saved target format than when the format did not conform to the pre-saved target format.
[0081] The electronic device also determines a language score based on whether the language corresponding to the output intuitive response and the output deep thinking response is consistent with the pre-saved target language. The language score is higher when the language is the target language than when the language is not the target language.
[0082] After determining the consistency score, format score, and language score, the electronic device can determine the target score based on the consistency score, format score, language score, and their respective weights.
[0083] After determining the target score, the electronic device trains the original target large model based on the deviation between the labeled intuitive response and the output intuitive response, the deviation between the labeled deep thinking response and the output deep thinking response, and the target score.
[0084] In this embodiment, during the GRPO reinforcement learning training process, the target large model continuously optimizes its thought chain reasoning based on its score. Simultaneously, the optimized thought chain reasoning continuously refines the intuitive output, ensuring that the target large model obtains a good intuitive output response that facilitates thought chain reasoning when it begins outputting. This training stage is the core step in optimizing the large model's intuitive output using reasoning training.
[0085] Example 5: To improve the accuracy of the responses, based on the above embodiments, in this embodiment, after training the target large model, the target large model is fine-tuned in the following way: The pre-saved input question is input into the target large model, and the output of the target large model is obtained as a response to be processed. The response to be processed includes an intuitive response to be processed and a deep thinking response to be processed. Send the pending response to the preset device; The system receives preference responses returned by the preset device, including intuitive responses and deep thinking responses. Based on the deviation between the intuitive preference response and the intuitive response to be processed, and the deviation between the deep thinking preference response and the deep thinking response to be processed, the system adjusts the target large model.
[0086] To improve the accuracy of responses, electronic devices can fine-tune the target large model based on human preferences.
[0087] Specifically, the electronic device can input a pre-saved input question into a target large-scale model, obtain the output of the target large-scale model as a pending response, including an intuitive response and a pending deep-thinking response. This pending response is then sent to a preset device. The user adjusts the pending intuitive and deep-thinking responses based on the preset device to obtain a preferred response, including a preferred intuitive response and a preferred deep-thinking response. The preset device then sends the preferred intuitive and deep-thinking responses back to the electronic device. The electronic device adjusts the target large-scale model based on the deviations between the preferred intuitive response and the pending intuitive response, as well as the deviations between the preferred deep-thinking response and the pending deep-thinking response. This makes the output of the target large-scale model more closely resemble human preferences.
[0088] In supervised fine-tuning training, the training dataset uses a pre-trained target large model to answer user inputs and is then manually modified to obtain a training dataset that conforms to human preferences. In human preference learning training, there are two types of reward functions: one is a reward model similar to traditional large model training, used for reward scoring in general question-answering scenarios; the other is a rule-based scoring method used for reward scoring in scenarios such as programming, mathematics, and logical reasoning. For the intuitive output and final answer of the large model, the reward model is used to determine user-friendliness. For all outputs of the large model, including intuitive output, thought chain reasoning, and the final answer, the reward model is used to determine harmlessness.
[0089] After being trained on human preferences, the intuitive output of the large model is further optimized for inference training in general scenarios. This allows the large model to output an initial response that is both user-friendly and intuitive in any dialogue scenario.
[0090] In this embodiment, model training is divided into four core steps, namely "cold start training dataset construction → supervised fine-tuning (cold start) training → inference training (reinforcement learning training) → human preference training".
[0091] To provide an accurate and effective response, based on the above embodiments, the method in this application embodiment further includes: The intuitive response to the target, which includes a first special marker, is displayed. After the intuitive response to the target is displayed, the deep thinking response to the target, which includes a second special marker, is displayed.
[0092] To provide accurate and effective responses, electronic devices can differentiate and display intuitive responses from deep, thoughtful responses using preset special markers. For example, a first special marker could be used, such as... <intuition>To identify the content of the target's intuitive response. Electronic devices can first display content containing... <intuition>The target intuitive response is marked. These responses are typically short, intuitive, and easy to understand, designed to quickly provide the user with the information they need. After the target intuitive response is displayed, the electronic device then displays a second special marker, such as... <reasoning>The goal is to respond with in-depth thinking. <reasoning>The content within the markers includes a detailed thought process and a final summary, providing users with a deeper understanding and background information.
[0093] For example, in practical applications, electronic devices can first identify and extract information containing... <intuition>The system identifies and displays the target intuitive response to the user. This section is concise and clear, quickly meeting the user's basic needs. After the target intuitive response is displayed, the electronic device continues to identify and extract information containing... <reasoning>The marked target requires in-depth thinking and response.
[0094] For the target large model, after inference training and optimization of the large model's intuitive output, its inference output after deployment also follows an autoregressive pattern. The specific pattern is as follows: the target large model first outputs... <intuition> This token then outputs an intuitive response. The target model, leveraging its knowledge and intuition, quickly generates a preliminary answer. Once the intuitive response is complete, the model outputs...< / intuition> At this point, this content can be directly displayed to the user as an intuitive and rapid response from the large model to the user's input. Subsequently, the target large model outputs... <think> This token marks the start of the thought chain reasoning process. In thought chain reasoning, the target large model spontaneously engages in thinking. When the thinking process concludes, the output is released.< / think> Finally, the target big model provides the final answer based on intuitive output and reasoning results from the thought process chain.
[0095] In the embodiments of this application, during the training process of the target large model, for example, <intuition> 、< / intuition> , <think> 、< / think> Using special tokens to strictly control and guide the model's thought process, ensuring that intuitive output and logical reasoning proceed in sequence, is key to realizing the unique function of the reasoning-based training optimization method for large-scale intuitive output. Through reasoning training and reinforcement learning, the initial responses of the large model possess good intuition, reasoning ability to reflect on intuitive output results, and reasoning capabilities. These are the core elements for improving the model's rapid response accuracy and logical thinking intelligence.
[0096] This unique process, based on special tokens, guides a large model to first generate intuitive outputs, then proceed with thought chain reasoning, and finally produce the final result. It includes the rules for using these special tokens and how they are integrated with the model generation process. It also involves specially constructed data in a specific format for supervised fine-tuning training, and a complete set of training algorithms and strategies for subsequently using reinforcement learning to improve the model's capabilities.
[0097] Compared with existing general large models and inference large models, the method and system proposed in this application for optimizing the intuitive output of large models based on inference training have the following significant advantages: Intuitive optimization: Existing general-purpose large models have only been trained on human preferences, without intuitive optimization training. Their training process lacks the key step of logical reasoning and thinking to correct intuition, which is essential for human intelligence.
[0098] Clear and controllable display of fast and slow thinking processes: Existing general-purpose models only show fast thinking results without slow thinking processes, and reasoning models only show slow thinking processes and final results without fast thinking processes. However, this model, guided by a special token, clearly shows users the process and results of fast and slow thinking, enhancing the responsiveness, transparency, and explainability of the interaction, and improving the user experience.
[0099] Good intuition and result judgment correction: The initial answer of the large model has good intuition and can reason about the intuitive output and reflect and correct it according to the situation. Compared with the general large model, which does not have good intuition and logical reasoning ability and the reasoning large model cannot quickly give an initial result, this model can provide more reliable answers in different scenarios.
[0100] Efficient training optimization: Through a specially designed training method that combines supervised fine-tuning and reinforcement learning, the training speed and accuracy of intuitive output of the model are improved in a targeted manner. Compared with the traditional training method of large models, it can improve the performance of the model more efficiently and adapt to more complex tasks.
[0101] Table 1 is a comparative table of the target large model, the general large model, and the inference large model provided in the embodiments of this application.
[0102]
[0103] Table 1
[0104] As shown in Table 1, the target large model provided in this application embodiment has a faster response speed, higher accuracy, and higher transparency of the thinking process than the general large model and the inference large model.
[0105] Example 6: Based on the same inventive concept, this application provides a large model recovery device, please refer to... Figure 3 The device includes: The input receiving module 301 is used to receive questions to be answered and input the questions to be answered and pre-saved prompt words into the target large model; wherein, the prompt words are used to prompt the target large model to output intuitive and deep thinking responses; Processing module 302 is used to receive the target intuitive response output by the target big model based on intuition when processing the question to be answered; and to receive the target deep thinking response output by the target big model; wherein, the target big model calls the database based on the target intuitive response and the question to be answered to obtain target information related to the keywords in the target intuitive response and the keywords in the question to be answered, and the target big model supplements and adjusts the target intuitive response based on the target information to obtain the target deep thinking response.
[0106] In one possible implementation, the processing module 302 is further configured to acquire any sample question from a pre-saved cold start dataset, and labeled responses saved for that sample question, wherein the labeled responses include labeled intuitive responses and labeled deep thinking responses; wherein the cold start dataset is used for supervised fine-tuning training of the target large model; input the sample question and the labeled responses into the original target large model to acquire the output responses output by the original target large model; wherein the output responses include output intuitive responses and output deep thinking responses; and train the original target large model based on the deviation between the labeled responses and the output responses.
[0107] In one possible implementation, the processing module 302 is further configured to: acquire any pre-saved input question; input the input question into a general large model; acquire a first response output by the general large model, wherein the first response is a response based on intuition from the general large model; input the input question into a reasoning large model; acquire a second response output by the reasoning large model, wherein the second response is a response based on deep thinking from the reasoning large model; concatenate the first response and the second response, and add a preset marker between the first response and the second response to obtain a labeled response; and add the labeled response and the input question together to a cold start dataset.
[0108] In one possible implementation, the processing module 302 is further configured to train the target large model based on the GRPO algorithm in mathematical, programming, logical reasoning, or other scenarios.
[0109] In one possible implementation, the processing module 302 is further configured to: determine the target answer corresponding to the sample question based on pre-saved recognition rules; determine a consistency score based on whether the output answer in the summary of the output response is consistent with the target answer; determine a format score based on whether the format of the output response conforms to a pre-saved target format; determine a language score based on whether the language corresponding to the output intuitive response and the output deep thinking response included in the output response is consistent with a pre-saved target language; and determine a target score based on the consistency score, the format score, the language score, and their respective weights. The processing module 302 is specifically used to train the original target large model based on the deviation between the labeled response and the output response and the target score.
[0110] In one possible implementation, the processing module 302 is further configured to input a pre-saved input question into the target large model, and obtain the output response to be processed from the target large model, wherein the response to be processed includes an intuitive response to be processed and a deep thinking response to be processed. The pending response is sent to a preset device; a preference response is received from the preset device, the preference response including an intuitive response and a preference deep thinking response; the target large model is adjusted based on the deviation between the preference intuitive response and the pending intuitive response, and the deviation between the preference deep thinking response and the pending deep thinking response.
[0111] In one possible implementation, the processing module 302 is further configured to display the target intuitive response containing a first special marker, and after the target intuitive response is displayed, to display the target deep thinking response containing a second special marker.
[0112] Example 7: Based on the same inventive concept, this application provides an electronic device that can realize the large model reply function described above. Please refer to... Figure 4 The device includes a processor 401, a memory 403, and a communication bus 404, wherein the processor 401, the communication interface 402, and the memory 403 communicate with each other through the communication bus 404.
[0113] The memory 403 stores a computer program, which, when executed by the processor 401, causes the processor 401 to perform the following steps: Receive questions to be answered, and input the questions to be answered and pre-saved prompts into the target large model; wherein, the prompts are used to prompt the target large model to output intuitive and deep thinking responses; Receive the target intuitive response output by the target large model based on intuition when processing the question to be answered; It also receives the target deep thinking response output by the target big model; wherein, the target big model calls the database based on the target intuitive response and the question to be answered, obtains target information related to the keywords in the target intuitive response and the keywords in the question to be answered, and the target big model supplements and adjusts the target intuitive response by slow thinking according to the target information to obtain the target deep thinking response.
[0114] In one possible implementation, the target large model is trained in the following manner: Obtain any sample question from a pre-saved cold start dataset, and the labeled responses saved for that sample question, wherein the labeled responses include labeled intuitive responses and labeled deep thinking responses; wherein the cold start dataset is used for supervised fine-tuning training of the target large model; The sample question and the labeled response are input into the original target large model to obtain the output response output by the original target large model; wherein, the output response includes an intuitive response and a deep thinking response; The original target large model is trained based on the deviation between the labeled response and the output response.
[0115] In one possible implementation, the cold start dataset is generated in the following manner: Obtain any pre-saved input question, input the input question into a general model, and obtain the first response output by the general model, wherein the first response is the response output by the general model based on intuition; then input the input question into a reasoning model, and obtain the second response output by the reasoning model, wherein the second response is the response output by the reasoning model based on deep thinking; concatenate the first response and the second response, and add a preset mark between the first response and the second response to obtain the labeled response; The labeled response is added to the cold start dataset along with the input question.
[0116] In one possible implementation, the method further includes: The target large model is trained based on the GRPO algorithm in mathematical, programming, logical reasoning or other scenarios.
[0117] In one possible implementation, training the target large model based on the GRPO algorithm includes: Based on pre-saved recognition rules, the target answer corresponding to the sample question is determined. A consistency score is determined based on whether the output answer in the summary of the output response is consistent with the target answer. A format score is determined based on whether the format of the output response conforms to the pre-saved target format. A language score is determined based on whether the language corresponding to the output intuitive response and the output deep thinking response contained in the output response is consistent with the pre-saved target language. The target score is determined based on the consistency score, the format score, the language score, and their respective weights. The step of training the original target large model based on the deviation between the labeled response and the output response includes: The original target large model is trained based on the deviation between the labeled response and the output response and the target score.
[0118] In one possible implementation, after training the target large model, the target large model is fine-tuned in the following way: The pre-saved input question is input into the target large model, and the output of the target large model is obtained as a response to be processed. The response to be processed includes an intuitive response to be processed and a deep thinking response to be processed. Send the pending response to the preset device; The system receives preference responses returned by the preset device, including intuitive responses and deep thinking responses. Based on the deviation between the intuitive preference response and the intuitive response to be processed, and the deviation between the deep thinking preference response and the deep thinking response to be processed, the system adjusts the target large model.
[0119] In one possible implementation, the method further includes: The intuitive response to the target, which includes a first special marker, is displayed. After the intuitive response to the target is displayed, the deep thinking response to the target, which includes a second special marker, is displayed.
[0120] The communication bus mentioned in the above server can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0121] Communication interface 402 is used for communication between the above-mentioned electronic device and other devices.
[0122] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0123] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0124] Example 8: Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium storing a computer program executable by an electronic device. When the program is run on the electronic device, the electronic device performs the following steps: Receive questions to be answered, and input the questions to be answered and pre-saved prompts into the target large model; wherein, the prompts are used to prompt the target large model to output intuitive and deep thinking responses; Receive the target intuitive response output by the target large model based on intuition when processing the question to be answered; It also receives the target deep thinking response output by the target big model; wherein, the target big model calls the database based on the target intuitive response and the question to be answered, obtains target information related to the keywords in the target intuitive response and the keywords in the question to be answered, and the target big model supplements and adjusts the target intuitive response by slow thinking according to the target information to obtain the target deep thinking response.
[0125] In one possible implementation, the target large model is trained in the following manner: Obtain any sample question from a pre-saved cold start dataset, and the labeled responses saved for that sample question, wherein the labeled responses include labeled intuitive responses and labeled deep thinking responses; wherein the cold start dataset is used for supervised fine-tuning training of the target large model; The sample question and the labeled response are input into the original target large model to obtain the output response output by the original target large model; wherein, the output response includes an intuitive response and a deep thinking response; The original target large model is trained based on the deviation between the labeled response and the output response.
[0126] In one possible implementation, the cold start dataset is generated in the following manner: Obtain any pre-saved input question, input the input question into a general model, and obtain the first response output by the general model, wherein the first response is the response output by the general model based on intuition; then input the input question into a reasoning model, and obtain the second response output by the reasoning model, wherein the second response is the response output by the reasoning model based on deep thinking; concatenate the first response and the second response, and add a preset mark between the first response and the second response to obtain the labeled response; The labeled response is added to the cold start dataset along with the input question.
[0127] In one possible implementation, the method further includes: The target large model is trained based on the GRPO algorithm in mathematical, programming, logical reasoning or other scenarios.
[0128] In one possible implementation, training the target large model based on the GRPO algorithm includes: Based on pre-saved recognition rules, the target answer corresponding to the sample question is determined. A consistency score is determined based on whether the output answer in the summary of the output response is consistent with the target answer. A format score is determined based on whether the format of the output response conforms to the pre-saved target format. A language score is determined based on whether the language corresponding to the output intuitive response and the output deep thinking response contained in the output response is consistent with the pre-saved target language. The target score is determined based on the consistency score, the format score, the language score, and their respective weights. The step of training the original target large model based on the deviation between the labeled response and the output response includes: The original target large model is trained based on the deviation between the labeled response and the output response and the target score.
[0129] In one possible implementation, after training the target large model, the target large model is fine-tuned in the following way: The pre-saved input question is input into the target large model, and the output of the target large model is obtained as a response to be processed. The response to be processed includes an intuitive response to be processed and a deep thinking response to be processed. Send the pending response to the preset device; The system receives preference responses returned by the preset device, including intuitive responses and deep thinking responses. Based on the deviation between the intuitive preference response and the intuitive response to be processed, and the deviation between the deep thinking preference response and the deep thinking response to be processed, the system adjusts the target large model.
[0130] In one possible implementation, the method further includes: The intuitive response to the target, which includes a first special marker, is displayed. After the intuitive response to the target is displayed, the deep thinking response to the target, which includes a second special marker, is displayed.
[0131] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0132] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0133] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0134] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0135] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.< / reasoning> < / intuition> < / reasoning> < / reasoning> < / intuition> < / intuition> < / summary> < / summary> < / summary> < / reasoning> < / intuition> < / summary> < / intuition>
Claims
1. A large-scale model response method, characterized in that, The method further includes: Receive questions to be answered, and input the questions to be answered and pre-saved prompts into the target large model; wherein, the prompts are used to prompt the target large model to output intuitive and deep thinking responses; Receive the target intuitive response output by the target large model based on intuition when processing the question to be answered; It also receives the target deep thinking response output by the target big model; wherein, the target big model calls the database based on the target intuitive response and the question to be answered, obtains target information related to the keywords in the target intuitive response and the keywords in the question to be answered, and the target big model supplements and adjusts the target intuitive response by slow thinking according to the target information to obtain the target deep thinking response.
2. The method according to claim 1, characterized in that, The target large model is trained in the following way: Obtain any sample question from a pre-saved cold start dataset, and the labeled responses saved for that sample question, wherein the labeled responses include labeled intuitive responses and labeled deep thinking responses; wherein the cold start dataset is used for supervised fine-tuning training of the target large model; The sample question and the labeled response are input into the original target large model to obtain the output response output by the original target large model; wherein, the output response includes an intuitive response and a deep thinking response; The original target large model is trained based on the deviation between the labeled response and the output response.
3. The method according to claim 2, characterized in that, The cold start dataset is generated in the following way: Obtain any pre-saved input question, input the input question into a general model, and obtain the first response output by the general model, wherein the first response is the response output by the general model based on intuition; then input the input question into a reasoning model, and obtain the second response output by the reasoning model, wherein the second response is the response output by the reasoning model based on deep thinking; concatenate the first response and the second response, and add a preset mark between the first response and the second response to obtain the labeled response; The labeled response is added to the cold start dataset along with the input question.
4. The method according to claim 2, characterized in that, The method further includes: In mathematical, programming, logical reasoning, or other scenarios, the target large model is trained using the group relative policy optimization GRPO algorithm.
5. The method according to claim 4, characterized in that, The training of the target large model based on the GRPO algorithm includes: Based on pre-saved recognition rules, the target answer corresponding to the sample question is determined. A consistency score is determined based on whether the output answer in the summary of the output response is consistent with the target answer. A format score is determined based on whether the format of the output response conforms to the pre-saved target format. A language score is determined based on whether the language corresponding to the output intuitive response and the output deep thinking response contained in the output response is consistent with the pre-saved target language. The target score is determined based on the consistency score, the format score, the language score, and their respective weights. The step of training the original target large model based on the deviation between the labeled response and the output response includes: The original target large model is trained based on the deviation between the labeled response and the output response and the target score.
6. The method according to any one of claims 2-5, characterized in that, After training the target large model, it is fine-tuned in the following way: The pre-saved input question is input into the target large model, and the output of the target large model is obtained as a response to be processed. The response to be processed includes an intuitive response to be processed and a deep thinking response to be processed. Send the pending response to the preset device; Receive preference responses returned by the preset device, including intuitive responses and preferences based on in-depth thinking; The target large model is adjusted based on the deviation between the preferred intuitive response and the intuitive response to be processed, as well as the deviation between the preferred deep thinking response and the deep thinking response to be processed.
7. The method according to claim 1, characterized in that, The method further includes: The intuitive response to the target, which includes a first special marker, is displayed. After the intuitive response to the target is displayed, the deep thinking response to the target, which includes a second special marker, is displayed.
8. A large model recovery device, characterized in that, The device includes: The input receiving module is used to receive questions to be answered and input the questions to be answered and pre-saved prompts into the target large model; wherein, the prompts are used to prompt the target large model to output intuitive and deep thinking responses; The processing module is used to receive the target intuitive response output by the target big model based on intuition when processing the question to be answered; and to receive the target deep thinking response output by the target big model; wherein, the target big model calls the database based on the target intuitive response and the question to be answered to obtain target information related to the keywords in the target intuitive response and the keywords in the question to be answered, and the target big model supplements and adjusts the target intuitive response by slow thinking based on the target information to obtain the target deep thinking response.
9. An electronic device, characterized in that, The electronic device includes at least a processor and a memory, the processor being used to execute a computer program stored in the memory to implement the steps of the large model recovery method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the large model recovery method as described in any one of claims 1-7.