Method, device and equipment for generating reply, medium and product
By adaptively identifying the complexity of a question and breaking it down into sub-questions, and combining a lightweight large language model and a reflective model, the problems of retrieval failure and excessive resource consumption in intelligent question answering systems when dealing with complex questions are solved, achieving efficient and low-cost solutions to complex questions.
Patent Information
- Application Number
- CN202511129286.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-10-17
AI Technical Summary
Existing intelligent question-answering systems suffer from retrieval failures, insufficient generation quality, and excessive computational resource consumption when dealing with complex questions, making large-scale deployment difficult.
By adaptively identifying the complexity of the problem, the problem is broken down into sub-problems, and combined with retrieval-enhanced generation techniques, a lightweight large language model and a reflective model are used to generate responses.
It improves the system's ability to solve complex problems, meets users' needs for obtaining different information, improves the user experience, and balances effectiveness and cost.
Smart Images

Figure CN120804268A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure generally relate to the field of information processing, and in particular, to methods, apparatuses, devices, media and products for generating a reply BACKGROUND
[0002] Large language models (LLMs) based on the transformer architecture have achieved breakthrough progress in open-domain natural language understanding and generation through large-scale unsupervised pre-training. Such models exhibit strong semantic reasoning capabilities in intelligent question-answering systems and can directly generate text replies that conform to grammar and logic, significantly improving the naturalness of human-computer interaction. The core advantage lies in the generalization of universal knowledge—without adjusting the model structure for specific tasks, it can be adapted to multiple question-answering scenarios through prompt engineering, becoming the standard technical foundation for modern dialogue systems.
[0003] To expand the coverage of dynamic knowledge by LLMs, retrieval-augmented generation (RAG) architecture has gradually become the mainstream of industry practice. This technology cascades an information retrieval system with an LLM generation module: first, it retrieves relevant document snippets from an external knowledge base in real time through a dense vector retrieval engine, and then inputs them as context into the LLM for answer synthesis, improving the LLM's ability to handle user questions. SUMMARY
[0004] Disclosed herein is a method, apparatus, device, medium and product for generating a reply.
[0005] According to a first aspect of the present disclosure, a method for generating a reply is provided. The method comprises determining a complexity type of a question in response to receiving the question. The method comprises obtaining a sub-question of the question by splitting the question in response to the complexity type of the question indicating that the question is a complex question. The method further comprises generating a reply corresponding to the question based on a retrieval result corresponding to the sub-question.
[0006] According to a second aspect of the present disclosure, an apparatus for generating a reply is provided, which comprises a complexity determination module configured to determine a complexity type of a question in response to receiving the question; a question splitting module configured to obtain a sub-question of the question by splitting the question in response to the complexity type of the question indicating that the question is a complex question; and a reply generation module configured to generate a reply corresponding to the question based on a retrieval result corresponding to the sub-question.
[0007] In a third aspect of the present disclosure, an electronic device is provided, comprising at least one processor; and a storage device for storing at least one program, wherein when the at least one program is executed by the at least one processor, the at least one processor implements the method according to the first aspect of the present disclosure.
[0008] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method according to the first aspect of the present disclosure is implemented.
[0009] In a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method according to the first aspect of the present disclosure when executed by a processor.
[0010] It should be understood that the content described in this content section is not intended to limit the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present disclosure.
[0012] Figure 1 A schematic diagram illustrating an example environment in which the apparatus and / or methods of some embodiments of the present disclosure may be implemented;
[0013] Figure 2 A schematic diagram illustrating an example method for generating a reply according to some embodiments of the present disclosure is illustrated;
[0014] Figure 3 illustrates a flow chart of adaptive complex question response generation according to some embodiments of the present disclosure;
[0015] Figure 4 FIGURE 1 illustrates a schematic diagram of problem complexity identification and routing according to some embodiments of the present disclosure;
[0016] Figure 5 A schematic diagram illustrating a solution path for a complex problem according to some embodiments of the present disclosure;
[0017] Figure 6 A schematic diagram illustrating an example of a framework for training a reflection model according to some embodiments of the present disclosure is illustrated;
[0018] Figure 7 A schematic block diagram of an apparatus for generating a reply according to some embodiments of the present disclosure is illustrated;
[0019] Figure 8 FIG. 1 illustrates a schematic block diagram of an example device suitable for use in implementing embodiments of the present disclosure. DETAILED DESCRIPTION
[0020] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0021] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the use range, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to the relevant laws and regulations.
[0022] For example, when receiving the active request of the user, the prompt information is sent to the user to explicitly prompt the user that the operation to be performed by the user will need to acquire and use the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the electronic device, the application program, the server or the storage medium, etc. software or hardware that performs the operation of the technical solutions of the present disclosure according to the prompt information.
[0023] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be the manner of a pop-up window, and the prompt information may, for example, be presented in the form of text in the pop-up window. In addition, the pop-up window may, for example, also carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0024] It can be understood that the above notification and acquisition of the authorization of the user are only illustrative, and do not limit the implementation manners of the present disclosure, and other manners that meet the relevant laws and regulations can also be applied to the implementation manners of the present disclosure.
[0025] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0026] In the description of embodiments of the disclosure, the term "includes" and its conjugates are to be construed as open-ended inclusive of items that follow the term, indicating that "comprising" the recited items but not excluding items not listed. The term "based on" is to be construed as "based at least in part on." The term "one embodiment" or "an embodiment" is to be construed as "at least one embodiment." The term "another embodiment" is to be construed as "at least one other embodiment." The terms "first," "second," etc. can refer to different or the same objects. Other explicit or implicit definitions can also be included below.
[0027] As described above, in existing intelligent question answering systems, RAG technology is widely used in basic information query scenarios. This technology directly generates a reply by matching user questions with a knowledge base, and has the advantages of high efficiency and low cost when dealing with simple problems (such as single entity query, definition explanation, etc.). However, when facing complex problems (such as problems requiring multi-dimensional analysis, multi-step reasoning, or cross-domain integration), the RAG technology has the following inherent defects: retrieval failure: the information needs of complex problems cannot be satisfied by a single retrieval, resulting in a mismatch between the retrieval results and the problem intent; generation quality limitation: lack of deep analysis ability for the problem, resulting in insufficient value of the reply information, and inability to meet the user's demand for obtaining complex information; cost-effectiveness imbalance: if a large language model (LLM) is used to analyze complex problems throughout the process, the computational resource consumption is too high, making it difficult to achieve large-scale deployment of services. To this end, embodiments of the present disclosure propose a method for generating a reply. In this method, the computing device determines the complexity type of a question in response to receiving the question. In response to the complexity type of the question indicating that the question is a complex question, the question is split to obtain a sub-question of the question. Then, based on the retrieval results corresponding to the sub-question, a reply corresponding to the question is generated. Through this method, the system's ability to solve complex problems can be improved, meeting the user's demand for obtaining different information and improving the user experience. In addition, it can also select different reply generation strategies according to the difference in the complexity of the user's question, improve the system's ability to solve simple problems and complex problems that need to be split, balance the effect and cost, and improve the user experience.
[0028] Embodiments of the present disclosure will be described in detail below with further reference to the accompanying drawings. Figure 1 An example environment in which devices and / or methods of embodiments of the present disclosure can be implemented is shown. In the environment 100, a computing device 102 can be used to generate a reply corresponding to a user question.
[0029] Examples of the computing device 102 include, but are not limited to, a personal computer, a server computer, a handheld or laptop device, a mobile device such as a mobile phone, a personal digital assistant (PDA), a media player, etc., a multi-processor system, a consumer electronic product, a minicomputer, a mainframe computer, a distributed computing environment including any of the above systems or devices, etc.
[0030] In the environment 100, the computing device 102 can be used to perform a question answering task that takes a user question 104 as input and outputs a question answer 108. The question answering process is accomplished by the question answering system 106 inside the computing device to determine the complexity, split the question into sub-questions, retrieve the results, and generate the question answer. The structure shown in the example 100 represents an end-to-end question answering framework that includes the analysis, splitting, and answering of questions of different complexity levels in the question answering system 106. Relevant information can be retrieved and corresponding answers can be generated by a large language model in the question answering system for questions of different complexity levels.
[0031] When receiving a user question 104, the question answering system 106 detects whether the user has enabled the complexity mode, in which case the question answering system 106 will split the received user question 104 into sub-questions and continue the subsequent retrieval task. If the user has not enabled the complexity mode, the question answering system 106 will analyze the complexity of the received user question 104, which is performed by a lightweight (e.g., 2B parameter) large language model. Based on the detection result, the question answering system 106 will classify the complexity type of the question into one of the following: simple question (e.g., “What is the first link?”), summary introduction (e.g., “What is your opinion on the increasing difficulty of immigration applications?”), parallel disassembly (e.g., “Chylomicron, pre-beta lipoprotein, beta lipoprotein, and alpha lipoprotein are located on the electrophoresis strip relative to the positive and negative electrodes.”), multiple conditional constraints (e.g., “In the golden age of cinema, which highly acclaimed films effectively used the montage technique?”), or multi-step retrieval (e.g., “What are the common characteristics of stocks that doubled in 2024? Where are they mainly concentrated?”). The complexity type can be used to adopt different question answering paths for simple questions and other complexity type questions. This can achieve adaptive recognition of the complexity of user questions and improve the answering ability of the question answering system.
[0032] When the user selects the complex mode or the complexity type of the user question 104 is classified into one of the following: summary introduction, parallel decomposition, multiple limiting conditions or multi-step retrieval, the question-answering system 106 will split the user question 104 into one or more sub-questions. The splitting of the question is performed by the reflective model. After receiving the user question or generating a sub-question for the user question, the reflective model will analyze the question and determine whether a new round of question splitting is needed to solve the problem. This process is accompanied by information retrieval corresponding to the sub-questions. When the retrieved information is judged by the reflective model to be sufficient to answer the user question 104, the question-answering system 106 will end the question splitting, organize the retrieved information and hand it over to the answer module to generate an answer. The splitting of complex questions improves the intelligent question-answering system's ability to understand and solve complex problems, and meets the user's needs for responding to questions of different complexities.
[0033] The reflective model for decomposing complex problems is trained using an outcome reward model (ORM) reinforcement learning mechanism. It also employs format-constrained and correctness-constrained rewards, enabling training using only open-source complex problem datasets. This eliminates the need for manual standard subtask data and reduces deployment costs. A complexity analysis model is also used to classify and identify the open-source complex problem dataset. A certain proportion of questions corresponding to simple questions, overviews, parallel decompositions, multiple constraints, or multi-step searches are retained as the training dataset for the reflective model, improving its comprehensive problem-solving capabilities.
[0034] In addition, the question-answering system 106 also includes the retrieval of simple questions or one or more sub-questions corresponding to complex questions, for example, by using search engines or some applications to search for information related to the sub-questions in data sources, and organizing it into a given format to facilitate use as input to the answer generation module to obtain a response to the user question 104.
[0035] This method improves the system's ability to answer complex questions, meeting users' diverse information needs and enhancing the user experience. Furthermore, because it can select different response generation strategies based on the complexity of user questions, it enhances the intelligent question-answering system's ability to answer both simple questions and complex questions that require decomposition, balancing effectiveness and cost, while improving the user experience.
[0036] Combined with the above Figure 1 A schematic diagram of an example environment in which the devices and / or methods of some embodiments of the present disclosure may be implemented is described below. Figure 2 A diagram depicting an example method for generating a reply according to some embodiments of the present disclosure. Figure 2 The method in can be Figure 1by the computing device 102 in the system 100 or any suitable device.
[0037] As shown in FIG. 2, an example method 200 describes a flow of a method for generating a reply. Such a method can improve the ability of a large language model to solve complex problems by adaptively identifying the complexity of a problem, breaking down the sub-questions of the problem step by step, and then combining retrieval enhancement generation. Figure 2 At block 202, the computing device 102 determines the complexity type of the question in response to the received question. When the computing device 102 receives the user question 104, it will assign the user question 102 to different question processing paths according to its complexity.
[0038] In some examples, the predefined complexity types include simple questions, review introductions, parallel disassembly, multiple qualification conditions, or multi-step retrieval. Simple questions correspond to questions that can obtain corresponding reference information directly through simple or single retrieval, and such questions can be generated by a lightweight RAG at low cost. Review introductions, parallel disassembly, multiple qualification conditions, or multi-step retrieval questions cannot obtain corresponding reference information through simple or single retrieval, and such complex questions are difficult to retrieve effective reference information through a conventional RAG, and need to be split to adapt to the conventional RAG process.
[0039] In some examples, the computing device 102 will configure an active instruction (e.g., user selection) indicating the start of the complex mode. In response to receiving the user question 104, the computing device 102 will detect whether there is an active instruction indicating the start of the complex mode. If the active instruction is detected, all user questions 104 will enter the subsequent complex question solving channel; if the active instruction is not detected, the complexity of the user question 104 will be determined to determine whether the question should enter the complex question solving channel. Active selection of the complex mode and adaptive detection of the complexity of the user question can balance the balance of user demand, configuration cost, and reply accuracy.
[0040] In some examples, the method used by the computing device 102 to determine the complexity of the user question is to submit the user question to a complexity analysis model for judgment. Such a complexity analysis model is trained using a question and difficulty type dataset, and a language model with a small number of parameters (e.g., 2B parameters) can improve the efficiency of complexity type determination.
[0041] In some examples, the questions in the question and difficulty type dataset come from a set of user questions in a question and answer system log, and in addition, based on the complexity type determination of the same question by multiple large language models, the complexity type corresponding to the user question is selected by voting, and the question and difficulty type dataset is constructed.
[0042] In some examples, the questions in the question and difficulty type dataset come from a set of user questions in a question and answer system log, and in addition, based on the complexity type determination of the same question by multiple large language models, the complexity type corresponding to the user question is selected by voting, and the question and difficulty type dataset is constructed.
[0043] For example, to obtain the question and difficulty type dataset, the computing device can obtain a first set of sample questions using log information. The sample questions can be obtained from the log of the question answering system, for example, to generate the first set of sample questions. Then, the computing device further determines a plurality of complexity types for a first sample question in the first set of sample questions using a plurality of models. For example, each model in the plurality of models can be used to process the first sample question to give a complexity type. Thus, for the sample question, a plurality of complexity types can be obtained using the plurality of models. Then, the computing device determines a sample complexity type for the first sample question using the plurality of complexity types. For example, the sample complexity type can be determined using a voting decision for the plurality of complexity types, in which the complexity type with the highest number of votes can be selected as the sample complexity type for the sample question. Thus, the question and difficulty type dataset can be obtained using this manner to train the complexity analysis model. For example, the first sample question and the sample complexity type can be used to train the model.
[0044] At block 204, the computing device 102 obtains sub-questions of the question by splitting the question in response to the complexity type of the question indicating that the question is a complex question. The complexity type indicates that the user question corresponds to a review introduction, a parallel disassembly, a multi-limit condition, or a multi-step retrieval, which will enter the step of complex question splitting, while simple questions will be solved through the RAG channel. In this way, the classification of user questions according to complexity helps the question answering system to improve the accuracy of complex question replies.
[0045] In some embodiments, for a complex question, the computing device can first determine whether the obtained retrieval results for the question are sufficient to answer the question. If the question is just received, there are no retrieval results available to answer the question. If there are some retrieval results, it is necessary to determine whether the retrieval results can answer the question. If the obtained retrieval results can answer the question, there is no need to split sub-questions from the question, but the obtained retrieval results can be provided to the final reply model to generate a reply to the question. If there are no retrieval results or the obtained retrieval results cannot answer the question, the question needs to be split to obtain sub-questions. Then, corresponding sub-retrieval tasks are generated for the sub-questions to retrieve corresponding retrieval results.
[0046] The above reflection reasoning process for complex problems can be performed by a reflection model, i.e., the splitting of complex problems and the generation of sub-tasks corresponding to the retrieval of sub-problems are performed by the reflection model. In this process, if there is a sub-task generation, the computing device 102 retrieves reference information related to the corresponding sub-problem from a pre-configured data source (e.g., web search, specific application, etc.). If the reflection reasoning process determines that no sub-task generation of the retrieval of sub-problems related to the user question is generated, the generation of new sub-tasks is stopped. The retrieved retrieval documents related to the user question 104 are then generated in a specific format and provided as input to the final reply model.
[0047] At block 206, the computing device 102 generates a reply corresponding to the question based on the retrieval results corresponding to the sub-questions. After obtaining the retrieval results of the sub-questions corresponding to the user question, the retrieval results can be provided to the reflection model to further determine whether the retrieved results are sufficient to answer the user question by the reflection model. If the existing retrieved results can answer the user question, the retrieved results can be provided to the final reply model to generate a reply to the question. For example, the user question 104 and the retrieval results can be provided as input to the final reply model to obtain a solution reply to the user question. Additionally, the final reply model can be a large language model.
[0048] By this method, the system's ability to answer complex questions can be improved, the user's demand for different information can be met, and the user experience can be improved. In addition, since it can also select different reply generation strategies according to the difference in complexity of the user question, the intelligent question answering system's ability to answer simple questions and complex questions that need to be split can be improved, the effect and cost are considered, and the user experience is improved.
[0049] The above describes an example method for generating a reply according to some embodiments of the present disclosure. The following describes a flowchart of adaptive complex question reply generation according to some embodiments of the present disclosure. Figure 2 The above describes an example method for generating a reply according to some embodiments of the present disclosure. The following describes a flowchart of adaptive complex question reply generation according to some embodiments of the present disclosure. Figure 3 The above describes an example method for generating a reply according to some embodiments of the present disclosure. The following describes a flowchart of adaptive complex question reply generation according to some embodiments of the present disclosure.
[0050] Example 300 shows a flowchart for adaptive complex question response generation. The computing device 102 receives a user question at box 302, and the process proceeds to box 304 to determine whether the current complex mode instruction indicates that the complex mode is started. If the complex mode instruction indicates that the complex mode is started, all user questions 104 will enter the complex problem solving channel as complex questions. If the complex mode instruction does not indicate that the complex mode is started, the process proceeds to box 306. At this time, the complexity analysis of the problem will introduce the complexity analysis model 308 to determine whether the user question is a complex question at box 310 (the complexity type corresponds to one of the summary introduction, parallel decomposition, multiple restrictions or multi-step retrieval). If the complexity analysis module determines that the current user question is a complex question, the process enters the complex problem solving channel. If the complexity analysis model determines that the current user question is not a complex question, the process proceeds to box 318 to generate an answer response through the original RAG channel.
[0051] When the complex mode instruction indicates that complex mode is enabled and / or the user question is determined to be a complex question, the response generation process will enter the complex problem solving channel. The user question will generate subtasks through reflective reasoning at box 312, and the reflective model 314 will be introduced at 312 to split the complex question. At box 316, it will be determined whether a subtask corresponding to the retrieval sub-question of the current user question has been generated. If a corresponding subtask has been generated, the process will proceed to box 320, where the search module will search for reference information related to the sub-question from a data source (e.g., web search 322, in-app search 324). If no corresponding subtask has been generated, the process will proceed to box 326, where the answer module will introduce the final response model 328. Using the searched information related to one or more sub-questions of the simple question or complex question as input, it will obtain the final response 330 corresponding to the user question and present it in a streaming manner on the user interface. In addition, if the user question is determined to be a complex question, the thinking content 332 corresponding to the question splitting and generating sub-questions will also be presented in a streaming / non-streaming manner on the user interface.
[0052] Through this adaptive complex question response generation process, the computing device can adaptively divide user questions into two different processing channels based on their complexity, ensuring efficient processing of simple questions while improving the ability to answer complex questions.
[0053] Combined with the above Figure 3 A flow chart of adaptive complex question response generation according to some embodiments of the present disclosure is described. Figure 4 A diagram illustrating problem complexity identification and routing in accordance with some embodiments of the present disclosure is provided.
[0054] like Figure 4 As shown, the example process 400 isFigure 3 At block 402, the computing device receives a user question 402, and Figure 3 Similarly, next at block 404, it is determined whether the user has selected the complex mode. If so, the complex problem solving channel 412 is performed. If the user has not selected the complex mode, the complexity analysis of the problem is further performed at block 406, at which point the complexity analysis model 408 is called to determine the type of problem.
[0055] Table 1 below shows the complexity types of the problems
[0056] Table 1: Problem complexity type table
[0057]
[0058] It can be seen that simple problems usually contain only one sub-problem. Other types of complexity usually include multiple sub-problems. For example, the overview introduction includes more content, including sub-problems that form many aspects of the problem. As for parallel decomposition, it usually contains sub-problems that can be solved into multiple parallel sub-problems. Multi-restriction conditions are problems formed by using multiple conditional restrictions, and multi-step retrieval is formed by multiple problems with dependencies. Then, at box 410, if it is determined to be a complex problem, it enters the complex problem solving channel 412. If it is not a complex problem, it is processed through the original RAG channel 414. Therefore, by classifying problems of different complexities and adopting different processing methods, computing resources can be saved when processing simple problems, and when processing complex problems, the corresponding results can be obtained quickly and accurately.
[0059] For the complexity analysis model 408, a set of questions can be obtained from the question answering system log, and then a number of different models can be used to determine the different complexity types of each question. These models can be large language models, and the following tips can be used:
[0060] The code block is as follows:
[0061] Please categorize the complexity of the following problems: Simple, General, Parallel, Multi-Constraint, or Multi-Step. Please refer to the examples provided for your categorization.
[0062] #Example
[0063] Simple Questions
[0064] Question: What does the first link mean?
[0065] Overview
[0066] Question: What do you think about the increasing difficulty of immigration applications?
[0067] Parallel disassembly
[0068] Question: Where are chylomicrons, pre-beta lipoproteins, beta lipoproteins, and alpha lipoproteins located on the electrophoresis strip relative to the positive and negative electrodes?
[0069] Multiple qualification conditions
[0070] Question: Which highly acclaimed films from the Golden Age of Cinema effectively used the rhythm montage technique?
[0071] Multi-step search
[0072] Question: What common characteristics do stocks that doubled in 2024 share? In which areas are they primarily concentrated?
[0073] Question: {user question}
[0074] Voting on the determination results of the question obtained by the model using the above prompt to determine the difficulty type of the user question, and constructing a training data set for training the complexity analysis model. Then, the complexity analysis model is trained using the training data set.
[0075] Figure 5 A schematic diagram of a complex question solving path is illustrated according to some embodiments of the present disclosure. Figure 5 Example 500 in is a further description of the complex question solving path in Figure 3 Example 500 in is a further description of the complex question solving path in
[0076] For complex problems 502, the computing device can refer to the reflection model 516 to perform reflection reasoning at 514 to generate sub-tasks. In this process, the computing device needs to analyze whether the search results available now can answer the complex problem. For example, for the multi-step search problem “What are the common characteristics of stocks that doubled in price in 2024? Where are they mainly concentrated in?” whether the search results obtained can answer the question completely. In the initial stage, there are usually no search results, and the question cannot be answered. At this time, the question in the first step of multi-step search can be split out, for example, “What are the stocks that doubled in price in 2024?” is split out to generate a sub-question, and then a corresponding sub-task is generated. Then, at block 506, if it is determined that there is a new task generated, for example, a new sub-task is generated, then at block 508, search is performed, for example, through web search 510 or in-application search 512 to obtain search results corresponding to the sub-task. Then, the search results are returned to the reflection model for further determination of whether the search results obtained can answer the complex problem. If it can be answered, the reflection model 516 will no longer generate a new task, and thus the search results obtained are provided to the answer module 514 to generate a final reply. If the search results obtained cannot answer the complex problem, further splitting is performed to generate a new sub-task, and finally after sufficient information is obtained, all search results are returned to the answer module 514 for processing.
[0077] Figure 6 FIG. 7 illustrates a schematic diagram of an example of a framework for training a reflection model according to some embodiments of the present disclosure. In example 700, the training of the reflection model is shown. Before describing the training of the reflection model, the construction of the training data is described first. For training data, the cost of manually annotating the data of sub-tasks is high, and the feasibility is low. Therefore, in the present disclosure, model training can be performed through reinforcement learning based on the outcome reward model (ORM), so that only using the open-source complex problem data set with standard answers can realize the construction of the reflection model.
[0078] By using the foregoing Figure 4The described complexity analysis model classifies and identifies open-source complex problem datasets, retains simple problems, summary introductions, parallel disassembly, multiple conditional constraints, and multi-step searches in a certain proportion, for example, a 1:3:5:10:20 ratio can be used to build, taking into account the diversity and coverage of complex problems, and improving the comprehensive ability of the reflection model. After obtaining the training data, the reflection model can be trained using group relative policy optimization (GRPO) for reinforcement learning. The training dataset constructed using the open-source complex problem dataset can be referred to as the second sample problem set.
[0079] During this process, the complex problem 602 can be provided to the training module 604 for training the reflection model. Then the reflection model 606 determines whether the retrieved information is sufficient to answer the complex problem. If the complex problem cannot be answered, a sub-problem can be split from the complex problem, thereby generating a corresponding sub-retrieval task. Then, the retrieval task for the sub-problem is performed at block 608. At this time, web search 610 and / or in-application search 612 in some applications can be performed. The search results are then returned to the reflection model. If the problem can be answered according to these returned search results, a result in the result set 614 can be generated. By inputting this complex problem multiple times to the reflection model, multiple different results can be generated. For example, results-1, results-2, and results-G included in the result set 614. Then, the results in the result set 614 can be input into the correctness reward model 618 and the format reward model 620 to determine the corresponding rewards. By inputting each result into the correctness reward model 618 and the format reward model 620 respectively, the correctness reward and the format reward for the result can be obtained, and then the two rewards are combined to obtain the final reward for the result, for example, the reward in the reward set 622. The reward set 622 includes multiple rewards corresponding to multiple results. Then, the rewards in the reward set are input into the advantage function to calculate the corresponding advantage, for example, the advantage set 626 includes the advantage for each result.
[0080] In addition, the multiple results can also be processed with the reference model 616 processing the results of the complex problem to calculate the corresponding KL divergence. Then, the KL divergence and the advantage in the advantage set 626 are used to adjust the reflection model, thereby completing the training of the reflection model 606.
[0081] The following is an example of a prompt for a reflection model:
[0082] The code block is as follows:
[0083]
[0084]
[0085]
[0086] As described above, the outcome reward model (ORM) includes two parts of rule-based rewards: one part is the format constraint reward, which is implemented through the format reward model 620, and the other part is the correctness reward of the generated answer, which is implemented through the correctness reward model 618. Both of the two models can be implemented through a large language model (LLM) and scored to improve generalization.
[0087] For the format reward model, which is used for format constraint, the code block used is as follows:
[0088] Code block
[0089] Determine whether the output format of the model is correct and output "correct" or "incorrect".
[0090] Requirements:
[0091] 1. If status_run is incomplete, it must contain a search query.
[0092] 2. It should contain <status_run>...
[0093] and the code block:
[0094] Format reward =
[0095] 0.2 if the format is correct
[0096] 0 if the format is incorrect
[0097] The correctness reward model for determining the correctness of the generated answer, the corresponding code block is as follows:
[0098] Code block:
[0099] For a given question and its corresponding gold answer, evaluate the accuracy of the predicted answer. If the predicted answer completely matches the meaning and basic details of the gold answer, it is considered correct. If the prediction is accurate, answer "true",
[0100] otherwise answer "false".
[0101] Question: {}
[0102] Gold answer: {}
[0103] Predicted answer: {}
[0104] And code block:
[0105] Correctness reward =
[0106] Reward 1.0 if result is correct
[0107] Reward 0 if result is incorrect
[0108] After the above result reward model reinforcement learning process, the reflection model is constructed, that is, the user's complex problem can be task decomposed, and effective reference information can be retrieved and obtained step by step until the user's problem can be solved, and then the final reply generation can be performed.
[0109] In the final reply generation, the retrieved information can be used for final reply generation. For example, the reference information sufficient to answer the user is obtained through the reflection model, and the above reference information and answer requirements are integrated into the prompt of the reply model to answer the user's question. The reply model can be any suitable large language model. For example, the prompt of the reply model is as follows,
[0110] Code block:
[0111] # The following content is the search result related to the user's message:
[0112] {search_results}
[0113] In the search results I provided to you, the format of each result is [doc X begin]...[doc X end], where X represents the numerical index of each document. Please quote the context at the end of the relevant sentence as appropriate.
[0114] When replying, please note the following points:
[0115] - Today is {cur_date}.
[0116] - The content in the search results may not all be closely related to the user's question. You need to evaluate and filter the search results according to the question.
[0117] - For list type questions (for example, list all flight information), try to limit the answer to within 10 key points, and inform the user that they can refer to the search source for complete information. Prefer to provide the most complete and relevant items in the list. Unless necessary, avoid mentioning content not provided in the search results.
[0118] - For creative tasks (for example, writing a paper), make sure to cite references in the text, such as
[0119] [citation:3][citation:5], rather than just at the end of the text. You need to interpret and summarize the user's needs, choose the appropriate format, make full use of the search results, extract the key information, and produce an insightful, creative and professional answer. Extend the length of the response as much as possible, elaborate on each point from multiple angles,
[0120] Make sure the content is rich and thorough.
[0121] If your response is long, please structure it appropriately and summarize it in sections. If you need to elaborate on a point by point basis, please limit your response to 5 points and consolidate related content.
[0122] -For objective questions and answers, if the answer is very short, you can add one or two relevant sentences to enrich the content.
[0123] - Choose an appropriate and visually appealing format for your response based on user needs and the content of your answer to ensure good readability.
[0124] - Your answer should synthesize information from multiple relevant web pages and avoid citing the same web page repeatedly.
[0125] - Your response should use the same language as the user's question, unless the user requests otherwise.
[0126] -#The user's message is:
[0127] {question}
[0128] Figure 7 FIG2 illustrates a schematic block diagram of an apparatus for generating a reply according to some embodiments of the present disclosure. Figure 7 As shown, the device 700 can be Figure 1 The apparatus 700 is implemented in a computing device 102, and the apparatus 700 includes a complexity determination module 702, configured to determine the complexity type of a question in response to receiving a question; a question splitting module 704, configured to obtain sub-questions of the question by splitting the question in response to the complexity type of the question indicating that the question is a complex question; and a reply generation module 706, configured to generate a reply corresponding to the question based on the retrieval results corresponding to the sub-questions.
[0129] In some embodiments, the complexity type is one of: simple question, summary introduction, parallel decomposition, multiple restrictions, or multi-step search.
[0130] In some embodiments, the complexity determination module 702 includes: a first type determination module configured to determine the complexity type of a question by inputting the question into a complexity analysis model.
[0131] In some embodiments, the training module of the complexity analysis model comprises: a first sample question set obtaining module configured to obtain a first sample question set using the log; a plurality of complexity type determining modules configured to determine a plurality of complexity types of a first sample question in the first sample question set by a plurality of models; a sample complexity type determining module configured to determine a sample complexity type for the first sample question based on the plurality of complexity types; and a first training module configured to train the complexity analysis model based on the first sample question and the sample complexity type.
[0132] In some embodiments, the sample complexity type determining module comprises a voting decision determining module configured to determine the sample complexity type using a voting decision based on the plurality of complexity types.
[0133] In some embodiments, the question splitting module 704 comprises: a question determining module configured to determine whether the obtained search result for the question is sufficient to answer the question; and a splitting module configured to split the question to obtain a sub-question in response to the obtained search result being unable to answer the question.
[0134] In some embodiments, the apparatus 700 further comprises a first generating module configured to generate a reply corresponding to the question based on the obtained search result in response to the obtained search result being able to answer the question.
[0135] In some embodiments, the question splitting module 704 comprises a sub-question obtaining module configured to obtain the sub-question by inputting the question into a reflection model.
[0136] In some embodiments, the training module of the reflection model comprises: a second sample question set obtaining module configured to obtain a second sample question set having standard answers; and a reinforcement learning training module configured to train the reflection model using reinforcement learning based on the second sample question set.
[0137] In some embodiments, the reinforcement learning training module comprises: a reward determining module configured to determine, for a second sample question in the second sample question set, a correctness reward and a format reward of a correctness of a predicted result corresponding to the second sample question using the reflection model; and a second training module configured to train the reflection model based on the correctness reward and the format reward.
[0138] In some embodiments, the apparatus 700 further comprises a thinking process providing module configured to provide a thinking process of the reflection model for display.
[0139] In some embodiments, the apparatus 700 further comprises a search result obtaining module configured to obtain a search result by performing a search sub-task corresponding to the sub-question.
[0140] Figure 8 A schematic block diagram of an example device 800 that can be used to implement embodiments of the present disclosure is shown. Figure 1 The computing device 102 in FIG. 1 can be implemented with the device 800. As shown, the device 800 includes a central processing unit (CPU) 801 that can perform various suitable actions and processes according to computer program instructions stored in a read-only memory (ROM) 802 or computer program instructions loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required by the device 800 for operation are also stored in the RAM 803. The CPU 801, the ROM 802, and the RAM 803 are connected to each other by a bus 804. An input / output (I / O) interface 807 is also connected to the bus 804.
[0141] Various components in the device 800 are connected to the I / O interface 807, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; the storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0142] The various processes and procedures described above, such as the method 200, can be performed by the processing unit 801. For example, in some embodiments, the method 200 can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the CPU 801, one or more actions of the example method 200 described above can be performed
[0143] The present disclosure can be a method, apparatus, system, and / or computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for performing various aspects of the present disclosure.
[0144] Computer readable storage media can be tangible storage media which can retain and store instructions for use by an instruction execution device. Computer readable storage media can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer readable storage media include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0145] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0146] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0147] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0148] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0149] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0150] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0151] Embodiments of the present disclosure have been described above, and the description is intended to be illustrative, and not restrictive, of the disclosed embodiments. Many modifications and variations of the described embodiments are possible, and all such modifications and variations are intended to be within the scope of the described embodiments. The description used herein is intended to be illustrative, and not restrictive, of the described embodiments. The scope of the described embodiments is not limited to the examples and / or embodiments described herein but only by the claims and their equivalents.
Claims
1. A method for generating a response, comprising: In response to receiving the question, determining a complexity type of the question; In response to the complexity type of the problem indicating that the problem is a complex problem, obtaining sub-problems of the problem by splitting the problem; Based on the retrieval results corresponding to the sub-questions, a reply corresponding to the question is generated.
2. The method according to claim 1, wherein the complexity type is one of the following: simple question, summary introduction, parallel decomposition, multiple restrictions or multi-step search.
3. The method according to claim 1 , wherein determining the complexity type of the problem comprises: The complexity type of the problem is determined by inputting the problem into a complexity analysis model.
4. The method according to claim 3, wherein the training of the complexity analysis model comprises: Using the log to obtain a first sample question set; Determining multiple complexity types of first sample problems in the first sample problem set by using multiple models; Determining a sample complexity type for the first sample problem based on the multiple complexity types; as well as The complexity analysis model is trained based on the first sample problem and the sample complexity type.
5. The method according to claim 4, wherein determining the sample complexity type for the first sample problem based on the plurality of complexity types comprises: Based on the multiple complexity types, a voting decision is made to determine the sample complexity type.
6. The method according to claim 1, wherein obtaining sub-problems of the problem by splitting the problem comprises: determining whether the retrieved search results for the question are sufficient to answer the question; In response to the obtained search results being unable to answer the question, the question is split to obtain the sub-questions.
7. The method according to claim 6, further comprising: In response to the obtained search results being able to answer the question, a reply corresponding to the question is generated based on the obtained search results.
8. The method according to claim 1, wherein obtaining sub-problems of the problem by splitting the problem comprises: The sub-questions are obtained by inputting the question into a reflective model.
9. The method according to claim 8, wherein the training of the reflection model comprises: Obtaining a second set of sample questions with standard answers; as well as Based on the second sample question set, the reflection model is trained using reinforcement learning.
10. The method according to claim 9, wherein based on the second sample question set, using reinforcement learning to train the reflection model comprises: For a second sample question in the second sample question set, determining a correctness reward and a format reward for a prediction result corresponding to the second sample question using the reflection model; as well as The reflection model is trained based on the correctness reward and the format reward.
11. The method according to claim 8, further comprising: The thought process of the reflective model is provided for display.
12. The method according to claim 1, further comprising: The retrieval result is obtained by executing the retrieval subtask corresponding to the subquestion.
13. An apparatus for generating a reply, comprising: a complexity determination module configured to determine a complexity type of the question in response to receiving the question; A question splitting module is configured to obtain sub-questions of the question by splitting the question in response to the complexity type of the question indicating that the question is a complex question; as well as The reply generation module is configured to generate a reply corresponding to the question based on the retrieval results corresponding to the sub-questions.
14. An electronic device comprising: at least one processor; as well as A storage device for storing at least one program, wherein when the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 12.
15. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 12 when executed by a processor.
16. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Question answering method and device, electronic equipment and computer readable storage medium
CN116244418A
Document question and answer method, system and equipment based on RAG and medium
CN119003725A
Multi-intelligent-agent physical examination enhancement generation method and device, equipment and storage medium
CN120256588A
Knowledge question and answer method and device, electronic equipment and storage medium
CN120353897A
Question answering system and method based on complex reasoning pipeline
US20250190741A1