Fine-tuning method of preset model, question and answer method based on preset model, and device

By rewriting and evaluating the retrieval results and answers of the pre-defined model, and combining process-outcome dual reinforcement learning, the problem of low response accuracy in the fine-tuning of large language models is solved, and higher response accuracy is achieved.

CN121009964BActive Publication Date: 2026-01-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511527072.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-01-27
Estimated Expiration
2045-10-24

AI Technical Summary

Technical Problem

Existing large language models lack reinforcement learning for internal reasoning processes during fine-tuning, leading to reduced response accuracy.

Method used

By rewriting and reflecting on the pre-set model, retrieval questions are generated and retrieval results and answers are evaluated. Fine-tuning is then performed using process-outcome dual reinforcement learning to optimize retrieval strategies and answer generation.

Benefits of technology

It improves the accuracy of the preset model in retrieval and answer generation, and achieves more comprehensive question understanding and answer generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009964B_ABST
    Figure CN121009964B_ABST
Patent Text Reader

Abstract

The application provides a preset model fine-tuning method, a preset model-based question and answer method and equipment, and relates to the technical field of artificial intelligence. The fine-tuning method comprises: rewriting a sample question by using a preset model to obtain a retrieval question; retrieving based on the retrieval question to obtain a retrieval result for the retrieval question; reflecting the retrieval result by using the preset model to obtain a reflection result for the retrieval result; generating an answer to the sample question based on the retrieval result by using the preset model in the case where the reflection result represents that the retrieval result meets a predetermined answer generation condition; evaluating the retrieval result and the answer respectively to obtain a process evaluation score for the retrieval result and a result evaluation score for the answer; and fine-tuning parameters of the preset model based on a training sample obtained from the sample question, the answer, the retrieval result, the process evaluation score and the result evaluation score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to a method for fine-tuning a preset model, a question-answering method based on a preset model, and an electronic device. Background Technology

[0002] With the rapid development of artificial intelligence, large language models have made significant breakthroughs in research. They have gained powerful general language understanding and generation capabilities through pre-training on massive amounts of data and have been widely used in various fields.

[0003] In practical applications, fine-tuning large language models is an important step in adapting them to specific domains or tasks. In related technologies, fine-tuning of large language models typically employs supervised fine-tuning methods, relying on "input-output" paired samples for training. This approach only performs reinforcement learning on the output results, lacking reinforcement learning of the internal reasoning process, leading to a decrease in the accuracy of the large language model's responses. Summary of the Invention

[0004] In view of this, this application provides a method for fine-tuning a preset model, a question-answering method based on a preset model, and an electronic device.

[0005] One aspect of this application provides a method for fine-tuning a pre-defined model, comprising: rewriting a sample question using the pre-defined model to obtain a retrieval question; performing a retrieval based on the retrieval question to obtain retrieval results for the retrieval question; reflecting on the retrieval results using the pre-defined model to obtain a reflection result for the retrieval results; generating an answer to the sample question using the pre-defined model based on the retrieval results, provided that the reflection result indicates that the retrieval results meet predetermined answer generation conditions; evaluating the retrieval results and the answer respectively to obtain a process evaluation score for the retrieval results and a result evaluation score for the answer; and fine-tuning the parameters of the pre-defined model using training samples constructed based on the sample question, the answer, the retrieval results, the process evaluation score, and the result evaluation score.

[0006] One aspect of this application provides a question-answering method based on a preset model, the method comprising: receiving a question input by an object, processing the question using a preset model, and generating an answer to the question; wherein the preset model is obtained by fine-tuning based on the aforementioned fine-tuning method.

[0007] Another aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0008] According to the technical solution of this application, by evaluating the retrieval results and answers separately, and constructing training samples based on sample questions, answers, retrieval results, process evaluation scores, and result evaluation scores, process-result dual reinforcement learning can be achieved using process evaluation scores and result evaluations. This allows the preset model to more comprehensively understand the problems it has in the two stages of retrieval and answer generation, and achieve deep optimization of the retrieval strategy and answer generation in both stages, thereby improving the accuracy of the preset model's response. Attached Figure Description

[0009] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments of this application with reference to the accompanying drawings.

[0010] Figure 1 An exemplary system architecture diagram is shown, illustrating a fine-tuning method for applying a preset model according to an embodiment of this application.

[0011] Figure 2 A flowchart of a fine-tuning method for a preset model according to an embodiment of this application is shown.

[0012] Figure 3 A schematic flowchart illustrating the process of determining search results according to an embodiment of this application is shown.

[0013] Figure 4 A flowchart illustrating the determination of reflection results according to an embodiment of this application is shown.

[0014] Figure 5 A flowchart of a model fine-tuning method according to another embodiment of this application is shown.

[0015] Figure 6 A flowchart of a question-and-answer method based on a preset model according to an embodiment of this application is shown.

[0016] Figure 7 The diagram illustrates the structure of a fine-tuning device for a preset model according to an embodiment of this application.

[0017] Figure 8 The diagram illustrates the structure of a question-and-answer device based on a preset model according to an embodiment of this application.

[0018] Figure 9 A block diagram illustrating a fine-tuning method suitable for implementing a preset model according to an embodiment of this application is shown schematically. Detailed Implementation

[0019] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or first components, but do not exclude the presence or addition of one or more other features, steps, operations, or first components.

[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0022] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0023] Figure 1 An exemplary system architecture diagram is shown, illustrating a fine-tuning method for applying a preset model according to an embodiment of this application.

[0024] It is important to note that Figure 1 The examples shown are merely examples of system architectures applicable to the embodiments of this application, intended to help those skilled in the art understand the technical content of this application. They do not imply that the embodiments of this application cannot be used in other devices, systems, environments, or scenarios. For instance, in another embodiment, an exemplary system architecture for applying a fine-tuning method for a preset model may include a terminal device. However, the terminal device can implement the fine-tuning method for the preset model provided in the embodiments of this application without interacting with a server.

[0025] like Figure 1As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0026] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).

[0027] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0028] Server 105 can be a server that provides various services, such as a backend management server that supports the content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0029] It should be noted that the fine-tuning method of the preset model provided in this application embodiment can generally be executed by the first terminal device 101, the second terminal device 102, and the third terminal device 103. Alternatively, the fine-tuning method of the preset model provided in this application embodiment can also generally be executed by the server 105. The fine-tuning method of the preset model provided in this application embodiment can also be executed by a server or server cluster that is different from the server 105 and is capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0030] It should be understood that Figure 1 The number of the first terminal device, second terminal device, third terminal device, network, and server is only a certain number. Depending on the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0031] In the technical solution of this application, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information all comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and they do not violate public order and good morals.

[0032] In the technical solution of this application, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0033] It should be noted that the sequence numbers of the operations in the following methods are for descriptive purposes only and should not be considered as indicating the execution order of the operations. Unless explicitly stated otherwise, the method does not need to be executed in the exact order shown.

[0034] Figure 2 A flowchart of a fine-tuning method for a preset model according to an embodiment of this application is shown.

[0035] like Figure 2 As shown, the fine-tuning method of the preset model includes operations S210 to S260.

[0036] In operation S210, the sample question is rewritten using a preset model to obtain the retrieval question.

[0037] In operation S220, a search is performed based on the search question, and search results are obtained for the search question.

[0038] In operation S230, the search results are reflected upon using a preset model to obtain the reflection results on the search results.

[0039] In operation S240, if the reflection result characterization of the retrieval results meets the predetermined answer generation conditions, the preset model is used to generate the answer to the sample question based on the retrieval results.

[0040] In operation S250, the search results and answers are evaluated separately to obtain process evaluation scores for the search results and result evaluation scores for the answers.

[0041] In operation S260, the parameters of the preset model are fine-tuned based on the training samples constructed from sample questions, answers, retrieval results, process evaluation scores, and result evaluation scores.

[0042] For example, the preset model can be a generative model. For instance, it can include a large language model, a large visual model, a multimodal large model, etc. The preset model involved in the embodiments of this application can be a general large model, or it can be an expert large model obtained by fine-tuning based on requirements; the embodiments of this application do not limit this.

[0043] The sample question can be a historical question input by the user into the preset model. The sample question can be relatively vague and complex. The preset model has question planning capabilities, which can be used to rewrite the sample question into a clear and concise retrieval question.

[0044] The pre-defined model can retrieve search results for a given query by calling a search engine. The search engine can be any one or more data source interfaces, such as internet search engine interfaces, text or vector knowledge bases, databases, etc. The search engine can retrieve a predetermined number of search results for the given query, such as 3, 5, 6, etc.

[0045] The pre-defined model also possesses reflective capabilities. It can reflect on the search results, compare them to determine if there are any flaws, and generate a reflective result. When the search results are free of flaws, it indicates that the reflective result meets the predetermined answer generation conditions.

[0046] For example, search results and answers can be input into an evaluation preset model respectively, and the search results and answers can be evaluated separately. The evaluation preset model can be the preset model itself or other large language models with evaluation capabilities. This application does not limit the type of evaluation preset model.

[0047] The process evaluation score can reflect whether the search results meet the requirements of the question. If the process evaluation score is low, it means that the search results may not be accurate or relevant enough. Based on the process evaluation score, when fine-tuning the preset model, it can guide the preset model to adjust the search strategy to obtain better search results.

[0048] The result evaluation score is a direct reflection of the answer quality. If the result evaluation score is low, it means that the answer is not accurate enough. Based on the result evaluation score, when fine-tuning the preset model, it can guide the preset model to adjust the answer generation strategy to improve the answer quality.

[0049] According to the embodiments of this application, by evaluating the retrieval results and answers respectively, and constructing training samples based on the sample questions, answers, retrieval results, process evaluation scores, and result evaluation scores, process-result dual reinforcement learning can be achieved using process evaluation scores and result evaluations. This allows the preset model to more comprehensively understand the problems it has in the two stages of retrieval and answer generation, and achieve deep optimization of the two stages of retrieval strategy and answer generation, thereby improving the accuracy of the preset model's response.

[0050] According to an embodiment of this application, rewriting a sample question using a preset model to obtain a retrieval question may include: inputting the sample question and a first prompt template representing a strategy for determining the retrieval question into the preset model to generate structured demand prompt information; generating multiple sub-sample questions based on the structured demand prompt information; and obtaining the retrieval question based on the multiple sub-sample questions.

[0051] For example, the first prompt template can include structured prompts, which may include, but are not limited to, prompts for demand type, core elements, analysis dimensions, implicit demand judgments, and follow-up question judgments. The demand type clarifies the question category, such as "industry trend analysis," "causal relationship decomposition," or "data comparison." Core elements may include multiple elements such as research object, time range, core indicators, and analysis dimensions. Implicit demand judgments indicate information that the user may be interested in but is not explicitly mentioned. Follow-up question judgments determine if there is any missing information in the question; if follow-up questions are needed, concise follow-up question statements are generated.

[0052] Based on structured prompts, the sample problem can be broken down into structured requirement prompts. For example, if the sample problem is "Analyze the core reasons for the decline in new energy vehicle sales in 2024," the structured requirement prompts could be:

[0053] Request type: Industry causal analysis;

[0054] Key elements: Research object = new energy vehicles, time range = the whole year of 2024, core indicator = year-on-year decline in sales, analysis dimensions = policy, supply chain costs, and consumer demand;

[0055] Implicit requirement: It must include differences between specific vehicle models (pure electric / hybrid);

[0056] Data preference: Prioritize authoritative institutions.

[0057] For example, structured demand prompts can be combined to identify multiple concise and clear sub-sample questions, and these sub-sample questions can be sorted according to their importance to obtain a list of retrieval questions composed of multiple sub-sample questions.

[0058] Continuing with the sample question "Analyze the core reasons for the decline in new energy vehicle sales in 2024", the generated list of search questions could be: "2024 New Energy Vehicle Sales Decline Due to Policy Withdrawal", "2024 Lithium Price and New Energy Vehicle Costs", "2024 New Energy Vehicle Consumer Willingness Survey".

[0059] According to embodiments of this application, by combining sample questions with a first prompt template and inputting them into a preset model, structured demand prompt information is generated. This transforms ambiguous or complex user queries into clear and actionable retrieval instructions, reducing semantic ambiguity. Generating multiple sub-sample questions based on structured prompts decomposes the dimensions of the sample questions, expands the retrieval coverage, and enhances the accuracy and robustness of the retrieval strategy.

[0060] Figure 3 A schematic flowchart illustrating the process of determining search results according to an embodiment of this application is shown.

[0061] like Figure 3 As shown, the sample question 301 and the first prompt template 302 are input into the preset model M303. The first prompt template 302 contains structured prompt words. Based on the structured prompt words and the sample question, the preset model M303 can generate structured demand prompt information 304. Using the structured demand prompt information 304, multiple sub-sample questions 305 are generated. The preset model can call the retrieval tool 306 to perform searches on the multiple sub-sample questions 305, obtaining multiple search results 307. Thus, multi-dimensional rewriting of the sample question can be achieved, expanding the search scope.

[0062] According to embodiments of this application, reflecting on the search results using a preset model to obtain a reflection result for the search results may include: inputting the search results and a second prompt template representing a strategy for determining the reflection result into the preset model to generate reflection prompt information; if the reflection prompt information indicates that the search results do not meet the predetermined answer generation conditions, performing a second search based on the reflection prompt information to obtain supplementary search results for the sample question; and reflecting on the search results and supplementary search results using the preset model to obtain a reflection result for the search results and supplementary search results.

[0063] For example, the second prompt template can include multiple reflection subtasks, each used to reflect on the search results from different dimensions. The preset model can execute multiple reflection subtasks in parallel based on the search results to generate reflection prompt information.

[0064] The predetermined answer generation conditions are used to determine whether the search results are good enough. For example, the predetermined answer generation conditions may be that the reflection prompt message shows that the search results are without defects or that the reflection prompt message shows that the number of supplementary searches has met the predetermined number.

[0065] Re-searching based on reflection prompts can involve generating new, more precise search keywords based on the reflection results, or redirecting to different data sources to obtain supplementary search results.

[0066] The search results and supplementary search results are input back into the preset model. The second prompt template guides the user to reflect on the search and supplementary results again. The reflection results can assess whether the supplementary search has addressed the initial deficiencies and whether the current information is sufficiently relevant, complete, and accurate. If the results are still unsatisfactory, iterative searches can be performed until the predetermined number of searches is reached.

[0067] For example, the search results can be filtered based on the reflection results of the search results and supplementary search results, removing search results that are erroneous or irrelevant, resulting in updated search results, and answers can be generated based on the updated search results.

[0068] According to embodiments of this application, by reflecting on the initial search results using a second prompt template, problems such as information gaps or insufficient relevance can be automatically diagnosed. When it is identified that the results do not meet the conditions, a targeted re-search can be triggered, making up for the blind spots of the initial search, effectively reducing the risks caused by incomplete or one-sided information, and upgrading the search process from a single search to a continuously optimized adaptive cycle.

[0069] According to embodiments of this application, the strategy for determining the reflection result includes a conflict detection strategy, a logic detection strategy, and an evidence completion strategy. The search results and a second prompt template representing the strategy for determining the reflection result are input into a preset model to generate reflection prompt information. This includes: performing information conflict detection on the search results according to the conflict detection strategy in the second prompt template, and generating conflict detection prompt information; performing logic vulnerability detection on the search results according to the logic detection strategy in the second prompt template, and generating logic detection prompt information; performing evidence integrity detection on the search results according to the evidence completion strategy in the second prompt template, and generating evidence completion prompt information; and obtaining reflection prompt information based on the conflict detection prompt information, logic detection prompt information, and evidence completion prompt information.

[0070] The second prompt template includes prompts for strategies such as conflict detection, logical fallacy detection, and evidence completion. Conflict detection prompts guide the pre-defined model to detect factual conflicts among multiple search results. Logical fallacy detection prompts guide the model to detect causal logical flaws in the search results. Evidence completion prompts guide the model to assess the sufficiency of the evidence chain in the search results.

[0071] For example, a conflict detection strategy may include multiple execution steps, guiding a pre-defined model to progressively reflect on each search result. For instance, the first step may involve guiding the large model to read the search results and extract factual statements into a fact list. The second step involves comparing each factual statement in the fact list pairwise, recording the conflict type, and scoring the conflict to obtain a conflict value. Conflict types include, but are not limited to, time conflicts, numerical conflicts, and opinion conflicts. Time conflicts indicate inconsistencies in the timing or timeliness of events; numerical conflicts indicate different values ​​for the same indicator; and opinion conflicts indicate inconsistent conclusions about the same event from search results from different data sources. The third step generates conflict detection prompts based on the conflict information and conflict values. When the conflict value meets a predetermined threshold, it indicates that the reflection information does not meet the predetermined answer generation conditions. Therefore, a secondary supplementary search can be performed based on the conflict information to obtain the latest or authoritative supplementary search results containing the conflict topic. Factual statements are extracted from the supplementary search results and added to the fact list for re-comparison, thereby obtaining updated conflict detection prompts. Multiple supplementary searches and reflections can be performed, and the number of executions can be recorded. The conflict detection task ends when the predetermined number of executions is met or the conflict value is lower than the preset value.

[0072] For example, a logic detection strategy can also include multiple execution steps. For instance, the first step can be to extract causal assertions from the fact list into causal chains. The second step is to detect whether each causal chain satisfies the premises sufficient to draw a conclusion, whether there are hidden hypotheses, and whether there are counterexamples or missing data causing jumps, etc., indicating causal vulnerabilities. The third step is to score the logical completeness of each causal chain, obtaining a causal score; if the causal score is lower than a preset value, the causal chain is marked as a vulnerability. The fourth step is to generate causal detection prompts based on the marked vulnerability causal chains. If a vulnerability is found, it indicates that the reflection information does not meet the predetermined answer generation conditions; therefore, a secondary supplementary search and reflection can be performed based on the marked vulnerability causal chains. Similarly, multiple supplementary searches and reflections can be performed until there are no causal vulnerabilities or the number of searches meets the predetermined number.

[0073] For example, an evidence completion strategy can also include multiple execution steps. For instance, the first step might be to list the sequence of key evidence needed to answer the question. The second step might be to compare the sequence of facts with the sequence of key evidence to identify any missing key evidence. The third step might be to generate evidence completion suggestions based on the missing key evidence. When key evidence is missing, a secondary supplementary search and reflection can be performed based on the missing key evidence. Similarly, multiple supplementary searches and reflections can be performed until the key evidence is complete or the predetermined number of searches is met.

[0074] According to embodiments of this application, by introducing multi-level reflection strategies, the quality and reliability of search results can be systematically improved. Conflict detection strategies can identify contradictions in search results, ensuring internal data consistency; logic detection strategies can analyze the rationality of reasoning chains, avoiding loopholes or jumps; and evidence completion strategies assess the completeness supporting the search results, strengthening the foundation of the conclusions. By integrating these strategies into a preset model and generating comprehensive reflection prompts, information can be fully verified, avoiding invalid output.

[0075] Figure 4 A flowchart illustrating the determination of reflection results according to an embodiment of this application is shown.

[0076] like Figure 4 As shown, when sample question 301, second prompt template 401, and search result 307 are input into preset model M303, preset model M303 can generate conflict detection prompt information 402 based on the conflict detection strategy in second prompt template 401, generate logic detection prompt information 403 based on the logic detection strategy in second prompt template 401, and generate evidence completion prompt information 404 based on the evidence completion strategy in second prompt template 401. Based on conflict detection prompt information 402, logic detection prompt information 403, and evidence completion prompt information 404, reflection result 405 is generated.

[0077] According to embodiments of this application, evaluating the search results and answers separately to obtain process evaluation scores for the search results and result evaluation scores for the answers may include: evaluating the matching degree between sample questions and answers using a preset evaluation model to obtain result evaluation scores, or responding to feedback information received from the recipient regarding the answer and obtaining result evaluation scores for the answer based on the feedback information; evaluating the relevance of each search result to the question to obtain a relevance evaluation result for each search result; and obtaining a process evaluation score based on the relevance evaluation results for each search result.

[0078] For example, the correctness of the answer can be evaluated using a pre-defined assessment model, with a value of 1 for a correct answer and 0 for an incorrect answer. Alternatively, the correctness can be assessed based on feedback from an object (such as a user). For instance, a user liking the generated answer is recorded as 1, and a user disliking the generated answer is recorded as 0. See formula (1):

[0079] (1).

[0080] The result evaluation score can be determined based on the correctness of the answer. See formula (2):

[0081] (2).

[0082] For example, the relevance between the search results and the question can be evaluated using a pre-defined evaluation model. The calculation of the relevance can be found in formula (3):

[0083] (3)

[0084] in, These are relevant keywords. This indicates the evaluation of the pre-set model. Let i represent the problem of the i-th subsample. Represents the problem of the i-th subsample. In step j, retrieve the recalled documents. This represents the correlation score, which ranges from 0 to 1, where 0 represents no correlation and 1 represents perfect correlation.

[0085] The process reward score can be determined based on the relevance of the search results to the question and the correctness of the answer. The calculation can be found in formula (4):

[0086] (4)

[0087] in, To determine the relevance between the i-th question and the document retrieved in step j, To adjust the coefficients, δ is positive if the final answer is correct, and negative if the final answer is incorrect.

[0088] According to embodiments of this application, a combination of user feedback (likes / dislikes) and evaluation scores based on a preset model is used to assess the correctness of answers. Correct answers are rewarded, while incorrect answers are penalized. The process evaluation score employs a relevance verification method based on the search results and the question, which can reward solutions that yield relevant search results and correct answers, while penalizing solutions that yield incorrect answers or irrelevant search results.

[0089] According to embodiments of this application, fine-tuning the parameters of a preset model based on training samples constructed from sample questions, answers, retrieval results, process evaluation scores, and result evaluation scores may include: constructing training samples by using sample questions, answers, and retrieval results as input information and the comprehensive evaluation score obtained from process evaluation scores and result evaluation scores as labels; obtaining multiple training samples, and, if the multiple training samples meet a predetermined quantity condition, fine-tuning the parameters of the preset model based on the multiple training samples.

[0090] For example, the comprehensive evaluation score The calculation can be found in formula (5):

[0091] (5).

[0092] For example, the preset model can construct multiple training samples based on collecting different sample questions, obtaining retrieval results, answers, and comprehensive evaluation scores. By continuously accumulating training samples, a model fine-tuning training is initiated when the number of training samples meets a predetermined condition. For example, a model fine-tuning training is initiated every time 1024 training samples are accumulated, thereby achieving continuous optimization of the preset model.

[0093] According to embodiments of this application, fine-tuning the parameters of a preset model based on multiple training samples includes: performing multiple rounds of iterative fine-tuning of the preset model based on multiple training samples; wherein, each round of iterative fine-tuning includes: determining the target loss based on the advantage value determined by the comprehensive evaluation score of the training samples in the previous round and the probability ratio of the retrieval strategy before and after the parameter update; the advantage value characterizes the quality of the retrieval strategy, and the retrieval strategy includes at least one of the strategies for determining the retrieval question, determining the retrieval result, and determining the reflection result; the probability of the retrieval strategy characterizes the probability that the retrieval strategy is adopted; updating the parameters of the preset model based on the average loss determined by the target losses of the multiple training samples to obtain the preset model after fine-tuning in the previous round; inputting the sample question into the preset model after fine-tuning in the current round to generate the retrieval result and answer for the current round; evaluating the retrieval result and answer for the current round to obtain a process evaluation score for the retrieval strategy for the current round and a result evaluation score for the answer for the current round.

[0094] By iteratively fine-tuning the preset model using multiple training samples in multiple rounds, the comprehensive reward score can be used as the optimization target. By using multiple training samples and iteratively fine-tuning in multiple rounds, the comprehensive reward score can be continuously maximized, thus completing one round of fine-tuning training.

[0095] For example, the advantage value of a search strategy can be determined based on the overall score and a preset baseline value. The baseline value can be represented as the average evaluation score of different search strategies, and the advantage value can be further represented as the degree to which the search strategy is superior to the average level. For example, if the overall evaluation score is 10 and the baseline value is 7, then the advantage value is 3, indicating that the search strategy performs 3 units better than the average level. The direction for optimizing the search strategy can be determined based on the advantage value.

[0096] The policy probability ratio reflects the probability change of the new detection policy relative to the old retrieval policy. It can be used to constrain the update magnitude of the retrieval policy to avoid gradient explosion or policy collapse.

[0097] By combining the advantage value and the policy probability ratio, the advantage value can guide the gradient update direction, while the policy probability ratio can control the parameter update step size. The two work together on the target loss, which can achieve a balance between effective exploration and stable updating of parameters.

[0098] The target loss of each of the multiple training samples can be determined separately, and the average loss of the multiple target losses can be obtained. The parameters of the preset model can be updated using the average loss. This ensures that the fine-tuning of the model is based on the optimization of the overall sample data, rather than optimization for individual cases.

[0099] After completing one round of parameter updates, updated search results and answers can be generated based on the updated model and sample questions, and evaluated to obtain process evaluation scores and result evaluation scores, which can then be used for the next round of parameter updates.

[0100] According to an embodiment of this application, determining the target loss based on the advantage value determined by the process evaluation score and result evaluation score of the training sample in the previous round, and the policy probability ratio of the retrieval policy before and after the parameter update, may include: determining a first loss based on the product of the advantage value and the policy probability ratio; determining a second loss based on preset pruning parameters, the advantage value, and the policy probability ratio, wherein the pruning parameters are used to limit the policy probability ratio to a target range; and determining the target loss based on the smaller value of the first loss and the second loss.

[0101] For example, the model parameters can be updated according to the following loss function:

[0102] (6)

[0103] in, This represents the average loss across multiple training samples; This represents the average of the target losses over multiple training samples; x represents the sample problem; y represents the retrieval strategy. This represents a set of multiple training samples.

[0104] Indicates the probability ratio of the strategies. = , These are the model parameters that need to be updated. These are the parameters of the old model before the update; The new model is in the problem The probability of choosing retrieval strategy y. The old model is the problem. The probability of choosing retrieval strategy y.

[0105] This represents the advantage value, used to determine the merits of search strategy y when the problem is x. If A(x,y)>0: Choosing search strategy y for problem x results in a higher overall evaluation score than the average. That is, the search strategy is good, and updating the parameters can increase the probability of using this strategy. Conversely, if... If the value is less than 0, the search strategy is a bad strategy. When updating the parameters, the probability of using this search strategy can be reduced.

[0106] This represents the clipping function. Indicates the clipping parameters. We can set it to 0.2 to limit the policy probability ratio to between 0.8 and 1.2, thus preventing the new model from being too far from the old model.

[0107] The function is a strategy to suppress extremes, choosing the smaller value from two options to avoid "over-favoring" or "over-rejecting" a certain search strategy.

[0108] For example, if =1.5, =2, If the ratio is 0.2, then the cropped ratio is 1.2, and the target loss is min(1.5×2,1.2×2)=2.4. This achieves a balance between the advantageous retrieval strategy and the constraint update step size.

[0109] This represents a regularization term used to prevent the pre-defined model from losing its original basic capabilities (such as understanding task generation and execution strategies, and simple reasoning logic) during updates. KL divergence measures the difference in probability distribution between the new model and the original model. Penalty weights control the intensity of the penalty for KL divergence. A value of 0.1 can be used. By using regularization terms to additionally constrain the overall change in the strategy distribution, it ensures that the updated retrieval strategy does not deviate too far from the old retrieval strategy, maintaining the continuity and exploratory capability of the strategy.

[0110] By minimizing the loss function, the pre-defined model can continuously iterate to develop an efficient and stable deep retrieval strategy. This means the retrieval strategy achieves a high overall evaluation score without deviating from its original capabilities or losing its core functionality.

[0111] Figure 5 A flowchart of a model fine-tuning method according to another embodiment of this application is shown.

[0112] like Figure 5As shown, sample question 301 is input into preset model M303. The sample question is rewritten using preset model M303 to obtain multiple sub-sample questions 305. Each sub-sample question 305 is searched to obtain search results 307. The search results are then reflected upon using preset model M303 to obtain reflection results 405. If reflection result 405 meets the answer generation conditions, answer 502 is generated based on search result 307. If reflection result 405 does not meet the predetermined answer generation conditions, a second search is performed based on reflection result 405 to obtain supplementary search results. The search results and supplementary search results are then reflected upon using preset model M303 to obtain reflection results for the search results and supplementary search results. This process is iterated until reflection result 405 meets the predetermined answer generation conditions or the number of iterations reaches a predetermined number. Finally, answer 502 is generated based on the search results and supplementary search results.

[0113] A comprehensive evaluation score is obtained by evaluating the process evaluation score obtained from assessing the search results and the answer evaluation score obtained from assessing the answers. Training samples 501 are constructed based on the sample questions, search results, answers, and comprehensive evaluation scores. These training samples 501 are then used to fine-tune the pre-defined model M303.

[0114] Figure 6 A flowchart of a question-and-answer method based on a preset model according to an embodiment of this application is shown.

[0115] like Figure 6 As shown, the method includes operations S610 to S620.

[0116] Problems with receiving object input when operating S610.

[0117] When operating the S620, a preset model is used to process the problem and generate an answer to the problem.

[0118] According to an embodiment of this application, the preset model can be obtained by fine-tuning according to the fine-tuning method of operations S210 to S260.

[0119] According to the embodiments of this application, since the preset model is obtained after being fine-tuned by the above-mentioned fine-tuning method, the generated answer can be more accurate.

[0120] Figure 7 The diagram illustrates the structure of a fine-tuning device for a preset model according to an embodiment of this application.

[0121] like Figure 7 As shown, the fine-tuning device 700 of the preset model includes a planning module 710, a retrieval module 720, a reflection module 730, a generation module 740, an evaluation module 750, and a fine-tuning module 760.

[0122] The planning module 710 is used to rewrite the sample question using a preset model to obtain the retrieval question.

[0123] The retrieval module 720 is used to perform retrieval based on the retrieval question and obtain retrieval results for the retrieval question.

[0124] The reflection module 730 is used to reflect on the search results using a preset model to obtain reflection results on the search results.

[0125] The generation module 740 is used to generate answers to sample questions based on the search results using a preset model, provided that the search results represent the reflection results and meet the predetermined answer generation conditions.

[0126] The evaluation module 750 is used to evaluate the search results and the answers separately, and obtain the process evaluation score for the search results and the result evaluation score for the answers.

[0127] The fine-tuning module 760 is used to fine-tune the parameters of the preset model based on the training samples constructed from sample questions, answers, retrieval results, process evaluation scores, and result evaluation scores.

[0128] According to an embodiment of this application, the planning module 710 includes a first generation submodule, a second generation submodule, and a determination submodule.

[0129] The first generation submodule is used to input the sample question and the first prompt template representing the strategy for determining the retrieval question into the preset model to generate structured requirement prompt information.

[0130] The second generation submodule is used to generate multiple sub-sample questions based on structured requirement prompts.

[0131] The determination submodule is used to obtain the retrieval question based on multiple subsample questions.

[0132] According to an embodiment of this application, the reflection module 730 includes a third generation submodule, a supplementary retrieval submodule, and a supplementary reflection submodule.

[0133] The third generation submodule is used to input the search results and the second prompt template representing the strategy for determining the reflection results into the preset model to generate reflection prompt information.

[0134] The supplementary retrieval submodule is used to perform a second retrieval based on the reflection prompt information when the retrieval results do not meet the predetermined answer generation conditions, so as to obtain supplementary retrieval results for the sample question.

[0135] The supplementary reflection submodule is used to reflect on the search results and supplementary search results using a preset model, and obtain reflection results on the search results and supplementary search results.

[0136] According to the embodiments of this application, the strategy for determining the reflection result includes a conflict detection strategy, a logic detection strategy, and an evidence completion strategy; the reflection module 730 includes: a first prompt information generation submodule, a second prompt information generation submodule, a third prompt information generation submodule, and a reflection prompt information generation submodule.

[0137] The first prompt information generation submodule is used to perform information conflict detection on the search results according to the conflict detection strategy in the second prompt template, and generate conflict detection prompt information.

[0138] The second prompt information generation submodule is used to perform logical vulnerability detection on the search results according to the logical detection strategy in the second prompt template, and generate logical detection prompt information.

[0139] The third prompt information generation submodule is used to perform evidence integrity checks on the search results based on the evidence completion strategy of the second prompt template, and generate evidence completion prompt information.

[0140] The reflection prompt information generation submodule is used to generate reflection prompt information based on conflict detection prompt information, logic detection prompt information, and evidence completion prompt information.

[0141] According to an embodiment of this application, the evaluation module 750 includes a first result evaluation submodule, a second result evaluation submodule, a first process evaluation submodule, and a second process evaluation submodule.

[0142] The first result evaluation submodule is used to evaluate the degree of matching between sample questions and answers using a preset evaluation model, and obtain the result evaluation score.

[0143] The second result evaluation submodule is used to respond to the feedback information received from the object regarding the answer, and to obtain the result evaluation score of the answer based on the feedback information.

[0144] The first process evaluation submodule is used to evaluate the relevance of each search result to the question, and obtain the relevance evaluation result of each search result.

[0145] The second process evaluation submodule is used to obtain a process evaluation score based on the relevance evaluation results of each search result.

[0146] According to an embodiment of this application, the fine-tuning module 760 includes a construction submodule and a fine-tuning submodule.

[0147] The construction submodule is used to construct training samples by taking sample questions, answers and retrieval results as input information, and the comprehensive evaluation score obtained by process evaluation score and result evaluation score as labels.

[0148] The fine-tuning submodule is used to acquire multiple training samples and, if the number of training samples meets a predetermined condition, to fine-tune the parameters of the preset model based on the multiple training samples.

[0149] According to an embodiment of this application, the fine-tuning submodule is used to perform multiple rounds of iterative fine-tuning on a preset model based on multiple training samples; the fine-tuning submodule includes a target loss determination unit, a parameter update unit, a generation unit, and an evaluation unit.

[0150] The target loss determination unit is used to determine the target loss based on the advantage value determined by the comprehensive evaluation score of the training samples in the previous round and the ratio of the probability of the retrieval strategy before and after the parameter update. The advantage value characterizes the quality of the retrieval strategy, which includes at least one of the following: strategy for determining the retrieval question, strategy for determining the retrieval result, and strategy for determining the reflection result. The probability of the retrieval strategy characterizes the probability that the retrieval strategy will be adopted.

[0151] The parameter update unit is used to update the parameters of the preset model based on the average loss determined by the target loss of multiple training samples, so as to obtain the preset model after fine-tuning in the previous round.

[0152] The generation unit is used to input sample questions into the preset model after fine-tuning in the current round, and generate the search results and answers for the current round.

[0153] The evaluation unit is used to evaluate the search results and answers in the current round, and obtains a process evaluation score for the search strategy in the current round and a result evaluation score for the answer in the current round.

[0154] According to an embodiment of this application, the target loss determination unit includes a first determination subunit, a second determination subunit, and a third determination subunit.

[0155] The first determining subunit is used to determine the first loss based on the product of the advantage value and the strategy probability ratio.

[0156] The second determining subunit is used to determine the second loss based on preset pruning parameters, advantage value, and policy probability ratio. The pruning parameters are used to limit the policy probability ratio to a target range.

[0157] The third determining subunit is used to determine the target loss based on the smaller of the first loss and the second loss.

[0158] Figure 8 The diagram illustrates the structure of a question-and-answer device based on a preset model according to an embodiment of this application.

[0159] like Figure 8 As shown, the question-and-answer device 800 includes a receiving module 810 and an answer generation module 820.

[0160] The receiving module 810 is used to receive object input.

[0161] The answer generation module 820 uses a preset model to process the question and generate an answer to the question.

[0162] Any one or more of the modules, submodules, units, and subunits according to the embodiments of this application, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, units, and subunits according to the embodiments of this application can be implemented by dividing them into multiple modules. Any one or more of the modules, submodules, units, and subunits according to the embodiments of this application can be at least partially implemented as hardware circuits, such as field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits (ASICs), or implemented by hardware or firmware in any other reasonable manner by integrating or packaging circuits, or implemented in any one of software, hardware, and firmware, or in a suitable combination of any of these. Alternatively, one or more of the modules, submodules, units, and subunits according to the embodiments of this application can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.

[0163] Figure 9 A block diagram illustrating a fine-tuning method suitable for implementing a preset model according to an embodiment of this application is shown schematically.

[0164] Figure 9 A block diagram of an electronic device suitable for implementing the methods described above, according to an embodiment of this application, is illustrated schematically. Figure 9 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0165] Electronic devices are intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The first components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described and / or claimed herein.

[0166] like Figure 9As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.

[0167] Multiple first components in electronic device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of displays, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0168] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the fine-tuning method. For example, in some embodiments, the fine-tuning method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the fine-tuning method described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform the fine-tuning method by any other suitable means (e.g., by means of firmware).

[0169] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0170] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to the processor or controller of a general-purpose computer, special-purpose computer, or other programmable test apparatus, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0171] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0172] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0173] The systems and technologies described herein can be implemented in computing systems that include a back-end first component (e.g., as a data server), or a computing system that includes a middleware first component (e.g., an application server), or a computing system that includes a front-end first component (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such a back-end first component, middleware first component, or front-end first component. The first components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0174] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.

[0175] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

[0176] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.

Claims

1. A method for fine-tuning a preset model, characterized in that, include: The sample question is rewritten using a pre-defined model to obtain the retrieval question; A search is performed based on the search question to obtain search results for the search question. The search results are analyzed using a preset model to obtain a reflection result on the search results; If the reflection result indicates that the retrieval result meets the predetermined answer generation conditions, then a preset model is used to generate the answer to the sample question based on the retrieval result. The search results and the answers are evaluated separately to obtain a process evaluation score for the search results and a result evaluation score for the answers. The parameters of the preset model are fine-tuned based on the training samples constructed from the sample questions, answers, retrieval results, process evaluation scores, and result evaluation scores. The preset model is iteratively fine-tuned in multiple rounds based on multiple training samples; The iterative fine-tuning in any round includes: The target loss is determined based on the advantage value determined by the comprehensive evaluation score of the training samples in the previous round and the probability ratio of the retrieval strategy before and after the parameter update. The advantage value represents the degree of superiority or inferiority of the retrieval strategy. The retrieval strategy includes at least one of the following: a strategy for determining the retrieval question, a strategy for determining the retrieval result, and a strategy for determining the reflection result. The probability of the retrieval strategy represents the probability that the retrieval strategy will be adopted. The parameters of the preset model are updated based on the average loss determined by the target loss of each of the multiple training samples, so as to obtain the preset model after fine-tuning in the previous round. The sample question is input into the preset model after fine-tuning for the current round to generate the search results and answers for the current round; The search results and answers for the current round are evaluated to obtain a process evaluation score for the search strategy for the current round and a result evaluation score for the answer for the current round.

2. The fine-tuning method according to claim 1, characterized in that, The parameters of the preset model are fine-tuned based on the training samples constructed from the sample questions, answers, retrieval results, process evaluation scores, and result evaluation scores, including: The sample questions, answers, and search results are used as input information, and the comprehensive evaluation score obtained from the process evaluation score and the result evaluation score is used as a label to construct training samples; Multiple training samples are obtained, and if the multiple training samples meet a predetermined number condition, the parameters of the preset model are fine-tuned based on the multiple training samples.

3. The fine-tuning method according to claim 1, characterized in that, Based on the advantage value determined by the comprehensive evaluation score of the training samples in the previous round, and the ratio of the probability of the retrieval strategy before and after the parameter update, the target loss is determined, including: The first loss is determined by the product of the advantage value and the strategy probability ratio; Based on preset pruning parameters, the advantage value, and the strategy probability ratio, a second loss is determined, wherein the pruning parameters are used to limit the strategy probability ratio to a target range; The target loss is determined based on the smaller of the first loss and the second loss.

4. The fine-tuning method according to claim 1, characterized in that, The sample question is rewritten using a pre-defined model to obtain the retrieval question, including: The sample question and the first prompt template representing the strategy for determining the retrieval question are input into the preset model to generate structured requirement prompt information; Based on the structured requirement prompts, multiple sub-sample questions are generated; The retrieval question is derived based on multiple sub-sample questions.

5. The fine-tuning method according to claim 1 or 4, characterized in that, The search results are analyzed using a preset model to obtain a reflection result on the search results, including: The search results and a second prompt template representing the strategy for determining the reflection results are input into a preset model to generate reflection prompt information; If the reflection prompt information indicates that the search results do not meet the predetermined answer generation conditions, a second search is performed based on the reflection prompt information to obtain supplementary search results for the sample question. The search results and supplementary search results are analyzed using a preset model to obtain the analysis results.

6. The fine-tuning method according to claim 5, characterized in that, The strategies for determining the results of reflection include conflict detection strategies, logic detection strategies, and evidence completion strategies. The search results and a second prompt template representing the strategy for determining reflection results are input into a preset model to generate reflection prompt information, including: Based on the conflict detection strategy in the second prompt template, information conflict detection is performed on the search results, and conflict detection prompt information is generated; Based on the logic detection strategy in the second prompt template, the search results are subjected to logic vulnerability detection, and logic detection prompt information is generated; Based on the evidence completion strategy in the second prompt template, the search results are subjected to evidence integrity detection, and evidence completion prompt information is generated; The reflection prompt information is obtained based on the conflict detection prompt information, logic detection prompt information, and evidence completion prompt information.

7. The fine-tuning method according to claim 1, characterized in that, The search results and the answers are evaluated separately to obtain a process evaluation score for the search results and a result evaluation score for the answers, including: The matching degree between the sample questions and answers is evaluated using a preset evaluation model to obtain a result evaluation score, or the result evaluation score of the answer is obtained based on the feedback information received from the object. The relevance of each search result to the search question is evaluated to obtain the relevance evaluation result of each search result. The process evaluation score is obtained based on the relevance evaluation results of each search result.

8. A question-answering method based on a preset model, characterized in that, The method includes: The problem of receiving input from the object; The problem is processed using a preset model to generate an answer to the problem; The preset model is obtained by fine-tuning based on the fine-tuning method as described in any one of claims 1 to 7.

9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Low-cost retrieval enhancement generation evaluation method based on branch splicing

    CN118966198A

  • Question and answer method and device based on large model, electronic equipment and medium

    CN119884300A

  • Large language model training method and device, electronic equipment and storage medium

    CN120633756A

  • Structured query statement generation method and system

    CN120780730A