Training sample generation method, training method and related device

By constructing a search tree containing questions, candidate answers, and standard answers, high-quality training samples are generated, solving the problem of high manual evaluation costs in Internet healthcare, realizing automated quality evaluation model training, and improving evaluation efficiency and consistency.

CN122065032APending Publication Date: 2026-05-19ALI HEALTH TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ALI HEALTH TECH CO LTD
Filing Date
2026-02-05
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In scenarios with extremely high risk control requirements, such as internet healthcare, existing technologies rely on manual evaluation to assess the quality of answers from large language models, resulting in high evaluation costs and limited scalability, making it difficult to efficiently obtain high-quality training samples.

Method used

By acquiring initial samples of questions, candidate answers, and standard answers, a search tree containing multiple reasoning paths is constructed to form target training samples, including questions, candidate answers, reasoning paths, and standard answers. The training data for the quality assessment model is then automatically constructed using a generation device.

Benefits of technology

It enables the automatic construction of high-quality training samples, reduces the reliance on manual writing of evaluation reasons, improves the generation efficiency of quality evaluation models and the consistency of evaluation, and ensures the security and reliability of Internet medical Q&A services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122065032A_ABST
    Figure CN122065032A_ABST
Patent Text Reader

Abstract

The invention provides a training sample generation method, a training method and a related device. The method comprises the following steps: acquiring a primary sample comprising a question, a candidate answer and a standard answer; wherein the candidate answers are to-be-evaluated answers for the question; the standard answer comprises standard content for answering the question; constructing a search tree based on the primary sample; wherein the search tree comprises a plurality of reasoning paths; the reasoning path comprises a thinking chain for obtaining quality evaluation of the candidate answers based on the questions and the standard answers; the thinking chain comprises a plurality of reasoning nodes, and each reasoning node comprises at least one reasoning sentence; and combining the primary sample with the reasoning path of the search tree to form a target training sample comprising the question, the candidate answer, the reasoning path and the standard answer. The sample generation quality can be improved to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method for generating training samples, a training method, and related apparatus. Background Technology

[0002] In scenarios with extremely high risk control requirements, such as internet healthcare, manual evaluation is typically used to assess the quality of dialogue answers generated by large language models. However, this manual evaluation heavily relies on domain experts who must manually judge the quality of each answer, its corresponding candidate answers, and the standard answer. This process is not only highly specialized but also extremely time-consuming and labor-intensive, resulting in persistently high evaluation costs.

[0003] To reduce reliance on manual annotation, the industry has begun exploring the use of large language models to automatically perform answer quality assessment. However, building a reliable and accurate automatic assessment model is highly dependent on high-quality training samples. Therefore, efficiently acquiring or synthesizing high-quality training samples has become a key challenge in improving automatic assessment capabilities and achieving scalable and low-cost quality assurance. Summary of the Invention

[0004] In view of this, one or more embodiments of this application provide a method for generating training samples, a training method, and related apparatus, which can improve the quality of generated training samples to a certain extent.

[0005] In a first aspect, one or more embodiments of this application propose a method for generating training samples, comprising: obtaining a primary sample including a question, candidate answers, and a standard answer; wherein the candidate answers are answers to be evaluated for the question; the standard answer includes standard content for answering the question; constructing a search tree based on the primary sample; wherein the search tree includes multiple reasoning paths; the reasoning path includes a thought chain for deriving a quality evaluation of the candidate answers based on the question and the standard answer; the thought chain includes multiple reasoning nodes, each reasoning node including at least one reasoning sentence; and combining the primary sample and the reasoning paths of the search tree to form a target training sample including the question, candidate answers, reasoning paths, and standard answer.

[0006] Secondly, one or more embodiments of this application propose a training method for a quality assessment model, comprising: training the quality assessment model based on target training samples generated by the aforementioned training sample generation method; wherein the quality assessment model is used to generate quality assessments for answers to user questions corresponding to online question-answering models.

[0007] Thirdly, one or more embodiments of this application propose a training sample generation apparatus, comprising: an acquisition module, a construction module, and a combination module; the acquisition module is used to acquire a primary sample including a question, candidate answers, and a standard answer; wherein the candidate answers are answers to be evaluated for the question; the standard answer includes standard content for answering the question; the construction module is used to construct a search tree based on the primary sample; wherein the search tree includes multiple reasoning paths; the reasoning path includes a thought chain that derives a quality evaluation of the candidate answers based on the question and the standard answer; the thought chain includes multiple reasoning nodes, each reasoning node including at least one reasoning sentence; the combination module is used to combine the primary sample and the reasoning paths of the search tree to form a target training sample including a question, candidate answers, reasoning paths, and a standard answer.

[0008] Fourthly, one or more embodiments of this application propose a training apparatus for a quality assessment model, comprising: a training module for training a quality assessment model based on the target training samples as described above; wherein the quality assessment model is used to generate a quality assessment for the answers to user questions corresponding to an online question-answering model.

[0009] Fifthly, one or more embodiments of this application provide a computer device including a memory and a processor, wherein the memory stores at least one computer program, which is loaded and executed by the processor to implement the method as described above.

[0010] In a sixth aspect, one or more embodiments of this application provide a computer program product including computer instructions that, when executed by a processor, implement the method as described above.

[0011] In a seventh aspect, one or more embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method as described above.

[0012] As can be seen from the above embodiments, multiple embodiments in this application first obtain a primary sample including a question, candidate answers, and a standard answer, and then construct a search tree containing multiple reasoning paths based on the primary sample. Each reasoning path records multiple reasoning nodes for quality evaluation of candidate answers based on the question and the standard answer in the form of a thought chain. Then, the primary sample is combined with the reasoning paths in the search tree to form a target training sample that simultaneously contains a question, candidate answers, reasoning paths, and a standard answer. This realizes the automatic construction of training data for the quality evaluation model and improves the generation quality of the training samples. Attached Figure Description

[0013] Figure 1 This is a schematic diagram illustrating the process of a training sample generation method provided in one embodiment of this application.

[0014] Figure 2 This is a flowchart illustrating a method for generating training samples according to an embodiment of this application.

[0015] Figure 3 This is a schematic diagram of a training sample generation device provided in one embodiment of this application.

[0016] Figure 4 This is a schematic diagram of a computer device provided in one embodiment of this application. Detailed Implementation

[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments.

[0018] In the description of the embodiments of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0019] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0020] In scenarios with extremely high risk control requirements, such as internet healthcare, large language models have begun to be used to generate candidate answers for patient questions. However, to ensure that these candidate answers are controllable in terms of safety, professionalism, and compliance, the industry typically still uses manual evaluation. Doctors or medical editors evaluate the quality of each model output and optimize the model or adjust security strategies accordingly. This manual evaluation model relies heavily on reviewers with professional backgrounds. The review process is meticulous and heavily dependent on individual experience, resulting in high overall evaluation costs and limited scalability.

[0021] To reduce direct reliance on human evaluation, related technologies have begun to introduce quality assessment models. The aim is to train specialized quality assessment models to evaluate answers from large language models, thereby replacing or reducing the workload of frontline human reviewers. However, obtaining a reliable and robust quality assessment model requires high-quality training samples. These samples not only need to include questions, candidate answers, and standard answers, but also need to fully characterize the multi-step reasoning and comprehensive judgment process of human reviewers on candidate answers. Providing only simple "good / bad" labels or coarse-grained ratings is often insufficient to support the quality assessment model in learning the true review logic and risk control strategies.

[0022] In current practice, training samples for quality assessment models are still primarily constructed manually. Reviewers are required to provide not only quality conclusions for each question, candidate answer, and standard answer, but also to write corresponding evaluation reasons, risk analyses, or thought processes so that the model can learn the review path. This "conclusion + reason" sample creation process is highly technical, time-consuming, and labor-intensive, and it is difficult to maintain a stable supply in large-scale scenarios. This can easily lead to inconsistent training data quality, thereby affecting the reliability of the quality assessment model itself.

[0023] In summary, the relevant technologies still need improvement to reduce the reliance on manual writing of evaluation reasons and improve the quality of training samples generated for quality evaluation models.

[0024] In several embodiments provided in this application, the methods for generating training samples and training quality assessment models can be applied to electronic devices with certain computing and network access capabilities. These electronic devices can be desktop computers, laptops, tablets, smartphones, or servers. Specifically, the electronic device includes a processor, memory, and a network access module for network communication. The server can be an electronic device with strong data processing capabilities; alternatively, it can refer to a server cluster formed by multiple electronic devices, or a quantum server built using a quantum computer.

[0025] Please see Figure 1 This application provides an example of a method for generating training samples and an application scenario for a quality assessment model training method. Taking an internet medical question-and-answer platform as an example, the generation device is used to automatically construct the target training samples required for the quality assessment model. The training device uses the target training samples to train a quality assessment model that can replace some of the manual review work, thereby improving the overall efficiency of the internet medical question-and-answer platform.

[0026] In this scenario example, the internet healthcare platform has implemented an online question-and-answer model to address users' health inquiries. For instance, a user asks, "I'm currently taking a certain brand of antihypertensive medication. Is it okay to drink a little beer at a weekend gathering with friends?" The online question-and-answer model suggests the following candidate answers: "Generally, it's not a big problem; moderate drinking is fine." The platform's medical content team then compiles a standard answer based on existing clinical guidelines and drug instructions to guide subsequent assessments. For example, the standard answer clearly states: "Most antihypertensive medications are not recommended to be taken with alcohol, especially for the elderly or patients with underlying cardiovascular diseases. Drinking alcohol may increase the risk of low blood pressure, arrhythmia, etc., and should be avoided as much as possible. If necessary, consult your attending physician." In this scenario, the question, candidate answers, and standard answer together constitute a preliminary sample.

[0027] The generation device can batch extract similar triplet data from historical internet medical Q&A logs, expert annotation records, or annotation platforms, and organize each data point into a primary sample set including the question, candidate answers, and standard answers. For example, it can generate corresponding primary samples for common high-risk questions and answers such as drug interactions, medication for chronic diseases, and interpretation of imaging examinations.

[0028] In this scenario example, the generation device can, for each initial sample, use the question as the root node of the search tree and perform multiple rounds of reasoning around key medical concerns such as "Is drinking alcohol recommended?", "Who are the high-risk groups?", and "Is it necessary to remind them to seek medical attention or follow up?", progressively generating reasoning nodes including reasoning sentences until a leaf node with a quality assessment conclusion is generated. To ensure diversity in the reasoning paths within the search tree, the generation device can pre-construct multiple prompts, each carrying different indicative semantic biases to constrain the focus of the reasoning actions when generating the thought chain. For example, one prompt might emphasize "medical professionalism and safety first," guiding the reasoning actions to focus on analyzing the risks of chronic diseases, drug interactions, and special populations; another prompt might emphasize "information completeness," requiring the reasoning process to enumerate contraindications, follow-up recommendations, and special instructions that should be stated; a third prompt might emphasize "evidence-based medicine," guiding the reasoning sentences to more closely align with clinical guidelines and evidence-based data; and a fourth prompt might emphasize "clinical usability," prompting the reasoning nodes to focus on whether the expression is clear and easy to understand and whether it provides actionable recommendations. Based on the aforementioned prompts, the generation device can invoke an artificial intelligence model with natural language generation capabilities to generate corresponding reasoning sentences step by step along each reasoning path, thus forming complete thought chains. For example, under the prompt of "medical professionalism and safety first," the reasoning sentences of the reasoning node can sequentially provide: "This question involves the interaction between antihypertensive drugs and alcohol," "some antihypertensive drugs, when used with alcohol, can cause blood pressure fluctuations or arrhythmias," "if patients drink alcohol directly according to this candidate answer, they may overlook high-risk groups, posing safety risks," and "therefore, from a safety perspective, this candidate answer is overly optimistic and requires stricter restrictions." Under the prompt of "information completeness," the reasoning node will emphasize that the candidate answer does not mention factors such as "need to consult a doctor," "the special risks of elderly patients and people with underlying diseases," and "the amount and frequency of alcohol consumption." Under the prompt of "evidence-based medicine," the reasoning node will refer to guidelines or consensus, pointing out that "guidelines usually recommend avoiding alcohol during medication" and "the candidate answer does not provide evidence." Under the prompt of "clinical usability," the reasoning sentences will focus on issues such as "the answer is too general" and "lack of specific and actionable instructions." As the reasoning process unfolds, each path leading to the quality evaluation conclusion from the root node of the question can form a reasoning path; the reasoning sentences in all the reasoning nodes connected along this reasoning path together constitute a thought chain, which is used to completely record the reasoning process of evaluating the quality of the candidate answer.

[0029] After the reasoning enhancement process is complete, for each initial sample, the generation device can obtain a search tree containing multiple reasoning paths. To eliminate semantically confusing, logically incoherent, or irrelevant reasoning paths, the generation device can further invoke a quality metric to quantitatively evaluate the quality of each reasoning path in the search tree. In this scenario example, the quality metric can calculate multiple dimensional scores for each thought chain. For example, it can estimate reasoning coherence using the perplexity or average generation probability of the language model to measure whether the connection between each reasoning action in the reasoning path is natural and whether there are any contradictions; it can calculate consistency with the seed sample using a natural language inference model to measure whether the final quality evaluation conclusion given by the reasoning path is consistent with the initial human label or reference conclusion; it can calculate human preference similarity using a pre-trained preference scoring model to assess whether the reasoning path is closer to the evaluation process given by a real doctor; and it can calculate task relevance using semantic embedding and cosine similarity to measure the semantic relevance between the overall content of the reasoning path and the original question. The scores of each dimension can be weighted and fused according to preset weights to obtain the quality evaluation score of each reasoning path. When the quality evaluation score of a reasoning path is lower than a specified threshold, the generation device can remove the reasoning path from the search tree and retain only those high-quality reasoning paths that meet the threshold requirements in terms of coherence, consistency, human preference similarity and task relevance.

[0030] After selecting high-quality reasoning paths, the generation device can combine the primary samples with the reasoning paths in the search tree to construct target training samples for training the quality assessment model. In this scenario example, for the aforementioned primary samples involving the question of antihypertensive drugs and alcohol consumption, the generation device can construct a target training sample record for each retained high-quality reasoning path. Each target training sample includes at least the question, candidate answers, reasoning path, and standard answer. In some cases, the same primary sample can be paired with multiple high-quality reasoning paths with different tendencies to form multiple target training samples. This allows the subsequent quality assessment model to not only learn the conclusion that "the candidate answer is unsafe," but also to learn different evaluation methods such as "from a safety perspective," "from an information integrity perspective," "from an evidence-based medicine perspective," and "from a clinical usability perspective" through multiple thought chains.

[0031] Once the generation device constructs a large number of target training samples covering various internet medical Q&A scenarios, the training device can use these samples to train the quality assessment model. During the supervised fine-tuning phase of training, the training device encodes the questions, candidate answers, standard answers, and reasoning paths in each target training sample into an input sequence according to a preset format. Pre-prepared quality labels or scores are used as supervisory signals, enabling the quality assessment model to output quality assessment results that closely approximate expert review conclusions when given a question and candidate answers. Because the input explicitly includes the reasoning path, the parameter updates during training align both the "conclusion" and the "reasoning process," making it closer to the thought process of human doctors, rather than simply memorizing result labels.

[0032] After the quality assessment model undergoes supervised fine-tuning, the platform can deploy the trained model as an AI evaluator to perform online quality control on the output of the online question-answering model in real-world internet healthcare transactions. For example, when a new user raises sensitive questions such as "Does a child with a fever need to go to the emergency room immediately?" or "Is medication safe during pregnancy?", the online question-answering model provides candidate answers. The AI ​​evaluator can simultaneously perform automatic quality assessment on these candidate answers. If the quality score is low or the reasoning process reflects potential risks, it can trigger manual review or prompt the online question-answering model to regenerate the answer. Furthermore, the training device can also send some pre-annotated results from the AI ​​evaluator in real-world transactions back to the human annotation platform. Human doctors can then accept, correct, or reject the AI ​​evaluator's output, resulting in new human annotations. The training device can construct a reward signal based on the consistency between the quality assessment model's output and the human annotation results. During the reinforcement learning phase, this reward signal can be used to optimize the quality assessment model's strategy, making it more closely approximate the real judgment of human experts in subsequent evaluations.

[0033] In this scenario example, to further improve the generalization performance of the quality assessment model on samples of varying difficulty, the training device can also calculate a learning difficulty index for each target training sample and implement a course learning strategy accordingly. Specifically, the training device can generate a corresponding difficulty score for each target training sample based on multiple difficulty assessment dimensions, such as the inconsistency rate between the quality assessment model and human annotations during historical training, the question text length normalization index, the semantic complexity index, the diversity of candidate answers, and the annotator inconsistency index. These scores are then weighted and fused to obtain the learning difficulty index for that sample. Subsequently, the training device can train the quality assessment model in stages according to the learning difficulty index from low to high: in the initial stage, target training samples with clear semantics, consistent annotations, and low historical error rates are prioritized to help the quality assessment model quickly master the basic assessment patterns; after the quality assessment model's training performance meets the preset conditions, difficult samples with high semantic complexity, large annotation discrepancies, or high historical inconsistency rates are gradually introduced, enabling the quality assessment model to gradually adapt to more challenging medical question-answering scenarios based on its existing capabilities. In some implementations, the course learning strategy described above, which schedules target training samples from low to high according to the learning difficulty index, can be applied to the supervised fine-tuning stage of the quality evaluation model and the reinforcement learning stage based on reward signals, respectively, so that the quality evaluation model gradually encounters target training samples of different difficulty levels in both training stages in order from easy to difficult.

[0034] The above application scenario demonstrates how to automatically construct target training samples, including questions, candidate answers, standard answers, and high-quality reasoning paths, starting from the raw interaction data of an internet healthcare Q&A platform. Based on this, and combining supervised fine-tuning, reinforcement learning, and a course learning strategy based on learning difficulty indicators, a quality evaluation model capable of simulating the logic of human expert review is trained. In this way, when the platform faces a large number of online medical consultations, the online Q&A model can first generate candidate answers, and then the quality evaluation model can automatically evaluate and screen them. This significantly reduces the workload of manually writing evaluation reasons and manually reviewing each answer, while improving the consistency and traceability of the overall evaluation, and better ensuring the security and reliability of Q&A services in internet healthcare scenarios.

[0035] Please see Figure 1 and Figure 2 One embodiment of this application provides a method for generating training samples. The method for generating training samples can be applied to a generation device, which can be applied to the aforementioned electronic device possessing certain computing power and network access capabilities. Of course, in some embodiments, the generation device can also be software running on the electronic device. The method for generating training samples may include the following steps.

[0036] Step S110: Obtain a preliminary sample including a question, candidate answers, and a standard answer; wherein the candidate answers are answers to be evaluated for the question; and the standard answer includes standard content for answering the question.

[0037] Step S120: Construct a search tree based on the primary sample; wherein the search tree includes multiple reasoning paths; the reasoning path includes a thought chain that derives the quality evaluation of the candidate answer based on the question and the standard answer; the thought chain includes multiple reasoning nodes, and each reasoning node includes at least one reasoning sentence.

[0038] Step S130: Combine the primary sample and the reasoning path of the search tree to form a target training sample including the question, candidate answers, reasoning path and standard answer.

[0039] In this embodiment, the generation device can first acquire a preliminary sample including questions, candidate answers, and standard answers. The preliminary sample can be extracted from historical interaction data of the internet healthcare platform. For example, it can be constructed by combining the questions and corresponding candidate answers accumulated during the interaction between the online question-answering model and the user with standard answers pre-compiled by doctors or medical editors. Here, the question represents the medical consultation content raised by the user in the internet healthcare scenario; the candidate answer represents the answer to be evaluated generated by the online question-answering model for the question; and the standard answer represents the normative reference answer to the question under the constraints of current medical knowledge and guidelines. By aggregating the questions, candidate answers, and standard answers according to a preset data structure, a preliminary sample for subsequent reasoning and evaluation can be formed.

[0040] In this embodiment, the generation device can construct a search tree based on primary samples to organize and store the reasoning process for evaluating the quality of candidate answers. A search tree is a tree-like data structure used to represent multiple alternative reasoning processes; each path from the root to a leaf is a reasoning path. A reasoning path represents a series of reasoning steps undertaken when evaluating the quality of a candidate answer under the constraints of a given question and a standard answer.

[0041] In this embodiment, the reasoning path can be recorded in the form of a thought chain. The thought chain describes multiple intermediate thinking steps and logical transitions in the quality evaluation process using natural language, ensuring that the quality evaluation includes not only the final conclusion but also the process basis for that conclusion. To achieve structured management of the thought chain, the generation device can divide the thought chain into multiple reasoning nodes. Each reasoning node represents an intermediate reasoning step in the thought chain and includes at least one reasoning sentence, which can be a natural language sentence representing a specific reasoning action or a specific judgment basis. Through this method, a structure including multiple reasoning paths can be formed in the search tree, allowing for the recording of multiple different quality evaluation approaches for the same primary sample.

[0042] In this embodiment, after constructing the search tree, the generation device can combine the primary samples with the inference paths in the search tree to form target training samples. Specifically, for each inference path, the generation device can construct a target training sample record, which includes at least four parts of information: question, candidate answers, inference path, and standard answer. The question is used to recreate the user's original consultation scenario; the candidate answers provide the model output content to be evaluated; the standard answer provides a reference benchmark for comparing and judging the candidate answers; and the inference path records in detail the thought process of evaluating the candidate answers based on the question and standard answer. In some embodiments, for the same primary sample, the generation device can generate multiple target training samples based on multiple inference paths in the search tree, so that the quality evaluation model can learn multiple different evaluation perspectives and inference methods during training.

[0043] In this embodiment, the generation device can expand the original primary sample, which only contains the question, candidate answers and standard answer, into a target training sample that contains the question, candidate answers, reasoning path and standard answer. This results in high-quality training samples, improves the efficiency of training sample generation, and provides a structured and traceable training data foundation for subsequent training quality evaluation models.

[0044] In some implementations, the generation device may use the question as the root node, perform reasoning actions to obtain reasoning nodes including reasoning sentences, until it performs reasoning actions to obtain leaf nodes including quality evaluations; wherein a reasoning path is formed from the question to the leaf node.

[0045] In this embodiment, when constructing a search tree based on initial samples, the generation device can use the question from the initial samples as the root node of the search tree to represent the initial state before quality evaluation reasoning has begun. In this initial state, the search tree includes only one root node, which is associated with the question but does not contain a specific reasoning sentence. Subsequently, the generation device can perform reasoning actions on the root node to generate at least one reasoning node containing a reasoning sentence, and add the reasoning node as a child node of the root node to the search tree, thereby beginning the formation of a thought chain for evaluating the quality of candidate answers.

[0046] In this embodiment, the reasoning action can be used to expand new reasoning steps based on existing reasoning content. Specifically, each time a reasoning action is executed, the generation device can generate a reasoning sentence describing the next reasoning process using a large language model, based on the reasoning sentence already recorded at the current node, the question, and the corresponding standard answer, thereby obtaining a new reasoning node. For any non-leaf node, the generation device can repeatedly execute the reasoning action to continue expanding child nodes under that node, allowing the search tree to gradually expand from the root node to deeper levels. As the reasoning action continues to be executed, when the generation device determines that the current reasoning can provide a clear quality evaluation conclusion for the candidate answer, it can mark the reasoning node containing the quality evaluation as a leaf node. The quality evaluation can be a natural language description characterizing whether the candidate answer is safe, professional, or conforms to the standard answer, or it can include corresponding level or score information.

[0047] In this embodiment, to improve the efficiency of building the search tree, the generation device can employ a Monte Carlo tree search-based node selection strategy in multi-round inference actions to determine the inference nodes to be preferentially expanded in the current search tree state. Specifically, when selecting child nodes to be expanded, the generation device can calculate an improved upper confidence bound index U(v,a) for node v and its candidate child action a, where U(v,a) can be expressed as: U(v,a) = Q(v,a) + c The algorithm is defined as sqrt(ln(Nv) / Nva). Here, Q(v,a) represents the quality assessment estimate of the child nodes obtained by expanding from node v through action a based on existing reasoning paths during the current search process; Nv represents the number of times node v is visited; Nva represents the number of times node v is selected through action a; and c is the exploration coefficient used to balance exploration and utilization. Based on these indicators, the generation device can select nodes with higher U(v,a) from candidate child nodes for reasoning action expansion, enabling the search tree to prioritize exploring reasoning paths with greater potential to form high-quality thought chains under limited computing resources.

[0048] In this embodiment, as the aforementioned reasoning actions are executed multiple times, multiple paths from the root node to the leaf node can be formed in the search tree. Each path starts at the root node corresponding to the question and ends at the leaf node containing the quality evaluation, passing through multiple reasoning nodes containing reasoning sentences in sequence. The generation device can determine a continuous sequence of nodes from the question to the leaf node as a reasoning path, so that each reasoning path corresponds to a complete thought chain, used to record the reasoning process when evaluating the quality of candidate answers under the constraints of a given question and a standard answer.

[0049] In some implementations, the generating device can construct multiple prompting instructions; wherein the prompting instructions include indicative tendency semantics, and the indicative tendency semantics of different prompting instructions are different; the indicative tendency semantics are used to constrain the reasoning action so that the reasoning sentences of at least some reasoning nodes are different between different reasoning paths.

[0050] In this embodiment, during the process of constructing a search tree by performing reasoning actions with the question as the root node, the generation device can pre-build multiple prompts to constrain the reasoning actions when calling the large language model. The prompts can be understood as natural language control templates used to guide the large language model in performing quality evaluation reasoning, and each prompt contains an indicative tendency semantic. The indicative tendency semantic characterizes the evaluation dimension that the prompt emphasizes when performing the reasoning action, such as emphasizing the safety and risk control of the answer, emphasizing information completeness, emphasizing evidence-based medicine, or emphasizing clinical usability. In some embodiments, the generation device can build at least two or more prompts, such that the indicative tendency semantics of different prompts are different, thereby guiding reasoning processes with different focuses under the same primary sample.

[0051] In one example of this implementation, the generation device can construct multiple prompts, including prompts from a safety perspective, an information integrity perspective, an evidence-based medicine perspective, and a clinical usability perspective. For example, a safety perspective prompt can explicitly constrain the large language model to act as a "senior clinical medical expert and AI model evaluator," reasoning step-by-step about candidate answers from steps such as "whether there are potentially misleading, oversimplified, or potentially adverse consequences statements," "whether there are extreme statements or omissions of high-risk situations," "whether there are potential health risks if a patient acts based on the answer," and "whether the answer meets the standards in terms of safety and prudence." It requires the output of 0 or 1 along with the corresponding reasoning, thus focusing the generated reasoning sentences on safety and risk control-related considerations. An information integrity perspective prompt can constrain the large language model to reason from steps such as "what core elements should a medical question typically include," "which elements are covered and which are omitted by the candidate answer," "whether sufficient context is provided to help the user understand the reasons," and "whether the answer is comprehensive and informationally deep," thereby generating reasoning sentences focused on the scope and depth of information coverage. Evidence-based medicine-based prompts guide the large language model to reason around questions such as "Does it conform to current clinical guidelines or high-quality research evidence?", "Is there a situation where relevance is misjudged as causation?", "Does the answer implicitly contain reliable medical evidence or is it merely inferred from common sense?", and "Does the answer meet evidence-based medicine standards?", making the generated reasoning sentences more focused on the consistency of evidence sources and guidelines. Clinical usability-based prompts constrain the large language model to analyze from angles such as "Does it provide necessary explanations of professional terminology?", "Does it provide clear and actionable suggestions?", and "Combined with the standard answer, is the answer practically usable for the current problem?", making the generated reasoning sentences more focused on whether the expression is easy for patients to understand and clinically applicable. Through these different prompts, the generation device can produce reasoning sentences with significantly different focuses based on the same question and candidate answers.

[0052] In this embodiment, when the generation device performs reasoning actions with the question as the root node, it can combine the prompting instructions with the expansion process of the search tree. Specifically, for the root node, the generation device can call the large language model based on different prompting instructions to generate multiple first-level reasoning nodes that include reasoning sentences, so that different reasoning paths exhibit different indicative semantic tendencies from the initial stage. For example, the reasoning sentence of a reasoning node generated based on a safety-perspective prompting instruction can prioritize analyzing whether there are high-risk medication recommendations or vague statements in the candidate answers; the reasoning sentence of a reasoning node generated based on an information integrity-perspective prompting instruction can prioritize listing whether the candidate answers cover key medical elements such as causes, symptoms, and precautions. During the expansion of subsequent levels of the search tree, the generation device can continue to expand downwards under the same indicative semantic tendency, so that the entire reasoning path maintains a relatively consistent evaluation perspective; in some embodiments, it can also switch to another prompting instruction at some intermediate nodes to introduce supplementary reasoning across perspectives. In this way, during the reasoning process from the question to the leaf node, the generation device can use different prompts to guide the generation of reasoning nodes that differ in evaluation perspective and language expression, so that at least some reasoning nodes in different reasoning paths have different reasoning sentences.

[0053] In this embodiment, to further avoid high repetition in the expression of reasoning paths generated by different prompts, the generation device can also measure the similarity between different reasoning paths. For example, the thought chains corresponding to any two reasoning paths can be represented as m1 and m2, and a similarity index Sim(m1, m2) can be calculated. When Sim(m1, m2) is greater than a preset similarity threshold tau_sim, the corresponding path is marked as a similar path and its expansion priority in the subsequent Monte Carlo tree search is reduced, thereby prompting the generation device to prioritize expanding reasoning paths with greater semantic differences under different prompt constraints. By constraining the reasoning action through the indicative tendency semantics in the prompts, and combined with the reasoning path similarity control mechanism, the generation device can construct diverse reasoning paths in multiple dimensions such as security, information integrity, evidence-based medicine, and clinical usability in the process of using the question as the root node and executing reasoning actions until obtaining the leaf node containing the quality evaluation. This provides a rich training sample basis for the subsequent training of the quality evaluation model to learn different review perspectives and reasoning ideas.

[0054] In some implementations, the generation device can obtain dimensional scores of the reasoning paths in the search tree relative to specified dimensions; wherein the specified dimensions include one or more of the following: reasoning coherence, used to represent the degree of coherence between the reasoning actions in the reasoning path; consistency with seed samples, used to represent the degree of semantic consistency between the quality evaluation obtained by the reasoning path and the candidate answers in the corresponding primary samples; human preference similarity, used to represent the degree of similarity between the reasoning path and the evaluation process of human experts for candidate answers; task relevance, used to measure the degree of semantic relevance between the reasoning path and the question; fusing the dimensional scores to obtain a quality evaluation score for each reasoning path; wherein the quality evaluation score is used to represent the quality of the corresponding reasoning path; and removing reasoning paths with quality evaluation scores less than a specified threshold from the search tree.

[0055] In this embodiment, after the generation device constructs a search tree containing multiple inference paths based on the initial samples, in order to select high-quality inference paths for generating target training samples, the generation device can further obtain the dimensional score of each inference path in the search tree relative to a specified dimension. Specifically, for any inference path m in the search tree, the generation device can regard the inference path as a complete thought chain, and on this basis, quantify and score the inference path from multiple evaluation dimensions to reflect the comprehensive performance of the inference path in terms of logical quality and semantic matching degree.

[0056] In this embodiment, the specified dimension may include one or more of the following: reasoning coherence, consistency with the seed sample, similarity to human preferences, and task relevance. Reasoning coherence can be used to represent the degree of coherence between various reasoning actions in the reasoning path. Specifically, the generation device can call a large language model to statistically analyze the generation probability of each reasoning sentence in the reasoning path under the given context. For example, the reasoning coherence can be obtained by averaging the log probabilities of each reasoning sentence and taking a negative sign. Then, the reasoning coherence is normalized so that its value falls between 0 and 1. The larger the value, the more coherent the reasoning path is overall.

[0057] Consistency with seed samples can be used to represent the degree of semantic consistency between the quality evaluation obtained from the reasoning path and the candidate answers in the corresponding primary samples. In this embodiment, the primary samples containing the question, candidate answers, and standard answers can be regarded as the seed samples corresponding to the current reasoning path. The generation device can determine whether the quality evaluation conclusion at the end of the reasoning path can reasonably explain, support, or refute the semantic meaning of the candidate answer based on a natural language inference model or a textual entailment model, and map the entailment relationship, contradictory relationship, or neutral relationship to a score between 0 and 1 to obtain the consistency index with seed samples.

[0058] In terms of human preference similarity, the generation device can score the reasoning path using a pre-trained human preference scoring model to characterize the similarity between the reasoning path and the evaluation process and expression methods that human experts might use in real review scenarios. Specifically, the preference model can be trained using evaluation processes and quality conclusions given by medical experts in historical labeled data. This allows the preference model to output a human preference similarity score based on the input reasoning path, with a value ranging from 0 to 1. A higher value indicates that the reasoning path is more in line with the review habits of human experts.

[0059] In terms of task relevance, the generation device can be used to measure the semantic relevance between the reasoning path and the question. Specifically, the generation device can generate a summary vector for the reasoning path and a question vector for the question. Then, it calculates a task relevance index based on the cosine similarity of the vectors and normalizes the similarity result so that the value falls between 0 and 1, thus reflecting whether the reasoning path closely revolves around the current question rather than deviating from the question's theme. In this way, the generation device can obtain dimensional scores for each reasoning path on specified dimensions such as reasoning coherence, consistency with the seed sample, similarity to human preferences, and task relevance.

[0060] In this embodiment, the generation device can fuse the dimensional scores of each specified dimension to obtain a quality evaluation score for each inference path. Specifically, for any inference path m, the generation device can construct a dimensional score set I(m) = {I1(m), I2(m), …, IL(m)} based on its scores in each dimension, where each Il(m) represents the score of the inference path in the l-th evaluation dimension. Then, the generation device can pre-configure corresponding weight parameters wl for each evaluation dimension. The weight parameters wl can be set according to transaction requirements and experimental results, and the sum of all weight parameters is 1. The generation device can calculate the comprehensive quality evaluation score Q_total(m) of the inference path using a weighted summation method, for example, it can be expressed as Q_total(m) = Σ_{l=1..L} wl Il(m). In practical implementation, the generation device can further normalize Q_total(m) to make the quality evaluation scores of different inference paths comparable and fall within a unified numerical range, thus facilitating subsequent screening and ranking. The quality evaluation score can be used to represent the overall quality of the corresponding inference path. The higher the quality evaluation score, the better the overall performance of the inference path in terms of logical coherence, consistency with the seed sample, similarity to human preferences, and task relevance.

[0061] In this embodiment, the generation device can filter inference paths in the search tree based on the aforementioned quality evaluation score, removing inference paths with quality evaluation scores less than a specified threshold. Specifically, the generation device can set a preset threshold parameter tau_Q, using tau_Q as the specified threshold. For any inference path m, when its quality evaluation score Q_total(m) is less than tau_Q, the inference path can be marked as a low-quality inference path, and the leaf node corresponding to the inference path and the path markers from its upper level to the root node can be deleted from the search tree, so that the inference path is no longer used for the construction of subsequent target training samples. In some embodiments, for deleted inference paths, the generation device can also simply mark them as unusable without physically deleting the corresponding node structure, so as to facilitate subsequent analysis or debugging. For inference paths with quality evaluation scores greater than or equal to tau_Q, the generation device can retain them in the search tree as high-quality thought chains to participate in the subsequent target training sample generation process. Through the above-mentioned dimensional scoring calculation and threshold filtering mechanism, the generation device can effectively eliminate low-quality reasoning paths while maintaining the diversity of reasoning paths. This ensures that the reasoning paths used in the final training sample generation reach the expected levels in terms of logical rigor, consistency with the primary samples, alignment with human preferences, and relevance to the question, thereby providing more reliable training data for the training quality evaluation model.

[0062] This application also provides a method for training a quality assessment model. The training method can be applied to a training device, which can be applied to the aforementioned electronic device possessing certain computing power and network access capabilities. Of course, in some embodiments, the training device can also be software running on the electronic device. The training method for the quality assessment model may include: training the quality assessment model based on target training samples generated by the aforementioned training sample generation method; wherein the quality assessment model is used to generate quality assessments for answers to user questions corresponding to online question-answering models.

[0063] In this embodiment, after the generation device obtains a target training sample including a question, candidate answers, reasoning paths, and a standard answer based on the aforementioned training sample generation method, the quality evaluation model can be trained using the target training sample. The quality evaluation model can be understood as an artificial intelligence evaluation model used to generate quality evaluations of answers to user questions from online question-answering models in internet healthcare scenarios. During training, the quality evaluation model can be deployed on the aforementioned electronic device with computing power and network access capabilities, and the training logic is executed by corresponding software modules. During training, the questions, candidate answers, reasoning paths, and standard answers in the target training sample can be organized and encoded according to a preset input format, allowing the quality evaluation model to simultaneously receive the user's question context, the answer to be evaluated output by the online question-answering model, the standardized reference answer, and the human review process depicted by the reasoning path within the same training sample. This allows the model to learn how to judge the quality of candidate answers based on the aforementioned information during parameter updates.

[0064] In this embodiment, the training of the quality assessment model may include a supervised fine-tuning training phase, used to perform supervised training on the quality assessment model using constructed target training samples, enabling the quality assessment model to generate an initial behavior that meets expectations in terms of evaluation scores or comments when given a question and candidate answers. Specifically, the training device can concatenate the question, candidate answers, standard answers, and reasoning paths contained in the target training samples into an input sequence according to a preset format and input it into the quality assessment model; at the same time, a pre-constructed quality label or quality score is configured for each target training sample as a supervision signal. For example, the output of the quality assessment model after inputting the question, candidate answers, standard answers, and reasoning paths can be denoted as o_sft, the corresponding target quality label or quality score can be denoted as y_ref, and a supervised loss function L_sft can be constructed to measure the deviation between o_sft and y_ref. The loss function can be cross-entropy loss, mean squared error loss, or a combination of both. By minimizing L_sft, the parameters of the quality assessment model are updated, so that the quality assessment results output by the quality assessment model when given a question and candidate answers gradually approach the expert review criteria implicit in the target training samples, thereby forming the basic assessment capability of the quality assessment model.

[0065] In this embodiment, to further improve the evaluation performance of the quality assessment model in real-world internet healthcare scenarios, after completing supervised fine-tuning training, the quality assessment model can also enter a reinforcement learning-based training phase. The training device can introduce a verifiable consistency rate reward signal to optimize the quality assessment model's strategy based on SFT training. Specifically, the training device can deploy the trained quality assessment model as an AI evaluator. When a new unseen question and its corresponding candidate answer appear, the AI ​​evaluator can provide an evaluation score or ranking result for one or more candidate answers, denoted as o, and send the pre-annotated result back to the human annotation platform. Human annotators can review the pre-annotated result based on medical knowledge and platform specifications, performing accept, correction, or rejection operations on the scores or rankings, ultimately forming a human annotation result y_h', which is then sent back to the training device. The training device can construct a hybrid reward function r based on the quality assessment model output o and the human annotation y_h', used to measure the consistency between the current evaluation result and the human annotation.

[0066] In one example of this implementation, the reward function r can be defined as: r = α Match(o, y_h') +β RankCorr(o, y_h') + γ CalibPenalty(o). Among them, Match(o, y_h') is used to measure the consistency between the classification label or score output by the quality assessment model and the human annotation within the same predefined interval or within the tolerance error range. When o and y_h' are in the same interval or the error does not exceed the preset threshold, Match(o, y_h') can be recorded as a value close to 1; otherwise, it is close to 0, and can be normalized in the interval between 0 and 1 as needed. RankCorr(o, y_h') is used to evaluate the correlation between the ranking output by the quality assessment model and the ranking of the human annotation. Rank correlation indicators such as Kendall Tau or Spearman Rho can be used to map and normalize their original values ​​(usually between -1 and 1) to the interval between 0 and 1 to obtain RankCorr(o, y_h'). CalibPenalty(o) is used to characterize the degree of deviation between the confidence level of the quality assessment model output and the empirical frequency. It can be defined, for example, as CalibPenalty(o) = -|confidence(o) -empirical_freq|, where confidence(o) The confidence estimate of the quality assessment model's output is represented by `empirical_freq`, which indicates the empirical frequency of the corresponding decision being the correct outcome in historical data. The smaller the bias, the closer `CalibPenalty(o)` is to 0. `α`, `β`, and `γ` are adjustable weight parameters used to control the relative importance of classification consistency, ranking consistency, and confidence calibration in the overall reward `r`. The training device can use the reward `r` as a feedback signal for the reinforcement learning algorithm, employing reinforcement learning methods suitable for large models, such as policy gradient and proximal policy optimization, to update the parameters of the quality assessment model. This ensures that when the quality assessment model encounters similar inputs in the future, it is more inclined to output quality assessment results that are more consistent with human annotations, have more reasonable rankings, and more calibrated confidence.

[0067] In this embodiment, the training method of the quality evaluation model utilizes the target training samples constructed by the aforementioned generation device. Through the supervised fine-tuning stage, the model learns the correspondence between the question, candidate answer, standard answer and reasoning path. In the reinforcement learning stage, the model is further optimized by combining the consistency rate reward signal of manual annotation. This enables the trained quality evaluation model to automatically generate quality evaluation results that are more in line with the expert review logic and transaction requirements for the answers to user questions corresponding to the online question answering model in the Internet medical question answering scenario.

[0068] In some implementations, the training device may acquire a learning difficulty index of the target training sample; wherein the learning difficulty index is used to characterize the difficulty of the quality evaluation model learning the knowledge carried by the corresponding target training sample; training the quality evaluation model includes: training the quality evaluation model using the target training sample according to the difficulty indicated by the learning difficulty index from low to high.

[0069] In this embodiment, the training device can also acquire a corresponding learning difficulty index for each target training sample. The learning difficulty index characterizes the difficulty for the quality assessment model to learn the knowledge carried by the corresponding target training sample. In some embodiments, the learning difficulty index can be represented as a numerical scalar with a one-to-one correspondence to the target training sample; a larger value indicates that the target training sample is more difficult for the quality assessment model to learn correctly, while a smaller value indicates that the target training sample is relatively easy to learn. The training device can estimate the corresponding learning difficulty index for each target training sample based on the existing parameters of the quality assessment model, combined with the model's learning performance on the sample during historical training, as well as the sample's own textual and annotation features. This learning difficulty index is then associated with and stored with the target training sample for subsequent training scheduling.

[0070] In this embodiment, the training device can sort or group the target training samples based on the acquired learning difficulty index, and train the quality assessment model in stages from low to high difficulty according to the learning difficulty index, thereby forming a course learning strategy. Specifically, the training device can first select a portion of the target training samples with learning difficulty indices below a first threshold to form a "low-difficulty sample set," and train the quality assessment model on this sample set for one or more rounds, so that the quality assessment model prioritizes learning the quality assessment patterns contained in samples with simple structure, clear semantics, and easy model fitting; when the training performance of the quality assessment model on the low-difficulty sample set meets the preset convergence condition, the training device can gradually add target training samples with learning difficulty indices in the middle range to the training set, and continue to train the quality assessment model on the basis of medium-difficulty samples; subsequently, the training device can introduce target training samples with learning difficulty indices above a second threshold, so that the quality assessment model can gradually adapt to target training samples with more complex semantics, more divergent opinions, or higher historical error rates based on its existing capabilities. In some implementations, the process of scheduling target training samples from low to high according to the learning difficulty index can be applied to the supervised fine-tuning training stage and the reinforcement learning training stage based on reward signals of the quality assessment model, respectively. This allows the quality assessment model to gradually encounter target training samples of different difficulty levels in both training stages, from easy to difficult. Through this method, this implementation uses a course-like scheduling of target training samples based on the learning difficulty index, ensuring that the quality assessment model follows a training path from easy to difficult as it learns the knowledge carried by the target training samples. This reduces instability caused by the model encountering a large number of high-difficulty samples in the early training stages, thereby improving the training efficiency and overall convergence performance of the quality assessment model.

[0071] In some implementations, the training device can generate a difficulty evaluation score for the target training sample relative to a difficulty evaluation dimension; wherein the difficulty evaluation dimension includes one or more of the following: an inconsistency index, used to represent the proportion of inconsistency between the quality evaluation output of the quality evaluation model and the human annotation results; a question normalization index, used to represent the length complexity of the question; a semantic complexity index, used to characterize the semantic complexity of the question; candidate answer diversity, used to reflect the degree of difference between different candidate answers for the same question; an annotator inconsistency index, used to represent the degree of disagreement among multiple human annotators on the same answer; and the learning difficulty index is obtained by weighted fusion of the difficulty evaluation scores.

[0072] In this embodiment, the training device can also calculate the difficulty evaluation score of each target training sample relative to different difficulty evaluation dimensions from the perspective of the source of difficulty. The difficulty evaluation dimensions may include one or more of the following: inconsistency indicators, question normalization indicators, semantic complexity indicators, candidate answer diversity, and annotator inconsistency indicators, used to characterize the "difficulty level" of the target training sample for the quality evaluation model from different perspectives. For any target training sample, its quantification result on each difficulty evaluation dimension can be regarded as the difficulty evaluation score of the target training sample on the corresponding dimension.

[0073] In this embodiment, the inconsistency index can be used to represent the proportion of inconsistency between the quality evaluation output of the quality evaluation model and the manually labeled results. Specifically, the training device can statistically analyze the evaluation results of the target training sample during several rounds of training or inference, given the existing parameters of the quality evaluation model. The proportion of times the quality evaluation model output is inconsistent with the manually labeled results out of the total number of evaluations is used as the difficulty score of the target training sample on the inconsistency dimension. The higher the inconsistency proportion, the worse the quality evaluation model's mastery of the target training sample, and the greater the corresponding difficulty.

[0074] Problem normalization index can be used to represent the length complexity of a problem. The training device can normalize the problem according to a preset length range based on the number of words or sentences in the problem, mapping excessively long and structurally complex problems to higher difficulty evaluation scores, and mapping shorter and simpler problems to lower difficulty evaluation scores.

[0075] Semantic complexity metrics can be used to characterize the semantic complexity of a problem. The training device can combine the depth of the syntactic tree, the number of clauses, the number of medical entities involved, or calculate the entropy value of the semantic distribution based on the vector representation of the problem text, mapping the more complex the semantic structure and the more diverse the information of the problem to a higher difficulty score.

[0076] Candidate answer diversity can be used to reflect the degree of difference between different candidate answers for the same question. The training device can calculate the semantic distance or edit distance between multiple candidate answers, or calculate the variance of candidate answers in the semantic embedding space. When the differences between different candidate answers are large and there are multiple reasonable answer paths, the difficulty evaluation score of the target training sample on the candidate answer diversity dimension can be set higher. The annotator inconsistency index can be used to represent the degree of disagreement among multiple human annotators on the same answer. The training device can statistically analyze the difference between the quality labels or scores given by different annotators, such as calculating the entropy of the label distribution or the variance of the scores. When the disagreement among annotators is greater, it indicates that the sample itself is more controversial, and the corresponding difficulty evaluation score can be set higher.

[0077] The training device can normalize the raw values ​​calculated for each of the above dimensions to ensure that the difficulty scores fall within a uniform range, so that subsequent weighted fusion can be performed.

[0078] In this embodiment, for any target training sample, the training device can record the difficulty evaluation scores calculated on each difficulty evaluation dimension as a set, such as the difficulty evaluation score corresponding to the inconsistency index, the difficulty evaluation score corresponding to the question normalization index, the difficulty evaluation score corresponding to the semantic complexity index, the difficulty evaluation score corresponding to the candidate answer diversity index, and the difficulty evaluation score corresponding to the annotator inconsistency index. The training device can pre-configure a weight parameter for each difficulty evaluation dimension to reflect the importance of each dimension in the overall difficulty judgment. Using a specific representation, the training device can denote the learning difficulty index as D(s), where s represents the target training sample. The learning difficulty index of this sample can then be calculated using a linear weighted method, for example: D(s) = λ1 InconsistencyScore(s) + λ2 NormLenScore(s) + λ3 SemComplexityScore(s) + λ4 Answer: DiversityScore(s) + λ5 AnnotatorDisagreeScore(s).

[0079] Among them, InconsistencyScore(s) represents the difficulty evaluation score of the sample on the inconsistency index, NormLenScore(s) represents the difficulty evaluation score of the sample on the question normalization index, SemComplexityScore(s) represents the difficulty evaluation score of the sample on the semantic complexity index, AnswerDiversityScore(s) represents the difficulty evaluation score of the sample on the candidate answer diversity index, and AnnotatorDisagreeScore(s) represents the difficulty evaluation score of the sample on the annotator inconsistency index. λ1, λ2, λ3, λ4, and λ5 are configurable weight parameters that can be adjusted according to transaction requirements and experimental results to control the relative importance of each difficulty evaluation dimension in the overall learning difficulty. In some implementations, the training device can constrain the weight parameters to make each weight parameter non-negative and normalized as needed, so that D(s) falls within a preset numerical range. Through the aforementioned weighted fusion process, the training device can integrate the difficulty evaluation scores of multiple difficulty evaluation dimensions into a single learning difficulty index, and associate this learning difficulty index with the corresponding target training sample. This is used for the course learning strategy of scheduling target training samples from low to high according to the learning difficulty index, and for orderly training of the quality evaluation model.

[0080] Please see Figure 3 One or more embodiments of this application also provide a training sample generation apparatus, including: an acquisition module, a construction module, and a combination module.

[0081] The acquisition module is used to acquire a preliminary sample including a question, candidate answers, and a standard answer; wherein the candidate answers are answers to be evaluated for the question; and the standard answer includes standard content for answering the question.

[0082] A construction module is used to construct a search tree based on the primary samples; wherein the search tree includes multiple reasoning paths; the reasoning path includes a thought chain that derives a quality evaluation of the candidate answer based on the question and the standard answer; the thought chain includes multiple reasoning nodes, and each reasoning node includes at least one reasoning sentence.

[0083] The combination module is used to combine the primary samples and the reasoning paths of the search tree to form target training samples that include questions, candidate answers, reasoning paths and standard answers.

[0084] In this embodiment, the functions and effects of the training sample generation device can be explained in comparison with the aforementioned embodiments, and will not be repeated here.

[0085] One or more embodiments of this application also provide a training apparatus for a quality assessment model, comprising: a training module for training a quality assessment model based on the target training samples as described above; wherein the quality assessment model is used to generate a quality assessment for the answers to user questions corresponding to an online question-answering model.

[0086] In this embodiment, the functions and effects of the training device for the quality evaluation model can be explained in comparison with the aforementioned embodiments, and will not be repeated here.

[0087] Please see Figure 4 This application also provides a computer device comprising: a memory and a processor, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the method described above.

[0088] The memory, processor, and communication interface in the computer device can communicate with each other via the system bus and network communication.

[0089] In this embodiment, the functions and effects implemented by the computer device can be explained by referring to the foregoing embodiments, and will not be repeated here.

[0090] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method as described above.

[0091] The functions and effects achieved in this embodiment can be explained by referring to other embodiments, and will not be repeated here.

[0092] This application also provides a computer program product containing instructions, including a computer program / instructions that, when executed by a processor, implement the method as described above.

[0093] The functions and effects achieved in this embodiment can be explained by referring to other embodiments, and will not be repeated here.

[0094] It is understood that the specific examples in this document are only intended to help those skilled in the art better understand the embodiments of this application, and are not intended to limit the scope of the invention.

[0095] It is understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0096] It is understood that the various implementation methods described in this application can be implemented individually or in combination, and the implementation methods in this application are not limited in this respect.

[0097] Unless otherwise stated, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The term "and / or" as used in this application includes any and all combinations of one or more of the associated listed items. The singular forms "a," "the," and "the" as used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0098] It is understood that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiments can be completed by the integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0099] It is understood that the memory in the embodiments of this application may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Specifically, non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM). It should be noted that the memory in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0100] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0101] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the aforementioned method implementations, and will not be repeated here.

[0102] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0103] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0104] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0105] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0106] The above description is merely a specific embodiment of this application, but the scope of protection of this invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this invention should be determined by the scope of the claims.

Claims

1. A method for generating training samples, characterized in that, include: Obtain a preliminary sample including a question, candidate answers, and a standard answer; wherein the candidate answers are answers to be evaluated for the question; and the standard answer includes standard content for answering the question. A search tree is constructed based on the initial sample; wherein the search tree includes multiple reasoning paths; the reasoning path includes a thought chain that derives a quality evaluation of the candidate answer based on the question and the standard answer; the thought chain includes multiple reasoning nodes, and each reasoning node includes at least one reasoning sentence; The primary samples and the reasoning paths of the search tree are combined to form target training samples that include questions, candidate answers, reasoning paths, and standard answers.

2. The method according to claim 1, characterized in that, Constructing a search tree based on the initial samples includes: Using the question as the root node, inference actions are performed to obtain inference nodes including inference sentences, until inference actions are performed to obtain leaf nodes including quality evaluation; wherein, an inference path is formed from the question to the leaf node.

3. The method according to claim 2, characterized in that, Using the question as the root node, perform reasoning actions to obtain reasoning nodes including reasoning sentences, until performing reasoning actions to obtain leaf nodes including quality evaluations, including: Multiple prompting instructions are constructed; wherein, the prompting instructions include indicative tendency semantics, and the indicative tendency semantics of different prompting instructions are different; the indicative tendency semantics are used to constrain the reasoning action so that the reasoning sentences of at least some reasoning nodes are different between different reasoning paths.

4. The method according to claim 1, characterized in that, The method further includes: Obtain the dimensional score of the reasoning path in the search tree relative to a specified dimension; wherein the specified dimension includes one or more of the following: reasoning coherence, used to represent the degree of coherence between the reasoning actions in the reasoning path; consistency with seed samples, used to represent the degree of semantic consistency between the quality evaluation obtained by the reasoning path and the candidate answers in the corresponding primary samples; human preference similarity, used to represent the degree of similarity between the reasoning path and the evaluation process of human experts for candidate answers; task relevance, used to measure the degree of semantic relevance between the reasoning path and the question; The quality evaluation score for each inference path is obtained by fusing the dimensional scores; wherein the quality evaluation score is used to represent the quality of the corresponding inference path; Remove inference paths from the search tree that have a quality rating score lower than a specified threshold.

5. A training method for a quality assessment model, characterized in that, include: The quality evaluation model is trained using the target training samples generated by the method as described in any one of claims 1 to 4; wherein the quality evaluation model is used to generate quality evaluations for the answers to user questions corresponding to the online question answering model.

6. The method according to claim 5, characterized in that, The method further includes: Obtain the learning difficulty index of the target training sample; wherein, the learning difficulty index is used to characterize the difficulty of the quality evaluation model in learning the knowledge carried by the corresponding target training sample; Training the quality assessment model includes: training the quality assessment model using target training samples, according to the difficulty level indicated by the learning difficulty index from low to high.

7. The method according to claim 6, characterized in that, The method further includes: Generate a difficulty evaluation score for the target training sample relative to the difficulty evaluation dimension; wherein the difficulty evaluation dimension includes one or more of the following: inconsistency index, used to represent the proportion of inconsistency between the quality evaluation output of the quality evaluation model and the human annotation result; question normalization index, used to represent the length complexity of the question; semantic complexity index, used to characterize the semantic complexity of the question; candidate answer diversity, used to reflect the degree of difference between different candidate answers for the same question; annotator inconsistency index, used to represent the degree of disagreement among multiple human annotators on the same answer; The learning difficulty index is obtained by weighted fusion of the difficulty evaluation scores.

8. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, causes the processor to implement the method as described in any one of claims 1 to 7.

9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to implement the method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, Includes computer instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 7.