Generative iterative optimization method and device for health consultation, equipment and medium
By receiving user questions, generating candidate answers, conducting multi-dimensional evaluations, and calibrating confidence, a weighted preference dataset is constructed. The dialogue generation model is then fine-tuned with confidence weights, solving the problem of high optimization costs for health consultation models and achieving efficient self-iterative optimization and performance improvement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-28
AI Technical Summary
The optimization of existing health consultation models is costly and inefficient, making it difficult to continuously improve performance, mainly due to the large-scale manual annotation required by doctors.
By receiving health consultation questions submitted by users, user profiles are obtained, multiple candidate answers are generated, multi-dimensional evaluation and confidence calibration are performed, answers that need to be corrected are selected, a structured weighted preference dataset is constructed, and the dialogue generation model is dynamically fine-tuned and optimized with confidence weighting to achieve self-iterative optimization.
It reduced optimization costs, improved optimization efficiency, provided high-quality and targeted training materials, and continuously improved the performance of the health consultation model.
Smart Images

Figure CN121938534A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a generative iterative optimization method, apparatus, device, and medium for health consultation. Background Technology
[0002] AI-assisted smart elderly care health consultation is an intelligent service technology that uses natural language processing, deep learning and other technologies to automatically generate questions and suggestions for the health consultation needs of the elderly. Its core value lies in providing convenient and professional health guidance for the elderly, making up for the gap in elderly care health service resources and improving the accessibility of services. It has been gradually promoted and applied in many fields.
[0003] For example, in the smart elderly care facility scenario within the medical field, AI-assisted health consultation systems can cover multiple sub-needs such as chronic disease management, daily health care, and medication consultation. They can quickly respond to the health questions of the elderly, providing standardized advice on diet, exercise, and medication, reducing immediate reliance on professional medical personnel. Another example is in the health insurance scenario within the financial sector. Insurance companies can use AI-assisted health consultations to intelligently assess the health status of elderly policyholders during underwriting, renewal, or claims processes. This helps identify potential health risks, verify the authenticity of insurance information, and thus optimize the underwriting and claims process.
[0004] High-quality training data is crucial for optimizing the performance of health consultation models. However, since health consultation involves medical expertise and requires doctors to participate in annotation, the labor costs are high and the efficiency is low. Traditional model optimization methods usually rely on continuous large-scale manual annotation, which makes the optimization of health consultation models costly and inefficient, thus affecting the improvement of model performance. Summary of the Invention
[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide a generative iterative optimization method, apparatus, equipment and medium for health consultation that can be applied to the financial field, medical field or other related fields. Its main purpose is to realize the self-optimization and iteration of the health consultation model, reduce optimization costs and improve optimization efficiency, thereby continuously improving the performance of the health consultation model.
[0006] The technical solution of the present invention is as follows: The first aspect of this invention provides a generative iterative optimization method for health consultation, comprising: Receive health consultation questions submitted by users and obtain user profiles; Based on the user profile, the current version of the dialogue generation model is driven to generate multiple candidate answers to the health consultation question using the corresponding generation strategy. The multiple candidate answers are evaluated and their confidence is calibrated in multiple dimensions. Based on the evaluation and calibration results, the answers that need to be reviewed and corrected by experts are selected, and corresponding corrected answer samples are obtained. The corrected response samples were re-evaluated and their confidence levels were calibrated in multiple dimensions, and a structured weighted preference dataset was constructed based on the evaluation and confidence data of the corrected response samples. The current version of the dialogue generation model is dynamically fine-tuned using confidence weighting through the structured weighted preference dataset to obtain a new version of the dialogue generation model for the next round of iterative optimization.
[0007] A second aspect of the present invention provides a generative iterative optimization apparatus for health consultation, comprising: The receiving and acquisition module is used to receive health consultation questions submitted by users and obtain user profiles. The generation module is used to drive the current version of the dialogue generation model to generate multiple candidate answers to the health consultation question according to the user profile and the corresponding generation strategy. The multi-dimensional evaluation module is used to perform multi-dimensional evaluation and confidence calibration on the multiple candidate answers, and to select the answers that need to be corrected by experts based on the evaluation and calibration results, and to obtain the corresponding corrected answer samples. The preference data construction module is used to re-evaluate and calibrate the confidence of the corrected answer samples in multiple dimensions, and to construct a structured weighted preference dataset based on the evaluation and confidence data of the corrected answer samples. The dynamic optimization module is used to perform confidence-weighted dynamic fine-tuning optimization of the current version of the dialogue generation model using the structured weighted preference dataset, so as to obtain a new version of the dialogue generation model for the next round of iterative optimization.
[0008] A third aspect of the present invention provides a computer device including at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform the generative iterative optimization method for health consultation described above.
[0009] A fourth aspect of the present invention provides a non-volatile computer-readable storage medium storing computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the above-described generative iterative optimization method for health consultation.
[0010] Beneficial Effects: This invention discloses a generative iterative optimization method, apparatus, device, and medium for health consultation. Compared to existing technologies, this invention receives health consultation questions submitted by users and obtains user profiles. Based on these user profiles, it drives the current version of the dialogue generation model to generate multiple candidate answers to the health consultation questions using corresponding generation strategies. The multiple candidate answers undergo multi-dimensional evaluation and confidence calibration. Based on the evaluation and calibration results, answers requiring expert review and correction are selected, and corresponding corrected answer samples are obtained. The corrected answer samples are then re-evaluated and recalibrated in multiple dimensions, and a structured weighted preference dataset is constructed based on the evaluation and confidence data of the corrected answer samples. The current version of the dialogue generation model is dynamically fine-tuned using the structured weighted preference dataset to obtain a new version of the dialogue generation model for the next round of iterative optimization. By evaluating multiple candidate answers generated in the current version in multiple dimensions and selecting some answers for correction to construct structured preference data with quality weights, high-quality and targeted training materials are provided for model optimization, reducing optimization costs and improving optimization efficiency. Attached Figure Description
[0011] To more clearly illustrate the solutions in this invention, the accompanying drawings used in the description of the embodiments of this invention will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0012] Figure 1 A schematic diagram of an application environment for the generative iterative optimization method for health consultation provided in an embodiment of the present invention; Figure 2 A flowchart of a generative iterative optimization method for health consultation provided in an embodiment of the present invention; Figure 3 A schematic diagram of the functional modules of the generative iterative optimization device for health consultation provided in an embodiment of the present invention; Figure 4 A schematic diagram of the hardware structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0013] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. The embodiments of the invention are described below in conjunction with the accompanying drawings.
[0014] The generative iterative optimization method for health consultation provided in this invention can be applied to, for example... Figure 1 In the application environment, it includes a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0015] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).
[0016] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0017] Server 105 can be a server providing various services, such as a backend server supporting the content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend server can analyze and process received user requests and other data, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices. Server 105 can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the shortcomings of traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"), such as high management difficulty and weak business scalability. Server 105 can also be a server for a distributed system or a server combined with blockchain.
[0018] It should be noted that the generative iterative optimization method for health consultation provided in this application embodiment can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the generative iterative optimization device for health consultation provided in this embodiment can also be located in the first terminal device 101, the second terminal device 102, or the third terminal device 103. Alternatively, the generative iterative optimization method for health consultation provided in this embodiment can generally be executed by the server 105. Correspondingly, the generative iterative optimization device for health consultation provided in this embodiment can generally be located in the server 105.
[0019] It should be understood that the number of terminal devices, networks, and servers listed above is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be used.
[0020] like Figure 2 As shown, the generative iterative optimization method for health consultation provided in this embodiment of the invention specifically includes the following steps: S201. Receive health consultation questions submitted by users and obtain user profiles.
[0021] In this embodiment, health consultation questions refer to inquiries raised by users regarding their own health conditions. The users are primarily elderly, but other groups such as young and middle-aged users can also receive and answer health consultation questions. Consultation content can include inquiries about disease symptoms, medication questions, dietary advice requests, and exercise guidance requests, such as "I have a 5-year history of hypertension, and my blood pressure was 155 / 95 mmHg this morning. What should I be careful about?" or "Can I eat grapefruit while taking valsartan?"
[0022] When receiving prompts for health consultation questions from users, the system also obtains the user profile of the user currently asking the question. The user profile is structured data that represents the user's health-related characteristics, including basic information (age, gender), health status (past medical history, current symptoms, allergy history), medication information (name of current medication, dosage, duration of medication), and lifestyle habits (dietary preferences, exercise frequency), so as to generate personalized answers for different user situations.
[0023] Specifically, users can submit health consultation questions through a health consultation app or mini-program on their mobile devices, choosing from input methods such as text input and voice-to-text (adapting to the operating habits of elderly users). When a question is submitted, the system automatically links to the user's historical profile data. For new users, a guided questionnaire collects profile information, such as "Do you have a history of chronic diseases such as hypertension or diabetes?" and "Are you currently taking any medication?", providing a basis for subsequent processing. By clearly defining the user's core needs and personalized characteristics, a foundation is provided for generating personalized answers, avoiding generic responses.
[0024] For example, a 70-year-old user, Mr. Li, submitted a consultation question through the Smart Elderly Care APP: "My blood pressure was 155 / 95 mmHg this morning. Do I need to adjust my medication? What should I pay attention to?" The system automatically retrieved Mr. Li's historical user profile, which showed that he was 70 years old, male, had a 5-year history of hypertension, no history of diabetes or allergies, was currently taking valsartan 80mg / day for 3 years, had a salty diet, and exercised 1-2 times a week. The system then linked the question to the profile for further processing.
[0025] S202. Based on the user profile, drive the current version of the dialogue generation model to generate multiple candidate answers to the health consultation question using the corresponding generation strategy.
[0026] In this embodiment, the current version of the dialogue generation model is a dedicated dialogue model for health consultation that has undergone initial training and multiple rounds of iterative optimization. It includes multiple sub-models for generating responses from different professional fields, enabling parallel generation of answers from various professional perspectives. The generation strategy is a dynamic rule that adjusts the focus and style of the response content based on the characteristics of the current user profile. For example, it optimizes the language for elderly users (simplifying sentence structure and avoiding technical terms), strengthens medication safety tips for users with chronic diseases, and avoids contraindications for users with allergies.
[0027] After receiving the bound health consultation questions and user profiles, the system can first perform semantic analysis on the consultation questions to clarify the core needs of the questions. Then, after analyzing key features of the profile (such as age, medical history, and medication use), strategy adjustment information is generated. For example, it might suggest strengthening medication safety tips or simplifying the expression style. This drives the current version of the dialogue generation model to adjust its generation focus and generate multiple candidate answers from different professional dimensions, ensuring coverage of the core needs of the questions. For adjusting the generation strategy, scenario-based strategy templates can also be preset, such as "elderly chronic disease user template" or "post-operative rehabilitation user template." The corresponding template is matched according to the user profile to quickly determine the generation focus; this embodiment does not limit this. By generating multiple candidate answers, user needs are covered from multiple dimensions, and the generation strategy is dynamically adjusted based on the user profile, improving the relevance and comprehensiveness of the answers.
[0028] For example, based on user Li's inquiry and profile, the current version of the dialogue generation model is driven to generate multiple candidate answers in parallel from different professional perspectives, including: Model A (Medication Management Focus): "Do not adjust the valsartan dosage on your own. It is recommended to monitor your blood pressure regularly and record it. Inform your doctor of today's data at your next follow-up visit." Model B (Dietary Focus): "Pay attention to a low-salt diet. Today, you can appropriately increase your intake of potassium-rich vegetables such as spinach and bananas." Model C (Exercise Focus): "Avoid strenuous exercise immediately when blood pressure is high; try a slow walk instead." S203. Perform multi-dimensional evaluation and confidence calibration on the multiple candidate answers, select the answers that need to be corrected by experts based on the evaluation and calibration results, and obtain the corresponding corrected answer samples.
[0029] In this embodiment, a comprehensive multi-dimensional evaluation system is constructed based on the core needs of health consultation. Specifically, it includes four key dimensions: medical accuracy (whether the answer conforms to medical common sense and professional standards), age-appropriate language (whether the sentence structure is concise, the expression is easy to understand, and whether it avoids professional jargon), practicality of the answer (whether the suggestions are actionable and relevant to the user's actual situation), and emotional support (whether it includes reassuring language and alleviates the user's anxiety). This multi-dimensional evaluation system independently scores each candidate answer across multiple dimensions, and performs confidence calibration based on the multi-dimensional scoring results. That is, it calculates the confidence score of each candidate answer by quantifying the reliability of the evaluation results, thereby judging the credibility of the answer quality. From all candidate answers, those requiring expert review and correction are selected.
[0030] The specific answers to be corrected can be candidate answers with confidence levels below a preset threshold or those posing potential risks. These require review and correction by medical experts to ensure accuracy. If the answers to be corrected are pushed to the expert terminal via an expert interface, the medical experts will review each received answer, correcting any imprecise wording, inaccurate suggestions, or risky content to create a corrected answer sample. Therefore, the corrected answer sample, after expert review and correction, represents high-quality answers confirmed by the experts. This provides high-quality, targeted preference data for the model's iterative optimization, thus providing core material for iterative optimization based on the dialogue generation model's own output, reducing optimization costs.
[0031] For example, for the candidate answers from the three professional perspectives mentioned above, each candidate answer is independently scored from four core dimensions: medical accuracy, age-appropriate language, answer practicality, and emotional support. For instance, (Medical accuracy: Model A -0.95, Model B -0.75, Model C -0.90; Practicality: Model A -0.90, Model B -0.85, Model C -0.80; Answer practicality: Model A -0.95, Model B -0.85, Model C -0.95; Emotional support: Model A -0.95, Model B -0.75, Model C -0.90). Then, based on the multi-dimensional scores, the uncertainty score (e.g., score variance) of the review result for each candidate answer is calculated. Finally, the confidence level of each candidate answer is calculated according to a preset formula, and candidate answers with low confidence levels (e.g., <0.6) are automatically identified as requiring expert review and correction. For example, the candidate answer generated by Model B has a high score variance and a confidence level below 0.6, therefore it is selected as a sample requiring correction.
[0032] By establishing an interface with the medical expert system, answers requiring correction are periodically pushed to medical experts for manual review and revision. For example, if a candidate answer from Model B is identified as a sample requiring correction and involves specific dietary advice, it is pushed to a nutritionist on the platform. The nutritionist corrects it to: "Pay attention to a low-salt diet. You can appropriately increase potassium-rich foods, such as spinach and bananas. However, if you are taking potassium-sparing diuretics or have renal insufficiency, you should be cautious about high-potassium diets and consult your attending physician." This corrected answer is marked as the "best answer," and the correction results and annotation information are fed back to the system to obtain the corresponding corrected answer sample.
[0033] S204. The corrected response samples are re-evaluated and their confidence levels are calibrated in multiple dimensions, and a structured weighted preference dataset is constructed based on the evaluation and confidence data of the corrected response samples.
[0034] In this embodiment, after receiving corrected answer samples from experts, a secondary evaluation is conducted using the same multi-dimensional evaluation system. This involves using the same dimensions and weights as the initial evaluation to calculate the corrected multi-dimensional score and confidence level. Since expert correction ensures answer accuracy, the confidence level is typically significantly improved after correction. By re-evaluating and calibrating the corrected answer samples, the quality improvement effect of expert correction is quantified, ensuring high reliability of the samples included in the preference dataset. Subsequently, a structured weighted preference dataset containing input, output, and quality weights is constructed based on the evaluation and confidence data of the corrected answer samples in a unified format for subsequent model fine-tuning. For example, the constructed standard preference data is a triple (input question, best answer, ranking score), where "best answer" is the highest-scoring answer or the answer corrected by the expert, and "ranking score" is the corresponding weighted average score. Further, an extended weighted preference data quadruple (input question, best answer, ranking score, confidence weight) is constructed, where the confidence weight is obtained by normalizing the final confidence level of the corrected answer samples to reflect the contribution of sample quality to model training, thereby supporting differentiated training. By constructing a structured weighted preference dataset, high-quality training materials with reliability weights are formed, providing clear objectives for model optimization, reducing interference from invalid data, and significantly reducing the amount of data labeled by experts, thus lowering labeling costs.
[0035] Preferably, after the model is built, each data point can be fully recorded with information such as model version, scores for each dimension of the review, confidence level, whether it has been manually reviewed, and expert identification. Furthermore, the data can be divided into three quality levels—high (≥0.8), medium (0.5~0.8), and low (<0.5)—based on the confidence level and stored separately to support subsequent differentiated training strategies and data asset management.
[0036] For example, the nutritionist's corrected answer sample in the above embodiment is evaluated a second time, and the multidimensional scores are: medical accuracy 0.98, language age-appropriateness 0.92, answer practicality 0.95, emotional support 0.85, weighted average score 0.94, confidence level 0.94, and confidence weight 0.9. The weighted preference data quadruple (input question, best answer, ranking score, confidence weight) constructed at this time is: "A 70-year-old hypertensive patient's blood pressure is 155 / 95 mmHg while taking valsartan. Should the medication be adjusted and what are the precautions?", "Your current blood pressure of 155 / 95 mmHg is slightly high. Please do not adjust the valsartan dosage yourself... (corrected complete answer)", 0.94, 0.9), and this quadruple is stored in the weighted preference dataset.
[0037] S205. The current version of the dialogue generation model is dynamically fine-tuned and optimized using the structured weighted preference dataset to obtain a new version of the dialogue generation model for the next round of iterative optimization.
[0038] In this embodiment, samples from the weighted preference dataset are used to adjust model parameters through a pre-constructed confidence-weighted loss function. This allows the model to prioritize learning the generation logic of high-confidence, high-quality samples, thereby achieving dynamic fine-tuning optimization based on confidence weighting. For example, the structured weighted preference dataset is divided into a training set (80%) and a validation set (20%). A confidence-weighted ranking loss function is used for model fine-tuning. This loss function includes at least a confidence weight term and may also include a safety constraint term, guiding the model to learn the optimal response generation pattern while suppressing the generation of risky content (such as "self-discontinuing medication" or "recommended folk remedies"). The confidence weights of different samples reflect sample quality, thus having varying degrees of impact on model parameter updates during fine-tuning. High-quality samples have a greater influence on model parameter updates, thereby achieving a dynamic adjustment of model parameter optimization. After fine-tuning, a new version of the dialogue generation model is generated, recording information such as the model version number, training data batch, and performance metrics. It is then deployed to a test environment for effect verification. Once verification is successful, it is used for the next round of iterative optimization.
[0039] For example, the weighted preference dataset constructed in the above embodiments is input into the current version of the dialogue generation model. Fine-tuning is performed using a weighted ranking loss function that includes confidence weights and safety constraints. Since the modified answer samples obtained after the nutritionist's correction have high confidence weights, they have a greater impact on updating the model parameters during fine-tuning. The weighted ranking loss function guides the model to learn and generate rigorous answers similar to the corrected ones, while the safety constraints suppress the risk of giving absolute dietary advice without mentioning contraindications. The new version of the dialogue generation model obtained after fine-tuning can improve the professionalism and rigor of its dietary advice.
[0040] In one embodiment, step S102 includes: Feature extraction is performed on the health consultation question to obtain an input question representation, and context-aware encoding is performed on the user profile to obtain the corresponding latent vector; The latent vector is fused with the input question representation to obtain a fused hint vector; The fused prompt vector and the input question representation are input into the current version of the dialogue generation model, which contains several professional generation sub-models that have been fine-tuned with knowledge from different professional domains. The generation strategies of the several professional generation sub-models are adjusted using the fusion prompt vector. Based on the adjusted generation strategies, the several professional generation sub-models decode the input question representation and generate several candidate answers from different professional perspectives.
[0041] In this embodiment, a pre-trained BERT model for the medical field can be used to extract features from health consultation questions. This involves preprocessing the questions, including word segmentation and stop word removal, before inputting them into the BERT model for feature extraction to obtain the corresponding input question representation. This vector fully preserves the semantic information of the question, thus transforming the natural language consultation question into a high-dimensional vector that the model can process. Simultaneously, context-aware encoding of the user profile is performed. Specifically, this can be achieved using a conditional variational autoencoder or attention context encoding network to encode the user profile (including age, past medical history, current medications, allergies, etc.) into a latent vector. This latent vector captures the correlation between various dimensions of the user profile through probability distribution modeling, achieving efficient representation of contextual information.
[0042] Then, through vector concatenation and attention weighting mechanisms, the latent vector and the input question representation are fused to generate a fused prompt vector that simultaneously incorporates the question's intent and the user context, providing a unified guiding signal for subsequent model adjustments and generation strategies. Specifically, the latent vector and the input question representation are first aligned in dimensions to ensure consistency. Then, an attention fusion mechanism is used, with the input question representation as the query vector and the mapped latent vector as the key vector, to calculate their attention weights. These weights reflect the relevance of each dimension of the user profile to the question's intent; for example, when the question involves medication adjustments, the attention weight of the "current medication" dimension is significantly increased. Finally, the input question representation and the weighted latent vector are added element-wise to obtain the fused prompt vector, which retains the core semantics of the question while incorporating user context features highly relevant to the question.
[0043] The fused prompt vectors and input question representations are then fed into the current version of the dialogue generation model. The dialogue generation model adopts a modular architecture of a main model and sub-models. The main model is responsible for overall coordination, while several specialized generation sub-models serve as core execution units. The specialized generation sub-models are pre-tuned on subdivided medical domain datasets (such as drug management datasets, diet and nutrition datasets, sports rehabilitation datasets, etc.) to possess the ability to generate professional knowledge in that domain and can independently output targeted answers.
[0044] Upon receiving the fused suggestion vector, the generation strategies of several specialized sub-models are dynamically adjusted by analyzing the user context features within it. Specifically, this involves adjusting generation parameters and probability distributions. For instance, the parameter responsible for "choosing a suggested direction" in the model might automatically increase the weight of "reminding users not to adjust their own medication" and decrease the weight of "recommending new drugs" due to the fused vector's mention of "70+ years old taking valsartan." Similarly, because the fused vector contains the "elderly user" attribute, the model will tend to generate "mild and easy-to-understand" expressions when selecting words and organizing sentences, while reducing the probability of generating "technical jargon and complex sentences." After the strategy adjustments are complete, each specialized sub-model performs semantic decoding on the input question representation based on the adjusted generation strategy, generating candidate answers in natural language from different professional perspectives to ensure that the candidate answers are both professional and relevant to the user's actual situation.
[0045] In one embodiment, after step S102, the method further includes: The multiple candidate answers are identified by using a pre-trained knowledge graph, and the professional knowledge entities and entity relationships in each candidate answer are identified. The professional knowledge entities and entity relationships are security-screened according to preset risk interception rules, and candidate answers that meet the preset interception rules are intercepted.
[0046] In this embodiment, the pre-trained knowledge graph is a specialized knowledge graph for the medical and health field, containing core nodes and relationships such as diseases, drugs, symptoms, foods, and contraindications. Examples include "hypertension - contraindications - high-salt foods," "valsartan - interaction - potassium-sparing diuretics," and "diabetes - recommendation - low-sugar diet." This pre-trained knowledge graph can be used to efficiently and reliably identify professional knowledge entities and their relationships. The pre-trained knowledge graph is used to identify entities and relationships in multiple candidate answers, recognizing professional knowledge entities and entity relationships in each candidate answer. For example, a professional generation sub-model generates the candidate answer: "You can try reducing your valsartan dosage and eating more grapefruit to help lower your blood pressure." Entity recognition using the pre-trained knowledge graph identifies the professional knowledge entities: "valsartan" (drug), "reducing dosage" (medical procedure), "grapefruit" (food), and "lowering blood pressure" (medical effect); and entity relationships such as: "reducing dosage - object - valsartan," "eating more - object - grapefruit," and "grapefruit - effect - lowering blood pressure."
[0047] Next, the system uses pre-set risk interception rules to screen professional knowledge entities for safety. These rules are a pre-defined set of prohibitive rules based on medical common sense and clinical guidelines, targeting high-risk behaviors and erroneous associations that may harm user health. Examples include "self-adjusting prescription drug dosages," "recommending foods / drugs that are contraindicated with current medications," "promoting folk remedies as alternatives to regular treatments," and "absolute promises of efficacy." The identified professional knowledge entities and their relationships are matched against the risk interception rules to determine if candidate answers contain high-risk content. For example, a risk interception rule could be "self-adjustment - prescription drugs - high risk." Therefore, if "reducing dosage - target - valsartan" is identified, this rule is matched, and the candidate answer is automatically blocked, not proceeding to the subsequent multi-dimensional evaluation stage. The reason for the block is recorded. If no risk rules are matched or only low-risk rules are matched, the candidate answer is considered safe and proceeds to the next evaluation stage.
[0048] This embodiment adds a pre-screening security step after candidate answers are generated. By using knowledge graph identification and risk rule screening, high-risk answers are intercepted in advance, reducing the health risks caused by incorrect suggestions from the source and ensuring the efficiency of subsequent evaluation and expert review.
[0049] In one embodiment, step S103 includes: Each candidate answer is scored for quality across multiple preset dimensions to obtain a corresponding multidimensional score. The confidence level is calculated based on the multidimensional score of each candidate answer, and then compared with a preset threshold. If the confidence level is lower than a preset threshold, the corresponding candidate answer is confirmed as an answer to be corrected, and the answer to be corrected is pushed to the expert terminal for review and correction to obtain the corresponding corrected answer sample.
[0050] In this embodiment, multiple preset dimensions are used, namely four core dimensions: medical accuracy, language age-appropriateness, answer practicality, and emotional support. Each dimension has clear scoring standards and rules to ensure the objectivity and consistency of the scoring. At the same time, each dimension is also configured with a corresponding weight to reflect its importance. For example, medical accuracy has a weight of 0.4, language age-appropriateness has a weight of 0.2, answer practicality has a weight of 0.3, and emotional support has a weight of 0.1.
[0051] When scoring the quality of each candidate answer across multiple preset dimensions, a combination of rule matching and model scoring can be used. First, a score is scored using preset rule matching, for example, deducting medical accuracy points directly for identifying "self-adjusting medication". Then, a pre-trained evaluation model scores each dimension, and the average of the two scores is taken as the final score for that dimension, outputting a multi-dimensional score vector, such as [0.95, 0.90, 0.85, 0.80].
[0052] The confidence score is calculated based on the multidimensional scores of each candidate answer. Specifically, a weighted average score can be calculated first, that is, the multidimensional score vector is weighted and summed according to the preset weights of each dimension to obtain the weighted average score. Then, Monte Carlo Dropout or a method that integrates multiple review sub-models is used to calculate the uncertainty score (such as the score variance) of the review result of each candidate answer. For example, three independently trained evaluation sub-models are deployed to score the same candidate answer, and the variance of the three sets of multidimensional scores is calculated. The larger the variance, the more unstable the evaluation result and the higher the uncertainty. The variance is normalized to the 0-1 range to obtain the uncertainty score. Finally, the confidence score of each candidate answer is calculated by the formula: Confidence score = (weighted average score) * (1 - uncertainty score). This confidence score comprehensively reflects the quality of the answer and the reliability of the evaluation result. The confidence score is compared with a preset threshold, such as 0.6, to determine whether a review is required.
[0053] If the confidence level is below a preset threshold, it indicates potential issues such as imprecise medical descriptions, insufficient practicality, or inappropriate language, but it does not meet the high-risk interception criteria. Therefore, manual review and correction by medical experts are required. Candidate answers with confidence levels below the preset threshold are automatically selected as answers requiring correction and categorized by professional field, then pushed to relevant experts for focused review and correction. For example, drug-related answers are pushed to internists, and diet-related answers to nutritionists. Upon receiving the answers, medical experts review them based on user profiles and their professional knowledge. If the answer is imprecise, details are added (e.g., adding contraindications to dietary recommendations); if practicality is insufficient, the operability of the recommendations is optimized (e.g., specifying exercise duration); if language is inappropriate, the expression style is adjusted (e.g., explaining professional terminology). After correction, a corrected answer sample is received from the expert's terminal. This corrected answer sample represents a high-quality answer after expert correction, facilitating targeted learning and optimization by the model. This ensures that the samples used for model fine-tuning are high-quality data, providing reliable material for model optimization while avoiding the high costs associated with full data annotation.
[0054] In one embodiment, step S104 includes: The quality of the corrected response sample is re-evaluated from multiple preset dimensions to obtain the corresponding corrected multidimensional score. The corrected multidimensional score is weighted according to the weights of multiple preset dimensions to obtain a weighted average score and the corrected confidence level of the corrected answer sample is calculated. The corrected confidence scores are normalized and used as the confidence weights of the corrected response samples. Based on the health consultation questions, corrected answer samples, weighted average scores, and confidence weights, a structured weighted preference dataset is constructed.
[0055] In this embodiment, the re-evaluation of the corrected answer sample uses the same dimensions, scoring criteria, and confidence calculation methods as the initial evaluation to ensure the comparability of the evaluation results. Specifically, during the re-evaluation, the multi-dimensional evaluation module uses the four dimensions of medical accuracy, linguistic age-appropriateness, answer practicality, and emotional support, along with their corresponding scoring criteria, to score the corrected answer sample, generating a corrected multi-dimensional score vector. Furthermore, using preset weights for each dimension (e.g., medical accuracy 0.4, linguistic age-appropriateness 0.2, answer practicality 0.3, emotional support 0.1), the corrected multi-dimensional score vector is weighted and summed to obtain a weighted average score. Similarly, using the same uncertainty score calculation method and confidence formula as the initial evaluation, combined with the weighted average score, the corrected confidence score is calculated to characterize the quality of the corrected sample.
[0056] When constructing the weighted preference dataset, the corrected confidence scores are first normalized, that is, the confidence scores of different samples are uniformly mapped to the 0-1 interval, which facilitates horizontal comparison and model processing. This normalized confidence score is used as the confidence score weight of the corrected response sample. This confidence score weight is an indicator that reflects the reliability of the corrected response sample. It is used to adjust the contribution of the sample to the parameter update when fine-tuning the model. That is, the higher the confidence score weight, the greater the influence of the sample on the model training, ensuring that high-quality and reliable samples play a dominant role.
[0057] The constructed structured weighted preference dataset is a standardized dataset in the form of quadruples. Each quadruple is defined as: (input question, best answer, ranking score, confidence weight), where the input question is the health consultation question submitted by the user, the "best answer" is the corrected answer sample after expert feedback, the ranking score is the weighted average score, and the confidence weight is the normalized value of the final confidence score. This ensures that each quadruple contains a complete relationship between "input-output-quality-reliability," facilitating model reading and learning, allowing model fine-tuning to accurately focus on high-quality, high-reliability samples, and improving training efficiency and optimization results.
[0058] In one embodiment, step S105 includes: Construct a weighted ranking loss function that includes confidence weight constraints; Based on the confidence weight of each sample in the structured weighted preference dataset, the corresponding weighted dynamic loss value is calculated using the weighted ranking loss function; The parameters of the current version of the dialogue generation model are fine-tuned and optimized based on the weighted dynamic loss value.
[0059] In this embodiment, the weighted ranking loss function is the basis for achieving confidence-weighted fine-tuning. A weighted ranking loss function containing confidence weight constraints is constructed so that it incorporates both sample confidence weights and ranking loss. Preferably, a safety constraint term can also be added when constructing the weighted ranking loss function, so that it not only guides the model to learn the ranking relationship between the optimal answer and other candidate answers, but also adjusts the training weights according to the reliability of the samples, and also suppresses the generation of risky content, ensuring the professionalism and safety of model optimization.
[0060] In practice, the weighted ranking loss function is constructed as follows:
[0061] in, For the sample The confidence weights reflect sample quality and allow for dynamic adjustment of their impact on parameter updates. Pairwise (such as RankNet) or Listwise (such as ListNet) sorting loss functions can be used; For security constraints; To balance hyperparameters, used to adjust the strength of safety constraints, safety constraint terms... By adding a suppression term for preset risk terms (such as "discontinuing medication on one's own" or "recommendation of folk remedies") to the loss function, the specific form is as follows:
[0062] in Indicates the first The term "risk" refers to specific risk-related terms; "context" refers to the contextual information. Given a context, the model generates the first... Risk words The probability of generating risky words by the model is reduced by this security constraint.
[0063] The weighted preference data quadruples from the weighted preference dataset are used as training samples and input into the current version of the dialogue generation model. The model generates predicted answers to the "input questions" in the training samples. For each sample, its confidence weight is extracted. Calculate the ranking loss for this sample. This involves comparing the ranking difference between the predicted answer and the "optimal answer" (corrected answer sample) in the training samples, and calculating the safety constraint term. This involves assessing the probability of generating risky words in the predicted response. , ,λ, Substituting into the weighted ranking loss function, the corresponding weighted dynamic loss value is calculated. This weighted dynamic loss value is the loss calculation result for each sample under the current model parameters. It reflects the degree of fit of the model to the sample. That is, the smaller the loss value, the closer the current answer generated by the model is to the optimal answer. The higher the confidence weight of the sample, the more significant the impact of the loss value on training.
[0064] Based on the weighted dynamic loss value, a gradient descent algorithm is used for parameter fine-tuning and optimization. Specifically, it can be done by traversing all samples in the weighted preference dataset, accumulating the total weighted dynamic loss value, and using it as the basis for updating model parameters. Then, the gradient of the loss value with respect to each model parameter is calculated through the backpropagation algorithm. The model parameters are updated according to the gradient direction and the learning rate, so that the total loss value gradually decreases. By minimizing the total weighted dynamic loss value, the model's generation parameters are adjusted, allowing the model to gradually learn the generation logic of the optimal answer. At the same time, safety constraints are strengthened, improving the quality and safety of the answers, achieving the dual goals of differentiated learning and safety reinforcement. This allows high-confidence samples to dominate the model parameter update while suppressing the generation of risky words.
[0065] In one embodiment, after step S105, the method further includes: The difference in performance of the dialogue generation model before and after iterative optimization is evaluated using a test set, and / or user feedback data on the new version of the dialogue generation model is collected. The parameters of the multi-dimensional assessment and confidence calibration are updated based on the effect difference data and / or the user feedback data for use in the next round of assessment and calibration.
[0066] In this embodiment, the effect difference data is used to quantify the optimization effect of model iteration, while user feedback data is used to reflect the model's performance in real-world application scenarios. Both, individually and in combination, can provide a basis for updating evaluation parameters. Specifically, when obtaining the effect difference data, the test set can be input into both the old version of the model before iteration and the new version of the model after iteration, generating two sets of responses. Using indicators consistent with the multi-dimensional evaluation (medical accuracy, age-appropriate language, response usability, and emotional support) and scoring criteria, the two sets of responses are scored separately, and the difference value of each indicator (new version score - old version score) is calculated to form the effect difference data, which may include, for example, the improvement rate of scores in each dimension, the overall accuracy improvement rate, and the decrease in the risk content generation rate.
[0067] When collecting user feedback data, the new version of the model can be deployed to some users for trial operation. After users use the model to get answers, feedback can be collected through questionnaires and other methods, such as collecting satisfaction ratings (1-5 points, 5 points being the most satisfactory), whether it is necessary to ask the question again (yes / no), automatic timer for consultation dwell time, etc., while allowing users to submit text feedback.
[0068] The parameters for multi-dimensional assessment and confidence calibration are updated using at least one of the following: performance difference data and user feedback data. These parameters include the weights of each assessment dimension and preset confidence thresholds. These parameters are dynamically adjusted based on performance difference data and / or user feedback data. For example, if user feedback shows the highest focus on "answer usability" and lower focus on "emotional support," the weight of answer usability is adjusted from 0.3 to 0.35, and the weight of emotional support is adjusted from 0.1 to 0.05. Alternatively, if the risk content generation rate of the new model significantly decreases (e.g., from 1.8% to 0.3%), and user satisfaction is high, indicating improved reliability of the model's generated answers, the preset confidence threshold can be appropriately increased from 0.7 to 0.75. This rigorously filters out low-quality answers, further improves expert review efficiency, and makes the multi-dimensional assessment system more aligned with the model's actual performance and user needs, ensuring the accuracy of the next round of assessments. Once the parameters are updated, the new evaluation parameters are stored in the system for multi-dimensional evaluation and confidence calibration in the next iteration, enabling the evaluation system and model optimization to iterate synchronously and form a closed loop of continuous optimization.
[0069] It should be noted that there is no necessary order between the above steps. Those skilled in the art will understand from the description of the embodiments of the present invention that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.
[0070] Further reference Figure 3 As a response to the above Figure 2 The present invention provides an embodiment of a generative iterative optimization device for health consultation, which is implemented in accordance with the method shown. Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0071] like Figure 3 As shown, the generative iterative optimization device 30 for health consultation described in this embodiment includes: The receiving and acquiring module 301 is used to receive health consultation questions submitted by users and acquire user profiles; The generation module 302 is used to drive the current version of the dialogue generation model to generate multiple candidate answers to the health consultation question according to the user profile and the corresponding generation strategy. The multi-dimensional evaluation module 303 is used to perform multi-dimensional evaluation and confidence calibration on the multiple candidate answers, select the answers that need to be corrected by experts based on the evaluation and calibration results, and obtain the corresponding corrected answer samples. The preference data construction module 304 is used to re-evaluate and calibrate the confidence of the corrected answer sample in multiple dimensions, and to construct a structured weighted preference dataset based on the evaluation and confidence data of the corrected answer sample. The dynamic optimization module 305 is used to perform confidence-weighted dynamic fine-tuning optimization of the current version of the dialogue generation model using the structured weighted preference dataset, so as to obtain a new version of the dialogue generation model for the next round of iterative optimization.
[0072] The module referred to in this invention is a series of computer program instruction segments that can perform specific functions. It is more suitable than a program for describing the generative iterative optimization execution process of health consultation. For specific implementation methods of each module, please refer to the corresponding method embodiments above, which will not be repeated here.
[0073] In one embodiment, the generation module 302 includes: The feature extraction and encoding module is used to extract features from the health consultation question to obtain an input question representation, and to perform context-aware encoding on the user profile to obtain the corresponding latent vector; The fusion module is used to fuse the latent vector with the input problem representation to obtain a fused hint vector; The input module is used to input the fused prompt vector and the input question representation into the current version of the dialogue generation model, which includes several professional generation sub-models that have been fine-tuned with knowledge from different professional domains. An adaptive generation module is used to adjust the generation strategies of the plurality of professional generation sub-models using the fusion prompt vector. The plurality of professional generation sub-models decode the input question representation based on the adjusted generation strategy to generate a plurality of candidate answers from different professional perspectives.
[0074] In one embodiment, the device 30 further includes: The entity recognition module is used to identify entities and relationships in the multiple candidate answers using a pre-trained knowledge graph, and to identify professional knowledge entities and entity relationships in each candidate answer. The security filtering module is used to perform security screening on the professional knowledge entities and entity relationships according to preset risk interception rules, and to intercept candidate answers that meet the preset interception rules.
[0075] In one embodiment, the multi-dimensional evaluation module 303 includes: The multi-dimensional evaluation unit is used to score the quality of each candidate answer from multiple preset dimensions to obtain the corresponding multi-dimensional score. A confidence calibration unit is used to calculate the corresponding confidence level based on the multidimensional score of each candidate answer and compare the confidence level with a preset threshold. The sample screening and correction unit is used to confirm the corresponding candidate answer as an answer to be corrected if the confidence level is lower than a preset threshold, and push the answer to be corrected to the expert terminal for review and correction to obtain the corresponding corrected answer sample.
[0076] In one embodiment, the preference data construction module 304 includes: The secondary evaluation unit is used to re-evaluate the quality of the corrected response sample from multiple preset dimensions to obtain the corresponding corrected multidimensional score. The scoring weighting unit is used to weight the corrected multidimensional score according to the weights of multiple preset dimensions, obtain a weighted average score, and calculate the corrected confidence level of the corrected answer sample. A normalization unit is used to normalize the corrected confidence level, which serves as the confidence weight for the corrected response sample. The data construction unit is used to construct a structured weighted preference dataset based on the health consultation questions, corrected answer samples, weighted average scores, and confidence weights.
[0077] In one embodiment, the dynamic optimization module 305 includes: Function building unit, used to construct a weighted ranking loss function that includes confidence weight constraints; The weighted loss calculation unit is used to calculate the corresponding weighted dynamic loss value using the weighted ranking loss function based on the confidence weight of each sample in the structured weighted preference dataset. The dynamic fine-tuning unit is used to fine-tune and optimize the parameters of the current version of the dialogue generation model based on the weighted dynamic loss value.
[0078] In one embodiment, the device 30 further includes: The testing and feedback collection module is used to evaluate the difference in performance of the dialogue generation model before and after iterative optimization through a test set, and / or collect user feedback data on the new version of the dialogue generation model. The parameter update module is used to update the parameters of the multi-dimensional evaluation and confidence calibration based on the effect difference data and / or the user feedback data, for use in the next round of evaluation and calibration.
[0079] In the above embodiments, this invention discloses a generative iterative optimization device for health consultation. It receives health consultation questions submitted by users and obtains user profiles. Based on the user profiles, it drives the current version of the dialogue generation model to generate multiple candidate answers to the health consultation questions using a corresponding generation strategy. The multiple candidate answers are evaluated and their confidence levels are calibrated in multiple dimensions. Based on the evaluation and calibration results, answers requiring expert review and correction are selected, and corresponding corrected answer samples are obtained. The corrected answer samples are re-evaluated and their confidence levels are calibrated in multiple dimensions, and a structured weighted preference dataset is constructed based on the evaluation and confidence data of the corrected answer samples. The current version of the dialogue generation model is dynamically fine-tuned using the structured weighted preference dataset to obtain a new version of the dialogue generation model for the next round of iterative optimization. By evaluating multiple candidate answers generated in the current version in multiple dimensions and selecting some answers for correction to construct structured preference data with quality weights, high-quality and targeted training materials are provided for model optimization, reducing optimization costs and improving optimization efficiency.
[0080] Another embodiment of the present invention provides a computer device, such as... Figure 4 As shown, the computer device 40 includes: One or more processors 401 and memory 402, Figure 4 The following section uses a processor 401 as an example. The processor 401 and the memory 402 can be connected via a bus or other means. Figure 4 Taking the example of a connection between China and Israel via a bus.
[0081] Processor 401 performs various control logic functions for computer device 40. It can be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), microcontroller, ARM (AcornRISC), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components. Furthermore, processor 401 can also be any conventional processor, microprocessor, or state machine. Processor 401 can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP, and / or any other such configuration.
[0082] The memory 402, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the generative iterative optimization method for health consultation in this embodiment of the invention. The processor 401 executes various functional applications and data processing of the computer device 40 by running the non-volatile software programs, instructions, and units stored in the memory 402, thereby implementing the generative iterative optimization method for health consultation in the above method embodiment.
[0083] Memory 402 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of computer device 40. Furthermore, memory 402 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 402 may optionally include memory remotely located relative to processor 401, and these remote memories may be connected to computer device 40 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof. One or more units stored in memory 402, when executed by one or more processors 401, perform the steps of the generative iterative optimization method for health consultation in any of the above method embodiments.
[0084] In the above embodiments, the present invention discloses a computer device that receives health consultation questions submitted by users and obtains user profiles; based on the user profiles, drives the current version of the dialogue generation model to generate multiple candidate answers to the health consultation questions using corresponding generation strategies; performs multi-dimensional evaluation and confidence calibration on the multiple candidate answers, and filters out answers that need expert review and correction based on the evaluation and calibration results, obtaining corresponding corrected answer samples; re-evaluates and recalibrates the corrected answer samples in multiple dimensions, and constructs a structured weighted preference dataset based on the evaluation and confidence data of the corrected answer samples; uses the structured weighted preference dataset to dynamically fine-tune and optimize the current version of the dialogue generation model with confidence weights, obtaining a new version of the dialogue generation model for the next round of iterative optimization. By performing multi-dimensional evaluation on the multiple candidate answers generated by the current version and filtering out some answers for correction to construct structured preference data with quality weights, high-quality and goal-oriented training materials are provided for model optimization, reducing optimization costs and improving optimization efficiency.
[0085] This invention provides a non-volatile computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are executed by one or more processors, they perform the steps of the generative iterative optimization method for health consultation in any of the above method embodiments.
[0086] In the above embodiments, the present invention discloses a non-volatile computer-readable storage medium that receives health consultation questions submitted by users and obtains user profiles. Based on the user profiles, the current version of the dialogue generation model is driven to generate multiple candidate answers to the health consultation questions using a corresponding generation strategy. The multiple candidate answers are evaluated and their confidence levels are calibrated in multiple dimensions. Based on the evaluation and calibration results, answers requiring expert review and correction are selected, and corresponding corrected answer samples are obtained. The corrected answer samples are re-evaluated and their confidence levels are calibrated in multiple dimensions, and a structured weighted preference dataset is constructed based on the evaluation and confidence data of the corrected answer samples. The current version of the dialogue generation model is dynamically fine-tuned and optimized using the structured weighted preference dataset to obtain a new version of the dialogue generation model for the next round of iterative optimization. By evaluating multiple candidate answers generated by the current version in multiple dimensions and selecting some answers for correction to construct structured preference data with quality weights, high-quality and targeted training materials are provided for model optimization, reducing optimization costs and improving optimization efficiency.
[0087] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0088] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0089] In summary, the generative iterative optimization method, apparatus, device, and medium for health consultation disclosed in this invention include: receiving a health consultation question submitted by a user and obtaining a user profile; driving a current version of the dialogue generation model to generate multiple candidate answers to the health consultation question using a corresponding generation strategy based on the user profile; performing multi-dimensional evaluation and confidence calibration on the multiple candidate answers, selecting answers requiring expert review and correction based on the evaluation and calibration results, and obtaining corresponding corrected answer samples; re-evaluating and re-calibrating the corrected answer samples using multi-dimensional evaluation and confidence calibration, and constructing a structured weighted preference dataset based on the evaluation and confidence data of the corrected answer samples; and dynamically fine-tuning the current version of the dialogue generation model using the structured weighted preference dataset to obtain a new version of the dialogue generation model for the next round of iterative optimization. By performing multi-dimensional evaluation on the multiple candidate answers generated by the current version and selecting some answers for correction to construct structured preference data with quality weights, high-quality and goal-oriented training materials are provided for model optimization, reducing optimization costs and improving optimization efficiency.
[0090] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The computer program can be stored in a non-volatile, computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The storage medium can be a memory, magnetic disk, floppy disk, flash memory, optical storage, etc.
[0091] It should be noted that any software tools or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. It should be understood that the application of this invention is not limited to the examples described above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A generative iterative optimization method for health consultation, characterized in that, include: Receive health consultation questions submitted by users and obtain user profiles; Based on the user profile, the current version of the dialogue generation model is driven to generate multiple candidate answers to the health consultation question using the corresponding generation strategy. The multiple candidate answers are evaluated and their confidence is calibrated in multiple dimensions. Based on the evaluation and calibration results, the answers that need to be reviewed and corrected by experts are selected, and corresponding corrected answer samples are obtained. The corrected response samples were re-evaluated and their confidence levels were calibrated in multiple dimensions, and a structured weighted preference dataset was constructed based on the evaluation and confidence data of the corrected response samples. The current version of the dialogue generation model is dynamically fine-tuned using confidence weighting through the structured weighted preference dataset to obtain a new version of the dialogue generation model for the next round of iterative optimization.
2. The generative iterative optimization method for health consultation according to claim 1, characterized in that, The step involves driving the current version of the dialogue generation model to generate multiple candidate answers to the health consultation question using a corresponding generation strategy based on the user profile, including: Feature extraction is performed on the health consultation question to obtain an input question representation, and context-aware encoding is performed on the user profile to obtain the corresponding latent vector; The latent vector is fused with the input question representation to obtain a fused hint vector; The fused prompt vector and the input question representation are input into the current version of the dialogue generation model, which contains several professional generation sub-models that have been fine-tuned with knowledge from different professional domains. The generation strategies of the several professional generation sub-models are adjusted using the fusion prompt vector. Based on the adjusted generation strategies, the several professional generation sub-models decode the input question representation and generate several candidate answers from different professional perspectives.
3. The generative iterative optimization method for health consultation according to claim 1, characterized in that, After the method drives the current version of the dialogue generation model to generate multiple candidate answers to the health consultation question in parallel according to the user profile and the corresponding generation strategy, the method further includes: The multiple candidate answers are identified by using a pre-trained knowledge graph, and the professional knowledge entities and entity relationships in each candidate answer are identified. The professional knowledge entities and entity relationships are security-screened according to preset risk interception rules, and candidate answers that meet the preset interception rules are intercepted.
4. The generative iterative optimization method for health consultation according to claim 1, characterized in that, The process involves multi-dimensional evaluation and confidence calibration of the multiple candidate answers, followed by screening out answers requiring expert review and correction based on the evaluation and calibration results, and obtaining corresponding corrected answer samples, including: Each candidate answer is scored for quality across multiple preset dimensions to obtain a corresponding multidimensional score. The confidence level is calculated based on the multidimensional score of each candidate answer, and then compared with a preset threshold. If the confidence level is lower than a preset threshold, the corresponding candidate answer is confirmed as an answer to be corrected, and the answer to be corrected is pushed to the expert terminal for review and correction to obtain the corresponding corrected answer sample.
5. The generative iterative optimization method for health consultation according to claim 1, characterized in that, The process involves re-evaluating and recalibrating the confidence levels of the corrected response samples across multiple dimensions, and constructing a structured weighted preference dataset based on the evaluation and confidence data of the corrected response samples, including: The quality of the corrected response sample is re-evaluated from multiple preset dimensions to obtain the corresponding corrected multidimensional score. The corrected multidimensional score is weighted according to the weights of multiple preset dimensions to obtain a weighted average score and the corrected confidence level of the corrected answer sample is calculated. The corrected confidence scores are normalized and used as the confidence weights of the corrected response samples. Based on the health consultation questions, corrected answer samples, weighted average scores, and confidence weights, a structured weighted preference dataset is constructed.
6. The generative iterative optimization method for health consultation according to claim 1, characterized in that, The step of dynamically fine-tuning and optimizing the current version of the dialogue generation model using the structured weighted preference dataset with confidence weights includes: Construct a weighted ranking loss function that includes confidence weight constraints; Based on the confidence weight of each sample in the structured weighted preference dataset, the corresponding weighted dynamic loss value is calculated using the weighted ranking loss function; The parameters of the current version of the dialogue generation model are fine-tuned and optimized based on the weighted dynamic loss value.
7. The generative iterative optimization method for health consultation according to claim 1, characterized in that, After the method involves dynamically fine-tuning the current version of the dialogue generation model using the structured weighted preference dataset to obtain a new version of the dialogue generation model for the next round of iterative optimization, the method further includes: The difference in performance of the dialogue generation model before and after iterative optimization is evaluated using a test set, and / or user feedback data on the new version of the dialogue generation model is collected. The parameters of the multi-dimensional assessment and confidence calibration are updated based on the effect difference data and / or the user feedback data for use in the next round of assessment and calibration.
8. A generative iterative optimization device for health consultation, characterized in that, include: The receiving and acquisition module is used to receive health consultation questions submitted by users and obtain user profiles. The generation module is used to drive the current version of the dialogue generation model to generate multiple candidate answers to the health consultation question according to the user profile and the corresponding generation strategy. The multi-dimensional evaluation module is used to perform multi-dimensional evaluation and confidence calibration on the multiple candidate answers, and to select the answers that need to be corrected by experts based on the evaluation and calibration results, and to obtain the corresponding corrected answer samples. The preference data construction module is used to re-evaluate and calibrate the confidence of the corrected answer samples in multiple dimensions, and to construct a structured weighted preference dataset based on the evaluation and confidence data of the corrected answer samples. The dynamic optimization module is used to perform confidence-weighted dynamic fine-tuning optimization of the current version of the dialogue generation model using the structured weighted preference dataset, so as to obtain a new version of the dialogue generation model for the next round of iterative optimization.
9. A computer device, characterized in that, Includes at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the generative iterative optimization method for health consultation as described in any one of claims 1-7.
10. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the generative iterative optimization method for health consultation as described in any one of claims 1-7.