Hospital guidance model training method and device, storage medium and program product
By introducing safety constraint rules and collaborative models, using pre-diagnosis dialogue and in-diagnosis data to build preferred data, the accuracy and safety of the guide model are solved, and the adaptability and user experience of the guide model are improved.
Patent Information
- Application Number
- CN202511046364.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-07-29
AI Technical Summary
The existing guide model has limited the accuracy and performance improvement of medical advice when user statements are inaccurate and data in the diagnosis are not fully utilized.
Introduce security constraint rules and collaborative models, build preference data by obtaining pre-diagnosis dialogue and in-diagnosis data, and use virtual QRA data sets for supervision and fine-tuning to ensure that the model output meets a safe and reliable expected model.
It improves the accuracy and safety of the guide model, enhances the adaptability and generalization ability to different situations, and optimizes user experience and medical efficiency.
Smart Images

Figure CN120561598A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of model training technology, and in particular to a training method, device, storage medium, and program product for a diagnosis guidance model. Background Art
[0002] In the medical field, the intelligent medical guidance system uses large models to obtain the user's main complaint information through natural language interaction, and provides medical advice based on this information. It aims to help users quickly find suitable departments and doctors, improve medical efficiency, reduce the workload of hospital guides, and optimize the user's medical experience.
[0003] However, current guidance models face numerous challenges in practical application. For one thing, users often lack medical knowledge, leading to inaccurate descriptions. This leads to biased information about the patient's complaint, which in turn affects the accuracy of the treatment recommendations it produces. Furthermore, while structured medical data generated during the consultation phase, such as electronic medical records and examination reports, contains a wealth of medical information, this data is not fully utilized to guide the optimization training of pre-diagnosis models. This makes it difficult for guidance models to learn and improve from actual diagnosis and treatment results, limiting their performance. Summary of the Invention
[0004] In view of this, the embodiments of the present disclosure provide a training method, device, storage medium and program product for a medical guidance model, which can use safety constraint rules to introduce clear "safety red line" rules in the reinforcement learning stage of medical guidance model training to ensure the security, credibility, controllability and interpretability of the model output. Compared with the traditional DPO training method that only uses the medical guidance model and the basic model to construct preference data, this solution introduces a collaborative model to construct preference data, which can make the preference data richer and more diverse, allowing the medical guidance model to be exposed to more diversified information during training, and then learn a wider range of human preference patterns.
[0005] In a first aspect, the present disclosure provides a method for training a diagnosis guidance model, which adopts the following technical solutions: Acquire pre-diagnosis conversation data and in-diagnosis data, and extract question data from the pre-diagnosis conversation data; Based on the diagnosis data, obtaining a reference standard answer to the question data; Input the question data into the collaborative model and the basic model respectively, and obtain a first answer content output by the collaborative model and a second answer content output by the basic model; constructing preference data based on the question data, the first answer content, the second answer content, and the reference standard answer; Based on the in-diagnosis data, a virtual QRA data set is obtained; Performing supervised fine-tuning on the basic model based on the virtual QRA dataset to obtain a reference model; Acquire safety constraint rules, and train the diagnosis guidance model based on the preference data, the safety constraint rules, and the reference model to acquire an expected model.
[0006] Optionally, obtaining a reference standard answer to the question data based on the diagnosis data includes: Preprocessing the diagnosis data to obtain key data; Based on the user's identification and the timeline of the consultation, align the problem data, the key data, and the in-consultation data; The data-aligned question data, key data and diagnosis data are input into the teacher's large language model to generate a reference standard answer for each question data.
[0007] Optionally, constructing preference data based on the question data, the first answer content, the second answer content, and the reference standard answer includes: Comparing the first answer content with the second answer content based on the question data and a reference standard answer corresponding to the question data, and determining high-quality answer content and low-quality answer content based on the comparison result; The question data, the high-quality answer content, and the low-quality answer content are combined into preference data.
[0008] Optionally, acquiring a virtual QRA dataset based on the in-diagnosis data includes: The QRA samples were constructed according to disease types; Splitting the diagnosis data to obtain diagnosis and treatment data sets for different disease types; Filling the QRA sample and the diagnosis and treatment data set into a preset template according to the disease type to form multiple prompt messages; Each prompt information is input into the teacher model respectively, and the virtual QRA data output by the teacher model constitutes a virtual QRA data set.
[0009] Optionally, the obtaining of the security constraint rules, training the diagnosis guidance model based on the preference data, the security constraint rules, and the reference model to obtain the expected model, includes: Construct an initially empty preference dataset; Determining current preference data through a traversal operation, and extracting high-quality answer content and low-quality answer content from the current preference data; Determining whether the high-quality reply content complies with security constraint rules; If the high-quality reply content complies with the security constraint rules, storing the current preference data in the preference data set; If the high-quality answer content does not comply with the security constraint rules, then determining whether the low-quality answer content complies with the security constraint rules; If the non-high-quality answer content meets the security constraint rule, the high-quality answer content and the non-high-quality answer content are subjected to preference inversion to form new preference data, and the new preference data is stored in the preference data set; If the non-high-quality reply content does not comply with the security constraint rules, the current preference data is deleted; After all preference data are traversed, the guidance model is trained based on the preference data set and the reference model to obtain the expected model.
[0010] Optionally, the obtaining of the security constraint rules, training the diagnosis guidance model based on the preference data, the security constraint rules, and the reference model to obtain the expected model, includes: Constructing an SC-DPO loss function based on the question data, high-quality answer content and non-high-quality answer content included in the preference data, and the security constraint rules included in the rule list; The guidance model is trained based on the reference model and the SC-DPO loss function to obtain the expected model.
[0011] Optionally, the expression of the SC-DPO loss function is: Where, represents the SC-DPO loss function; Represents the distribution of preference datasets; Represents all preference data samples sampled from distribution D ; Represents problem data; Indicates high-quality reply content; Indicates non-high-quality reply content; Represents the sigmoid function; represents a hyperparameter; Represents the guidance model; Represents the parameters of the guidance model; represents the reference model; represents the parameters of the reference model; express expectations; Indicates that the guidance model is given an input Generate a response under the conditions probability; Indicates that the reference model is given an input Generate a response under the conditions probability; Indicates that the guidance model is given an input Generate a response under the conditions probability; Indicates that the reference model is given an input Generate a response under the conditions probability; represents the safety penalty weight hyperparameter; Represents a list of rules; Represents a security penalty function. If the guidance model is given an input The response generated under the conditions If the security constraint rules in the rule list are not met, then , if the guidance model is given input The response generated under the conditions If the security constraint rules in the rule list are met, .
[0012] In a second aspect, the present disclosure also provides a training system for a diagnosis guidance model, which adopts the following technical solutions: Prepare a data acquisition module for acquiring pre-diagnosis conversation data and in-diagnosis data, and extract question data from the pre-diagnosis conversation data; A standard answer acquisition module, configured to acquire a reference standard answer to the question data based on the diagnosis data; a reply content acquisition module, configured to input the question data into the collaborative model and the basic model respectively, and obtain a first reply content output by the collaborative model and a second reply content output by the basic model; a preference data construction module, configured to construct preference data based on the question data, the first answer content, the second answer content, and the reference standard answer; A virtual data acquisition module, configured to acquire a virtual QRA data set based on the in-diagnosis data; A model supervision fine-tuning module, configured to perform supervised fine-tuning on the basic model based on the virtual QRA dataset to obtain a reference model; The diagnosis guidance model training module is used to obtain safety constraint rules, and train the diagnosis guidance model based on the preference data, the safety constraint rules and the reference model to obtain the expected model.
[0013] In a third aspect, the embodiments of the present disclosure further provide a computer device that adopts the following technical solution: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any of the training methods for the guidance model described above.
[0014] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute any of the above-mentioned training methods for the guidance model.
[0015] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, comprising a computer program / instruction, which implements the steps of any of the above methods when executed by a processor.
[0016] The training method of the guidance model provided by the embodiment of the present disclosure obtains pre-diagnosis conversation data and in-diagnosis data, comprehensively collects information about patients at different stages of the medical process, extracts question data from the pre-diagnosis conversation data, can accurately locate the patient's questions and needs, and provide clear input content for the model, ensuring that the model learns and optimizes for actual patient problems. The reference standard answers to the question data obtained based on the in-diagnosis data provide a reliable benchmark for the training of the guidance model. These standard answers are based on the actual diagnosis process and have high accuracy and authority, so that the guidance model has clear goals in the learning process and can be optimized in the right direction, thereby improving the accuracy of the guidance results. The collaborative model and the basic model are used to process the question data to obtain different response contents. The two models have different characteristics and advantages, and the output content can complement each other, providing a variety of choices for the subsequent construction of preference data, which increases the richness and comprehensiveness of the data. Preference data is constructed by combining question data, responses from different models, and reference standard answers. This data fully reflects the differences and advantages and disadvantages between different responses and the standard answers. This preference data provides rich learning material for the guidance model, enabling it to learn a wider range of human preference patterns during training, improving the model's adaptability and generalization to different situations. A virtual QRA dataset, derived from in-consultation data, further enriches the type and scale of training data. The virtual QRA dataset simulates a variety of possible pre-consultation and in-consultation scenarios, providing more learning samples for the base model, helping it better understand and handle complex medical problems and improving its performance and accuracy. Supervised fine-tuning of the base model using the virtual QRA dataset allows it to better adapt to the needs of medical guidance scenarios. Through fine-tuning, the reference model learns the patterns and regularities in the virtual dataset, providing more accurate reference and guidance for subsequent guidance model training. Safety constraints are captured and incorporated into guidance model training, setting clear boundaries and specifications for the guidance model's output, ensuring the security and reliability of guidance results. Combining preference data and reference models to train the guidance model enables the guidance model to learn human preferences while following safety rules, thereby obtaining an expected model that is both in line with user preferences and safe and reliable, improving the quality and efficiency of guidance and better meeting user needs.
[0017] The above description is only an overview of the technical solution of the present disclosure. In order to more clearly understand the technical means of the present disclosure, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the following specifically cites preferred embodiments and describes them in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 A flowchart of a training method for a diagnosis guidance model provided in an embodiment of the present disclosure; Figure 2 A flowchart of a method for generating a reference standard answer provided in an embodiment of the present disclosure; Figure 3 A flowchart of a method for constructing preference data provided in an embodiment of the present disclosure; Figure 4 A flowchart of a method for acquiring a virtual QRA dataset provided by an embodiment of the present disclosure; Figure 5 A flowchart of a method for obtaining security constraint rules provided in an embodiment of the present disclosure; Figure 6 A flowchart of a diagnosis guidance model training method provided in an embodiment of the present disclosure; Figure 7 Another flowchart of the diagnosis guidance model training method provided by the embodiment of the present disclosure; Figure 8 A block diagram of the principle of the training system for the medical guidance model provided in the embodiment of the present disclosure; Figure 9 A schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0020] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0021] It should be clear that the following embodiments of the present disclosure are described through specific concrete examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other in the absence of conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.
[0022] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this device and / or practice this method.
[0023] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The illustrations only show components related to the present disclosure and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.
[0024] Additionally, in the following description, specific details are provided to provide a thorough understanding of the examples. However, one skilled in the art will appreciate that the aspects described can be practiced without these specific details.
[0025] Reference Figure 1 The present disclosure provides a training method for a diagnosis guidance model, comprising the following steps: S1: Obtain pre-diagnosis conversation data and in-diagnosis data, and extract question data from the pre-diagnosis conversation data; S2: Based on the in-diagnosis data, obtain the reference standard answers to the question data; S3: Input the question data into the collaborative model and the basic model respectively, and obtain the first answer content output by the collaborative model and the second answer content output by the basic model; S4: Construct preference data based on question data, first answer content, second answer content and reference standard answers; S5: Based on the in-diagnosis data, a virtual QRA dataset is obtained; S6: Perform supervised fine-tuning on the base model based on the virtual QRA dataset to obtain a reference model; S7: Obtain security constraint rules, train the guidance model based on preference data, security constraint rules and reference model, and obtain the expected model.
[0026] The training method of the guidance model provided by the present invention comprehensively collects information about patients at different stages of the medical process by acquiring pre-diagnosis conversation data and in-diagnosis data, extracts question data from the pre-diagnosis conversation data, and can accurately locate the patient's questions and needs, provide clear input content for the model, and ensure that the model learns and optimizes for actual patient problems. The reference standard answers to the question data obtained based on the in-diagnosis data provide a reliable benchmark for the training of the guidance model. These standard answers are based on the actual diagnosis process and have high accuracy and authority, so that the guidance model has clear goals in the learning process and can be optimized in the right direction, thereby improving the accuracy of the guidance results. The collaborative model and the basic model are used to process the question data to obtain different response contents. The two models have different characteristics and advantages, and the output content can complement each other, providing a variety of choices for the subsequent construction of preference data, which increases the richness and comprehensiveness of the data. Preference data is constructed by integrating question data, response content from different models, and reference standard answers, which can fully reflect the differences and advantages and disadvantages between different responses and standard answers. These preference data provide rich learning materials for the guidance model, enabling it to learn a wider range of human preference patterns during training, thereby improving the model's adaptability and generalization capabilities to different situations.
[0027] The virtual QRA dataset, derived from in-diagnosis data, further enriches the type and scale of training data. The virtual QRA dataset can simulate a variety of possible pre-diagnosis and in-diagnosis scenarios, providing more learning samples for the underlying model. This helps the underlying model better understand and handle complex medical issues, improving its performance and accuracy. Using the virtual QRA dataset for supervised fine-tuning of the underlying model allows it to better adapt to the needs of medical guidance scenarios. Through fine-tuning, the reference model can learn the patterns and regularities in the virtual dataset, providing more accurate reference and guidance for subsequent training of the guidance model.
[0028] Obtaining safety constraints and incorporating them into guidance model training sets clear boundaries and specifications for the guidance model's output, ensuring the security and reliability of guidance results. Combining preference data with a reference model for guidance model training enables the guidance model to learn from human preferences while adhering to safety rules. This results in a desired model that is both user-friendly and secure, improving guidance quality and efficiency and better meeting user needs.
[0029] In S1, pre-diagnosis conversation data is extracted from the natural language interaction data between users and the pre-diagnosis large model. This data includes natural language question-and-answer pairs formed by users' proactive descriptions of their chief complaints, symptom questions and answers, medical history, and personal information. Because pre-diagnosis conversation data is relatively free-form, contains preliminary information, and reflects the user's perspective, it is necessary to pre-process and analyze these question-and-answer pairs using natural language processing technologies, such as text parsing and semantic understanding algorithms, to extract the question data, such as symptom descriptions and medical history-related questions raised by users.
[0030] Acquire intra-consultation data from hospital information systems. This data exists in structured or semi-structured formats and includes electronic medical records, physician-written medical history notes, laboratory test reports, and diagnostic conclusions. This data is characterized by accuracy, complete structure, and expert perspective. Pre-consultation conversation data and intra-consultation data are integrated, and the accuracy and completeness of the inter-consultation data are further verified and supplemented by comparing the inter-consultation data with the diagnostic conclusions, medical orders, and other information contained in the inter-consultation data.
[0031] To ensure data security, user identity information (such as name, etc.) in pre-diagnosis conversation data and in-diagnosis data is replaced with a unique identifier automatically generated by the system. This identifier is only used for internal system statistics and data association and does not contain any personally identifiable information, thereby effectively avoiding security risks.
[0032] Reference Figure 2 The flowchart of the reference standard answer generation method shown in the figure, "Obtaining reference standard answers to question data based on in-diagnosis data," includes the following steps: S21: Preprocess the diagnosis data to obtain key data; S22: Based on the user's ID and consultation timeline, align the problem data, key data, and in-consultation data; S23: Input the data-aligned question data, key data, and diagnosis data into the teacher's large language model to generate a reference standard answer for each question data.
[0033] In S21, preprocessing includes initial data acquisition, data cleaning and standardization, where the initial data includes chief complaint, current medical history, past medical history, examination results and medical advice; data cleaning and standardization refers to removing redundant information, unifying medical terminology and processing missing values.
[0034] In the diagnosis data, we use text matching to search for specific identifiers, such as "chief complaint." Following the identifier, we extract a continuous section of text as the chief complaint. For example, if the text contains "chief complaint: cough for 3 days," "cough for 3 days" will be extracted as the chief complaint. For complex text structures, we can combine regular expressions for precise matching to ensure accurate extraction of the chief complaint information.
[0035] Using the same method of word matching, search for the "Current Medical History" tag. The extracted text should then describe the symptom progression in detail, including the time of onset, changes in symptoms, and factors that alleviate or aggravate them. For example, "Current Medical History: The user developed a cough without apparent cause three days ago. Initially, it was a single cough that was ignored. However, the cough has worsened over the past two days and is accompanied by sputum production." This entire segment should be extracted as the current medical history information.
[0036] Medical history includes various aspects of information, such as allergy history and surgical history. We search for the "medical history" tag in the medical record text and further identify different types of medical history information. For allergy history, if the text mentions "Medical history: history of penicillin allergy," we extract "penicillin allergy history" as the allergy history information. For surgical history, such as "Previous appendectomy," we extract "appendectomy" as the surgical history information.
[0037] Examination results include imaging reports, laboratory test results, and other content. Search for identifiers such as "examination results" and "examination" to extract relevant information. An imaging report may include content such as "Chest X-ray shows thickened lung markings," while a laboratory test result may include "Routine blood test indicates elevated white blood cell count." Accurately extract this information.
[0038] Medical order information typically includes medication regimens and examination recommendations. Retrieve relevant content by searching for the "Medical Order" tag. A medication regimen might include "Oral amoxicillin capsules, 0.5g three times daily," while an examination recommendation might include "Further chest CT scan recommended." Extract and organize these contents separately.
[0039] Redundant information removal involves examining the extracted key information and removing redundant content such as duplicate records and process notes. For example, if the same symptom description or procedure record appears repeatedly in different locations, only one copy will be retained. Furthermore, irrelevant procedural text, such as the doctor's thought process and notes during the recording process, will be deleted to ensure data conciseness.
[0040] Unifying medical terminology involves creating a standardized dictionary of medical terms, unifying different medical terms into standard terms. For example, "hypertension" can be converted to "essential hypertension," and "diabetes" can be standardized into "type 2 diabetes." By traversing the extracted key information, identifying non-standard terms, and replacing them according to the dictionary, we ensure consistency and accuracy in medical terminology.
[0041] Handling missing values involves uniformly marking missing information. For example, if allergy history information is missing, it can be marked as "not recorded." If part of a test result is missing, it can also be marked accordingly, such as "no relevant test result record." This facilitates subsequent data statistics and analysis, avoiding errors caused by missing values.
[0042] Store preprocessed and standardized key data in an appropriate storage medium, such as a database or a new CSV file. The stored data should be clearly structured to facilitate subsequent querying, analysis, and use. Different types of key information can be stored in separate fields or tables to improve data organization and manageability.
[0043] In S22, the problem data, key data and diagnosis data all contain the user's unique identification information. By comparing the user identification, the problem data, key data and diagnosis data belonging to the same user are associated to form a preliminary data set with the user as the core.
[0044] After completing the association based on the unique identifier, the key element of the medical consultation timeline is introduced. The specific time when the user asks the question is recorded in the question data, and the consultation start time, examination time, diagnosis time and other information are obtained in the key data and the in-diagnosis data. The question data, key data and in-diagnosis data are sorted in chronological order, and the question data is preliminarily matched with the key data and the in-diagnosis data within the corresponding time period. For example, if the user asks a question about a symptom at a certain moment before the diagnosis, the relevant diagnosis, examination and other information before and after that moment is searched in the key data and the in-diagnosis data, so that the data can be preliminarily matched in the time dimension.
[0045] Based on the initial time matching, entity extraction is performed on the question data, key data, and in-diagnosis data. Key entity information such as symptoms, diseases, and body parts is identified from the question data. For example, for the question "I've had a severe headache these past few days. What's going on?", the entity extracted is "headache." The same entity extraction operation is performed on the key data and in-diagnosis data. For example, if the in-diagnosis report mentions "migraine causing head pain," entities such as "migraine" and "head pain" are extracted. Through entity extraction, core information is extracted from the data, preparing for subsequent semantic analysis.
[0046] Semantic similarity analysis is used to calculate the similarity between entities extracted from the problem data and entities extracted from the key data and the diagnosis data. For example, "headache" and "headache" have a high semantic similarity and can be determined to be strongly associated. While "headache" and "migraine" are not exactly the same, they are closely related in the medical context. Based on the semantic similarity results, the matching relationship between the problem data and the key data and the diagnosis data is further adjusted and determined. For data associated with entities with high similarity, they are more accurately bound together to enhance the accuracy of data alignment and ensure that the problem data, key data, and diagnosis data are highly consistent at the semantic level.
[0047] Although the correspondence between question data, key data, and in-diagnosis data is established through unique identification, timeline matching, entity extraction, and semantic similarity analysis, some inaccuracies or ambiguities may exist. Therefore, manual review is performed, with professional medical personnel or data auditors reviewing the automated matching results to determine whether the association between question data, key data, and in-diagnosis data is reasonable, and whether there are any mismatches or missed associations. Any issues discovered are promptly corrected to ensure the accuracy of data alignment and provide a reliable data foundation for the subsequent generation of reference standard answers.
[0048] In S23, the aligned question data, key data and diagnosis data are sorted and organized according to the input format required by the teacher's large language model, that is, the question data is input as the question part, the key data is input as important prompt information, and the diagnosis data is input as background knowledge into the model at the same time. After receiving the input data, the teacher's large language model uses its pre-trained knowledge and algorithms to reason, comprehensively considering the important information provided by the question data and key data and the background knowledge of the diagnosis data to generate a reference standard answer for each question data. The LLM language model is used to summarize the reference standard answers output by the teacher's large language model and convert them into clear and standardized medical expressions. By removing some colloquial and vague expressions, the reference standard answers are made more professional and accurate. For example, some popular symptom descriptions are converted into standard medical terms, and unclear medical instructions are converted into standardized medication plans and dosage instructions, so that the final reference standard answers meet the standards and requirements of the medical field.
[0049] The obtained reference standard answers are reviewed by reviewers who possess sufficient medical knowledge to determine the rationality and accuracy of the answers, verifying whether the answers conform to medical logic and are consistent with the in-diagnosis data and key data. Any answers that do not meet the requirements are corrected to improve the quality of the generated reference standard answers.
[0050] In S3, each question is fed into the collaborative model and the basic model. The two models analyze and output corresponding responses. The response output by the collaborative model is defined as the first response, and the response output by the basic model is defined as the second response. Deepseek-R1 can be used as the collaborative model, and Qwen3 or Qwen / Qwen1.5 - 7B - Chat can be used as the basic model.
[0051] In S4, refer to Figure 3 The flowchart of the preference data construction method shown in the figure, "Constructing preference data based on question data, first answer content, second answer content, and reference standard answer," includes the following steps: S41: Based on the question data and the reference standard answer corresponding to the question data, the first answer content and the second answer content are compared, and based on the comparison result, the high-quality answer content and the low-quality answer content are determined; S42: Combine the question data, the high-quality answer content, and the low-quality answer content into preference data.
[0052] In the above steps, the user's question data before diagnosis is collected, for example: "Which clinic should I go to if I have a headache, fever, and sore throat?"; the reference standard answer corresponding to the question data is obtained, for example: "Otolaryngology"; and the response content generated by the collaborative model and the basic model for the question data is obtained, for example: the first response content output by the collaborative model: "Otolaryngology", and the second response content output by the basic model: "Internal Medicine".
[0053] The responses generated by the collaborative model and the base model are compared for consistency with the reference standard answer. Each response is evaluated for accuracy, completeness, and logic to determine which response is closer to the reference standard answer. For example, if the first response is completely consistent with the reference standard answer, while the second response is relevant but less accurate, the first response generated by the collaborative model is superior to the second response generated by the base model.
[0054] Based on the comparative analysis results, a comparative answer sample is constructed in the following format: completion_a: Model b generates the result, e.g.: "Otolaryngology" completion_b: Model b generates the result, e.g.: "Internal Medicine" preference: a(chosen)>b(rejected) Among them, a refers to the collaborative model; b refers to the basic model.
[0055] The comparative answer sample has accurately specified the quality attributes of the first answer content and the second answer content, so it is possible to extract high-quality answer content and low-quality answer content from the comparative answer sample, and combine the question data, high-quality answer content and low-quality answer content into preference data.
[0056] In the above method, preference data is constructed by combining question data, reference standard answers, and the output results of the collaborative model and the basic model. This can accurately capture the user's preference patterns in scenarios such as medical consultation. These preference data can enable the model to provide answers that are more in line with user needs in subsequent interactions, greatly improving the user experience and the accuracy and effectiveness of model services.
[0057] In S5, refer to Figure 4 The flowchart of the virtual QRA dataset acquisition method is shown in the figure. "Acquiring a virtual QRA dataset based on in-diagnosis data" includes the following steps: S51: Construct QRA samples according to disease types; S52: Split the diagnosis data to obtain diagnosis and treatment data sets of different disease types; S53: According to the disease type, the QRA sample and the diagnosis and treatment data set are filled into the preset template to form multiple prompt information; S54: Input each prompt information into the teacher model respectively, and form a virtual QRA data set with the virtual QRA data output by the teacher model.
[0058] Among them, Q refers to Question, which is used to describe the user's symptoms and intentions, etc.; R refers to the analytical reasoning process (Reasoning). The detailed analytical reasoning process includes but is not limited to: preliminary judgment of symptoms, a list of possible causes or diseases, the logic of excluding or confirming certain diseases, consideration of differential diagnosis, and the basis for the final conclusion; A refers to Conclusion (Action), that is, the response content, which gives the final conclusion based on the reasoning process, such as the recommended department, possible diagnostic direction, or preliminary suggestions.
[0059] Based on established disease classification standards, we collect common problem patterns, analysis and derivation process patterns, and conclusion patterns associated with each disease and integrate them into QRA samples for each disease. Based on the diagnosis conclusion field in the in-diagnosis data, we categorize and split the in-diagnosis data by disease type. For example, we group all in-diagnosis data from users diagnosed with "diabetes" into one category, and data from users diagnosed with "hypertension" into another, and so on, to form a collection of diagnosis and treatment data for different disease types.
[0060] The QRA samples and diagnosis and treatment data sets belonging to the same disease category are filled into the preset template to form prompt information for different disease categories. Taking the prompt information of respiratory medicine as an example, the following is displayed: You are now an experienced chief physician in respiratory medicine and a skilled medical instructor. Your task is to construct a SFT sample for training the AI guidance model based on the structured medical record information provided below.
[0061] This sample needs to include three parts: a "question" that simulates the user's tone, a detailed "analysis process", and a clear "final conclusion".
[0062] Please strictly follow the following requirements: 1. Simulate user questions: Rewrite the "chief complaint" and "current medical history" in the medical record into natural, colloquial user questions.
[0063] 2. Reasoning It must be in a structured, point-by-point format.
[0064] The first step is "Symptom Breakdown and Preliminary Assessment", summarizing key positive and negative signs.
[0065] The second step is "differential diagnosis", which lists at least 2-3 possible diseases and explains the reasons for supporting or excluding them.
[0066] The third step is "information synthesis and decision-making". Based on the above analysis, the most likely diagnostic direction and the reasons for choosing the treatment department are obtained.
[0067] The entire process must be logically rigorous and reflect clinical thinking.
[0068] 3. Final Conclusion (Answer): Based on the analysis process, give a clear and easy-to-understand medical advice (such as a recommended department).
[0069] Please return your output strictly in the following JSON format without any extra text.
[0070] { "question": "(the simulated user question you generated)", "reasoning": "(the analysis process you generated)", "answer": "(the final conclusion you generated)" }
Input medical data set
[0071] In S6, the optimization objective is to minimize a token-level loss function, which measures the difference between the base model's predictions and the true labels, thereby improving model accuracy. AdamW is selected as the optimizer, which adaptively adjusts the learning rate of each parameter, helping the base model converge more stably and quickly. During training, a warm-up and cosine decay parameter tuning strategy is employed, gradually increasing the learning rate early in training to stabilize learning and then gradually decreasing it later, allowing for more refined parameter tuning of the base model. The virtual QRA dataset is used as training data, and a high-quality sft_qra_data.jsonl file is generated and loaded. This file contains the organized and processed training data for subsequent supervised fine-tuning. The base model and corresponding tokenizer are loaded from the pretrained model using relevant tools. Using a dedicated dataset loading tool, the sft_qra_data.jsonl file is loaded as a JSON dataset for base model training. Determine a series of training parameters, including the output directory, the training batch size for each device, the number of gradient accumulation steps, the learning rate, the number of training epochs, the number of logging steps, and the number of saved steps. If supported by the hardware, enable mixed-precision training to improve training efficiency. Define a formatter function to convert a JSON object (referring to a JSON dataset) into a conversational string formatted for the base model. This function processes each piece of training data during SFT (supervised fine-tuning) training to ensure that the training data meets the input requirements of the base model. Finally, initialize SFTTrainer, passing in the base model, training parameters, training dataset, tokenizer, and formatter, along with a maximum sequence length, and then start training. During training, the base model will use the AdamW optimizer and the warm-up + cosine decay parameter tuning strategy to gradually optimize its parameters by minimizing the token-level loss function, ultimately obtaining a reference model. These steps effectively perform supervised fine-tuning on the base model, improving its performance and accuracy on specific tasks.
[0072] In S7, refer to Figure 5 The flowchart of the method for obtaining security constraint rules is shown. "Obtaining security constraint rules" includes the following steps: S71: Obtaining initial rules input by medical experts in different fields, and determining triggers, constraint types, and matching methods that match the initial rules; S72: Based on the trigger, constraint type, and matching method, the initial rule is converted into an initial security constraint rule; S73: Review and verify the initial security constraint rules to obtain the security constraint rules.
[0073] In S71, medical experts from different fields, such as cardiology, neurology, and pediatrics, are organized to hold seminars or use online questionnaires to collect initial rules summarized by them based on their professional knowledge and clinical experience. These rules usually focus on common high-risk medical scenarios, misdiagnosis risks, safe use of medications, etc., such as identifying symptoms such as acute myocardial infarction, stroke, and high fever convulsions in children, and avoiding giving incorrect treatment recommendations.
[0074] If the initial rule involves complex semantic understanding, such as determining whether the user input belongs to a specific high-risk disease scenario, such as the example rule "SR-001" below, which detects symptom combinations highly associated with acute myocardial infarction, you should choose the semantic_classifier trigger. This is because this type of trigger can process semantic information and more accurately identify risks using a trained model. The example rule "SR-001" is as follows: rule_id: "SR-001" name: "Suspected heart attack symptoms must be directed to the emergency room" description: "Detects symptom combinations in user input that are highly correlated with acute myocardial infarction and ensures that the model's response includes mandatory urgent medical attention instructions." trigger: type: "semantic_classifier" # trigger type model_path: "models / classifiers / cardiac_emergency_detector.pt" # Trigger model path threshold: 0.95 # trigger activation threshold constraint: type: "MUST_INCLUDE" # constraint type elements: # Constraint elements - "Call emergency number immediately" - "Go to the emergency room immediately" - "Don't go there alone" match_method: "semantic" # element matching method For situations where the initial rule trigger condition is clearly a simple keyword, such as the following example rule "SR-002" that detects high-risk symptoms such as "coma" and "convulsions," the keyword_list trigger is suitable. It is simple and direct and can quickly match. The example rule "SR-002" is as follows: rule_id: "SR-002" name: "Recommendation of observing high-risk symptoms at home is prohibited" description: "When any symptom defined as 'high risk' is detected, the model is prohibited from recommending non-active interventions such as 'stay at home' or 'drink more water'." trigger: type: "keyword_list" keywords: ["coma", "convulsion", "blindness", "respiratory arrest"] constraint: type: "MUST_NOT_CONTAIN" elements: - "Home Observation" - "Get more rest" - "Handle on your own" - "Let's see in a few days." match_method: "fuzzy" # fuzzy matching If the triggering condition for the initial rule is neither suitable for simple keyword matching nor requires complex semantic classification, but instead requires the identification of semantically similar expressions, the embedding_similarity trigger can be used. This trigger, as a technical solution between keyword matching and semantic classifiers, offers unique advantages. For example, a pre-trained word embedding model (such as Word2Vec or BERT) is used to convert the initial rule trigger condition (such as "chest squeezing pain") into a vector embedding. After receiving user input, the same model is used to convert it into a vector. The cosine similarity between this vector and the trigger condition vector is calculated. A similarity threshold is set, and if the calculated result exceeds the threshold, the rule is triggered. This trigger can achieve a balance between semantic understanding and matching efficiency.
[0075] When the initial rule's goal is to ensure that the model-generated content includes specific key information, select the MUST_INCLUDE constraint type. In this case, the key information required by the initial rule must be clearly defined and listed as an elements list. The model's completion (generated content) must include at least one element from this list. For example, in the example rule "SR-001," the initial rule, "Suspected MI symptoms must be directed to the emergency room," aims to ensure that the model's response includes urgent medical instructions. Therefore, the constraint type is MUST_INCLUDE, and "Call 911 immediately," "Go to the emergency room immediately," and "Do not go alone" are listed as an elements list. This ensures that the model-generated content must include at least one element from this elements list.
[0076] If the initial rule prohibits the model from generating certain content, select the MUST_NOT_CONTAIN constraint type. List the prohibited content in the elements list to ensure that the model's completion does not include any elements from this list. For example, in the example rule "SR-002," the initial rule "Do not recommend home observation for high-risk symptoms" aims to prevent the model from recommending no proactive intervention. Therefore, select the MUST_NOT_CONTAIN constraint type and list "Home observation," "Get plenty of rest," "Treat yourself," and "Wait a few days" as the elements list to ensure that the model-generated content does not include any elements from this list.
[0077] If the initial rules have high semantic requirements, that is, they must accurately match semantics, such as in the example rule "SR-001" that requires the model's response to contain content semantically related to emergency medical instructions, choose semantic similarity matching to ensure that the model output content meets the semantic requirements. If the initial rules allow for a certain degree of flexibility and do not require exact matching, such as in the example rule "SR-002" that prohibits recommendations such as "stay at home for observation" that do not actively intervene, use fuzzy matching to handle differences in expression. If the initial rules require high matching accuracy, choose exact matching.
[0078] In S72, assign a unique rule_id (rule number) to the initial rule, name it, and add a description to clearly explain its function and purpose, such as "SR-001" in the example rule, "Suspected myocardial infarction symptoms must be directed to the emergency department." Configure trigger parameters. For the semantic_classifier trigger, specify the model path (model_path) and activation threshold (threshold), as configured in the example rule "SR-001." For the keyword_list trigger, list the keywords (keywords), as configured in the example rule "SR-002." Based on the constraint type, enter the constraint elements (elements) and match method (match_method). For the MUST_INCLUDE constraint type, list the elements that must be included; for the MUST_NOT_CONTAIN constraint type, list the elements that cannot be included, and specify whether the match method is "semantic," "fuzzy," or "exact." Based on the above, the rule number, initial rule name, and description assigned to the initial rule are integrated with the configured trigger parameters (model path, activation threshold, or keyword list set according to the trigger type), as well as the constraint elements and matching methods filled in according to the constraint type to form a security constraint rule. This rule can effectively regulate the output content of the model to ensure that it meets specific security and business requirements.
[0079] In S73, medical experts and technical personnel will review the initial safety constraint rules. Medical experts will examine the logic of the rules from a professional perspective, verifying that the trigger conditions and constraints are consistent with medical knowledge and clinical practice. Technical personnel will verify the correct configuration of the rules and the reasonableness of the trigger and constraint parameters. The rules will be validated using a test dataset that includes a variety of possible user inputs and model-generated content, simulating real-world scenarios. For each initial safety constraint rule, the triggers will be checked for correct triggering and the constraints will be checked for effective model output. For example, for rule "SR-001," the trigger will be verified to activate when the input contains symptoms associated with acute myocardial infarction, and the model response will include the required emergency medical instructions. For rule "SR-002," the model output will be verified to exclude prohibited content when the input contains high-risk symptoms. Based on the review and validation results, the initial safety constraint rules will be adjusted and optimized. If trigger misjudgments or omissions are detected, the trigger parameters, such as the threshold or keyword list, will be adjusted. If the constraints fail to effectively constrain the model output, the constraint elements or matching methods will be adjusted. After multiple adjustments and verifications, until the rules achieve satisfactory results, the security constraint rules are finally obtained and stored in the rule list.
[0080] After obtaining the safety constraint rules, DPO training is performed according to the safety constraint rules. The present disclosure provides two possible implementation plans. In one implementation plan, refer to Figure 6 The flowchart of the guidance model training method is shown. "Training the guidance model based on preference data, safety constraints, and reference models to obtain the desired model" includes the following steps: S74: construct an initially empty preference dataset; S75: determining current preference data through a traversal operation, and extracting high-quality answer content and low-quality answer content from the current preference data; S76: Determine whether the high-quality reply content complies with the security constraint rules; if so, execute S77; if not, execute S78; S77: storing the current preference data into the preference data set; S78: Determine whether the non-quality reply content complies with the security constraint rules; if so, execute S79; if not; execute S710; S79: performing preference reversal on the high-quality answer content and the low-quality answer content to form new preference data, and storing the new preference data in the preference data set; S710: Delete the current preference data; S711: Determine whether all preference data have been traversed; if so, execute S712; if not, return to S75 to determine the new current preference data; S712: Based on the preference data set and the reference model, the guidance model is trained to obtain the expected model.
[0081] In the above content, all preference data is traversed, and the preference data currently being traversed is the current preference data. High-quality and low-quality response content is extracted from it. By matching the high-quality and low-quality response content with security constraints, the processing method for the current preference data can be determined. If the high-quality response content meets the security constraints, the current preference data has training value and is retained. If the high-quality response content does not meet the security constraints, while the low-quality response content does, it indicates that the quality attributes of the high-quality and low-quality response content are incorrectly defined. The original high-quality response content is corrected to low-quality response content, and the original low-quality response content is corrected to high-quality response content. Through this preference reversal, new preference data is formed and saved. If both the high-quality and low-quality response content do not meet the security constraints, the current preference data has no training value and is deleted.
[0082] The method for matching high-quality and low-quality responses with security constraints is as follows: A rule matching function, check_safety, is constructed to check whether the response content (including high-quality and low-quality responses) complies with the security constraint rules. An empty list, violated_rules, is created to store the rule numbers of the security constraint rules that the response content does not comply with. Each security constraint rule in the rule list is then iterated over. The trigger in the security constraint rule is used to determine whether the response content meets the trigger condition. If so, the constraint condition is further determined. If so, the response content is deemed to comply with the security constraint rule, and the traversal continues to the next security constraint rule. If not, the response content is deemed to comply with the security constraint rule, and the rule number of that security constraint rule is stored in violated_rules before the traversal continues. If not, the response content is deemed to have not triggered the security constraint rule, indicating that the response content is safe and complies with the security constraint rule, and the traversal continues directly to the next security constraint rule. After the traversal, the violated_rules list is returned, and the violated_rules list can be used to identify the security constraint rule that the response content does not comply with.
[0083] The guidance model is used as the initial model, and the reference model is loaded. The reference model provides a reference standard and comparison basis for training. The loss function is defined according to the DPO method. This loss function is called the DPO loss function, which is used to measure the consistency between the guidance model output and the preference data. During training, the guidance model generates the third reply content based on the question data in the preference data, and compares the third reply content with the high-quality reply content and non-high-quality reply content in the preference data. The DPO loss function is calculated, and the optimizer updates the parameters of the guidance model according to the DPO loss function. Repeat the training and update steps until the guidance model converges or reaches the preset number of rounds. Use the validation set to evaluate the performance of the trained guidance model. After completion, the expected model is obtained. Among them, the calculation formula of the DPO loss function is: Where, represents the DPO loss function; Represents the guidance model; Represents the parameters of the guidance model; represents the reference model; represents the parameters of the reference model; express expectations; Represents the distribution of preference datasets; Represents all preference data samples sampled from distribution D ; Represents problem data; Indicates high-quality reply content; Indicates non-high-quality reply content; represents the sigmoid function, , Indicates the intermediate quantity, , represents the natural constant, which is the base of natural logarithms; Represents a hyperparameter, specifically a hyperparameter that controls the amplification factor of the logit scale term, which is used to adjust the slope of the sigmoid function; Indicates that the guidance model is given an input Generate a response under the conditions probability; Indicates that the reference model is given an input Generate a response under the conditions probability; Indicates that the guidance model is given an input Generate a response under the conditions probability; Indicates that the reference model is given an input Generate a response under the conditions probability.
[0084] This method modifies preference data through security constraint rules and then uses the modified preference data to perform DPO training on the guidance model. This not only improves the training effect, but also is simple to implement as this static processing method does not intrude or modify the core code of the DPO training framework (Transformers, etc.).
[0085] In another embodiment, referring to Figure 7 Another flowchart of the guidance model training method is shown. "Training the guidance model based on preference data, safety constraints, and a reference model to obtain the desired model" includes the following steps: S713: Constructing an SC-DPO loss function based on the question data, high-quality answer content and low-quality answer content included in the preference data, and the security constraint rules included in the rule list; S714: Train the guidance model based on the reference model and the SC-DPO loss function to obtain the expected model.
[0086] In the above steps, preference data is loaded, and the guidance model (i.e., policy model) and reference model forward propagate these answers to obtain their logits (raw scores output by the model). These logits are then used to calculate the Determined Point-of-View (DPO) loss function, which aims to optimize the guidance model so that it can more effectively distinguish between selected and rejected answers. When calculating the DPO loss function, a penalty term for safety constraints is introduced. This modification combines the standard DPO loss with the safety penalty to form the SC-DPO loss function, resulting in a more accurate total loss. Based on this total loss, backpropagation is performed to calculate the gradient, and then the optimizer is used to update the model weights. This process ensures that the guidance model not only considers the preferences of the answers during optimization, but also strictly adheres to safety constraints.
[0087] Among them, the calculation formula of SC-DPO loss function is: Where, represents the SC-DPO loss function; Represents the safety penalty weight hyperparameter, which is used to control the intensity of the safety penalty; Represents a list of rules; Represents a security penalty function. If the guidance model is given an input The response generated under the conditions If the security constraint rules in the rule list are not met, then , if the guidance model is given input The response generated under the conditions If the security constraint rules in the rule list are met, .
[0088] This method significantly improves the traditional DPO loss function by introducing a penalty term for safety constraints, resulting in a more comprehensive SC-DPO loss function. This improvement not only encourages the guidance model to understand the importance of safety through learning data but also provides a clear and enforced "safety brake" for the guidance model through a manually adjustable external constraint system. Guided by the SC-DPO loss function, the guidance model's optimization direction shifts from simply "becoming more human-like" to "pursuing optimal human preferences while ensuring safety." This shift ensures that the guidance model strictly adheres to safety red lines when generating responses, while also encouraging it to pursue optimal human preferences while adhering to key safety rules. This approach significantly improves the safety and reliability of the guidance model in practical applications, enabling direct and precise control over the AI's behavioral boundaries, thereby ensuring safety while optimizing the user experience.
[0089] Existing technologies usually only involve a reference model and a guidance model for DPO training. However, this solution also introduces a collaborative model. The collaborative model replaces the guidance model and jointly constructs preference data with the reference model before supervised fine-tuning. By using these preference data in the reinforcement learning process, the guidance model can learn a wider range of human preference patterns during training. This application breaks through the limitations of existing technologies. In reinforcement learning, preference data serves as important guiding information. When the guidance model makes a decision, it will compare it with the preference data. If the decision is in line with the human preferences reflected in the preference data, it will be given a positive reward; otherwise, it will be given negative feedback. With this reward and punishment mechanism, the guidance model can gradually adjust its own strategy and continuously learn and adapt to a wider range of human preference patterns. With the help of the preference data constructed by the reference model and the collaborative model before supervised fine-tuning, the guidance model can be exposed to more diverse preference information during training, thereby improving its generalization ability and adaptability. It can better meet the guidance needs of various users and provide more accurate and user-friendly guidance suggestions in actual applications, thereby improving the quality and efficiency of guidance.
[0090] Reference Figure 8 The present disclosure provides a training system for a medical guidance model, comprising: Prepare a data acquisition module 101 for acquiring pre-diagnosis conversation data and in-diagnosis data, and extract question data from the pre-diagnosis conversation data; The standard answer acquisition module 102 is used to obtain reference standard answers to question data based on the diagnosis data; The answer content acquisition module 103 is used to input the question data into the collaborative model and the basic model respectively, and obtain the first answer content output by the collaborative model and the second answer content output by the basic model; A preference data construction module 104 is used to construct preference data based on the question data, the first answer content, the second answer content and the reference standard answer; A virtual data acquisition module 105 is used to acquire a virtual QRA data set based on the in-diagnosis data; The model supervision fine-tuning module 106 is used to perform supervised fine-tuning on the basic model based on the virtual QRA dataset to obtain a reference model; The diagnosis guidance model training module 107 is used to obtain safety constraint rules, and train the diagnosis guidance model based on the preference data, safety constraint rules and reference model to obtain the expected model.
[0091] The various variations and specific examples of the training method of the medical guidance model provided above are also applicable to the training system of the medical guidance model provided in the present disclosure. Through the above detailed description of the training method of the medical guidance model, those skilled in the art can clearly know the implementation method of the training system of the medical guidance model. For the sake of brevity of the specification, it will not be described in detail here.
[0092] A computer device according to an embodiment of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc.
[0093] The processor may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In one embodiment of the present disclosure, the processor is configured to execute the computer-readable instructions stored in the memory, causing the computer device to execute all or part of the steps of the aforementioned method for training a diagnosis guidance model in each embodiment of the present disclosure.
[0094] Those skilled in the art should understand that in order to solve the technical problem of how to obtain a good user experience, this embodiment may also include well-known structures such as a communication bus and an interface, and these well-known structures should also be included in the scope of protection of this disclosure.
[0095] like Figure 9The present invention provides a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Figure 9 The computer device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0096] like Figure 9 As shown, a computer device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) or programs loaded from a storage device into a random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0097] Typically, the following devices can be connected to the I / O interface: input devices such as sensors or visual information acquisition devices; output devices such as display screens; storage devices such as tapes and hard disks; and communication devices. The communication device can allow the computer device to communicate with other devices (such as edge computing devices) wirelessly or by wire to exchange data. Figure 9 A computer device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0098] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the training method of the guidance model of the embodiment of the present disclosure are executed.
[0099] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.
[0100] According to the computer-readable storage medium of the embodiment of the present disclosure, non-transitory computer-readable instructions are stored thereon. When the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the training method of the diagnosis guidance model of each embodiment of the present disclosure are executed.
[0101] The above-mentioned computer-readable storage media include, but are not limited to, optical storage media (e.g., CD-ROMs and DVDs), magneto-optical storage media (e.g., MOs), magnetic storage media (e.g., magnetic tapes or mobile hard disks), media with built-in rewritable non-volatile memory (e.g., memory cards), and media with built-in ROM (e.g., ROM cartridges).
[0102] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.
[0103] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0104] In the present disclosure, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. The block diagrams of the devices, devices, equipment, and systems involved in the present disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "including," "comprising," "having," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0105] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.
[0106] It should also be noted that in the system and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0107] Various changes, substitutions, and modifications may be made to the technology described herein without departing from the teachings defined by the appended claims. Moreover, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of things, means, methods, and actions described above. Currently existing or later developed processes, machines, manufactures, compositions of things, means, methods, or actions that perform substantially the same function or achieve substantially the same results as the corresponding aspects described herein may be utilized. Accordingly, the appended claims include within their scope such processes, machines, manufactures, compositions of things, means, methods, or actions.
[0108] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0109] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A training method for a medical guidance model, characterized in that: include: Acquire pre-diagnosis conversation data and in-diagnosis data, and extract question data from the pre-diagnosis conversation data; Based on the diagnosis data, obtaining a reference standard answer to the question data; Input the question data into the collaborative model and the basic model respectively, and obtain a first answer content output by the collaborative model and a second answer content output by the basic model; constructing preference data based on the question data, the first answer content, the second answer content, and the reference standard answer; Based on the in-diagnosis data, a virtual QRA data set is obtained; Performing supervised fine-tuning on the basic model based on the virtual QRA dataset to obtain a reference model; Acquire safety constraint rules, and train the diagnosis guidance model based on the preference data, the safety constraint rules, and the reference model to acquire an expected model.
2. The training method of the medical guidance model according to claim 1, characterized in that: The step of obtaining a reference standard answer to the question data based on the diagnosis data includes: Preprocessing the diagnosis data to obtain key data; Based on the user's identification and the timeline of the consultation, align the problem data, the key data, and the in-consultation data; The data-aligned question data, key data and diagnosis data are input into the teacher's large language model to generate a reference standard answer for each question data.
3. The training method of the medical guidance model according to claim 1, characterized in that: The constructing of preference data based on the question data, the first answer content, the second answer content, and the reference standard answer includes: Comparing the first answer content with the second answer content based on the question data and a reference standard answer corresponding to the question data, and determining high-quality answer content and low-quality answer content based on the comparison result; The question data, the high-quality answer content, and the low-quality answer content are combined into preference data.
4. The training method of the medical guidance model according to claim 1, characterized in that: The acquiring of a virtual QRA data set based on the diagnosis data includes: The QRA samples were constructed according to disease types; Splitting the diagnosis data to obtain diagnosis and treatment data sets for different disease types; Filling the QRA sample and the diagnosis and treatment data set into a preset template according to the disease type to form multiple prompt messages; Each prompt information is input into the teacher model respectively, and the virtual QRA data output by the teacher model constitutes a virtual QRA data set.
5. The training method of the medical guidance model according to claim 3, characterized in that: The obtaining of the security constraint rules, training the diagnosis guidance model based on the preference data, the security constraint rules, and the reference model to obtain the expected model, includes: Construct an initially empty preference dataset; Determining current preference data through a traversal operation, and extracting high-quality answer content and low-quality answer content from the current preference data; Determining whether the high-quality reply content complies with security constraint rules; If the high-quality reply content complies with the security constraint rules, storing the current preference data in the preference data set; If the high-quality answer content does not comply with the security constraint rules, then determining whether the low-quality answer content complies with the security constraint rules; If the non-high-quality answer content meets the security constraint rule, the high-quality answer content and the non-high-quality answer content are subjected to preference inversion to form new preference data, and the new preference data is stored in the preference data set; If the non-high-quality reply content does not comply with the security constraint rules, the current preference data is deleted; After all preference data are traversed, the guidance model is trained based on the preference data set and the reference model to obtain the expected model.
6. The training method of the medical guidance model according to claim 3, characterized in that: The obtaining of the security constraint rules, training the diagnosis guidance model based on the preference data, the security constraint rules, and the reference model to obtain the expected model, includes: Constructing an SC-DPO loss function based on the question data, high-quality answer content and non-high-quality answer content included in the preference data, and the security constraint rules included in the rule list; The guidance model is trained based on the reference model and the SC-DPO loss function to obtain the expected model.
7. The training method of the medical guidance model according to claim 6, characterized in that: The expression of the SC-DPO loss function is: Where, represents the SC-DPO loss function; Represents the distribution of preference datasets; Represents all preference data samples sampled from distribution D ; Represents problem data; Indicates high-quality reply content; Indicates non-high-quality reply content; Represents the sigmoid function; represents a hyperparameter; Represents the guidance model; Represents the parameters of the guidance model; represents the reference model; represents the parameters of the reference model; express expectations; Indicates that the guidance model is given an input Generate a response under the conditions probability; Indicates that the reference model is given an input Generate a response under the conditions probability; Indicates that the guidance model is given an input Generate a response under the conditions probability; Indicates that the reference model is given an input Generate a response under the conditions probability; represents the safety penalty weight hyperparameter; Represents a list of rules; Represents a security penalty function. If the guidance model is given an input The response generated under the conditions If the security constraint rules in the rule list are not met, then , if the guidance model is given input The response generated under the conditions If the security constraint rules in the rule list are met, .
8. A computer device, characterized in that: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the training method of the guidance model described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the training method for the diagnosis guidance model described in any one of claims 1-7.
10. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Model training method and device, nonvolatile storage medium and electronic equipment
CN118051775A
Large language model training method and device
CN118349852A
Enhanced self-training-based medical scene dialogue generation method and system
CN118569381A
Medical treatment guide model training method and system, terminal and medium
CN119578497A
Large medical model based on DPO and application thereof
CN119650033A