Intelligent diagnostic medical model training method and model training device
By integrating large language models and medical expert knowledge, a multi-grained medical model is constructed, and the problems of diagnostic errors and insufficient professionalism in the existing technology are solved, achieving higher diagnostic accuracy and efficiency.
Patent Information
- Application Number
- CN202510133332.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-09
AI Technical Summary
Existing medical diagnostic technologies have the risk of misdiagnosis in the diagnosis of daily health problems and common diseases, are insufficient in professionalism, and common large language models are difficult to achieve high accuracy in the diagnosis of specific diseases.
By integrating the basic large language models and the professional knowledge of medical experts, we can build a general medical model and a specialized model for specific diseases, and combine the disease discrimination model to achieve the integration of multi-grained medical models and conduct intelligent diagnosis.
It improves the professionalism and accuracy of disease diagnosis, can better balance the diagnostic needs of different types of diseases, and improves the efficiency and accuracy of diagnosis.
Smart Images

Figure CN119964783A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and medical models, and in particular to an intelligent diagnosis medical model training method and a model training device. Background Art
[0002] With the development of economy and society, people's average life expectancy has gradually increased, bringing greater demand for medical and health resources. However, the existing medical resources are unevenly distributed in space and level, and cannot well meet the needs of the general public in terms of daily health, common diseases, chronic diseases, difficult and complicated diseases, etc. With the emergence and rapid promotion of large language models represented by ChatGPT, some professional models have also begun to appear in the medical field. Most of these models are used in scenarios such as triage, health information consultation and preliminary diagnosis of diseases, and their professionalism is usually not at a high level.
[0003] The research on intelligent technology in the medical field began when the concept of artificial intelligence was born. The simulation decision-making system represented by Mycin uses the expert system route to diagnose specific diseases and select treatment plans based on multiple manually written rules. This type of system has high accuracy in specific fields, but poor scalability. With the development of artificial intelligence technology, auxiliary diagnosis systems for specific diseases or specific categories have been developed in many different disease categories and different departments, but they basically have the characteristics of poor versatility. The large language model trained based on a large amount of medical data has a relatively wide knowledge coverage and has played a certain role in consulting on some daily health issues, but it lacks professionalism and is prone to misjudgment of some diseases with similar symptoms.
[0004] For example, the technical solutions in the prior art include: 1. Professional doctors provide consultation on the condition, arrange tests and examinations as needed, and then make a diagnosis and prescribe medicine based on the results. 2. Based on the patient's symptom description and preliminary disease classification, an expert system is used to diagnose and formulate a treatment plan. 3. Communicate with the general medical model, collect symptoms, diagnose, and give treatment advice.
[0005] The above technical solutions have their own shortcomings or usage limitations:
[0006] 1. There is no need to go to the hospital to consult a professional doctor for daily health problems; for common diseases, the diagnosis period is long and the cost is high; the level of doctors in some primary medical institutions is limited, misdiagnosis is easy to occur, and delayed disease treatment leads to more serious consequences.
[0007] 2. Most expert systems have strict scope of application, and doctors need to make preliminary judgments on whether they are applicable to a specific field. Moreover, most expert systems are aimed at professional physicians, with high thresholds for use, and many require additional testing and examinations.
[0008] 3. Due to its limited training data and professional limitations, the general large language model can only make rough judgments based on the patient's description and cannot make accurate diagnoses. In addition, for some common diseases with similar symptoms, it is easy to be misled by the patient's initial description, and there is a certain probability of misdiagnosis. Summary of the invention
[0009] In view of this, the problem to be solved by the present invention is to integrate the capabilities of the basic large language model and the professional knowledge provided by medical experts to build an intelligent diagnosis medical model with a high level of professionalism to assist doctors from different institutions in disease diagnosis and treatment. To this end, the present invention provides an intelligent diagnosis medical model training method, which includes:
[0010] Step 1: Train a general medical model for a variety of diseases, including:
[0011] Step 1.1: Obtain a basic large language model that meets the requirements of the usage scenario;
[0012] Step 1.2: Collect various text data in the medical field, including medical textbooks, medical literature, clinical trial data, electronic health records, professional terminology dictionaries, case reports, public health datasets, online medical communities and forum data;
[0013] Step 1.3: Fine-tune the base large language model using the data collected in step 1.2.
[0014] Step 2: Construct a major disease map to test the diagnostic accuracy of the general medical model trained in step 1 for major diseases.
[0015] Step 3: Training a disease-specific model (dedicated model) for a specific disease, which includes:
[0016] Step 3.1: Determine the base model for training. Use the general medical model trained in step 1, or use knowledge distillation to form a smaller professional model based on it.
[0017] Step 3.2: Filter the general data in step 1 to obtain data specific to the disease;
[0018] Step 3.3: Collect the latest clinical guidelines, medical literature, research reports, disease surveillance reports, epidemiological data, and prevention and control strategy data for the disease;
[0019] Step 3.4: Collect more diagnostic data of the disease, including clinical records, standard cases and negative cases of model dialogue, and review and modify them by medical experts in the field;
[0020] Step 3.5: Use the above data to further train the basic model to improve the model's diagnostic accuracy for the disease;
[0021] Step 3.6: Evaluate the new model to see if it meets the requirements. If not, return to step 3.3 until it meets the qualification standards or the indicators cannot be improved after multiple rounds of optimization.
[0022] Step 4: Train the disease discrimination model, including:
[0023] Step 4.1: Select an appropriate classifier model, which includes decision trees, naive Bayes, Bayesian networks in the field of traditional machine learning, and convolutional neural networks, recurrent neural networks, and Transformer models in the field of deep learning;
[0024] Step 4.2: Prepare training data, extract the main symptoms and diagnosis results from previous case data, organize the input into discrete feature labels or text descriptions, and output the correct diagnosis labels;
[0025] Step 4.3: Train the classifier and require the output to be a probability distribution over all disease labels, that is, the sum of the probabilities of all diseases is 1.
[0026] Step 5: After the model training is completed, diagnose the new cases. The specific steps are as follows:
[0027] Step 5.1: Use the dialogue model to communicate with the patient and collect their main symptoms;
[0028] Step 5.2: Input the patient's self-report into the disease discrimination model trained in step 4, and output the probability distribution of the case among all diseases;
[0029] Step 5.3: If the maximum value in the probability distribution is higher than a predetermined threshold, the disease-specific model corresponding to the disease is selected to diagnose it; if there is no corresponding disease-specific model for the disease, it is diagnosed by the general medical model;
[0030] Step 5.4: If the maximum value in the probability distribution does not exceed the predetermined threshold, select at least the top N diseases with the highest probability and a cumulative probability higher than the threshold. Both conditions must be met at the same time. Combine them into a hybrid expert model, predict the probability of generating the next token at the same time, and perform weighted summation of the outputs of multiple models according to the disease probability distribution to jointly complete the diagnosis.
[0031] Step 6: Full model evaluation and tuning, which includes:
[0032] Step 6.1: The model's diagnosis is reviewed by medical experts in the field. Each diagnosis needs to be reviewed in the initial operation, and then random inspections are carried out according to the set ratio;
[0033] Step 6.2: For the results that experts judged incorrectly during diagnosis, classify their error natures and classify them into corresponding case libraries. After reaching a certain scale, retrain the corresponding steps;
[0034] Step 6.3: For disease identification errors, correct the case data and add them to the disease identification model in step 4 for training;
[0035] Step 6.4: In the case where the disease identification is basically accurate but the diagnosis and medication are inaccurate, the expert will rewrite the diagnosis report and input the disease-specific model in step 3 or the general medical model in step 1 for retraining.
[0036] Preferably, step 2 specifically includes:
[0037] Step 2.1: Check the latest version of the health service statistical survey report to determine the spectrum of common diseases;
[0038] Step 2.2: Construct standard cases based on the typical symptoms of each common disease and generate diagnostic records through dialogue with the medical model;
[0039] Step 2.3: Evaluate the diagnostic quality of the medical model based on multiple pre-defined evaluation indicators;
[0040] Step 2.4: Conduct multiple evaluations for each disease and take the average index as the evaluation result;
[0041] Step 2.5: Based on the pre-set qualified indicators, determine whether the current model can meet the diagnostic needs of the current disease. If not, proceed to step 3, otherwise proceed to step 4.
[0042] In addition, the present invention also provides a model training device, the device comprising:
[0043] The first training unit is used to train a general medical model applicable to various diseases, which includes: obtaining a basic large language model that meets the requirements of the usage scenario; collecting various text data in the medical field, including medical textbooks, medical literature, clinical trial data, electronic health records, professional terminology dictionaries, case reports, public health data sets, online medical communities and forum data; using the collected data to fine-tune the basic large language model;
[0044] A model evaluation unit, used to construct a main disease map to test the diagnostic accuracy of the general medical model trained in the first training unit for the main diseases;
[0045] The second training unit is used to train disease-specific models for specific diseases, which includes: determining the basic model for training; filtering out data specific to the disease from the general data in the first training unit; collecting the latest clinical guidelines, medical literature, research reports, disease monitoring reports, epidemiological data and prevention and control strategy data for the disease; collecting more diagnostic data for the disease, including clinical medical records, negative cases of standard cases and model dialogues, and reviewing and modifying them by medical experts in the field; using the above data to further train the basic model to improve the model's diagnostic accuracy for the disease; and evaluating the indicators of the new model to determine whether it meets the requirements;
[0046] The third training unit is used to train the disease discrimination model, which includes: selecting an appropriate classifier model; preparing training data, extracting the main symptoms and diagnosis results from previous case data, organizing the input into discrete feature labels or text descriptions, and outputting the correct diagnosis label; training the classifier, requiring the output to be a probability distribution over all disease labels, that is, the sum of the probabilities of all diseases is 1;
[0047] The case diagnosis module is used to diagnose new cases after the model training is completed, which includes: using the dialogue model to communicate with the patient and collect his main symptoms; inputting the patient's self-report into the disease discrimination model trained in the third training unit, and outputting the probability distribution of the case among all diseases; if the maximum value in the probability distribution is higher than the predetermined threshold, selecting the disease-specific model corresponding to the disease to diagnose it; if the disease does not have a corresponding disease-specific model, it is diagnosed by the general medical model; if the maximum value in the probability distribution does not exceed the predetermined threshold, selecting at least the top N diseases with the highest probability and a cumulative probability higher than the threshold, and both conditions must be met at the same time, and they are combined into a hybrid expert model, and the probability of generating the next token is predicted at the same time, and the outputs of multiple models are weighted and summed according to the disease probability distribution to jointly complete the diagnosis.
[0048] Preferably, the device further comprises a full model tuning module for full model evaluation and tuning, which comprises:
[0049] The model's diagnosis is reviewed by medical experts in the field. Each diagnosis needs to be reviewed in the initial stage of operation, and random checks are carried out according to the set ratio later.
[0050] For the results that experts judged incorrectly during diagnosis, the nature of the errors was classified and they were classified into corresponding case libraries. After reaching a certain scale, the corresponding steps were retrained;
[0051] For errors in disease identification, the case data is corrected and added to the disease identification model for training;
[0052] In cases where the disease identification is basically accurate but the diagnosis and medication are inaccurate, experts will rewrite the diagnosis report and input the disease-specific model or general medical model for retraining.
[0053] Compared with the prior art, the present invention has the following beneficial effects:
[0054] The present invention provides a solution for the fusion of multi-granularity medical models. For different diseases, according to their diagnostic difficulty, a universal medical model and a disease-specific model for a single disease are used for diagnosis. At the same time, a complete evaluation solution and a technical route for classifying and identifying diseases are provided. Compared with the existing technology, this solution can better balance the diagnostic needs of different types of diseases, use multiple language models of different granularity and scale to jointly realize intelligent diagnosis functions, improve the professionalism and accuracy of diagnosis, and have higher operating efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 Schematic diagram of the process of training a smart diagnosis medical model in an embodiment of the present invention;
[0056] Figure 2 Schematic diagram of a decision tree of a disease discrimination model in an embodiment of the present invention. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessary confusion of the concept of the present invention.
[0058] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "the" used in the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0059] like Figure 1 As shown, in one embodiment, the intelligent diagnosis medical model training method of the present invention includes the following steps:
[0060] Step 1: Train a general medical model applicable to various diseases, which is divided into the following steps:
[0061] Step 1.1: Obtain a basic large language model that meets the requirements of the usage scenario (graphics memory usage, inference speed, etc.).
[0062] Step 1.2: Collect various text data in the medical field, including but not limited to medical textbooks and other teaching materials, medical literature, clinical trial data, electronic health records, professional terminology dictionaries, case reports, public health datasets, online medical communities and forums, etc.
[0063] Step 1.3: Use the collected data to fine-tune the basic large language model to give it professional capabilities in the medical field.
[0064] Step 2: Build a map of major diseases and test the diagnostic accuracy of the general medical model trained in step 1 for these diseases. The specific steps are as follows:
[0065] Step 2.1: Consult the latest version of the health service statistical survey report to determine the spectrum of common diseases.
[0066] Step 2.2: Based on the typical symptoms of each common disease, construct a standard case and communicate with the medical model to generate a diagnosis record.
[0067] Step 2.3: Evaluate the diagnostic quality of the medical model based on multiple pre-defined evaluation indicators.
[0068] Step 2.4: Perform multiple evaluations for each disease and take the average index as the evaluation result.
[0069] Step 2.5: Based on the pre-set qualified indicators, determine whether the current model can meet the diagnostic needs of the current disease. If not, proceed to step 3, otherwise proceed to step 4.
[0070] Step 3: Train a disease-specific model for a specific disease, which is a dedicated model. The specific steps are as follows:
[0071] Step 3.1: Determine the basic model for training. You can use the general medical model trained in step 1, or you can use knowledge distillation and other methods on its basis to form a smaller-scale professional model.
[0072] Step 3.2: Filter the general data in step 1 to obtain data specific to the disease.
[0073] Step 3.3: Collect the latest clinical guidelines, medical literature, research reports on the disease, as well as authoritative information such as disease surveillance reports, epidemiological data, and prevention and control strategies issued by official health agencies.
[0074] Step 3.4: Collect more diagnostic data of the disease, including clinical medical records, standard cases, and negative cases of model dialogue (no correct diagnosis results), etc., and review and modify them by medical experts in the field.
[0075] Step 3.5: Use the above data to further train the basic model to improve the diagnostic accuracy of the model for the disease.
[0076] Step 3.6: Evaluate the new model to see if it meets the requirements. If not, return to step 3.3 until it meets the qualification standards or the indicators cannot be improved after multiple rounds of optimization.
[0077] Step 4: Train the basic disease discrimination model. The specific steps are as follows:
[0078] Step 4.1: Select an appropriate classifier model. Alternatives include decision trees, naive Bayes, Bayesian networks, etc. in the field of traditional machine learning, as well as convolutional neural networks, recurrent neural networks, and Transformer models in the field of deep learning.
[0079] Step 4.2: Prepare training data, extract the main symptoms and diagnosis results from previous case data, organize the input into discrete feature labels (traditional machine learning methods) or text descriptions (deep learning methods), and output the correct diagnosis labels.
[0080] Step 4.3: Train the classifier and require the output to be a probability distribution over all disease labels, that is, the sum of the probabilities of all diseases is 1.
[0081] Step 5: After the model training is completed, diagnose the new cases. The specific steps are as follows:
[0082] Step 5.1: Use a conversational model (either a general medical model or a separate guiding model with certain medical capabilities) to communicate with the patient and collect their main symptoms.
[0083] Step 5.2: Input the patient's self-report into the disease discrimination model trained in step 4 (if it is a traditional machine learning model, feature extraction at the symptom level is required), and output the probability distribution of the case among all diseases.
[0084] Step 5.3: If the maximum value in the probability distribution is higher than a predetermined threshold (such as 80%), select a dedicated model corresponding to the disease to diagnose it. If there is no dedicated model for the disease, it will be completed by the general medical model.
[0085] Step 5.4: If the maximum value in the probability distribution does not exceed the predetermined threshold, select at least the top N diseases (typical value is 3) with the highest probability and the cumulative probability higher than the threshold (such as 95%). Both conditions must be met at the same time. They are combined into a hybrid expert model to predict the generation probability of the next token at the same time, and the outputs of multiple models are weighted and summed according to the disease probability distribution to jointly complete the diagnosis.
[0086] Step 6: Full model evaluation and tuning. The specific steps are as follows:
[0087] Step 6.1: The model's diagnosis is reviewed by field medical experts. Each diagnosis needs to be reviewed in the initial stage of operation, and random checks can be carried out at a certain ratio later.
[0088] Step 6.2: For the results that experts judge to be incorrect during diagnosis, classify the nature of the errors and place them into corresponding case libraries. After reaching a certain scale, retrain the corresponding steps.
[0089] Step 6.3: For disease identification errors, correct the case data and add it to the disease identification model in step 4 for training.
[0090] Step 6.4: In cases where the disease identification is basically accurate but the diagnosis and medication are inaccurate, the expert will rewrite the diagnosis report and input the disease-specific model of the disease in step 3 (if a professional model is used) or the general model in step 1 (if no professional model has been established for the disease) for retraining.
[0091] Specifically, the construction method of the standard case includes:
[0092] Collect maps of common diseases and corresponding symptoms, and construct a comparison table between diseases and symptoms, for example:
[0093] Insect bites
[0094] Local redness and swelling: The bite site becomes red and swollen, often with a pin-sized bruise in the center and a red halo around it.
[0095] · Itching: The bite site will itch significantly. Scratching may cause skin damage and secondary infection.
[0096] Heat stroke
[0097] Fever: The body temperature rises to over 38°C and in severe cases can exceed 40°C.
[0098] Dizziness and headache: Insufficient blood supply and lack of oxygen to the body can cause dizziness, head swelling, and head tumors, which may be accompanied by nausea and vomiting.
[0099] Thirst and excessive sweating: In the early stage, there is a lot of sweating and obvious thirst, but then there is no sweating due to the failure of sweat gland function.
[0100] · Palpitation and fatigue: The heartbeat speeds up, the body becomes weak and powerless, and in severe cases, blood pressure may drop and shock may occur.
[0101] Insufficient blood flow to the brain
[0102] Dizziness: Repeated dizziness and vertigo, feeling groggy and unclear, which may be aggravated by sudden head turning or body position change.
[0103] Headache: usually dull pain or distending pain, which may be accompanied by memory loss and inattention.
[0104] Blurred vision: Insufficient blood supply to the brain affects the function of the optic nerve, which may cause temporary vision loss.
[0105] Sleep disorders: Sleep problems such as insomnia, frequent dreams, and easy awakening are common.
[0106] Sequelae of stroke
[0107] Limb movement disorder: Hemiplegia of one limb, manifested as inability to lift or clench a fist with the upper limb, difficulty walking with the lower limb, dragging on the floor, limited joint movement, muscle atrophy, etc.
[0108] Speech disorders: Aphasia, including motor aphasia (can understand others’ speech but have difficulty expressing oneself), sensory aphasia (cannot understand others’ speech but can express oneself fluently but with errors), mixed aphasia, etc.
[0109] Cognitive impairment: memory loss, lack of concentration, decreased calculation ability, disorientation, etc., which seriously affect the patient's daily life and social function.
[0110] ·Dysphagia: Difficulty swallowing and choking when eating can lead to complications such as malnutrition and lung infection.
[0111] Then a real person can play the role of the patient and explain the symptoms one by one in a natural conversation. Another option is to use a large language model, let the model play the role of a patient with a disease based on the typical symptoms of the disease, and explain the symptoms according to certain rules (1-2 typical symptoms per round, some other confusing symptoms can be added, and the description of the symptoms is rewritten in a colloquial style), and have a conversation with the model to generate a conversation record of a typical case.
[0112] Evaluation indicators of model output: The evaluation indicators include both general indicators for evaluating the conversational capabilities of large language models, such as text fluency and security, as well as special requirements in the medical field, such as professionalism and output style. They also include initiative and friendliness required to ensure user experience and completeness of information collection during intelligent diagnosis and treatment. The following table includes a description of some indicators, and the specific evaluation methods and scoring standards are omitted.
[0113] Table 1:
[0114]
[0115] The structure of the disease discrimination model:
[0116] The main function of the disease discrimination model is to input a text T of length n = {t1, t2, …, t n}, a probability distribution D = (p1, p2, ..., p m ), so that For a model with parameter θ, the training target is P(D|T,θ). For different types of models, the training ideas are different. Figure 2 As shown, taking the decision tree as an example, an example is given, and its training method can adopt C4.5 or other decision tree algorithms.
[0117] With the rapid development of large language models, their applications in the medical field are also gradually increasing. Existing intelligent medical models are either based on traditional expert systems with limited application scope and high threshold for use, or are trained based on general medical data, and are not highly professional in diagnosing some specific diseases. The present invention provides a solution for the fusion of multi-granularity medical models. For different diseases, general models and professional models for single diseases are used for diagnosis according to their diagnostic difficulty. At the same time, a complete evaluation solution and a technical route for classifying and identifying diseases are provided. Compared with the existing technology, this solution can better balance the diagnostic needs of different types of diseases, use multiple language models of different granularities and scales to jointly realize intelligent diagnostic functions, improve the professionalism and accuracy of diagnosis, and have higher operating efficiency.
[0118] In addition, the present invention also provides a model training device, which includes:
[0119] The first training unit is used to train a general medical model applicable to various diseases, which includes: obtaining a basic large language model that meets the requirements of the usage scenario; collecting various text data in the medical field, including medical textbooks, medical literature, clinical trial data, electronic health records, professional terminology dictionaries, case reports, public health data sets, online medical communities and forum data; using the collected data to fine-tune the basic large language model;
[0120] A model evaluation unit, used to construct a main disease map to test the diagnostic accuracy of the general medical model trained in the first training unit for the main diseases;
[0121] The second training unit is used to train disease-specific models for specific diseases, which includes: determining the basic model for training; filtering out data specific to the disease from the general data in the first training unit; collecting the latest clinical guidelines, medical literature, research reports, disease monitoring reports, epidemiological data and prevention and control strategy data for the disease; collecting more diagnostic data for the disease, including clinical medical records, negative cases of standard cases and model dialogues, and reviewing and modifying them by medical experts in the field; using the above data to further train the basic model to improve the model's diagnostic accuracy for the disease; and evaluating the indicators of the new model to determine whether it meets the requirements;
[0122] The third training unit is used to train the disease discrimination model, which includes: selecting an appropriate classifier model; preparing training data, extracting the main symptoms and diagnosis results from previous case data, organizing the input into discrete feature labels or text descriptions, and outputting the correct diagnosis label; training the classifier, requiring the output to be a probability distribution over all disease labels, that is, the sum of the probabilities of all diseases is 1;
[0123] The case diagnosis module is used to diagnose new cases after the model training is completed, which includes: using the dialogue model to communicate with the patient and collect his main symptoms; inputting the patient's self-report into the disease discrimination model trained in the third training unit, and outputting the probability distribution of the case among all diseases; if the maximum value in the probability distribution is higher than the predetermined threshold, selecting the disease-specific model corresponding to the disease to diagnose it; if the disease does not have a corresponding disease-specific model, it is diagnosed by the general medical model; if the maximum value in the probability distribution does not exceed the predetermined threshold, selecting at least the top N diseases with the highest probability and a cumulative probability higher than the threshold, and both conditions must be met at the same time, and they are combined into a hybrid expert model, and the probability of generating the next token is predicted at the same time, and the outputs of multiple models are weighted and summed according to the disease probability distribution to jointly complete the diagnosis.
[0124] Preferably, the aforementioned device further includes a full model tuning module for full model evaluation and tuning, which includes:
[0125] The model's diagnosis is reviewed by medical experts in the field. Each diagnosis needs to be reviewed in the initial stage of operation, and random checks are carried out according to the set ratio later.
[0126] For the results that experts judged incorrectly during diagnosis, the nature of the errors was classified and they were classified into corresponding case libraries. After reaching a certain scale, the corresponding steps were retrained;
[0127] For errors in disease identification, the case data is corrected and added to the disease identification model for training;
[0128] In cases where the disease identification is basically accurate but the diagnosis and medication are inaccurate, experts will rewrite the diagnosis report and input the disease-specific model or general medical model for retraining.
[0129] It is understandable that if the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a corresponding computer-readable storage medium. Based on such an understanding, the present invention implements all or part of the processes in the above-mentioned corresponding embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0130] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0131] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0132] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0133] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for training an intelligent diagnostic medical model, characterized in that: The method comprises: Step 1: Train a general medical model for a variety of diseases, including: Step 1.1: Obtain a basic large language model that meets the requirements of the usage scenario; Step 1.2: Collect various text data in the medical field, including medical textbooks, medical literature, clinical trial data, electronic health records, professional terminology dictionaries, case reports, public health datasets, online medical communities and forum data; Step 1.3: Use the data collected in step 1.2 to fine-tune the basic large language model; Step 2: Construct a major disease map to test the diagnostic accuracy of the general medical model trained in step 1 for major diseases; Step 3: Training a disease-specific model for a specific disease, which includes: Step 3.1: Determine the base model for training. Use the general medical model trained in step 1, or use knowledge distillation to form a smaller professional model based on it. Step 3.2: Filter the general data in step 1 to obtain data specific to the disease; Step 3.3: Collect the latest clinical guidelines, medical literature, research reports, disease surveillance reports, epidemiological data, and prevention and control strategy data for the disease; Step 3.4: Collect more diagnostic data of the disease, including clinical records, standard cases and negative cases of model dialogue, and review and modify them by medical experts in the field; Step 3.5: Use the above data to further train the basic model to improve the model's diagnostic accuracy for the disease; Step 3.6: Evaluate the new model to see if it meets the requirements. If not, return to step 3.3 until it meets the qualification standards or the indicators cannot be improved after multiple rounds of optimization. Step 4: Train the disease discrimination model, including: Step 4.1: Select an appropriate classifier model, which includes decision trees, naive Bayes, Bayesian networks in the field of traditional machine learning, and convolutional neural networks, recurrent neural networks, and Transformer models in the field of deep learning; Step 4.2: Prepare training data, extract the main symptoms and diagnosis results from previous case data, organize the input into discrete feature labels or text descriptions, and output the correct diagnosis labels; Step 4.3: Train the classifier and require the output to be a probability distribution over all disease labels, that is, the sum of the probabilities of all diseases is 1; Step 5: After the model training is completed, diagnose the new cases. The specific steps are as follows: Step 5.1: Use the dialogue model to communicate with the patient and collect their main symptoms; Step 5.2: Input the patient's self-report into the disease discrimination model trained in step 4, and output the probability distribution of the case among all diseases; Step 5.3: If the maximum value in the probability distribution is higher than a predetermined threshold, the disease-specific model corresponding to the disease is selected to diagnose it; if there is no corresponding disease-specific model for the disease, it is diagnosed by the general medical model; Step 5.4: If the maximum value in the probability distribution does not exceed the predetermined threshold, select at least the top N diseases with the highest probability and a cumulative probability higher than the threshold. Both conditions must be met at the same time. Combine them into a hybrid expert model, predict the probability of generating the next token at the same time, and perform weighted summation of the outputs of multiple models according to the disease probability distribution to jointly complete the diagnosis.
2. The method according to claim 1, characterized in that It also includes step 6: full model evaluation and tuning, which includes: Step 6.1: The model's diagnosis is reviewed by medical experts in the field. Each diagnosis needs to be reviewed in the initial operation, and then random inspections are carried out according to the set ratio; Step 6.2: For the results that experts judged incorrectly during diagnosis, classify their error natures and classify them into corresponding case libraries. After reaching a certain scale, retrain the corresponding steps; Step 6.3: For disease identification errors, correct the case data and add them to the disease identification model in step 4 for training; Step 6.4: In the case where the disease identification is basically accurate but the diagnosis and medication are inaccurate, the expert will rewrite the diagnosis report and input the disease-specific model in step 3 or the general medical model in step 1 for retraining.
3. The method according to claim 2, characterized in that Step 2 specifically includes: Step 2.1: Check the latest version of the health service statistical survey report to determine the spectrum of common diseases; Step 2.2: Construct standard cases based on the typical symptoms of each common disease and generate diagnostic records through dialogue with the medical model; Step 2.3: Evaluate the diagnostic quality of the medical model based on multiple pre-defined evaluation indicators; Step 2.4: Conduct multiple evaluations for each disease and take the average index as the evaluation result; Step 2.5: Based on the pre-set qualified indicators, determine whether the current model can meet the diagnostic needs of the current disease. If not, proceed to step 3, otherwise proceed to step 4.
4. The method according to claim 3, characterized in that: In step 5.1, the conversation model uses a general medical model, or a separate guiding model with certain medical capabilities.
5. A model training device, characterized in that: The device comprises: The first training unit is used to train a general medical model applicable to various diseases, which includes: obtaining a basic large language model that meets the requirements of the usage scenario; collecting various text data in the medical field, including medical textbooks, medical literature, clinical trial data, electronic health records, professional terminology dictionaries, case reports, public health data sets, online medical communities and forum data; using the collected data to fine-tune the basic large language model; A model evaluation unit, used to construct a main disease map to test the diagnostic accuracy of the general medical model trained in the first training unit for the main diseases; The second training unit is used to train disease-specific models for specific diseases, which includes: determining the basic model for training; filtering out data specific to the disease from the general data in the first training unit; collecting the latest clinical guidelines, medical literature, research reports, disease monitoring reports, epidemiological data and prevention and control strategy data for the disease; collecting more diagnostic data for the disease, including clinical medical records, negative cases of standard cases and model dialogues, and reviewing and modifying them by medical experts in the field; using the above data to further train the basic model to improve the model's diagnostic accuracy for the disease; and evaluating the indicators of the new model to determine whether it meets the requirements; The third training unit is used to train the disease discrimination model, which includes: selecting an appropriate classifier model; preparing training data, extracting the main symptoms and diagnosis results from previous case data, organizing the input into discrete feature labels or text descriptions, and outputting the correct diagnosis label; training the classifier, requiring the output to be a probability distribution over all disease labels, that is, the sum of the probabilities of all diseases is 1; The case diagnosis module is used to diagnose new cases after the model training is completed, which includes: using the dialogue model to communicate with the patient and collect his main symptoms; inputting the patient's self-report into the disease discrimination model trained in the third training unit, and outputting the probability distribution of the case among all diseases; if the maximum value in the probability distribution is higher than the predetermined threshold, selecting the disease-specific model corresponding to the disease to diagnose it; if the disease does not have a corresponding disease-specific model, it is diagnosed by the general medical model; if the maximum value in the probability distribution does not exceed the predetermined threshold, selecting at least the top N diseases with the highest probability and a cumulative probability higher than the threshold, and both conditions must be met at the same time, and they are combined into a hybrid expert model, and the probability of generating the next token is predicted at the same time, and the outputs of multiple models are weighted and summed according to the disease probability distribution to jointly complete the diagnosis.
6. The model training device according to claim 5, characterized in that: The device also includes a full model tuning module, which is used for full model evaluation and tuning, and includes: The model's diagnosis is reviewed by medical experts in the field. Each diagnosis needs to be reviewed in the initial stage of operation, and random checks are carried out according to the set ratio later. For the results that experts judged incorrectly during diagnosis, the nature of the errors was classified and they were classified into corresponding case libraries. After reaching a certain scale, the corresponding steps were retrained; For errors in disease identification, the case data is corrected and added to the disease identification model for training; In cases where the disease identification is basically accurate but the diagnosis and medication are inaccurate, experts will rewrite the diagnosis report and input the disease-specific model or general medical model for retraining.
Citation Information
Cited By
Case recognition model training method and device, electronic equipment and storage medium
CN121257774A