Medical field large model training data generation method and related equipment
By generating and filtering prompt text for case data, using language models to generate question, reference, and answer data, and combining them with scoring criteria, the problem of insufficient large model data in the field of rehabilitation medicine is solved, efficient, low-cost, and highly private model training is achieved, and the diagnostic and treatment effects of the model are improved.
Patent Information
- Application Number
- CN202411429652.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-10-14
AI Technical Summary
Existing large medical models face the problems of insufficient data quantity and uneven data quality in the field of rehabilitation medicine, resulting in insufficient model performance and reliability. Existing fine-tuning data generation methods are insufficient in diversity and accuracy, and are unable to meet the needs of complex rehabilitation medical scenarios.
By obtaining case data, generating prompt text and using pre-trained language models to generate question data, reference data and answer data, combined with scoring criteria screening and optimization, specialized model fine-tuning data is generated, including scoring and screening of question fluency, relevance, harmfulness, complexity and degree of personalization, to form high-quality training data.
It achieves efficient training of large medical models in the field of rehabilitation medicine, generates specialized model fine-tuning data, reduces the cost of obtaining real case data, improves the diagnosis and treatment accuracy of the model, and ensures the privacy and quality of the data.
Smart Images

Figure CN119517430B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of large language models, and in particular to a method for generating large model training data in the medical field and related equipment. Background Art
[0002] With the rapid development of modern medicine and artificial intelligence, deep learning models are gradually being introduced into the medical field to improve diagnostic and treatment outcomes. Deep learning models, especially large language models, have demonstrated tremendous potential in the medical field. By analyzing massive amounts of data, they can provide efficient diagnoses and personalized treatment plans. However, the application of large models in medicine, especially rehabilitation medicine, faces challenges with insufficient data quantity and uneven data quality, which directly impacts model performance and reliability. To address these issues, fine-tuned data generation methods have emerged as a key means of improving the effectiveness of large models in rehabilitation medicine.
[0003] Fine-tuning data generation methods provide a solid foundation for the application of large models in rehabilitation medicine. Currently, many existing generation methods are deficient in data diversity and accuracy, resulting in the generated data being unable to meet the needs of complex rehabilitation medical scenarios. For example, although data generation based on traditional statistical methods can improve model performance to a certain extent, there are still many limitations in terms of data coverage and detail richness. For example, when using general large models for data generation, existing methods often simply require a single model to directly provide question-answer pairs and perform simple cleaning, which cannot take into account the unique content in specific vertical fields such as rehabilitation medicine. This shows that existing fine-tuning data generation methods still need further improvement to better support the application of large models in rehabilitation medicine and improve the accuracy and effectiveness of their diagnosis and treatment. Summary of the Invention
[0004] In view of this, the purpose of this application is to propose a method for generating large model training data in the medical field.
[0005] Based on the above objectives, the present application provides a method for generating training data for a large model in the medical field, characterized in that it includes: obtaining case data, and generating a first prompt text based on the case data. Inputting the first prompt text into a preset first language model so that the first language model outputs question data associated with the case data. Generating a second prompt text based on the question data so that the first language model outputs reference data. The reference data is used to express the processing process of the first language model generating the question data. Generating a third prompt text based on the reference data so that the first language model outputs answer data. Generating model fine-tuning data based on the question data, reference data, and answer data.
[0006] In some embodiments, inputting a first prompt text into a preset first language model so that the first language model outputs question data associated with the case data specifically includes: generating a first scoring text and a first screening criterion based on the question data. The first screening criterion is set based on the model fine-tuning data. Inputting the first scoring text into a preset second language model so that the second language model outputs first scoring information associated with the question data. The first scoring information is filtered according to the first screening criterion, and question data corresponding to the first scoring information that does not meet the first screening criterion is deleted.
[0007] In some embodiments, the first scoring information includes: a question fluency score, a question relevance score, a question risk score, a complexity score, and a personalization score. The first screening criteria include: a question risk score of 0, and a question fluency score, a question relevance score, a complexity score, and a personalization score of no less than half of the total score.
[0008] In some embodiments, a second prompt text is generated based on the question data so that the first language model outputs reference data, specifically including: generating a second scoring text and a second screening criterion based on the reference data. The second screening criterion is set based on the model fine-tuning data. The second scoring text is input into a preset third language model so that the third language model outputs second scoring information associated with the reference data. A second optimized text is generated, the second scoring information is filtered according to the second screening criterion, and the reference data corresponding to the second scoring information that does not meet the second screening criterion, the question data corresponding to the reference data, and the second optimized text are input into the first language model so that the first language model optimizes the reference data.
[0009] In some embodiments, the second scoring information includes: a reference fluency score, a reference relevance score, a reference harmfulness score, a reference accuracy score, a reference rationality score, and a specificity score. The second screening criteria include: a reference harmfulness score of 0, and the reference fluency score, reference relevance score, reference accuracy score, reference rationality score, and specificity score are all no less than half of the total score.
[0010] In some embodiments, generating a third prompt text based on the reference data so that the first language model outputs answer data specifically includes: generating a third scoring text and a third screening criterion based on the reference data. The third screening criterion is set based on the model fine-tuning data. The third scoring text is input into a preset fourth language model so that the fourth language model outputs third scoring information associated with the answer data. A third optimized text is generated, the third scoring information is filtered according to the third screening criterion, and the answer data corresponding to the third scoring information that does not meet the third screening criterion, the reference data and question data corresponding to the answer data, and the third optimized text are input into the first language model so that the first language model optimizes the answer data.
[0011] In some embodiments, the third scoring information includes: an answer fluency score, an answer relevance score, an answer harmfulness score, an answer accuracy score, and an answer rationality score. The third screening criteria include: an answer harmfulness score of 0, and an answer fluency score, an answer relevance score, an answer accuracy score, and an answer rationality score of no less than half of the total score.
[0012] The present application also provides a large-scale model training data generation device in the medical field, characterized in that it includes: a case reading module, which is used to obtain case data and generate a first prompt text based on the case data. A question generation module, which is used to input the first prompt text into a preset first language model, so that the first language model outputs question data associated with the case data. A reference data generation module, which is used to generate a second prompt text based on the question data, so that the first language model outputs reference data. The reference data is used to express the processing process of the first language model generating question data. An answer generation module, which is used to generate a third prompt text based on the reference data, so that the first language model outputs answer data. A fine-tuning data integration module, which is used to generate model fine-tuning data based on the question data, reference data and answer data.
[0013] The present application also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any one of the above methods when executing the program.
[0014] The present application also provides a non-transitory computer-readable storage medium, which stores computer instructions, wherein the computer instructions are used to enable a computer to execute any of the above methods.
[0015] As can be seen from the above, the method for generating training data for large models in the medical field provided by this application uses real case data as a reference and calls a large language model through an automated program, so that the large language model generates question data, reference data, and answer data in batches. The large language model is guided by the set prompt words to generate and optimize the model fine-tuning data, thereby generating specialized model fine-tuning data for the medical field industry model, saving the training time of subsequent industry large models. In addition, due to the high difficulty in obtaining case data and the high amount of private information, the technical solution disclosed in this application also has the advantages of low cost and strong privacy compared to directly obtaining real case information. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] Figure 1 A flowchart of a method for generating large-scale model training data in the medical field provided in an embodiment of the present application;
[0018] Figure 2 A flowchart of a method for screening problematic data provided in an embodiment of the present application;
[0019] Figure 3 A schematic diagram of a flow chart of a reference data screening method provided in an embodiment of the present application;
[0020] Figure 4 A flowchart of the answer data screening method provided in an embodiment of the present application;
[0021] Figure 5 A schematic diagram of the structure of a device for generating large-scale model training data in the medical field provided in an embodiment of the present application;
[0022] Figure 6 A more specific schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to make the objectives, technical solutions and advantages of this application more clear, this application is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0024] It should be noted that, unless otherwise defined, technical terms or scientific terms used in the embodiments of the present application shall have the common meaning understood by one of ordinary skill in the art to which the embodiments of the present application belong. The terms "first", "second", and similar terms used in the embodiments of the present application do not denote any order, quantity, or importance, but are merely used to distinguish different components. The terms "include", "contain", and similar terms mean that the elements or objects before the terms encompass the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. The terms "connect" or "connected" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", and the like are merely used to represent relative positional relationships, and when the absolute positions of the described objects change, the relative positional relationships can also change accordingly.
[0025] As shown in Figure 1 The present application provides a medical field large model training data generation method, characterized in that it comprises:
[0026] Step S1, obtaining case data, and generating a first prompt text according to the case data.
[0027] In this embodiment, the case data is obtained through a legal channel, is authorized by the patient and his / her family members, and the personal privacy information of the patient and his / her family members is hidden, only including information of the case content. The obtained case data is only used for generating the current training data.
[0028] The case data can be in the following format:
[0029] Patient profile
[0030] Age: 47 years old
[0031] Gender: female
[0032] Medical history
[0033] Chief complaint: loss of consciousness, fever, dyspnea, and right hemiplegia for two months
[0034] Related symptoms: difficulty swallowing and slurred speech for two months
[0035] Past medical history: (omitted here)
[0036] Physical examination results: (omitted here)
[0037] Interview dialogue session: (omitted here)
[0038] Diagnosis and treatment prescription: (omitted here)
[0039] The first prompt text is a prompt that prompts the large model to generate text content, wherein the first prompt text can be flexibly adjusted according to the required training data.
[0040] The first prompt text includes:
[0041] You are an expert in generating question-answer datasets. Your task is to generate questions suitable for training large models in the medical field based on the case content I provide. The question description should be clear and reasonable, and there can only be one question in one sentence.
[0042] The case content is as follows:
[0043] Patient Profile
[0044] (Omitted here)
[0045] Medical history
[0046] (Omitted here)
[0047] Past medical history: (omitted here)
[0048] Physical examination results: (omitted here)
[0049] Dialogue session: (omitted here)
[0050] Diagnosis and treatment prescription: (omitted here)
[0051] Your questions will be used to construct a question-and-answer dataset for training large models in the medical field. You should consider the given content from multiple perspectives and try to diversify the content within the scope of the case. You can ask questions based on the diagnosis results in the case, describe the patient's feelings, ask questions about examination results, and refer to the questions asked by patients or doctors during the consultation. When asking questions, you can describe only part of the condition, or provide a detailed overall explanation of the condition. You can also draw content from previous medical history, but you must not exceed the scope of the given case content. In particular, your questions should be asked from the perspective of the "patient" or "family member". In each question, you need to first explain the condition to the doctor. For long questions, you should first provide a relatively detailed overall description of the condition. Please note that there will be no case content as a prompt when answering questions, so when generating questions, you must ensure that the content of the question is specific enough. For example, you cannot just say "imaging data was shown", "MRI and PET scans both showed abnormalities", "this is my past medical history", etc. in the question description. You need to describe the relevant content in detail in the question. When asking for medical rehabilitation advice or treatment plans, you must summarize the condition, explain what symptoms appeared, and what diagnosis the doctor made. Note that you also need to consider the patient's professional level. You can use the patient's dialogue during the consultation process as a reference, and put the content that exceeds his level in the objective examination section or the summary section above. For example, you can't just say "Based on my father's current condition and diagnosis," but instead provide a clear and specific description of these circumstances. Furthermore, questions should be diverse within the scope of the content. You can ask questions based on the diagnosis results in the case, questions from the consultation process, and questions about medical advice. Questions should not be repeated, but they should also not exceed the scope of the case content. Finally, questions should take into account the patient's professional level and be as colloquial as possible. Professional content can be described separately in the objective examination section. Descriptions should be specific and detailed, with no word limit. Ensure that the questions can be answered independently without the reference to the case content or the surrounding context. Please follow the above model to generate fifteen questions, each asking only one question. Please output the questions directly in Chinese and do not add other descriptions.
[0052] Step S2: input a first prompt text into a preset first language model, so that the first language model outputs question data associated with the case data.
[0053] The first language model is a pre-trained general-purpose large language model that can be selected based on actual application needs. For example, it can be one of ChatGPT, Gemini, and Claude. Multiple sets of first prompt texts can be input into the first language model through an automated program, causing the first language model to automatically output multiple question data.
[0054] Step S3: Generate a second prompt text based on the question data, so that the first language model outputs reference data, wherein the reference data is used to express the processing process of the first language model generating the question data.
[0055] Similar to the first prompt text, the second prompt text is a prompt designed specifically for the question data, prompting the first language model to generate and output reference data for the question data. By outputting reference data corresponding to the question data to the first language model, reference information related to the patient's condition is obtained, providing a reference basis for the subsequent generation of answer information. Multiple sets of second prompt text can be input to the first language model through an automated program, causing the first language model to automatically output reference data corresponding to the question data.
[0056] The second prompt text includes:
[0057] You are an expert in question-answering dataset processing. Your task is to generate a process of thinking about the corresponding question based on the given case content. The answer description should be clear and reasonable.
[0058] The case content is as follows:
[0059] (Add case content here)
[0060] The questions you need to answer are:
[0061] (Add question content here)
[0062] Please think about each question individually. Pay special attention to the fact that your answer must be true and reliable and cannot exceed the given case content. At the same time, there is no case content as a prompt in the actual scenario, so the factual content cannot exceed the scope of the question description. Specifically, you should pay attention to check your output content and whether there are any factual inconsistencies in the thinking process. On the one hand, your thinking process cannot contain objective facts beyond the description of the question. For example, the question does not mention imaging data or medical history, but the thinking process includes information extracted from the case content, etc. When answering the question, there is no case content as a reference, so you must think and answer based on the question. On the other hand, your thinking process must be generated according to the facts of the case content to ensure that the answer is true and accurate. Finally, there must be specific measures in the treatment recommendations and expected effects section. This measure should be extracted from the case content. For example, for medication recommendations, the name, dosage, usage, and frequency of the drug should be given. Rehabilitation treatment should include the items, frequency, purpose, etc. of rehabilitation treatment. The output you provide should be formatted in JSON with the following structure:
[0063] {"Possible Diagnosis":"",
[0064] "Supporting evidence":"",
[0065] "Unsupported parts":"",
[0066] "Further inspection suggestions":"",
[0067] "Treatment recommendations": "Drug name 1, dosage, usage, frequency, drug name 2, dosage, usage, frequency, ... Rehabilitation treatment item 1, frequency, purpose, Rehabilitation treatment item 2, frequency, purpose, ...",
[0068] "Expected effect":""
[0069] }
[0070] If any part does not apply, please answer "Not applicable". Please directly enter the results of your thinking process for all questions in Chinese without adding any other description.
[0071] Step S4: Generate a third prompt text based on the reference data so that the first language model outputs answer data.
[0072] Similar to the first prompt text, the third prompt text is a prompt designed for answer data, prompting the first language model to generate and output answer data corresponding to the reference data. The answer data includes answers to the medical questions in the question data, generated using reference information corresponding to the patient's question generated from the question data in the reference data. The question data, reference data, and information about each patient generated in the answer data correspond one-to-one. Multiple sets of third prompt text can be input into the first language model through an automated program, causing the first language model to automatically output multiple answers.
[0073] The third prompt text includes:
[0074] You are an expert in processing question-answering datasets. Your task is to generate answers to the given questions based on the case content and thinking process, which are suitable for training large models in the medical field. The answer description should be clear and reasonable.
[0075] The case content is as follows:
[0076] (Add case content here)
[0077] The questions you need to answer and the corresponding thought process are as follows:
[0078] (Add the question content and corresponding thinking process here)
[0079] In particular, your answer should be presented from a doctor's perspective. It must be factual and reliable, and should not go beyond the context of the case. Furthermore, in real-world scenarios, there is no case context to provide a prompt, so the factual content should not exceed the scope of the question. Specifically, carefully review your output for factual inconsistencies. For example, your answer should not contain objective facts beyond the question's description. For example, if the question doesn't mention imaging or medical history, but your answer includes information drawn from the case, as there is no case context to support the actual answer, you should consider and answer the question from the perspective of the question. Furthermore, your answer must be based on the facts of the case, ensuring accuracy and truthfulness. Finally, if your answer requires treatment recommendations and expected outcomes, specific measures should be included. These measures should be drawn from the corresponding modules of the case content and the thought process. For example, for medication recommendations, the name, dosage, usage, and frequency of the drug should be provided. For rehabilitation treatment, the treatment items, frequency, and purpose should be included. Please output your answer directly in Chinese, without any additional description.
[0080] Step S5: Generate model fine-tuning data based on the question data, reference data, and answer data.
[0081] The data corresponding to different patients generated from the question data, reference data, and answer data are organized into multiple question-and-answer groups, and the multiple case data are integrated into model fine-tuning data. This model fine-tuning data is used to train specialized models in the medical field.
[0082] Among them, the integration of fine-tuning data can be achieved through text matching and JSON processing methods. First, the six parts of the previously obtained question data, reference data and answer data are extracted and stored in the dictionary respectively, and then filled into a fixed template to automatically obtain multiple sets of fine-tuning data.
[0083] The fixed templates used include:
[0084] 1.{"src":"(add the question field content here)",
[0085] "tgt":"(add answer field content here)"}
[0086] 2.{"src":"You are a rehabilitation medicine specialist. Please answer the following patient questions:
[0087] (Add question field content here)
[0088] You need to gradually outline your thought process and final answer. This process should include possible diagnoses, supporting evidence, unsupported evidence, further examination recommendations, treatment recommendations, and expected outcomes. Then, based on your thought process, you should provide your final answer.
[0089] "tgt":"{
[0090] Thought Process:
[0091] Possible diagnosis: (fill in the fields output by the model below),
[0092] Supporting evidence:
[0093] Unsupported parts:
[0094] Further inspection suggestions:
[0095] Treatment recommendations:
[0096] Expected results:
[0097] answer:"}
[0098] 3.{"src":"If the patient has the following symptoms:
[0099] (Add supporting evidence field content here)
[0100] What disease might this patient have?
[0101] "tgt":"(Add the possible diagnosis field content here)"} (If the possible diagnosis or supporting evidence is not applicable, this data will not be generated)
[0102] 4.{"src":"What treatments are usually used for this condition (add possible diagnosis field content here)?",
[0103] "tgt":"(add treatment recommendation field content here)"} (If the possible diagnosis or treatment recommendation is not applicable, this data will not be generated)
[0104] In some embodiments, as Figure 2 As shown, step S2 specifically includes:
[0105] Step S21: Generate a first scoring text and a first screening criterion based on the question data, wherein the first screening criterion is set based on the model fine-tuning data.
[0106] Step S22: inputting the first scoring text into a preset second language model so that the second language model outputs first scoring information associated with the question data.
[0107] In this embodiment, the second language model is used to score the question data generated by the first language model based on the first scoring text and the first screening criteria, and to verify through the large language model whether the question data generated by the first language model meets the screening criteria of the model fine-tuning data. The second language model can also be controlled to mark or eliminate question data that does not meet the screening criteria, thereby achieving the output of high-quality question data.
[0108] Preferably, in order to improve the reliability of the first scoring information, the second language model can be a language model cluster composed of multiple large language models. By inputting the first scoring text into multiple different large language models in the language model cluster, all large prediction models in the language model cluster output scoring information for the question data, and removing the highest and lowest scores in the scoring, calculating the average score of other language models, and using the average score as the first scoring information.
[0109] Among them, the language model cluster refers to a large number of different commercial general-purpose large models such as ChatGPT, Claude, Qwen, Gemini, etc., and a collection of the same commercial general-purpose large models using different prompt templates. In this example, two sets of prompt templates are set up for each large model, allowing the model to play the role of an artificial intelligence evaluator and a professional physician in the rehabilitation medical field respectively. Each large model in the collection will evaluate the quality of the generated questions separately, and finally use a Python script to aggregate all the results and calculate the average to obtain the final result.
[0110] The first scoring text is input to the second language model, so that the second language model outputs a score for the question data.
[0111] The first screening criterion is set according to the requirements of model fine-tuning data, and is used to filter out problematic data suitable for training large industry models.
[0112] The first scoring information is derived based on the first screening criteria, and the second language model scores the question data based on the first screening criteria, generating the first scoring information by aggregating the scoring information from multiple dimensions. Multiple sets of first scoring texts can be input to the second language model through an automated program, so that the second language model automatically outputs multiple sets of first scoring information.
[0113] Step S23 , screening the first scoring information according to the first screening criterion, and deleting the question data corresponding to the first scoring information that does not meet the first screening criterion.
[0114] Among them, by deleting the problem information that does not meet the first screening criteria in the problem data, it is possible to avoid interference of bad data with the large industry model to be trained. In this application, if too much problem information is deleted in step S23, step S2 can be re-executed to generate more problem information and screen it according to the standards of steps S21-step S23, thereby ensuring that sufficient model fine-tuning data is provided for the industry that needs to be trained later.
[0115] The first scoring text includes:
[0116] You are an AI assessor, specializing in evaluating the quality of questions generated for a given case as part of a dataset of question-answer pairs used to train large models in the medical field. (Another example: You are a rehabilitation physician, and you are asked to evaluate a set of questions based on a patient's case.) Your primary goal is to evaluate questions based on their fluency, relevance to the case, harmfulness, complexity, and personalization. Please use the following scale to evaluate each criterion:
[0117] (Add rating scale here)
[0118] The case you are referring to is:
[0119] (Add the case content here)
[0120] The questions you need to assess are:
[0121] (Add question content here)
[0122] The output you provide should be formatted in JSON and have the following structure:
[0123] {"fluency":"",
[0124] "Relevance":"",
[0125] "Hazard":"",
[0126] "Complexity":"",
[0127] "Level of Personalization":""}
[0128] For each part, directly fill in the corresponding score of the question. Please directly output a piece of Json format data without adding any other description.
[0129] In some embodiments, the first scoring information includes: a question fluency score, a question relevance score, a question harmfulness score, a complexity score, and a personalization score.
[0130] The first screening criteria include: the question harmfulness score is 0, and the question fluency score, question relevance score, complexity score and personalization score are not less than half of the total score.
[0131] In this embodiment, to ensure that the answers output by the subsequent trained industry model do not contain obvious errors, this application provides a screening criterion for question data, wherein the first scoring information includes:
[0132] The question information is scored for fluency, relevance, harmfulness, complexity, and personalization. The scoring criteria are as follows:
[0133] Question Fluency Scoring:
[0134] 1: The sentences are completely unreadable and incomprehensible.
[0135] 2: There are some coherent phrases in the sentence, but the whole sentence is difficult to understand.
[0136] 3: The sentence is generally fluent, but contains a few difficult-to-understand parts.
[0137] 4: The sentences are relatively fluent, easy to understand, and the expression is natural.
[0138] 5: The sentences are very fluent, the expression is clear and easy to understand, without any difficult parts.
[0139] Question relevance score:
[0140] 1: Completely irrelevant. The question has no obvious connection with the case content and may involve different topics.
[0141] 2: Partially related. The question slightly involves case information, but deviates from the core content and has a weak connection with the field of rehabilitation medicine.
[0142] 3: Mostly relevant. The question covers some important information in the case, but ignores some key content.
[0143] 4: Highly relevant, the questions accurately reflect the key points in the case and cover the main rehabilitation issues.
[0144] 5: Completely relevant. The questions accurately capture the core content of the case and reflect details and in-depth thinking.
[0145] Problem severity rating:
[0146] 1: The problem has obvious ethical and safety issues;
[0147] 0: The problem does not pose any ethical safety issues;
[0148] Problem Complexity Rating:
[0149] 1: The question is very simple, does not involve complex factors of rehabilitation treatment, and is only a basic statement or factual question.
[0150] 2: The question is slightly in-depth, but still stays on the surface and lacks guidance for rehabilitation decisions.
[0151] 3: The question is moderately complex, covering major factors in rehabilitation treatment, can guide reasonable discussion and decision-making, and is somewhat challenging but easy to understand.
[0152] 4: The question covers multiple complex variables and requires in-depth reasoning, suitable for expert discussion or special rehabilitation cases.
[0153] 5: The question involves highly complex logic or interdisciplinary content, which is difficult to answer.
[0154] Personalization score:
[0155] 1: Completely lack of personalization, the question is general and not related to the specific patient.
[0156] 2: Mild personalization, the question partially involves the patient's background or needs, but is more general.
[0157] 3: Moderate personalization, the question takes into account some individual factors of the patient, but does not delve into the patient's special needs.
[0158] 4: High personalization, the question fully reflects the individual differences, medical history and rehabilitation needs of the patient.
[0159] 5: Complete personalization, the question deeply considers the patient's special background and unique challenges in rehabilitation.
[0160] Question coverage: The proportion of the content covered by the question in the case content.
[0161] Based on the current scoring criteria, when the question danger score is not 0, it is determined that the question data at this time has ethical and safety issues for the patient's diagnosis, and the current question may pose a threat to the patient or others. At this time, all question danger scores that are not 0 should be deleted or re-input prompt to optimize them. When the question fluency score and question relevance score are less than 2.5 points, the current question provides less help to the training process of the required industry large model, and should be directly deleted. For question data that meets the requirements, add it to the question data set.
[0162] Preferably, the language model cluster can also be inputted with a first comparison text form to make the language model cluster evaluate the content amount of the question data set in the content amount of the case data, and the first comparison text specifically includes:
[0163] You are an AI assessor, specializing in calculating the proportion of the content covered by a given set of questions in a case. Your main goal is to think about the meaning of the question content and calculate the proportion of the information contained in the set of questions to the given case content based on the content of all questions:
[0164] The case you are referring to is:
[0165] (Add the case content here)
[0166] The set of questions you need to assess is:
[0167] (Add the current question set content here)
[0168] Please directly output an integer from 0 to 100 to indicate the percentage of the content covered by the given question set in the content of the case. No other description is required.
[0169] After the comparison, questions can be further generated by inputting a first optimized text into the first language model and conducting multiple rounds of dialogue with the first language model based on the previous dialogue content, so that the first language model continues to output the required question data. The first optimized text specifically includes:
[0170] The questions you've currently generated cover approximately ()% of the case content. Following the above ideas, generate fifteen more questions in different directions to address the remaining areas. Please generate questions that are less complex, easier to answer, more personalized, and fully reflect the individual differences of the patients. You should still adhere to the principles of question generation mentioned previously. Question descriptions should be specific and detailed, and should not require reliance on the case content to understand. Please directly input the questions in Chinese, without any additional descriptions.
[0171] Among them, "Please generate some questions that are less complex, easier to answer, more personalized, and fully reflect the individual differences of patients" will be automatically adjusted by the script according to the average complexity and average personalization of the current question set. For example, when the average complexity of the current question set is lower than 3, "less complex and easier to answer" will be adjusted to "more complex, requiring in-depth reasoning based on professional knowledge to answer."
[0172] The above process is then repeated, and the optimized question data and the first scored text are re-input into the language model cluster. The optimized question data is scored and the question data that does not meet the standards is removed. The question data that meets the requirements is added to the question data set. When the content of the question set accounts for more than 95%, the content in the question set is input as question information into the first language model for the next step.
[0173] In some embodiments, as Figure 3 As shown, step S3 specifically includes:
[0174] Step S31: Generate a second scoring text and a second screening criterion based on the reference data, wherein the second screening criterion is set based on the model fine-tuning data.
[0175] Step S32: inputting the second scoring text into the preset third language model so that the third language model outputs second scoring information associated with the reference data.
[0176] In this embodiment, the third language model is used to score the reference data generated by the first language model based on the second scoring text and the second screening criteria. The large language model is used to verify whether the reference data generated by the first language model meets the screening criteria for the model fine-tuning data. The third language model can also be controlled to mark or eliminate reference data that does not meet the screening criteria, thereby achieving the output of high-quality reference data.
[0177] Preferably, in order to improve the reliability of the second scoring information, the third language model can be a language model cluster composed of multiple large language models. By inputting the second scoring text into multiple different large language models in the language model cluster, all large prediction models in the language model cluster output scoring information for the reference data, and removing the highest and lowest scores in the scoring, calculating the average score of other language models, and using the average score as the second scoring information.
[0178] Among them, the language model cluster refers to a collection of a large number of different commercial general-purpose large models such as ChatGPT, Claude, Qwen, Gemini, etc., and the same commercial general-purpose large model using different prompt templates. In this example, two sets of prompt templates are set up for each large model, allowing the model to play the role of an artificial intelligence evaluator and a professional physician in the rehabilitation medicine field respectively. Each large model in the collection will evaluate the quality of the generated questions separately, and finally use an automated script to summarize all the results and calculate the average to obtain the final result.
[0179] The second scoring text is input to the third language model, so that the third language model outputs a score for the reference data.
[0180] The second screening criterion is set according to the requirements of model fine-tuning data, and is used to screen out reference data suitable for training large industry models.
[0181] The second scoring information is derived based on the second screening criteria. The third language model scores the reference data based on the second screening criteria, generating the second scoring information by aggregating the scoring information from multiple dimensions. Multiple sets of second scoring texts can be input to the third language model through an automated program, thereby causing the third language model to automatically output multiple sets of second scoring information.
[0182] Step S33: Generate a second optimized text, filter the second scoring information according to the second filtering criteria, and input the reference data corresponding to the second scoring information that does not meet the second filtering criteria, the problem data corresponding to the reference data, and the second optimized text into the first language model so that the first language model optimizes the reference data.
[0183] Among them, by optimizing the reference information in the reference data that does not meet the second screening criteria, it is possible to avoid interference of bad data on the large industry model to be trained.
[0184] The second scoring text includes:
[0185] You are an AI assessor, specializing in evaluating the quality of the thought process generated in response to a given problem. Your primary goal is to assess the thought process based on its fluency, relevance to the given problem, harmfulness, accuracy, rationality, and specificity, using the following scale to evaluate each criterion:
[0186] (Add scale here)
[0187] The case you are referring to is:
[0188] (Add the case content here)
[0189] The given question is:
[0190] (Add question content here)
[0191] The thought processes you need to assess are:
[0192] (Add thought process here)
[0193] The output you provide should be formatted in JSON and have the following structure:
[0194] {"fluency":"",
[0195] "Relevance":"",
[0196] "Hazard":"",
[0197] "accuracy":"",
[0198] "reasonableness":"",
[0199] "Specificity":""}
[0200] For each part, directly fill in the corresponding score of the question. Please directly output a piece of Json format data without adding any other description.
[0201] In some embodiments, the second scoring information includes: a reference fluency score, a reference relevance score, a reference harmfulness score, a reference accuracy score, a reference rationality score, and a specificity score.
[0202] The second screening criteria include: the reference harmfulness score is 0, and the reference fluency score, reference relevance score, reference accuracy score, reference rationality score and specificity score are not less than half of the total score.
[0203] In this embodiment, to ensure that the answers output by the subsequent industry model training do not contain obvious errors, this application provides a screening standard for reference data, wherein the second scoring information includes:
[0204] The reference information is scored for reference fluency, relevance, harmfulness, accuracy, rationality, and specificity respectively. The scoring criteria are as follows:
[0205] Fluency:
[0206] 1: The sentences are completely unreadable and incomprehensible.
[0207] 2: There are some coherent phrases in the sentence, but the whole sentence is difficult to understand.
[0208] 3: The sentence is generally fluent, but contains a few difficult-to-understand parts.
[0209] 4: The sentences are relatively fluent, easy to understand, and the expression is natural.
[0210] 5: The sentences are very fluent, the expression is clear and easy to understand, without any difficult parts.
[0211] Dependencies:
[0212] 1: Completely irrelevant. The generated thought process has no obvious connection to the question content and may involve an entirely different topic.
[0213] 2: Partially relevant, but overall deviates from the key content of the question, has only a slight connection with the peripheral information of the question, and does not answer the core of the question.
[0214] 3: Mostly relevant. The thinking process answers some important information in the question, but ignores some key points.
[0215] 4: Highly relevant. The thinking process accurately reflects the key content of the question and covers the main rehabilitation issues.
[0216] 5: Completely relevant. The thought process accurately captures the core content of the question and reflects details and in-depth thinking.
[0217] Hazards:
[0218] 1: There are obvious ethical safety issues in the thinking process.
[0219] 0: There are no ethical safety issues in the thinking process.
[0220] accuracy:
[0221] 1: The thinking process contains obvious factual errors that conflict with the case content and cannot be used as a reference for answering questions.
[0222] 2: There are no factual errors in the thinking process as a whole, but there are a few logical inference problems.
[0223] 3: The overall thinking process contains no factual errors or logical problems, but it is difficult to give effective suggestions for the unresolved parts of the problem.
[0224] 4: The thinking process is basically an accurate thinking about the problem, and it can also provide effective inspection suggestions for the parts of the problem that cannot be solved at present, but some content is not accurate enough.
[0225] 5: The thinking process is very good at thinking about the problem, and it can be used as a reference to solve the problem well.
[0226] Rationality:
[0227] 1. The "Possible Diagnosis" and "Supporting Evidence" sections contain a large amount of factual content beyond the problem description, which cannot be derived from the problem content in actual scenarios;
[0228] 2: The "Possible Diagnosis" and "Supporting Evidence" sections contain a small amount of factual content beyond the problem description, which is difficult to derive from the problem content in real scenarios;
[0229] 3: There is almost no factual content in the "Possible Diagnosis" and "Supporting Evidence" sections beyond the problem description, and a certain amount of content must be inferred based on the problem content;
[0230] 4: There is no factual content in the "Possible Diagnosis" and "Supporting Evidence" sections beyond the problem description, and a small amount of content must be inferred based on the problem content;
[0231] 5: The content in the "Possible Diagnosis" and "Supporting Evidence" sections can be directly derived from the question content;
[0232] Specificity:
[0233] 1. The "Further Examination Recommendations" and "Treatment Recommendations" sections are general and cannot be used as a guide;
[0234] 2: The "further examination recommendations" and "treatment recommendations" sections only provide a rough outline, making it difficult to come up with specific measures;
[0235] 3: The "Further Examination Recommendations" and "Treatment Recommendations" sections provide some general guidance and are of certain reference value;
[0236] 4: The "Further Examination Recommendations" and "Treatment Recommendations" sections provide explanations of the examination plan, rehabilitation treatment items, and medication regimen, but the specific frequency and duration of treatment are unclear.
[0237] 5. The "Further Examination Recommendations" and "Treatment Recommendations" sections clearly describe the examination plan, rehabilitation treatment items, and medication plan, down to the precise examination items, medication usage and dosage, and the frequency and purpose of rehabilitation programs.
[0238] In particular, the scoring help here has a small number of variations, and adjustments need to be made for the parts of the thinking process that are "not applicable". For example, if the "further inspection suggestions" in the thinking process are "not applicable", then the specific degree part of the scale does not need to evaluate the "further inspection suggestions".
[0239] The reference data is scored based on the current scoring criteria. If the reference hazard score is non-zero, it is determined that the reference data presents ethical safety issues for patient diagnosis and may pose a threat to the patient or other individuals. All reference data with a non-zero reference hazard score should be re-entered into the prompt for optimization. If the reference fluency score, reference relevance score, reference accuracy score, reference rationality score, and specificity score are below 2.5, the current reference data is of little use in the subsequent training of the industry's large model and should be optimized through multiple rounds of dialogue with the first language model.
[0240] Preferably, since the problem data has undergone a round of screening, the screening criteria for the reference data can be further improved. For example, the screening criteria for the reference relevance score can be increased to 3 points, and the screening criteria for the reference accuracy score, reference rationality score and specificity can be increased to 3.5 points, thereby improving the quality of the model fine-tuning data.
[0241] In particular, it is also possible to provide a method in which, based on the reference data generated by the first language model, a second optimized text is input into the first language model for multiple rounds of dialogue so that the first language model optimizes the reference data. The second optimized text specifically includes:
[0242] Please review your thought process. Your current thought process sentence contains some consecutive phrases, making it difficult to read. The "Further Examination Suggestions" and "Treatment Suggestions" sections only provide a general plan, making it difficult to come up with specific measures. Please try to revise your thought process to make your sentences smooth and clear without any difficult-to-understand content. In the "Further Examination Suggestions" and "Treatment Suggestions" sections, clearly describe the examination plan, rehabilitation treatment items, and medication plan, down to the examination items, medication usage and dosage, and the frequency and purpose of the rehabilitation program. Your output should still be formatted in JSON, with the following structure:
[0243] {"Possible Diagnosis":"",
[0244] "Supporting evidence":"",
[0245] "Unsupported parts":"",
[0246] "Further inspection suggestions":"",
[0247] "Treatment recommendations": "Drug name 1, dosage, usage, frequency, drug name 2, dosage, usage, frequency, ... Rehabilitation treatment item 1, frequency, purpose, Rehabilitation treatment item 2, frequency, purpose, ...",
[0248] "Expected effect":""
[0249] }
[0250] If any part does not apply, please answer "Not applicable". Please directly enter the results of your thinking process for all questions in Chinese without adding any other description.
[0251] In particular, the second optimized text has a certain amount of variations, which will be automatically and dynamically adjusted according to the score of the reference data. For example, when the specific degree score meets the requirements, the texts such as "In the "Further examination suggestions" and "Treatment suggestions" sections, only a general plan is given, and it is difficult to come up with specific measures" and "In the "Further examination suggestions" and "Treatment suggestions" sections, the inspection plan, rehabilitation treatment items, and medication plan are clearly described, accurate to the inspection items, usage and dosage of drugs, and the frequency and purpose of rehabilitation items" should be removed. In this application, when reference data that meets the standards cannot be obtained after five rounds of dialogue, the reference data with the highest average score and a harmfulness score of 0 generated during the dialogue will be selected as the result.
[0252] In some embodiments, as Figure 4 As shown, step S4 specifically includes:
[0253] Step S41: Generate a third scoring text and a third screening criterion based on the reference data, wherein the third screening criterion is set based on the model fine-tuning data.
[0254] Step S42: inputting the third scoring text into the preset fourth language model so that the fourth language model outputs third scoring information associated with the answer data.
[0255] In this embodiment, the fourth language model is used to score the answer data generated by the first language model based on the third scoring text and the third screening criteria. The large language model is used to verify whether the answer data generated by the first language model meets the screening criteria of the model fine-tuning data. The fourth language model can also be controlled to mark or eliminate answer data that does not meet the screening criteria, thereby achieving the output of high-quality answer data.
[0256] Preferably, in order to improve the reliability of the third scoring information, the fourth language model can be a language model cluster composed of multiple large language models. By inputting the third scoring text into multiple different large language models in the language model cluster, all large prediction models in the language model cluster output scoring information for the answer data, and removing the highest and lowest scores in the scoring, calculating the average score of other language models, and using the average score as the third scoring information.
[0257] Among them, the language model cluster refers to a collection of a large number of different commercial general-purpose large models such as ChatGPT, Claude, Qwen, Gemini, etc., and the same commercial general-purpose large model using different prompt templates. In this example, two sets of prompt templates are set up for each large model, allowing the model to play the role of an artificial intelligence evaluator and a professional physician in the rehabilitation medicine field respectively. Each large model in the collection will evaluate the quality of the generated questions separately, and finally use an automated script to summarize all the results and calculate the average to obtain the final result.
[0258] The third scoring text is input to the fourth language model, so that the fourth language model outputs a score for the answer data.
[0259] The third screening criterion is set according to the requirements of model fine-tuning data, and is used to filter out answer data suitable for training large industry models.
[0260] The third scoring information is derived based on the third screening criteria, and the fourth language model scores the answer data based on the third screening criteria, generating the third scoring information by aggregating the scoring information from multiple dimensions. Multiple sets of third scoring texts can be input to the fourth language model through an automated program, thereby causing the fourth language model to automatically output multiple sets of third scoring information.
[0261] In step S43, a third optimized text is generated, the third score information is screened according to a third screening standard, and the answer data corresponding to the third score information that does not meet the third screening standard, the reference data corresponding to the answer data, the question data, and the third optimized text are input to the first language model to optimize the answer data by the first language model.
[0262] Among them, by optimizing the answer information in the answer data that does not meet the third screening standard, the interference of bad data on the industry large model to be trained is avoided.
[0263] Among them, the third score text includes:
[0264] You are an artificial intelligence evaluator who specializes in evaluating the quality of answers to given problems based on the thinking process provided. Your main goal is to assess the thinking process based on the fluency, relevance, harmfulness, accuracy, and reasonableness of the answer content, referring to the given case content. Use the following scale to evaluate each criterion:
[0265] (Add the scale here)
[0266] The case content you refer to is:
[0267] (Add the case content here)
[0268] The given question and thinking process are:
[0269] (Add the question content and thinking process content here)
[0270] The answer you need to evaluate is:
[0271] (Add the answer content here)
[0272] The output you provide should be in JSON format, with the following structure:
[0273] {"fluency": "",
[0274] "relevance": "",
[0275] "harmfulness": "",
[0276] "accuracy": "",
[0277] "reasonableness": ""}
[0278] Where each part directly fills in the score corresponding to the question. Directly output a JSON format data, without adding any other description.
[0279] In some embodiments, the third scoring information includes: an answer fluency score, an answer relevance score, an answer harmfulness score, an answer accuracy score, and an answer rationality score.
[0280] The third screening criteria include: the answer harmfulness score is 0, and the answer fluency score, answer relevance score, answer accuracy score and answer rationality score are not less than half of the total score.
[0281] In this embodiment, to ensure that the answers output by the subsequent industry model training do not contain obvious errors, this application provides a screening standard for answer data, wherein the third scoring information includes:
[0282] The answer fluency score, answer relevance score, answer harmfulness score, answer accuracy score and answer rationality score of the question information are scored separately. The scoring criteria are as follows:
[0283] Fluency:
[0284] 1: The sentences are completely unreadable and incomprehensible.
[0285] 2: There are some coherent phrases in the sentence, but the whole sentence is difficult to understand.
[0286] 3: The sentence is generally fluent, but contains a few difficult-to-understand parts.
[0287] 4: The sentences are relatively fluent, easy to understand, and the expression is natural.
[0288] 5: The sentences are very fluent, the expression is clear and easy to understand, without any difficult parts.
[0289] Dependencies:
[0290] 1: Completely irrelevant. The generated answer has no obvious connection with the question content and may involve a completely different topic.
[0291] 2: Partially relevant, but overall deviates from the key content of the question, has only a slight connection with the peripheral information of the question, and does not answer the core of the question.
[0292] 3: Mostly relevant, the generated answer explains some important information in the question, but ignores some key points.
[0293] 4: Highly relevant. The generated answers accurately reflect the key content of the questions and cover the main rehabilitation issues.
[0294] 5: Completely relevant. The generated answer accurately captures the core content of the question and provides a positive response to the question.
[0295] Hazards:
[0296] 1: There are obvious ethical safety issues in the generated answers.
[0297] 0: There are no ethical safety issues in the generated answers.
[0298] accuracy:
[0299] 1: The generated answer contains obvious factual errors that conflict with the case content.
[0300] 2: The generated answers are generally free of factual errors, but have a few logical inference problems.
[0301] 3: The generated answer is generally free of factual errors and logical problems, but does not fully capture the content of the thought process.
[0302] 4: The generated answer is basically a feasible solution to the problem, which basically contains the content obtained during the thinking process.
[0303] 5: The generated answer provides a good answer to the question, comprehensively covers the content of the thinking process, and well solves the patient's problem or provides a clear examination plan.
[0304] Rationality:
[0305] 1: The answer contains a lot of factual content beyond the question description, which cannot be derived from the question content in actual scenarios.
[0306] 2: The answer contains a small amount of factual content beyond the question description, which is difficult to draw conclusions based on the question content in actual scenarios.
[0307] 3: There is almost no factual content in the answer beyond the question description, and a certain amount of content must be inferred based on the question content.
[0308] 4: The answer contains no factual content beyond the question description, and a small amount of content needs to be inferred based on the question content.
[0309] 5: The content of the answer can be directly derived from the content of the question.
[0310] The answer data is scored based on the current scoring criteria. If the answer's hazard score is non-zero, it is determined that the answer data presents ethical safety issues for the patient's diagnosis and that the current answer may pose a threat to the patient or others. All answers with a non-zero hazard score should be re-entered into the prompt for optimization. If the answer fluency score, answer relevance score, answer accuracy score, and answer rationality score are below 2.5, the current answer is of little help in the subsequent training of the large industry model required and should be optimized through multiple rounds of dialogue with the first language model.
[0311] In particular, a method may be provided in which, based on the answer data generated by the first language model, a third optimized text is inputted into the first language model for multiple rounds of dialogue so that the first language model optimizes the answer data. The third optimized text specifically includes:
[0312] Please review your response. Your current answer contains some run-on phrases, making it difficult to read. While your overall answer is factually sound and logically sound, it doesn't fully capture your thought process. Please review your response to ensure it flows smoothly, contains no incomprehensible content, and fully captures your thought process. It should address the patient's concerns or provide a clear examination plan. Please directly output your answer in Chinese, without adding any additional descriptions.
[0313] In particular, the third optimized text has a certain amount of variation and will be automatically and dynamically adjusted based on the score of the answer data. For example, if the accuracy score is only 2 points, the statement "The answer is generally free of factual errors and logical problems, but does not fully encompass the content of the thinking process" should be revised to "The generated answer contains obvious factual errors that conflict with the case content." In this application, if no reference data that meets the standards can be obtained after five rounds of dialogue, the reference data generated during the dialogue with the highest average score and a harmfulness score of 0 will be selected as the result.
[0314] As can be seen from the above embodiments of this application, this application uses real case data as a reference and calls a large language model through an automated program, so that the large language model generates question data, reference data, and answer data in batches, and guides the large language model to generate and optimize model fine-tuning data through the set prompt words, thereby generating specialized model fine-tuning data for the medical field industry model, saving the training time of subsequent industry large models. In addition, due to the high difficulty in obtaining case data and the high amount of private information, the technical solution disclosed in this application also has the advantages of low cost and strong privacy compared to directly obtaining real case information.
[0315] This application also provides screening criteria for question data, reference data, and answer data. By setting specific prompts, it guides the large language model to score and screen question data, reference data, and answer data, thereby obtaining more accurate model fine-tuning data for subsequent industry large model training.
[0316] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0317] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the disclosed technical solution based on the prompt message.
[0318] As an optional but non-limiting implementation, in response to a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0319] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0320] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario and performed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the method.
[0321] It should be noted that the above description is limited to some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0322] Based on the same technical concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a device for generating large-scale model training data in the medical field.
[0323] refer to Figure 5 , a device for generating large model training data in the medical field, including:
[0324] The case reading module 601 is used to obtain case data and generate a first prompt text according to the case data.
[0325] The question generating module 602 is configured to input a first prompt text into a preset first language model so that the first language model outputs question data associated with the case data.
[0326] The reference data generation module 603 is used to generate a second prompt text based on the question data so that the first language model can output reference data. The reference data is used to express the processing process of the first language model generating the question data.
[0327] The answer generation module 604 is configured to generate a third prompt text based on the reference data, so that the first language model outputs answer data.
[0328] The fine-tuning data integration module 605 is used to generate model fine-tuning data based on the question data, reference data and answer data.
[0329] For the convenience of description, the above devices are described as being divided into various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0330] The device of the above embodiment is used to implement the corresponding medical field large model training data generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0331] Based on the same technical concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the method for generating large model training data in the medical field described in any of the above embodiments.
[0332] Figure 6 10 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0333] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0334] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0335] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0336] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0337] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0338] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0339] The electronic device of the above embodiment is used to implement the corresponding medical field large model training data generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0340] Based on the same technical concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the method for generating large model training data in the medical field as described in any of the above embodiments.
[0341] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0342] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the method for generating large model training data in the medical field as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0343] It should be noted that the embodiments of the present application can be further described in the following manner:
[0344] A method for generating large-scale model training data in the medical field, characterized by comprising:
[0345] Obtain case data and generate a first prompt text based on the case data.
[0346] A first prompt text is input into a preset first language model so that the first language model outputs question data associated with the case data.
[0347] A second prompt text is generated based on the question data so that the first language model outputs reference data, wherein the reference data is used to express the processing process of the first language model generating the question data.
[0348] A third prompt text is generated based on the reference data so that the first language model outputs answer data.
[0349] Generate model fine-tuning data based on question data, reference data and answer data.
[0350] Preferably, inputting a first prompt text into a preset first language model so that the first language model outputs question data associated with the case data specifically includes:
[0351] A first scoring text and a first screening criterion are generated based on the question data, wherein the first screening criterion is set based on the model fine-tuning data.
[0352] The first scoring text is input into a preset second language model, so that the second language model outputs first scoring information associated with the question data.
[0353] The first scoring information is screened according to the first screening criteria, and the problem data corresponding to the first scoring information that does not meet the first screening criteria is deleted.
[0354] Preferably, the first scoring information includes: question fluency score, question relevance score, question harmfulness score, complexity score and personalization degree score.
[0355] The first screening criteria include: the question harmfulness score is 0, and the question fluency score, question relevance score, complexity score and personalization score are not less than half of the total score.
[0356] Preferably, generating a second prompt text based on the question data so that the first language model outputs reference data specifically includes:
[0357] A second scoring text and a second screening criterion are generated based on the reference data, wherein the second screening criterion is set based on the model fine-tuning data.
[0358] The second scoring text is input into the preset third language model, so that the third language model outputs second scoring information associated with the reference data.
[0359] Generate a second optimized text, filter the second scoring information according to the second filtering criteria, and input the reference data corresponding to the second scoring information that does not meet the second filtering criteria, the problem data corresponding to the reference data, and the second optimized text into the first language model so that the first language model optimizes the reference data.
[0360] Preferably, the second scoring information includes: a reference fluency score, a reference relevance score, a reference harmfulness score, a reference accuracy score, a reference rationality score and a specificity score.
[0361] The second screening criteria include: the reference harmfulness score is 0, and the reference fluency score, reference relevance score, reference accuracy score, reference rationality score and specificity score are not less than half of the total score.
[0362] Preferably, generating a third prompt text based on the reference data so that the first language model outputs answer data specifically includes:
[0363] A third scoring text and a third screening criterion are generated based on the reference data, wherein the third screening criterion is set based on the model fine-tuning data.
[0364] The third scoring text is input into a preset fourth language model, so that the fourth language model outputs third scoring information associated with the answer data.
[0365] A third optimized text is generated, the third scoring information is screened according to a third screening criterion, and the answer data corresponding to the third scoring information that does not meet the third screening criterion, the reference data and question data corresponding to the answer data, and the third optimized text are input into the first language model, so that the first language model optimizes the answer data.
[0366] Preferably, the third scoring information includes: answer fluency score, answer relevance score, answer harmfulness score, answer accuracy score and answer rationality score.
[0367] The third screening criteria include: the answer harmfulness score is 0, and the answer fluency score, answer relevance score, answer accuracy score and answer rationality score are not less than half of the total score.
[0368] A device for generating large-scale model training data in the medical field, characterized by comprising:
[0369] The case reading module is used to obtain case data and generate a first prompt text based on the case data.
[0370] The question generation module is used to input a first prompt text into a preset first language model so that the first language model outputs question data associated with the case data.
[0371] The reference data generation module is used to generate a second prompt text based on the question data so that the first language model outputs reference data. The reference data is used to express the processing process of the first language model generating the question data.
[0372] The answer generation module is used to generate a third prompt text based on the reference data so that the first language model outputs answer data.
[0373] The fine-tuning data integration module is used to generate model fine-tuning data based on question data, reference data and answer data.
[0374] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements any one of the above methods when executing the program.
[0375] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute any one of the methods described above.
[0376] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. Within the scope of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0377] In addition, for simplicity of description and discussion, and in order not to make the embodiment of the application difficult to understand, the known power supply / ground connection with integrated circuit (IC) chip and other components may or may not be shown in the accompanying drawings provided. In addition, the device can be shown in the form of a block diagram to avoid making the embodiment of the application difficult to understand, and this also takes into account the following fact, that is, the details of the embodiment of these block diagram devices are highly dependent on the platform to be implemented in the embodiment of the application (that is, these details should be fully within the scope of understanding of those skilled in the art). When specific details (for example, circuit) are set forth to describe exemplary embodiments of the application, it will be apparent to those skilled in the art that the embodiment of the application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.
[0378] Although the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may utilize the embodiments discussed.
[0379] The embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of this application.
Claims
1. A method for generating large model training data in the medical field, characterized in that: include: Acquire case data, and generate a first prompt text based on the case data; Inputting the first prompt text into a preset first language model so that the first language model outputs question data associated with the case data; Generate a first scoring text and a first screening criterion based on the question data; wherein the first screening criterion is set based on the required model fine-tuning data; Inputting the first scoring text into a preset second language model so that the second language model outputs first scoring information associated with the question data; Filtering the first scoring information according to the first screening criteria, and deleting the problem data corresponding to the first scoring information that does not meet the first screening criteria; generating a second prompt text based on the processed question data, so that the first language model outputs reference data; wherein the reference data is used to express the processing process of the first language model generating the question data; Generate a second scoring text and a second screening criterion based on the reference data; wherein the second screening criterion is set based on the model fine-tuning data; inputting the second scoring text into a preset third language model so that the third language model outputs second scoring information associated with the reference data; generating a second optimized text, filtering the second scoring information according to the second filtering criteria, and inputting the reference data corresponding to the second scoring information that does not meet the second filtering criteria, the question data corresponding to the reference data, and the second optimized text into the first language model, so that the first language model optimizes the reference data; generating a third prompt text based on the reference data so that the first language model outputs answer data; Generating a third scoring text and a third screening criterion based on the optimized reference data; wherein the third screening criterion is set based on the model fine-tuning data; inputting the third scoring text into a preset fourth language model so that the fourth language model outputs third scoring information associated with the answer data; generating a third optimized text, filtering the third scoring information according to the third filtering criteria, and inputting the answer data corresponding to the third scoring information that does not meet the third filtering criteria, the reference data and question data corresponding to the answer data, and the third optimized text into the first language model, so that the first language model optimizes the answer data; The model fine-tuning data is generated according to the processed question data, the optimized reference data and the optimized answer data.
2. A method for generating large-scale model training data in the medical field according to claim 1, characterized in that: The first scoring information includes: question fluency score, question relevance score, question harmfulness score, complexity score and personalization score; The first screening criteria include: the question harmfulness score is 0, and the question fluency score, the question relevance score, the complexity score and the personalization score are not less than half of the total score.
3. The method for generating large-scale model training data in the medical field according to claim 1, characterized in that: The second scoring information includes: a reference fluency score, a reference relevance score, a reference harmfulness score, a reference accuracy score, a reference rationality score, and a specificity score; The second screening criterion includes: the reference hazard score is 0, and the reference fluency score, the reference relevance score, the reference accuracy score, the reference rationality score and the specificity score are not less than half of the total score.
4. A method for generating large-scale model training data in the medical field according to claim 1, characterized in that: The third scoring information includes: answer fluency score, answer relevance score, answer harmfulness score, answer accuracy score and answer rationality score; The third screening criterion includes: the answer harmfulness score is 0, and the answer fluency score, the answer relevance score, the answer accuracy score and the answer rationality score are not less than half of the total score.
5. A device for generating large-scale model training data in the medical field, characterized in that: include: A case reading module, configured to obtain case data and generate a first prompt text according to the case data; a question generating module, configured to input the first prompt text into a preset first language model, so that the first language model outputs question data associated with the case data; Generate a first scoring text and a first screening criterion based on the question data; wherein the first screening criterion is set based on the required model fine-tuning data; Inputting the first scoring text into a preset second language model so that the second language model outputs first scoring information associated with the question data; Filtering the first scoring information according to the first screening criteria, and deleting the problem data corresponding to the first scoring information that does not meet the first screening criteria; a reference data generation module, configured to generate a second prompt text based on the processed question data, so that the first language model outputs reference data; wherein the reference data is used to express the processing process of the first language model generating the question data; Generate a second scoring text and a second screening criterion based on the reference data; wherein the second screening criterion is set based on the model fine-tuning data; inputting the second scoring text into a preset third language model so that the third language model outputs second scoring information associated with the reference data; generating a second optimized text, filtering the second scoring information according to the second filtering criteria, and inputting the reference data corresponding to the second scoring information that does not meet the second filtering criteria, the question data corresponding to the reference data, and the second optimized text into the first language model, so that the first language model optimizes the reference data; an answer generation module, configured to generate a third prompt text based on the reference data, so that the first language model outputs answer data; Generating a third scoring text and a third screening criterion based on the optimized reference data; wherein the third screening criterion is set based on the model fine-tuning data; inputting the third scoring text into a preset fourth language model so that the fourth language model outputs third scoring information associated with the answer data; generating a third optimized text, filtering the third scoring information according to the third filtering criteria, and inputting the answer data corresponding to the third scoring information that does not meet the third filtering criteria, the reference data and question data corresponding to the answer data, and the third optimized text into the first language model, so that the first language model optimizes the answer data; The fine-tuning data integration module is used to generate the model fine-tuning data based on the processed question data, the optimized reference data and the optimized answer data.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.
7. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Question and answer pair generation method and system based on large language model
CN118332086A
Method and device for generating fine tuning data of large language model in target domain
CN118536604A