Model training method and device, equipment and storage medium

By training medical report interpretation large language models, using thinking chains to construct prompt words and fine-tune supervision, the problem of automated interpretation of medical reports is solved and the accuracy and interpretability of interpretation is improved.

CN120541529AActive Publication Date: 2025-08-26ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511046848.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-08-26
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

There are a large number of medical terms and professional indicators in medical reports, which leads to medical staff needing to manually interpret and translate them into content that patients can understand, and lack of automated interpretation solutions.

Method used

By training medical reports to interpret large language models, using thinking chains to construct prompt words to guide the large language models to generate reference thinking chains and interpretation results, combined with supervision fine-tuning and reinforcement learning, output interpretation results and reasoning process.

Benefits of technology

It realizes automatic interpretation of medical reports, improves the accuracy and interpretability of interpretation results, and enhances users' trust in interpretation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541529A_ABST
    Figure CN120541529A_ABST
Patent Text Reader

Abstract

One or more embodiments of the invention provide a model training method and device, equipment and a storage medium, which are used for training a medical report interpretation large language model. Therefore, the trained medical report interpretation large language model can output a target medical report interpretation result corresponding to the input target medical report and the target to-be-answered question; and a reasoning process used for describing a target medical report interpretation result corresponding to the target to-be-answered question obtained based on the text content in the target medical report, namely a target thinking chain, can be output. In a specific training process, a first training sample including a medical report, a question to be answered and a first medical report interpretation result is acquired; outputting a reference thinking chain corresponding to the first training sample through the target large language model under the supervision of the first medical report interpretation result; and then adding the reference thinking chain to the first training sample to generate a second training sample, and training the medical report interpretation large language model through the second training sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of artificial intelligence technology, and in particular, to a model training method, apparatus, device, and storage medium. Background Art

[0002] In healthcare settings, medical professionals typically conduct interviews or medical examinations to assess a patient's health. Medical examinations can be categorized into various types, including imaging tests like X-rays, computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound; laboratory tests like blood tests, urine tests, and stool tests; and functional tests like electrocardiograms (ECGs), pulmonary function tests, and endoscopy. After the interview or medical examination, medical professionals typically issue a medical report documenting the patient's health status, examination conclusions, diagnosis, and treatment recommendations.

[0003] Because medical reports contain numerous medical terms, specialized indicators, and clinical judgment logic, they typically require medical personnel with specialized medical knowledge to interpret them. This requires translating these terms, indicators, and clinical judgment logic into easily understandable descriptions, enabling patients to understand their health status and treatment information. Therefore, how to automatically translate medical reports into understandable content for patients has become a pressing issue. Summary of the Invention

[0004] One or more embodiments of this specification provide a model training method, apparatus, device, and storage medium for training a medical report interpretation model capable of automatically interpreting medical reports.

[0005] In a first aspect, one or more embodiments of this specification provide a model training method, the method comprising: Obtaining a first training sample, the first training sample comprising: a medical report, a question to be answered, and a first medical report interpretation result obtained by answering the question to be answered based on text content in the medical report; Inputting the prompt words for constructing a thought chain containing the first training sample into a target large language model, so that the target large language model outputs a reference thought chain and a second medical report interpretation result under the supervision of the first medical report interpretation result, wherein the reference thought chain is used to describe the reasoning process of obtaining the second medical report interpretation result corresponding to the question to be answered based on the text content in the medical report; If it is determined that the second medical report interpretation result is correct according to the first medical report interpretation result, the reference thought chain is added to the first training sample to generate a second training sample; The large language model for medical report interpretation is trained using the second training sample. The large language model for medical report interpretation is used to output a target medical report interpretation result and a target thinking chain based on the target medical report and target question to be answered input by the user. The target thinking chain is used to describe the reasoning process of obtaining the target medical report interpretation result corresponding to the target question to be answered based on the text content in the target medical report.

[0006] In a second aspect, one or more embodiments of this specification provide a model training device, the device comprising: a sample acquisition module, configured to acquire a first training sample, the first training sample comprising: a medical report, a question to be answered, and a first medical report interpretation result obtained by answering the question to be answered based on text content in the medical report; a sample construction module, configured to input a thought chain construction prompt word containing the first training sample into a target large language model, so that the target large language model outputs a reference thought chain and a second medical report interpretation result under the supervision of the first medical report interpretation result, wherein the reference thought chain is used to describe the reasoning process of obtaining the second medical report interpretation result corresponding to the question to be answered based on the text content in the medical report; if the second medical report interpretation result is determined to be correct based on the first medical report interpretation result, the reference thought chain is added to the first training sample to generate a second training sample; The model training module is used to train a large language model for interpreting medical reports using the second training sample. The large language model for interpreting medical reports is used to output a target medical report interpretation result and a target thinking chain based on the target medical report and target question to be answered input by the user. The target thinking chain is used to describe the reasoning process of obtaining the target medical report interpretation result corresponding to the target question to be answered based on the text content in the target medical report.

[0007] In a third aspect, one or more embodiments of this specification provide an electronic device comprising: a memory, a processor, and a communication interface; wherein a computer program is stored on the memory, and when the computer program is executed by the processor, the processor can at least implement the model training method described in the first aspect.

[0008] In a fourth aspect, one or more embodiments of this specification provide a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor of an electronic device, the processor can at least implement the model training method described in the first aspect.

[0009] In a fifth aspect, one or more embodiments of this specification provide a computer program product, comprising: a computer program or instructions, which, when executed by a processor of an electronic device, enables the processor to at least implement the model training method described in the first aspect.

[0010] One or more embodiments of this specification provide a model training scheme for training a large language model for medical report interpretation. The training goal is to enable the trained large language model to output a target medical report interpretation result corresponding to the input target medical report and target question to be answered, and also to output a target thought chain describing the reasoning process of obtaining the target medical report interpretation result corresponding to the target question to be answered based on the text content in the target medical report. Thus, on the one hand, the target medical report interpretation result can be used to answer the target question to be answered in a user-friendly manner; on the other hand, the target thought chain can be used to explain the multiple reasoning steps of the large language model for medical report interpretation from the target medical report and the target question to be answered to obtain the target medical report interpretation result, thereby enhancing the user's perceived logical relationship between the target medical report interpretation result and the target medical report and the target question to be answered, improving the interpretability of the target medical report interpretation result and the user's trust in the target medical report interpretation result. In addition, the large language model for medical report interpretation obtains the target medical report interpretation result based on the multiple reasoning steps contained in the target thought chain, which can also improve the accuracy of the reasoning of the large language model for medical report interpretation. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in one or more embodiments of this specification, a brief introduction will be given below to the drawings required for the description of the embodiments. Obviously, the drawings described below are some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 A schematic diagram of a use scenario of a large language model for interpreting medical reports provided in one or more embodiments of this specification; Figure 2 A schematic diagram of a medical report interpretation interface provided in one or more embodiments of this specification; Figure 3 A flowchart of a model training method provided for one or more embodiments of this specification; Figure 4 A scenario diagram of a model training process provided in one or more embodiments of this specification Figure 1 ; Figure 5A scenario diagram of a model training process provided in one or more embodiments of this specification Figure 2 ; Figure 6 A scenario diagram of a model training process provided in one or more embodiments of this specification Figure 3 ; Figure 7 A schematic diagram of the structure of a model training device provided in one or more embodiments of this specification; Figure 8 A schematic diagram of the structure of an electronic device provided in one or more embodiments of this specification. DETAILED DESCRIPTION

[0013] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.

[0014] It should be noted that when one or more embodiments of this specification involve user information, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals must be provided for users to choose to authorize or refuse. In addition, the various models involved in this specification (including but not limited to language models or large models) comply with relevant laws and standards.

[0015] In addition, the step sequence in the following method embodiments is only an example and not a strict limitation.

[0016] The following first describes the concepts involved in one or more embodiments of this specification.

[0017] Large Language Models (LLMs) are AI models learned using a large-scale data corpus and a deep learning framework. They can be used to handle various tasks such as computer vision, speech recognition, machine translation, text summarization, and intelligent question answering. This specification does not limit the number of model parameters supported by LLMs, but aims to meet practical needs.

[0018] A multimodal large language model is one that can process and understand multiple types (modalities) of data. Specifically, in the fields of natural language processing (NLP) and computer vision (CV), multimodal large language models can process not only text data but also other forms of data, such as images, audio, and video.

[0019] Chain-of-Thought (CoT) is a model reasoning method that decomposes the complex task to be solved by a large language model into a series of logically related intermediate steps, guiding the large language model to gradually reason out the answer, thereby improving the model's reasoning ability and the interpretability of the reasoning results, that is, clarifying the model's reasoning logic.

[0020] Prompts are information related to the task to be performed that is fed into the large language model during user interaction with it. Prompts can include information such as the task execution context, execution rules, and input and output formats. Prompts guide the large language model to perform the task and produce the expected results.

[0021] Supervised fine-tuning (SFT) refers to the process of further training a model that has been pre-trained on a certain amount of unlabeled data using a labeled dataset. The goal of supervised fine-tuning is to make the pre-trained model better suited to the execution requirements of a specific task and improve its performance on that task.

[0022] Reinforcement learning (RL), similar to supervised and unsupervised learning, is a machine learning method. In RL, a large language model executes a corresponding decision-making strategy and continuously optimizes it based on feedback to maximize feedback rewards. The decision strategy corresponding to the maximum feedback reward is the target decision-making strategy learned by the large language model. Based on this learned target decision-making strategy, the large language model can achieve high task execution accuracy.

[0023] As mentioned above, in medical scenarios, there is a need to interpret medical reports into content that patients can understand. Based on this, the embodiments of this specification provide a model training method for training a large language model for medical report interpretation that can automatically interpret medical reports.

[0024] To facilitate understanding, this manual first explains the use process of the large language model for interpreting medical reports, and then explains the training process of the large language model for interpreting medical reports.

[0025] It is understood that the large language model for interpreting medical reports provided in the embodiments of this specification can be used not only by patients but also by others (e.g., medical staff) to interpret various medical reports for applications such as clinical auxiliary diagnosis or medical education. For ease of description, in the following embodiments, users of the large language model for interpreting medical reports will be referred to as users.

[0026] Figure 1 This is a schematic diagram of a scenario for using a large language model for interpreting medical reports provided in one or more embodiments of this specification. Figure 1 As shown, in the use scenario of a large language model for medical report interpretation, a client device 110 corresponding to a user and a server device 120 deployed with the large language model for medical report interpretation are involved. The client device 110 and the server device 120 are in communication with each other. Optionally, the client device 110 can be a terminal device such as a laptop, tablet, PC, smartphone, or robot; the server device 120 can be a physical server including an independent host, a virtual server, a cloud server, or a server cluster.

[0027] In an optional embodiment, the process of using the large language model for interpreting medical reports may be: The client device 110 calls the medical report interpretation service provided by the server device 120 and sends a medical report interpretation request to the server device 120. The medical report interpretation request includes: a target medical report to be interpreted and a target question to be answered.

[0028] Server device 120 receives the medical report interpretation request and inputs the target medical report and target question into the medical report interpretation large language model. The large language model is then used to generate: 1) a target thought chain describing the reasoning process by which the large language model, based on the textual content of the target medical report, arrives at a target medical report interpretation result corresponding to the target question; and 2) a target medical report interpretation result corresponding to the input target medical report and target question. The target medical report interpretation result and target thought chain are then fed back to client device 110 for display, thereby completing the interpretation of the target medical report.

[0029] Figure 1 The use of the large language model for interpreting medical reports in the embodiment of this specification is described from the perspective of the device interaction between the client device 110 and the server device 120. Figure 2 , from the perspective of human-computer interaction between the user and the client device 110, the use process of the large language model for interpreting medical reports in the embodiment of this specification is further explained.

[0030] Figure 2 This is a schematic diagram of a medical report interpretation interface provided by one or more embodiments of this specification. It is worth noting that Figure 2 The medical report interpretation interface shown is for illustrative purposes only and is not intended to be limiting.

[0031] like Figure 2 As shown, the medical report interpretation interface provides interactive controls for users to input a target medical report and a target question to be answered. For example, a user can use the "Take a photo" control to call the camera of the client device 110 to take a photo of the target medical report to be interpreted, thereby obtaining an image containing the target medical report, that is, inputting the target medical report. The user can also enter a desired health question in the input box, such as "Please help me interpret this medical report?", thereby inputting the target question to be answered.

[0032] In an optional embodiment, the user can also input the target medical report by, for example, uploading a scanned image or taking a photo of the target medical report. Of course, in addition to the image data format, the target medical report can also be input in a document format, such as Word or PDF. The embodiments of this specification do not limit the method for inputting the target medical report, nor do they limit the data format of the target medical report.

[0033] In response to the target medical report and target question to be answered input by the user, the client device 110 sends a medical report interpretation request containing the target medical report and the target question to be answered to the server device 120, so that the server device 120 outputs the target medical report interpretation result and the target thinking chain through the medical report interpretation large language model according to the medical report interpretation request.

[0034] After receiving the target medical report interpretation results and target thinking chain fed back by the server device 120 in response to the medical report interpretation request, the client device 110 displays the target medical report interpretation results and target thinking chain in the medical report interpretation interface in a preset display format for the user to view.

[0035] In terms of display format, the display format of the target medical report interpretation results and target thought chain can be customized. For example, Figure 2As shown in the figure, the target medical report interpretation result and the target thought chain are displayed in the form of conversation messages, where "AI" is used to represent the server device 120 deployed with the large language model for medical report interpretation. For example, the target medical report interpretation result and the target thought chain can be displayed differently using different visual effects.

[0036] In terms of displayed content, the target thought chain can optionally include at least one continuous reasoning step, which constitutes the reasoning process for deriving the target medical report interpretation result corresponding to the target unanswered question from the text content in the target medical report. The target medical report interpretation result includes, but is not limited to, at least one of the following interpretation results: an explanation of the text content contained in the target medical report, a health problem indicated in the target medical report, and a suggested solution to the health problem. The type of interpretation result included in the target medical report interpretation result can be customized.

[0037] For ease of understanding, let's assume the target medical report is a "Hepatitis B Five-item Test Report" and the target question is "Please help me interpret this medical report?" The text of the "Hepatitis B Five-item Test Report" contains basic information such as the patient's name and the report date (Ms. Li, February 17, 2025), as well as the specific values ​​and reference ranges for the five indicators: HBsAg, HBsAb, HBeAg, HBeAb, and HBcAb.

[0038] Correspondingly, the target thinking chain could be: This is a five-item hepatitis B test report for patient Li, dated February 17, 2025. The report lists five indicators: hepatitis B surface antigen (HBsAg), hepatitis B surface antibody (HBsAb), hepatitis B e antigen (HBeAg), hepatitis B e antibody (HBeAb), and hepatitis B core antibody (HBcAb). Each indicator has a specific value and reference range. First, HBsAg is negative, indicating that she is not currently infected with the hepatitis B virus. Next, HBsAb is also negative, indicating that she has no hepatitis B surface antibodies, meaning she has no immunity to the virus. Further down, HBeAg, HBeAb, and HBcAb are all negative, further confirming that she has neither active hepatitis B infection nor signs of previous infection. In summary, all indicators are within normal ranges, with no abnormalities found. However, given her lack of hepatitis B antibodies, this means she has no protection against the virus and could be infected if exposed. To sum up, there is no abnormality in this hepatitis B five-item test report, but the lack of hepatitis B antibodies is something that needs attention. It is recommended to get vaccinated and maintain good living habits.

[0039] Correspondingly, the target medical report interpretation result could read: Based on the five hepatitis B test reports you provided, all indicators are negative, indicating the following: 1) Negative hepatitis B surface antigen (HBsAg): Indicates no current hepatitis B virus infection. 2) Negative hepatitis B surface antibody (HBsAb): Indicates the lack of antibodies to the hepatitis B virus, possibly due to lack of hepatitis B vaccination or an ineffective immune response. 3) Negative hepatitis B e antigen (HBeAg), hepatitis B e antibody (HBeAb), and hepatitis B core antibody (HBcAb): Further rule out previous infection or occult infection. Overall, you are currently HBV-free, but lack the protection of hepatitis B antibodies, making you susceptible. Future exposure to hepatitis B (e.g., blood contact, sexual contact, etc.) poses a risk of infection. Vaccination and maintaining healthy lifestyle habits are recommended.

[0040] The above describes the process of using the large language model for interpreting medical reports provided in the embodiments of this specification.

[0041] Based on the above description, it can be seen that the large language model for medical report interpretation obtained by the model training method provided in the embodiment of this description can, on the one hand, output the target medical report interpretation result corresponding to the input target medical report and the target question to be answered; on the other hand, it can output the target thinking chain, wherein the target thinking chain is used to describe the reasoning process of the large language model for medical report interpretation to obtain the target medical report interpretation result corresponding to the target question to be answered based on the text content in the target medical report.

[0042] To achieve the aforementioned training objectives, the large language model for medical report interpretation is equipped with the ability to infer a target medical report interpretation result through a target thought chain, and to output the target medical report interpretation result and the target thought chain. One or more embodiments of this specification provide a model training method that automatically constructs thought chains using a target large language model, thereby generating training samples containing a medical report, a question to be answered, a medical report interpretation result, and a thought chain, for training the large language model for medical report interpretation.

[0043] In the examples of this specification, the data format of medical reports involved is an image. In this case, the target large language model and the medical report interpretation large language model are both multimodal large language models capable of processing both the text data corresponding to the question to be answered and the image data corresponding to the medical report.

[0044] The following describes the model training method provided in one or more embodiments of this specification.

[0045] Figure 3A flowchart of a model training method provided in one or more embodiments of this specification, such as Figure 3 As shown, it includes steps 310 to 380: Step 310: Obtain a first training sample, where the first training sample includes: a medical report, a question to be answered, and a first medical report interpretation result obtained by answering the question to be answered based on the text content in the medical report.

[0046] As an optional method for obtaining the first training sample, you can first collect several patients' medical reports, questions that patients have asked medical staff regarding these medical reports, and answers provided by medical staff to patients based on these medical reports. The collected medical reports, questions, and answers are all data that has been desensitized to protect the patient's privacy.

[0047] The medical report contains text content, which is used to record the patient's basic information, medical information, examination conclusions, diagnosis conclusions and treatment recommendations, etc., and can provide a reference for medical staff to answer patients' questions.

[0048] For example, basic patient information, such as name, date of birth, contact information, and report number. Medical information, such as the name of the hospital or department visited, the name of the examination item (e.g., "Hepatitis B Five Items," "Hematology Routine," "CT Scan," etc.), the name of the doctor (the doctor who issued the report or the diagnostician), and the time of the visit / examination. Test results, such as the name of the test indicator (e.g., "HBsAg," "White Blood Cell Count," "Blood Sugar Level"), the measured value (e.g., "Negative," "Positive," "5.6 mmol / L"), the normal range of the indicator (e.g., "0.00-5.00 pg / mL"), and whether the indicator is outside the normal range (e.g., "↑Increased," "↓Decreased," "Normal"), etc. Diagnostic conclusions, such as the presence of a disease or abnormal condition (e.g., "No abnormality found," "Indicates absence of hepatitis B surface antibodies"), possible etiology analysis, or condition assessment. Treatment recommendations, such as the doctor's advice based on the results (such as "hepatitis B vaccination is recommended" and "liver function review"), follow-up treatment recommendations (such as medication, surgery, lifestyle adjustments, etc.), and precautions (such as diet control, avoiding high-risk behaviors, etc.).

[0049] Afterwards, the collected data is preprocessed to verify its integrity and correctness and remove interfering data. For example, in an image containing a medical report, the clarity and completeness of the medical report's content are determined. If unclear or incomplete, the data is discarded. Another example is determining whether the collected questions are complete and relevant to the medical report. If incomplete or irrelevant, the data is discarded. For example, questions like "What's the weather like today?" are completely irrelevant to the medical report and can be discarded.

[0050] Optionally, the data preprocessing process can be implemented using a target large language model. Specifically, a prompt word containing the data to be preprocessed and a description of the preprocessing task can be input into the target large language model, so that the target large language model, guided by the prompt word, can complete verification of the integrity and correctness of the data to be preprocessed.

[0051] Among them, by preprocessing the collected data, the accuracy of the data used for model training can be effectively guaranteed, avoiding the problem of poor model reasoning effect caused by inaccurate data used for model training.

[0052] Next, based on the data preprocessing results, a first training sample is constructed. The first training sample includes: a medical report, a question to be answered, and the interpretation of the first medical report. The question to be answered is the question that the patient has previously asked the medical staff regarding the medical report in the collected data, and the interpretation of the first medical report is the answer provided by the medical staff to the patient based on the text content of the medical report.

[0053] The first training sample is used to construct a reference thought chain using the target large language model, and to generate a second training sample for training the large language model for medical report interpretation. The second training sample includes: a medical report, a question to be answered, the interpretation result of the first medical report, and the reference thought chain.

[0054] In step 320, the prompt words for constructing the thought chain containing the first training sample are input into the target large language model, so that the target large language model outputs a reference thought chain and a second medical report interpretation result under the supervision of the first medical report interpretation result. The reference thought chain is used to describe the reasoning process of obtaining the second medical report interpretation result corresponding to the question to be answered based on the text content in the medical report.

[0055] In the embodiment of this specification, the target large language model is a large language model that performs reasoning through thought chaining.

[0056] It is understandable that the accuracy of the inference results obtained through thought chain reasoning can reflect the accuracy of the thought chain. Based on this, while generating a reference thought chain corresponding to the first training sample through the target large language model, the target large language model can also generate a second medical report interpretation result corresponding to the first training sample.

[0057] Among them, the interpretation result of the second medical report is used to verify the correctness of the reference thinking chain; the reference thinking chain is used to describe the reasoning process of obtaining the interpretation result of the second medical report corresponding to the question to be answered based on the text content in the medical report of the first training sample.

[0058] For ease of description, the prompt words used to guide the target large language model to generate a reference thought chain and the second medical report interpretation results are referred to as: thought chain construction prompt words.

[0059] In an optional embodiment, a thought chain construction prompt word template can be pre-built. This thought chain construction prompt word template includes thought chain construction task description information, output content description information, etc., and has a slot reserved for filling in the first training sample. Therefore, after obtaining the first training sample, the first training sample can be directly filled into the preset slot in the thought chain construction prompt word template to obtain the thought chain construction prompt word.

[0060] For example, the prompt words for constructing a thought chain can be: { “First training sample: medical report, question to be answered, first medical report interpretation result” "Thought Chain Building Task Description: Please use the thought chain reasoning method to answer the question to be answered. Your answer should include multiple steps, each step consisting of three types of actions: 1) Inner Thinking: This is the thinking step. Note that describing comprehensive reasoning requires multiple "Inner Thinking" steps. Each step should begin with a brief title.

[0061] 2) Final Conclusion: Summarize the correct reasoning from the previous "Inner Thinking" step and provide the final answer.

[0062] 3) Verification: Verify the conclusion in the "Final Conclusion" step. If the conclusion is true, end the process. If not, return to the previous step. Output content description: The output format follows the following JSON structure: "COT": {"action":"Inner Thoughts", "title":"...", "content":"..."}, {“action”: “Final Conclusion”, “content”: “…”}, {“action”: “verify”, “content”: “…”}” }.

[0063] It is worth noting that, in actual applications, the thinking chain construction task description information, output content description information, etc. in the thinking chain construction prompt words can be customized and are not limited to the above examples.

[0064] After obtaining the thought chain construction prompt words, the thought chain construction prompt words are input into the target large language model, so that the target large language model generates a reference thought chain and a second medical report interpretation result under the guidance of the thought chain construction prompt words. Specifically, under the supervision of the first medical report interpretation result contained in the first training sample in the thought chain construction prompt words, the target large language model generates a corresponding reference thought chain and a second medical report interpretation result based on the text content and the questions to be answered in the medical report. Among them, the reference thought chain corresponds to the "COT" in the output content of the thought chain construction prompt words in the above example, and the second medical report interpretation result corresponds to the "final conclusion" in the output content of the thought chain construction prompt words in the above example.

[0065] Step 330: If it is determined that the second medical report interpretation result is correct based on the first medical report interpretation result, the reference thought chain is added to the first training sample to generate a second training sample.

[0066] Figure 4 A scenario diagram of a model training process provided in one or more embodiments of this specification Figure 1 .like Figure 4 As shown, after the target large language model outputs the reference thought chain and the second medical report interpretation result, it is further possible to verify whether the second medical report interpretation result is correct.

[0067] Optionally, the first report interpretation result may be used as a reference to verify whether the second medical report interpretation result is correct.

[0068] If the second medical report interpretation is correct based on the first medical report interpretation, then the reference thought chain generated by the target large language model is correct. The reference thought chain can be added to the first training sample to generate a second training sample. The second training sample includes: a medical report, a question to be answered, the first medical report interpretation, and the reference thought chain.

[0069] If the interpretation result of the second medical report is determined to be incorrect based on the interpretation result of the first medical report, it means that the reference thought chain generated by the target large language model is incorrect. A new reference thought chain and a new second medical report interpretation result can be regenerated by the target large language model, and the correctness of the new second medical report interpretation result can be re-verified until the second medical report interpretation result generated by the target large language model is correct, or the number of verifications reaches a set threshold (for example, 4 times). Optionally, for the second medical report interpretation result corresponding to a first training sample, if the number of verifications reaches a set threshold and the verification result of the second medical report interpretation result is still incorrect, the first training sample is discarded.

[0070] In the embodiment of this specification, in the case where the interpretation result of the second medical report is incorrect, a new reference thought chain and a new interpretation result of the second medical report can be regenerated using the target large language model in at least one of the following ways.

[0071] In an optional embodiment, if it is determined that the interpretation result of the second medical report is incorrect based on the interpretation result of the first medical report, the thought chain construction prompt words can be re-input into the target large language model so that the target large language model can re-output a new reference thought chain and a new second medical report interpretation result under the supervision of the first medical report interpretation result.

[0072] In this scheme, by re-inputting the thought chain construction prompt words into the target large language model, the target large language model is guided to try new reasoning logic, and the reference thought chain and the second medical report interpretation results are regenerated with new reasoning thinking, thereby exploring the reference thought chain that can obtain the correct second medical report interpretation results for constructing the second training sample.

[0073] In another optional embodiment, when verifying the correctness of the second medical report interpretation result using the first report interpretation result as a reference, the target large language model may first output difference information between the second medical report interpretation result and the first medical report interpretation result, as well as a verification result of the second medical report interpretation result indicated by the difference information. The difference information may be understood as incorrect information in the second medical report interpretation result relative to the first medical report interpretation result.

[0074] Afterwards, if the distinguishing information indicates that the interpretation result of the second medical report is incorrect, the distinguishing information and the reference thinking chain are filled into the thinking chain construction prompt words; the thinking chain construction prompt words containing the first training sample, the distinguishing information and the reference thinking chain are input into the target large language model, so that the target large language model adjusts the reference thinking chain according to the distinguishing information under the supervision of the first medical report interpretation result, so as to re-output a new reference thinking chain and a new second medical report interpretation result.

[0075] In this scheme, by adding distinguishing information and the generated reference thought chain to the original thought chain construction prompt words, the target large language model is helped to trace back possible errors in the previous reasoning logic, thereby correcting possible errors in the generated reference thought chain to regenerate a reference thought chain that can obtain the correct second medical report interpretation result for constructing the second training sample.

[0076] Figure 4 In the illustrated scenario, the correctness verification of the interpretation result of the second medical report after the target large language model outputs the reference thinking chain and the interpretation result of the second medical report is taken as an example to illustrate.

[0077] In actual applications, optionally, if the thought chain construction prompt contains information for guiding the target large language model to verify the correctness of the second medical report interpretation result, such as the "verification" information contained in the thought chain construction prompt in the above example, then in addition to generating a reference thought chain and the second medical report interpretation result, the target large language model can further verify the correctness of the second medical report interpretation result based on the first medical report interpretation result, and its verification result corresponds to the "verification" in the output content of the thought chain construction prompt in the above example. Thus, the target large language model can output the verification result while outputting the reference thought chain and the second medical report interpretation result. The embodiment of this specification does not limit the timing and method of verifying the second medical report interpretation result.

[0078] In the above embodiment, the correctness of the reference thought chain is verified by using the first medical report interpretation result as a reference to verify the second medical report interpretation result.

[0079] To further ensure the accuracy of the reference thought chain, in an optional embodiment, after determining that the interpretation result of the second medical report is correct based on the interpretation result of the first medical report, the following steps may be performed: The target large language model outputs the semantic match between the reference thought chain and the text content of the medical report of the first training sample, the question to be answered, and the interpretation result of the first medical report. If the semantic match is greater than a set threshold, the reference thought chain is added to the first training sample to generate a second training sample. If the semantic match is less than or equal to the set threshold, the target large language model re-outputs a new reference thought chain and a new interpretation result of the second medical report.

[0080] The semantic matching degree can be the semantic matching degree between the reference thought chain and the text content in the medical report of the first training sample, the question to be answered, and the interpretation result of the first medical report, respectively. For example, the first semantic matching degree between the reference thought chain and the text content in the medical report of the first training sample, the second semantic matching degree between the reference thought chain and the question to be answered, and the third semantic matching degree between the reference thought chain and the interpretation result of the first medical report. Optionally, different semantic matching degrees can correspond to the same or different set thresholds, and the size of the set threshold can be customized.

[0081] In this scheme, by determining the semantic matching degree between the reference thinking chain and the text content in the medical report of the first training sample, the questions to be answered and the interpretation results of the first medical report, the correlation between the reference thinking chain and the text content in the medical report, the questions to be answered and the interpretation results of the first medical report can be ensured, thereby ensuring the correctness of the second training sample obtained based on the first training sample and the reference thinking chain.

[0082] Step 340: Train the large language model for medical report interpretation using the second training sample. The large language model for medical report interpretation is used to output the target medical report interpretation result and the target thinking chain based on the target medical report and the target question to be answered input by the user. The target thinking chain is used to describe the reasoning process of obtaining the target medical report interpretation result corresponding to the target question to be answered based on the text content in the target medical report.

[0083] After obtaining the second training sample, the second training sample is used to perform supervised fine-tuning on the medical report interpretation large language model to be trained. Optionally, in order to improve training efficiency, the medical report interpretation large language model can be a pre-trained model.

[0084] It's understandable that the training and usage processes of the large language model for medical report interpretation are similar. In conjunction with the previous description of the large language model's usage, when training the large language model using the second training sample, its input is the medical report and question to be answered contained in the second training sample, and its output is the thought chain and the medical report interpretation result. For ease of description, the outputs of the large language model during training will be referred to as the predicted thought chain and the predicted medical report interpretation result.

[0085] In actual applications, the large language model for medical report interpretation performs corresponding tasks under the guidance of prompt words. In the embodiments of this specification, the prompt words used to guide the large language model for medical report interpretation to generate predictive thinking chains and predict medical report interpretation results are called: report interpretation prompt words.

[0086] In an optional embodiment, a report interpretation prompt word template can be pre-built. This template includes information describing the report interpretation task and output content, and has reserved slots for filling in the medical reports and unanswered questions in the second training sample. Thus, after obtaining the second training sample, the medical reports and unanswered questions in the second training sample can be filled into the preset slots in the report interpretation prompt word template to obtain the report interpretation prompt words.

[0087] Figure 5 A scenario diagram of a model training process provided in one or more embodiments of this specification Figure 2 .like Figure 5As shown, after obtaining the report interpretation prompt words, the report interpretation prompt words containing the medical report and the question to be answered can be input into the medical report interpretation large language model, so that the medical report interpretation large language model outputs the predicted medical report interpretation result and the predicted thought chain. Afterwards, with the goal of improving the first similarity between the predicted medical report interpretation result and the first medical report interpretation result and the second similarity between the predicted thought chain and the reference thought chain, a loss function is constructed. Optionally, the first similarity, the second similarity, etc. can be, for example, cosine similarity, Euclidean distance, etc. Finally, according to the loss function, the medical report interpretation large language model is trained. For example, when the loss value corresponding to the loss function is less than a certain value, the training of the medical report interpretation large language model is terminated.

[0088] In an optional embodiment, after the medical report interpretation large language model is trained using the second training sample, the medical report interpretation large language model can continue to be trained during the use of the medical report interpretation large language model to further improve its reasoning accuracy.

[0089] Figure 6 A scenario diagram of a model training process provided in one or more embodiments of this specification Figure 3 ,like Figure 6 As shown, during the use of the large language model for medical report interpretation, feedback data corresponding to the target medical report interpretation results and target thinking chains from users can also be obtained; then, the feedback data is used to retrain the large language model for medical report interpretation.

[0090] Among them, feedback data refers to the data generated when performing medical report interpretation tasks due to failure to provide correct or expected target medical report interpretation results and target thinking chains, which is commonly known as bad cases, referring to bad cases or error cases.

[0091] During the specific implementation process, the feedback data can be analyzed to determine the reasons why the medical report interpretation large language model fails to give the correct or expected target medical report interpretation results and target thinking chain results when performing the medical report interpretation task, so as to guide the further training of the medical report interpretation large language model.

[0092] Optionally, the feedback data may also include: reasons why the large language model for interpreting medical reports reported by the user fails to provide the correct or expected target medical report interpretation results and target thought chain results, etc.

[0093] As an optional method of training a large language model for medical report interpretation based on feedback data, the target medical report interpretation results and target thinking chain results actually output by the large language model for medical report interpretation can be corrected according to the reasons why the large language model for medical report interpretation fails to give the correct or expected target medical report interpretation results and target thinking chain results, such as: correcting a certain reasoning decision result in the target thinking chain, etc.; then, based on the corrected target medical report interpretation results and target thinking chain results, as well as the corresponding target medical report and target question to be answered, a third training sample for training the large language model for medical report interpretation is constructed, so as to train the large language model for medical report interpretation through the third training sample.

[0094] In this solution, the feedback data generated by the medical report interpretation large language model is used to train the medical report interpretation large language model, and more training samples are introduced, thereby better improving the reasoning ability of the medical report interpretation task.

[0095] Understandably, the medical report interpretation language model continuously generates feedback data during user use. If each new piece of feedback triggers training of the medical report interpretation language model, this would increase model training costs and hinder the improvement of the medical report interpretation language model's capabilities.

[0096] Optionally, during the use of the large language model for medical report interpretation, the data volume of the currently generated feedback data can be counted. In response to the data volume of the feedback data being greater than the set data volume threshold, the large language model for medical report interpretation can be trained based on the feedback data, for example: guiding the large language model for medical report interpretation to perform reinforcement learning.

[0097] Optionally, during the use of the large language model for medical report interpretation, the time interval between the historical time point corresponding to the previous model training of the large language model for medical report interpretation based on feedback data and the current time point can also be counted. In response to the time interval between the historical time point corresponding to the previous reinforcement learning of the large language model for medical report interpretation and the current time point being greater than the set time threshold, the large language model for medical report interpretation is trained according to the feedback data, for example: guiding the large language model for medical report interpretation to perform reinforcement learning.

[0098] In this solution, by counting the amount of feedback data, or the time interval between the time point corresponding to the previous model training of the medical report interpretation large language model guided by the feedback data and the current time point, the timing of training the medical report interpretation large language model through the feedback data can be reasonably controlled. This can not only reduce the training cost of the medical report interpretation large language model, but also help improve the model capabilities of the medical report interpretation large language model.

[0099] In summary, in the model training scheme provided by one or more embodiments of this specification, the target large language model can automatically construct a reference thought chain based on the first training sample, thereby generating a second training sample containing a medical report, a question to be answered, a medical report interpretation result and a reference thought chain, for training the medical report interpretation large language model. The medical report interpretation large language model trained based on the second training sample not only learns relevant medical knowledge, but also masters the ability to reason through thought chains. Therefore, whether it is a single target medical report or multiple target medical reports, the medical report interpretation large language model can integrate the logical relationship between the text content in the target medical report and the target question to be answered through thought chains, ensuring the integrity of information integration and the accuracy of reasoning logic. Based on this, the accuracy of the target medical report interpretation result obtained by using the reasoning steps corresponding to the target thought chain can also be improved. In addition, the output content of the medical report interpretation large language model in the embodiment of this specification includes both the target medical report interpretation results and the target thinking chain. Therefore, the user can not only obtain the answers to the questions they want to know, but also know the reasoning ideas of the medical report interpretation large language model. This not only enhances the interpretability of the target medical report interpretation results and the user's trust in the target medical report interpretation results, but also provides users with corresponding teaching functions, expanding the use scenarios of the medical report interpretation large language model.

[0100] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0101] The following will describe in detail the model training device of one or more embodiments of this specification. Those skilled in the art will understand that these devices can be configured using commercially available hardware components through the steps taught in this solution.

[0102] Figure 7A schematic diagram of the structure of a model training device provided in one or more embodiments of this specification, such as Figure 7 As shown, the device includes: a sample acquisition module 710, a sample construction module 720, and a model training module 730.

[0103] The sample acquisition module 710 is configured to acquire a first training sample, where the first training sample includes: a medical report, a question to be answered, and a first medical report interpretation result obtained by answering the question to be answered based on text content in the medical report.

[0104] The sample construction module 720 is used to input the thought chain construction prompt words containing the first training sample into the target large language model, so that the target large language model outputs a reference thought chain and a second medical report interpretation result under the supervision of the first medical report interpretation result. The reference thought chain is used to describe the reasoning process of obtaining the second medical report interpretation result corresponding to the question to be answered based on the text content in the medical report; if it is determined that the second medical report interpretation result is correct based on the first medical report interpretation result, the reference thought chain is added to the first training sample to generate a second training sample.

[0105] The model training module 730 is used to train the medical report interpretation large language model through the second training sample. The medical report interpretation large language model is used to output the target medical report interpretation result and the target thinking chain based on the target medical report and the target question to be answered input by the user. The target thinking chain is used to describe the reasoning process of obtaining the target medical report interpretation result corresponding to the target question to be answered based on the text content in the target medical report.

[0106] In an optional embodiment, the sample construction module 720 is also used to: if it is determined that the second medical report interpretation result is incorrect based on the first medical report interpretation result, the thought chain construction prompt word is re-input into the target large language model, so that the target large language model re-outputs a new reference thought chain and a new second medical report interpretation result under the supervision of the first medical report interpretation result.

[0107] In an optional embodiment, the sample construction module 720 is also used to: output the difference information between the second medical report interpretation result and the first medical report interpretation result through the target large language model, and the verification result of the second medical report interpretation result indicated by the difference information; if the difference information indicates that the second medical report interpretation result is incorrect, the difference information and the reference thinking chain are filled into the thinking chain construction prompt word; the thinking chain construction prompt word containing the first training sample, the difference information and the reference thinking chain is input into the target large language model, so that the target large language model adjusts the reference thinking chain according to the difference information under the supervision of the first medical report interpretation result, so as to re-output a new reference thinking chain and a new second medical report interpretation result.

[0108] In an optional embodiment, the sample construction module 720 is also used to: output the semantic matching degree between the reference thought chain and the text content in the medical report, the question to be answered and the first medical report interpretation result through the target large language model; if the semantic matching degree is greater than a set threshold, add the reference thought chain to the first training sample to generate a second training sample; if the semantic matching degree is less than or equal to the set threshold, re-output a new reference thought chain and a new second medical report interpretation result through the target large language model.

[0109] In an optional embodiment, the model training module 730 is specifically used to: input the report interpretation prompt words containing the medical report and the question to be answered into the medical report interpretation large language model, so that the medical report interpretation large language model outputs a predicted medical report interpretation result and a predicted thinking chain; construct a loss function with the goal of improving the first similarity between the predicted medical report interpretation result and the first medical report interpretation result and the second similarity between the predicted thinking chain and the reference thinking chain; and train the medical report interpretation large language model according to the loss function.

[0110] In an optional embodiment, the model training module 730 is further used to: obtain feedback data corresponding to the target medical report interpretation result and the target thinking chain from the user during the use of the medical report interpretation large language model; in response to the data volume of the feedback data being greater than a set data volume threshold, perform reinforcement learning on the medical report interpretation large language model according to the feedback data; or, in response to the time interval between the historical time point corresponding to the previous reinforcement learning of the medical report interpretation large language model and the current time point being greater than a set time threshold, perform reinforcement learning on the medical report interpretation large language model according to the feedback data.

[0111] In an optional embodiment, the data format of the medical report is an image.

[0112] Figure 7 The illustrated apparatus can execute the steps described in the aforementioned embodiments. The various embodiments in this specification are described in a progressive manner, and similar portions between the various embodiments can be referenced. Each embodiment focuses on the differences from the other embodiments. In particular, the apparatus embodiments are generally similar to the method embodiments, so their description is relatively simple. The detailed execution process and technical effects are described in the aforementioned method embodiments and will not be repeated here.

[0113] In one possible design, the above Figure 7 The structure of the model training device shown can be realized as an electronic device, such as Figure 8 As shown, the electronic device may include: a memory 810, a processor 820, and a communication interface 830. The memory 810 stores a computer program, and when the computer program is executed by the processor 820, the processor 820 can at least implement the model training method provided in the aforementioned embodiment.

[0114] The memory 810 may be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0115] Accordingly, one or more embodiments of the present specification further provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-mentioned method embodiments. The computer-readable storage medium may be volatile or non-volatile, or a combination thereof, and may be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices, or any other non-transmission medium. Accordingly, one or more embodiments of this specification also provide a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the processor is enabled to implement the various steps in the above-mentioned method embodiments. It should be understood that each process or a combination of multiple processes in the above-mentioned method flow can be implemented by a computer program or instruction. In addition, these computer programs or instructions can be applied to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable model training device, so that the processor of the general-purpose computer, the special-purpose computer, the embedded processor or other programmable model training device can be implemented as a device for implementing the corresponding functions in the above-mentioned method embodiments.

[0116] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Those skilled in the art can understand and implement the present invention without inventive effort.

[0117] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented by adding the necessary general-purpose hardware platform, or of course, by combining hardware and software. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a computer product. This specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0118] Finally, it should be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a..." does not preclude the presence of additional identical elements in the process, method, commodity, or apparatus that includes the element.

[0119] The above are merely examples of the present invention and are not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A model training method, characterized in that: The method comprises: Obtaining a first training sample, the first training sample comprising: a medical report, a question to be answered, and a first medical report interpretation result obtained by answering the question to be answered based on text content in the medical report; Inputting the prompt words for constructing a thought chain containing the first training sample into a target large language model, so that the target large language model outputs a reference thought chain and a second medical report interpretation result under the supervision of the first medical report interpretation result, wherein the reference thought chain is used to describe the reasoning process of obtaining the second medical report interpretation result corresponding to the question to be answered based on the text content in the medical report; If it is determined that the second medical report interpretation result is correct according to the first medical report interpretation result, the reference thought chain is added to the first training sample to generate a second training sample; The large language model for medical report interpretation is trained using the second training sample. The large language model for medical report interpretation is used to output a target medical report interpretation result and a target thinking chain based on the target medical report and target question to be answered input by the user. The target thinking chain is used to describe the reasoning process of obtaining the target medical report interpretation result corresponding to the target question to be answered based on the text content in the target medical report.

2. The method according to claim 1, characterized in that The method further comprises: If it is determined that the interpretation result of the second medical report is incorrect based on the interpretation result of the first medical report, the thought chain construction prompt words are re-input into the target large language model, so that the target large language model re-outputs a new reference thought chain and a new second medical report interpretation result under the supervision of the interpretation result of the first medical report.

3. The method according to claim 1, characterized in that The method further comprises: outputting, through the target large language model, difference information between the second medical report interpretation result and the first medical report interpretation result, and a verification result of the second medical report interpretation result indicated by the difference information; If the distinguishing information indicates that the interpretation result of the second medical report is incorrect, the distinguishing information and the reference thinking chain are filled into the thinking chain construction prompt word; The thought chain construction prompt words including the first training sample, the distinction information and the reference thought chain are input into the target large language model, so that the target large language model adjusts the reference thought chain according to the distinction information under the supervision of the first medical report interpretation result, so as to re-output a new reference thought chain and a new second medical report interpretation result.

4. The method according to claim 1, wherein After determining that the second medical report interpretation result is correct according to the first medical report interpretation result, the method further includes: Outputting, through the target large language model, the semantic matching degree between the reference thought chain and the text content in the medical report, the question to be answered, and the interpretation result of the first medical report; If the semantic matching degree is greater than a set threshold, adding the reference thought chain to the first training sample to generate a second training sample; If the semantic matching degree is less than or equal to the set threshold, a new reference thought chain and a new second medical report interpretation result are re-output through the target large language model.

5. The method according to claim 1, wherein The step of training the large language model for interpreting medical reports using the second training sample includes: Inputting report interpretation prompt words including the medical report and the question to be answered into a medical report interpretation large language model, so that the medical report interpretation large language model outputs a predicted medical report interpretation result and a predicted thought chain; constructing a loss function with the goal of improving a first similarity between the predicted medical report interpretation result and the first medical report interpretation result and a second similarity between the predicted thought chain and the reference thought chain; The medical report interpretation large language model is trained according to the loss function.

6. The method according to claim 1, characterized in that The method further comprises: During the use of the large language model for medical report interpretation, obtaining feedback data from the user corresponding to the target medical report interpretation result and the target thought chain; In response to the amount of the feedback data being greater than a set data amount threshold, performing reinforcement learning on the large language model for interpreting medical reports based on the feedback data; or In response to the time interval between the historical time point corresponding to the previous reinforcement learning of the medical report interpretation large language model and the current time point being greater than a set time threshold, reinforcement learning is performed on the medical report interpretation large language model based on the feedback data.

7. The method according to any one of claims 1 to 6, characterized in that The data format of the medical report is an image.

8. A model training device, characterized in that: The device comprises: a sample acquisition module, configured to acquire a first training sample, the first training sample comprising: a medical report, a question to be answered, and a first medical report interpretation result obtained by answering the question to be answered based on text content in the medical report; a sample construction module, configured to input a thought chain construction prompt word containing the first training sample into a target large language model, so that the target large language model outputs a reference thought chain and a second medical report interpretation result under the supervision of the first medical report interpretation result, wherein the reference thought chain is used to describe the reasoning process of obtaining the second medical report interpretation result corresponding to the question to be answered based on the text content in the medical report; if the second medical report interpretation result is determined to be correct based on the first medical report interpretation result, the reference thought chain is added to the first training sample to generate a second training sample; The model training module is used to train a large language model for interpreting medical reports using the second training sample. The large language model for interpreting medical reports is used to output a target medical report interpretation result and a target thinking chain based on the target medical report and target question to be answered input by the user. The target thinking chain is used to describe the reasoning process of obtaining the target medical report interpretation result corresponding to the target question to be answered based on the text content in the target medical report.

9. An electronic device, characterized in that: include: A memory, a processor, and a communication interface; wherein a computer program is stored on the memory, and when the computer program is executed by the processor, the processor executes the model training method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor of an electronic device, causes the processor to execute the model training method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Language model reinforcement learning training extension method and device based on reflection

    CN120146142A

  • Method and device for multi-stage generation of medical image question and answer thinking chain data

    CN120317386A

  • Interpreting Natural Language Comparisons During Visual Analysis

    US20240362261A1

  • Tuning generative models using latent-variable inference

    US20240386202A1

  • Optimizing large language models with meta learning and chain of thought

    US20250200387A1